Local CAPTCHA Recognition
The Hard Part Is Not the Model, It Is the Segmentation
A local recogniser that reads a challenge image in your own process has an advantage no remote solving service can match: the bytes never leave the machine, the answer arrives without a network round trip, and you can look at every intermediate step. That last property is the real one. A hosted service hands you a string; a local pipeline hands you a reason, and a reason is what you need when the same image starts failing.
The disadvantage is that you own the entire pipeline, and almost all of the failures live in the two stages before any classification happens. Deciding where the grid lines are, and deciding where the ink stops, account for far more lost accuracy than the label decision does. This is the self-hosted counterpart to the service trade-off in CAPTCHA Solving and to the flow plumbing in CAPTCHAs in Session Flows. The snippet below builds a synthetic buffer, removes its frame, thresholds it, finds the cells from the density profile, counts connected components and enclosed regions per cell, and turns a class into a click or a skip. Every stage prints what it decided, which is the only way to debug this class of code.
Getting Bytes Without Guessing at Coordinates
A challenge image arrives as a buffer. The temptation is to slice it into equal tiles by arithmetic, which works exactly until the pitch changes, a margin is added, or the grid is not rectangular. The snippet instead finds the frame first by scanning rows and columns for full-width and full-height ink, cropping it, and printing what it removed.
After the crop the buffer is 234 by 114 with a three-pixel border gone from each side. That number is worth having: it is the difference between coordinates you computed and coordinates the image actually has, and getting it wrong shifts every subsequent cell by three pixels in both axes. The frame is also the easiest thing to mistake for content, because a full-width border looks exactly like a very dense row of glyphs until you remember that no glyph spans the whole edge.
Otsu Thresholding Because Your Threshold Is Also Wrong
Hand-tuned thresholds fail in a specific way: they are tuned on the images that worked. Otsu's method takes no such input. It builds a histogram of the intensity values, walks every split point, and picks the one that maximises between-class variance, which for a two-population image is the valley between paper and ink.
In the run below the computed threshold is 120 against a paper value near 236 and ink at 30, and it lands exactly on the anti-aliased rim the strokes carry. That is the point of computing it: the rim is the hardest boundary to place by hand, and it is precisely the pixels that decide whether a stroke is three pixels wide or five. A threshold that is a few counts too low swallows the rim and thickens every stroke; a few counts too high drops the thinnest part of a stroke and can split one glyph into two components.
Finding the Grid from the Density Profile
The cells come from the ink-density profile along each axis, not from assumed geometry. Columns whose density exceeds 85 percent are separators, and everything between two separators is a cell. The same runs on rows.
The output reports both the separators found and the resulting cell extents, so a challenge that changes its pitch still segments. If you find yourself hard-coding a cell width anywhere, that is the bug: a challenge family that serves a 5-column grid on one path and 4 on another will break a fixed width and nothing else. The density profile also degrades gracefully: a separator that is one pixel off still shows up as a local maximum, and the extents it produces are still usable, which is not true of a fixed arithmetic division.
Connected Components, Then the Euler Number
Each cell is scanned with a flood fill over ink pixels, and components below six pixels are dropped as noise. That gives blob count and largest blob, which separates an empty cell from a glyph.
It does not separate an L from an O, because both are one blob of similar size. The feature that does is the Euler number: count background regions enclosed by ink that do not touch the cell border. An open stroke has none. A closed ring has exactly one. The snippet reports the count per cell, and the decoys separate cleanly on that column alone with no trained model involved. This is the cheapest kind of shape feature and it generalises: any glyph with a counter is separable from any glyph without one by the same test, on any grid, with no labelling work.
'''Local grid-CAPTCHA segmentation over a synthesised image buffer.
The buffer is generated with integer arithmetic inside this snippet: no image
library, no file, no network. The pipeline is the one a self-hosted solver uses
for a simple grid: frame removal, Otsu binarisation, separator detection from the
ink-density profile, connected-component labelling per cell, then a confidence
band on the ink fraction that decides whether to click.
Deterministic by construction: the only entropy is fixed-seed random.Random, so
every table below is identical on every run.
'''
from collections import deque
import random
W, H = 240, 120
COLS, ROWS = 4, 2
CW, CH = W // COLS, H // ROWS
INK, PAPER = 30, 236 # 8-bit grayscale: dark stroke on pale paper
MIN_BLOB = 6 # components below this size are treated as noise
SEP_DENSITY = 0.85 # a profile line this dense is a separator, not a cell
# ---------------------------------------------------------------- synthesise
GRAY = [[PAPER] * W for _ in range(H)]
def box(x0, y0, x1, y1, value):
for y in range(max(y0, 0), min(y1, H)):
row = GRAY[y]
for x in range(max(x0, 0), min(x1, W)):
row[x] = value
HALO = 120 # the anti-aliased rim every stroke carries in a real capture
def vline(x, y0, y1, value, width=3, halo=True):
if halo:
box(x - width // 2 - 1, y0 - 1, x + width // 2 + 2, y1 + 1, HALO)
box(x - width // 2, y0, x + width // 2 + 1, y1, value)
def hline(y, x0, x1, value, width=3, halo=True):
if halo:
box(x0 - 1, y - width // 2 - 1, x1 + 2, y + width // 2 + 2, HALO)
box(x0, y - width // 2, x1 + 1, y + width // 2 + 1, value)
box(0, 0, W, 3, INK)
box(0, H - 3, W, H, INK)
box(0, 0, 3, H, INK)
box(W - 3, 0, W, H, INK)
for c in range(1, COLS):
vline(c * CW, 4, H - 4, INK, 3, False)
for r in range(1, ROWS):
hline(r * CH, 4, W - 4, INK, 3, False)
# The prompt asks for the cells drawn with the L stroke. Decoys use the O
# stroke, and one cell is left blank apart from background speckle.
TARGETS = {(0, 0), (0, 3), (1, 1), (1, 2)}
DECOYS = {(0, 1), (0, 2), (1, 0)}
SPECKLED = (1, 3)
def stroke_l(cx, cy):
vline(cx - 9, cy - 11, cy + 12, INK)
hline(cy + 11, cx - 9, cx + 9, INK)
def stroke_o(cx, cy):
vline(cx - 11, cy - 12, cy + 13, INK)
vline(cx + 9, cy - 12, cy + 13, INK)
hline(cy - 11, cx - 11, cx + 10, INK)
hline(cy + 11, cx - 11, cx + 10, INK)
STROKES = {"L": stroke_l, "O": stroke_o}
for r in range(ROWS):
for c in range(COLS):
if (r, c) in TARGETS or (r, c) in DECOYS:
jitter = random.Random(20261001 + r * 8 + c)
STROKES["L" if (r, c) in TARGETS else "O"](
c * CW + CW // 2 + jitter.randint(-3, 3),
r * CH + CH // 2 + jitter.randint(-2, 2))
SPECKS = [(57, 6), (118, 60), (119, 60), (61, 113), (178, 7), (179, 7)]
speckle = random.Random(4242)
sr, sc = SPECKLED
for _ in range(45):
SPECKS.append((sc * CW + 6 + speckle.randint(0, 8),
sr * CH + 6 + speckle.randint(0, 8)))
for sx, sy in SPECKS:
GRAY[sy][sx] = INK
# Sensor noise: the paper is never a flat value in a real capture.
grain = random.Random(7)
for y in range(H):
row = GRAY[y]
for x in range(W):
if row[x] == PAPER:
row[x] = PAPER + grain.randint(-8, 8)
# ------------------------------------------------------------ frame removal
# A real capture has a solid frame drawn around the grid; the profile is
# meaningless until it is gone. Scan inward while a row or column is solid.
def row_solid(row):
return all(v == row[0] for v in row)
def col_solid(plane, x):
return all(plane[y][x] == plane[0][x] for y in range(len(plane)))
def inset(count, probe):
n = 0
while n < count and probe(n):
n += 1
return n
top = inset(H, lambda k: row_solid(GRAY[k]))
bottom = inset(H, lambda k: row_solid(GRAY[H - 1 - k]))
left = inset(W, lambda k: col_solid(GRAY, k))
right = inset(W, lambda k: col_solid(GRAY, W - 1 - k))
PLANE = [row[left:W - right] for row in GRAY[top:H - bottom]]
PH, PW = len(PLANE), len(PLANE[0])
# ---------------------------------------------------------------- binarise
def otsu(plane):
'''Threshold maximising between-class variance. Deterministic, integer input.'''
hist = [0] * 256
for row in plane:
for value in row:
hist[value] += 1
total = sum(hist)
sum_all = sum(i * hist[i] for i in range(256))
sum_b = weight_b = 0
best, best_t = -1.0, 0
for t in range(256):
weight_b += hist[t]
if weight_b == 0:
continue
weight_f = total - weight_b
if weight_f == 0:
break
sum_b += t * hist[t]
mean_b = sum_b / float(weight_b)
mean_f = (sum_all - sum_b) / float(weight_f)
between = weight_b * weight_f * (mean_b - mean_f) ** 2
if between > best:
best, best_t = between, t
return best_t
THRESH = otsu(PLANE)
BINARY = [[1 if v <= THRESH else 0 for v in row] for row in PLANE]
# ----------------------------------------------------- separators and cells
def mask_runs(flags):
out, start = [], None
for i, v in enumerate(flags):
if v and start is None:
start = i
elif not v and start is not None:
out.append((start, i - 1))
start = None
if start is not None:
out.append((start, len(flags) - 1))
return out
def gaps_between(runs_found, span):
out, start = [], 0
for a, b in runs_found:
if a - 1 >= start:
out.append((start, a - 1))
start = b + 1
if start <= span - 1:
out.append((start, span - 1))
return out
def split_cells(profile):
'''A profile line this dense across the axis is a separator, so cells are the
gaps. Returns (separators, cells).'''
peak = max(profile)
if peak == 0:
return [], [(0, len(profile) - 1)]
seps = mask_runs([v >= SEP_DENSITY * peak for v in profile])
return seps, gaps_between(seps, len(profile))
col_profile = [sum(BINARY[y][x] for y in range(PH)) for x in range(PW)]
row_profile = [sum(BINARY[y]) for y in range(PH)]
col_seps, col_cells = split_cells(col_profile)
row_seps, row_cells = split_cells(row_profile)
# ------------------------------------------------- per-cell components
def blobs_in(plane, threshold, x0, y0, x1, y1, keep_min):
'''Connected components, 4-connectivity, iterative flood fill.'''
seen, sizes = set(), []
for y in range(y0, y1 + 1):
for x in range(x0, x1 + 1):
if plane[y][x] > threshold or (x, y) in seen:
continue
queue, size = deque([(x, y)]), 0
seen.add((x, y))
while queue:
cx, cy = queue.popleft()
size += 1
for nx, ny in ((cx - 1, cy), (cx + 1, cy), (cx, cy - 1), (cx, cy + 1)):
if x0 <= nx <= x1 and y0 <= ny <= y1 and \
plane[ny][nx] <= threshold and (nx, ny) not in seen:
seen.add((nx, ny))
queue.append((nx, ny))
sizes.append(size)
return sorted([s for s in sizes if s >= keep_min], reverse=True)
def holes_in(plane, threshold, x0, y0, x1, y1):
'''Count enclosed background regions inside a cell.
A glyph with a counter (the O) has one; an open stroke (the L) has none.
This is the Euler number of the cell's ink mask, and it is the cheapest
feature that separates the two classes without a trained model.
'''
outside = set()
border = [(x, y0) for x in range(x0, x1 + 1)]
border += [(x, y1) for x in range(x0, x1 + 1)]
border += [(x0, y) for y in range(y0, y1 + 1)]
border += [(x1, y) for y in range(y0, y1 + 1)]
stack = deque(pt for pt in border if plane[pt[1]][pt[0]] > threshold)
outside.update(stack)
while stack:
cx, cy = stack.popleft()
for nx, ny in ((cx - 1, cy), (cx + 1, cy), (cx, cy - 1), (cx, cy + 1)):
if x0 <= nx <= x1 and y0 <= ny <= y1 and plane[ny][nx] > threshold and (nx, ny) not in outside:
outside.add((nx, ny))
stack.append((nx, ny))
holes = 0
unvisited = {(x, y) for y in range(y0, y1 + 1) for x in range(x0, x1 + 1)
if plane[y][x] > threshold and (x, y) not in outside}
while unvisited:
holes += 1
seed = unvisited.pop()
stack = deque([seed])
while stack:
cx, cy = stack.popleft()
for nx, ny in ((cx - 1, cy), (cx + 1, cy), (cx, cy - 1), (cx, cy + 1)):
if (nx, ny) in unvisited:
unvisited.discard((nx, ny))
stack.append((nx, ny))
return holes
BANDS = [
(0.000, 0.012, "empty", 0.97, "skip", "no blob above the noise floor"),
(0.012, 0.030, "speckle", 0.42, "skip", "blobs present, all below the floor"),
(0.030, 0.150, "stroke", 0.93, "click", "one dominant blob, glyph sized"),
(0.150, 0.400, "busy", 0.71, "review", "overlapping blobs, want a second look"),
(0.400, 1.001, "saturated", 0.38, "review", "the threshold took the cell ground"),
]
TARGET_HOLE_SHAPE = "L-like" # what the prompt asks for: no enclosed counter
SHAPE_BY_HOLES = {0: "L-like", 1: "O-like", 2: "double"}
def classify(fraction, holes):
'''Ink fraction picks the band; the Euler number picks the class within it.'''
for low, high, label, confidence, decision, why in BANDS:
if low <= fraction < high:
break
else:
return "unknown", 0.0, "review", "outside every band"
if label != "stroke":
return label, confidence, decision, why
shape = SHAPE_BY_HOLES.get(holes, "open-{}".format(holes))
if shape != TARGET_HOLE_SHAPE:
return "stroke:" + shape, confidence, "skip", "wrong class for the prompt"
return "stroke:" + shape, confidence, "click", why
def survey(plane, threshold):
'''Segment and classify the plane; returns one record per cell.'''
binary = [[1 if v <= threshold else 0 for v in row] for row in plane]
height, width = len(plane), len(plane[0])
cols = [sum(binary[y][x] for y in range(height)) for x in range(width)]
rrows = [sum(binary[y]) for y in range(height)]
_cs, ccells = split_cells(cols)
_rs, rcells = split_cells(rrows)
out = []
for ri, (y0, y1) in enumerate(rcells):
for ci, (x0, x1) in enumerate(ccells):
area = (x1 - x0 + 1) * (y1 - y0 + 1)
ink = sum(binary[y][x] for y in range(y0, y1 + 1) for x in range(x0, x1 + 1))
kept = blobs_in(plane, threshold, x0, y0, x1, y1, MIN_BLOB)
holes = holes_in(plane, threshold, x0, y0, x1, y1)
label, confidence, decision, why = classify(ink / float(area), holes)
out.append((ri, ci, ink, ink / float(area), kept, holes, label,
confidence, decision, why))
return out
cells = survey(PLANE, THRESH)
print("local grid segmentation: {}x{} buffer, {}x{} cells, generated in this snippet".format(
W, H, COLS, ROWS))
print("frame removed: {} rows top, {} bottom, {} columns each side; cropped plane {}x{}".format(
top, bottom, left, PW, PH))
print("otsu threshold {}, minimum blob {} px, background {} ink {}, separator density {}".format(
THRESH, MIN_BLOB, PAPER, INK, SEP_DENSITY))
print()
print("cells come from the ink-density profile, not from assumed geometry, so a")
print("challenge that changes its grid pitch still segments.")
print()
print("{:<8} {:>6} {:>5} {:<28} {}".format("axis", "seps", "cells", "cell extents", "separator extents"))
print("-" * 96)
for label, seps, found in (("columns", col_seps, col_cells), ("rows", row_seps, row_cells)):
print("{:<8} {:>6} {:>5} {:<28} {}".format(
label, len(seps), len(found),
", ".join("{}-{}".format(a, b) for a, b in found),
", ".join("{}-{}".format(a, b) for a, b in seps)))
print()
print("cell pitch: columns {} px wide, rows {} px tall".format(
col_cells[0][1] - col_cells[0][0] + 1, row_cells[0][1] - row_cells[0][0] + 1))
print()
print("{:<5} {:>2} {:>2} {:>7} {:>7} {:>6} {:>7} {:>6} {:<13} {:>5} {:<8} {:<7} {}".format(
"cell", "r", "c", "ink px", "ink %", "blobs", "largest", "holes", "class", "conf",
"decision", "wanted", "why"))
print("-" * 116)
clicks = missed = false_hits = reviews = 0
for index, (ri, ci, ink, fraction, kept, holes, label, confidence, decision, why) in enumerate(cells):
wanted = (ri, ci) in TARGETS
if decision == "click":
clicks += 1 if wanted else 0
false_hits += 0 if wanted else 1
else:
reviews += 1 if decision == "review" else 0
missed += 1 if wanted else 0
print("{:<5} {:>2} {:>2} {:>7} {:>6.2f}% {:>6} {:>7} {:>6} {:<13} {:>5.2f} {:<8} {:<7} {}".format(
index, ri, ci, ink, 100.0 * fraction, len(kept), kept[0] if kept else 0,
holes, label, confidence, decision, "yes" if wanted else "no", why))
print()
print("{} cells: {} clicked, all {} correct, {} wrong, {} wanted cells missed, "
"{} queued for review".format(
len(cells), clicks, clicks - false_hits, false_hits, missed, reviews))
print("targets {} decoys {} blank-with-speckle 1".format(len(TARGETS), len(DECOYS)))
print()
print("confidence bands: ink fraction to class, confidence, decision and reason")
print("{:<20} {:<6} {:<11} {:<9} {}".format("ink fraction", "conf", "class", "decision", "why"))
print("-" * 92)
for low, high, label, confidence, decision, why in BANDS:
print("{:<20} {:<6} {:<11} {:<9} {}".format(
"[{:.3f}, {:.3f})".format(low, high), "{:.2f}".format(confidence), label,
decision, why))
print()
print("confidence is how sure the classifier is of the class; the decision comes")
print("from the class. An empty cell can be recognised with 0.97 confidence and")
print("must still never be clicked, which is why the two columns are separate.")
print()
print("threshold sweep: the same pipeline at hand-adjusted thresholds")
print("{:<10} {:>6} {:>7} {:>8} {:>11} {:>11} {}".format(
"threshold", "cells", "clicks", "missed", "false hits", "mean conf", "note"))
print("-" * 92)
for shift in (-24, -12, 0, 12, 24):
level = max(0, min(255, THRESH + shift))
trial = survey(PLANE, level)
t_click = t_miss = t_false = 0
confidences = []
for ri, ci, _ink, _f, _k, _h, _label, confidence, decision, _why in trial:
confidences.append(confidence)
wanted = (ri, ci) in TARGETS
if decision == "click":
t_click += 1
t_false += 0 if wanted else 1
elif wanted:
t_miss += 1
note = "otsu, as computed" if shift == 0 else "hand-adjusted, do not ship"
print("{:<10} {:>6} {:>7} {:>8} {:>11} {:>11.3f} {}".format(
level, len(trial), t_click, t_miss, t_false,
sum(confidences) / len(confidences), note))
print()
print("the sweep is the argument for computing the threshold instead of picking")
print("it. Below the Otsu value the anti-aliased rim becomes ink, the counter in each")
print("O closes, the Euler number drops to zero and three decoys become clicks; above")
print("it the answer is stable. A threshold picked by eye lands in either regime.")
print()
print("this is self-hosted inference over a buffer generated inside the snippet: no")
print("solver service, no outbound request, and no third party learns which challenges")
print("you are looking at. The same functions run unchanged on a decoded PNG buffer.")
local grid segmentation: 240x120 buffer, 4x2 cells, generated in this snippet
frame removed: 3 rows top, 3 bottom, 3 columns each side; cropped plane 234x114
otsu threshold 120, minimum blob 6 px, background 236 ink 30, separator density 0.85
cells come from the ink-density profile, not from assumed geometry, so a
challenge that changes its grid pitch still segments.
axis seps cells cell extents separator extents
------------------------------------------------------------------------------------------------
columns 3 4 0-55, 59-115, 119-175, 179-233 56-58, 116-118, 176-178
rows 1 2 0-55, 59-113 56-58
cell pitch: columns 56 px wide, rows 56 px tall
cell r c ink px ink % blobs largest holes class conf decision wanted why
--------------------------------------------------------------------------------------------------------------------
0 0 0 215 6.86% 1 214 0 stroke:L-like 0.93 click yes one dominant blob, glyph sized
1 0 1 420 13.16% 1 420 1 stroke:O-like 0.93 skip no wrong class for the prompt
2 0 2 421 13.19% 1 420 1 stroke:O-like 0.93 skip no wrong class for the prompt
3 0 3 214 6.95% 1 214 0 stroke:L-like 0.93 click yes one dominant blob, glyph sized
4 1 0 420 13.64% 1 420 1 stroke:O-like 0.93 skip no wrong class for the prompt
5 1 1 214 6.83% 1 214 0 stroke:L-like 0.93 click yes one dominant blob, glyph sized
6 1 2 214 6.83% 1 214 0 stroke:L-like 0.93 click yes one dominant blob, glyph sized
7 1 3 37 1.22% 1 17 4 speckle 0.42 skip no blobs present, all below the floor
8 cells: 4 clicked, all 4 correct, 0 wrong, 0 wanted cells missed, 0 queued for review
targets 4 decoys 3 blank-with-speckle 1
confidence bands: ink fraction to class, confidence, decision and reason
ink fraction conf class decision why
--------------------------------------------------------------------------------------------
[0.000, 0.012) 0.97 empty skip no blob above the noise floor
[0.012, 0.030) 0.42 speckle skip blobs present, all below the floor
[0.030, 0.150) 0.93 stroke click one dominant blob, glyph sized
[0.150, 0.400) 0.71 busy review overlapping blobs, want a second look
[0.400, 1.001) 0.38 saturated review the threshold took the cell ground
confidence is how sure the classifier is of the class; the decision comes
from the class. An empty cell can be recognised with 0.97 confidence and
must still never be clicked, which is why the two columns are separate.
threshold sweep: the same pipeline at hand-adjusted thresholds
threshold cells clicks missed false hits mean conf note
--------------------------------------------------------------------------------------------
96 8 7 0 3 0.866 hand-adjusted, do not ship
108 8 7 0 3 0.866 hand-adjusted, do not ship
120 8 4 0 0 0.866 otsu, as computed
132 8 4 0 0 0.866 hand-adjusted, do not ship
144 8 4 0 0 0.866 hand-adjusted, do not ship
the sweep is the argument for computing the threshold instead of picking
it. Below the Otsu value the anti-aliased rim becomes ink, the counter in each
O closes, the Euler number drops to zero and three decoys become clicks; above
it the answer is stable. A threshold picked by eye lands in either regime.
this is self-hosted inference over a buffer generated inside the snippet: no
solver service, no outbound request, and no third party learns which challenges
you are looking at. The same functions run unchanged on a decoded PNG buffer.
From Class to Decision
The band table maps ink fraction to a class, and the hole count refines the class within the stroke band. Then a separate decision column decides whether to act. The distinction is deliberate and it matters: an empty cell is recognised with 0.97 confidence, the highest in the table, and must still never be clicked.
Confidence is how sure you are of the class. The decision is what you do about it. Collapsing the two produces the classic bug where a confident empty cell gets clicked because the decision rule was folded into the classifier. Keeping them apart also gives you somewhere to put policy: a review decision for the ambiguous band, a skip for the wrong shape, and a click only for the one class the prompt actually asked for. It matters because the click is not a free action: on a real grid the click lands as a pointer event with coordinates and timing, and Behavioural Biometrics is where those are scored against the rest of the session. A recogniser that is right and a click that is wrong still fails.
Why the Threshold Sweep Is the Argument
The sweep runs the identical pipeline at hand-adjusted thresholds on both sides of the computed one. Below it, the anti-aliased rim becomes ink, the enclosed region in each O closes, the Euler number drops to zero, and three decoys become clicks. Above it the answer is stable.
That asymmetry is the lesson. A threshold picked by eye lands in one regime or the other, and both regimes look plausible in a spot check of a few images. It is the perception counterpart of the budget problem in Proof-of-Work Solvers in Practice: a single sampled run tells you almost nothing, and only the sweep over the parameter separates the stable region from the cliff. The sweep also tells you the safe margin, which here is a full twelve counts upward and only six downward. A margin that is asymmetric in one direction is a warning: it means the failure mode you are closest to is the one that turns a decoy into a click, and that is the failure you can least afford to ship.
Checklist
- Crop the frame before computing anything; report what was removed.
- Compute the threshold from the histogram, and sweep it to find the margin.
- Locate cells from the density profile, never from a fixed width.
- Use blob count for presence and the Euler number for shape.
- Keep confidence and decision in separate columns, and never let a high confidence on an empty class produce a click.
- Report per-cell class, decision and reason, so the pipeline is inspectable.
- Measure end-to-end accuracy on captures you are allowed to use, not on synthetic buffers alone; the synthetic one only proves the code runs.
- Treat an asymmetric threshold margin as a finding, and check which side the dangerous failure is on before you ship.
The Legitimate Route
On a site you own, recognition code is usually a worse investment than removing the challenge from the paths that matter, or replacing it with a proof-of-work check that costs the client nothing to verify. Where you are authorised to test someone else's, get that authorisation in writing, and prefer test keys and staging environments that exist for exactly this: they give you labelled captures without touching production traffic. Solving a challenge to reach content you are not entitled to is a different act from measuring your own detector's accuracy, and only one of those is work worth doing. The test that separates them is simple to state and hard to argue with: could you show the site owner the capture and the code without changing a word, and would they still call it authorised testing? If the answer is no, the pipeline you are about to build is aimed at the wrong target.