Testing Stealth Pipelines
"It Worked Yesterday" Is Not a Test
Every stealth pipeline has a half-life, and nothing in it announces the day it expires. A UA that was current last month is a curiosity. A patch that closed one leak silently unpatched another in the same init script. A challenge added a step and your extractor now returns a shell. All of these ship to production as a green build, because no test asserts the thing that broke. A stealth pipeline needs its own regression suite, and most of it has to run without touching the target.
The Four Kinds of Test
Split by what they can catch, because they have different costs and different tolerances.
- Static invariants over a profile. Zero cost, runs in CI, catches the majority of regressions. Fields that contradict each other in a stored bundle: a UA that says Windows against a font list only macOS has.
- Golden comparison against recorded traffic. No live target, catches target-side changes you have absorbed and then broken: a header set that changed shape, a cookie that stopped being set, a redirect that appeared.
- Canary against a host you control. Small cost, catches what static checks cannot see, because it exercises the real network path with real timing: your egress, your DNS, your TLS stack, your proxy.
- Live smoke against the real target, sampled. Expensive, risky, rate-limited to a handful of requests per hour, and the only one of the four that can tell you the truth about the target's current state. Everything else exists so you can keep this small.
The discipline is in the ratio: the cheap three should catch ninety percent, so the fourth can run at a rate that does not matter. A pipeline whose only real test is the live job has no test, it has a production environment you are watching.
Cross-Field Invariants, and Why They Are the Highest Value
A contradiction is a signal a server can act on even when every individual field is plausible. navigator.userAgent claiming macOS while the WebGL renderer string names an Intel GPU is not detectable by looking at either value; it is detectable by comparing them, and it is exactly the comparison a defender runs.
Write the checks as invariants over the stored bundle, not as assertions in the runtime code, so they run before a context is ever created. Seventeen of them is a realistic number for a desktop profile:
| Field pair | The contradiction | Why it matters |
|---|---|---|
UA string vs navigator.platform |
Macintosh UA, Win32 platform | trivially cross-checked |
Sec-CH-UA-Platform vs UA token |
"Windows" against a Mac UA |
the header is sent before any script runs |
Sec-CH-UA major vs UA Chrome major |
131 vs 130 | version-rotation drift in a stored profile |
| screen vs viewport | viewport wider than screen | a headless default |
devicePixelRatio vs platform |
3.0 on a desktop macOS build | impossible on that hardware |
| primary language vs exit country | en-US behind a German residential exit |
one of the cheapest signals there is |
| timezone vs geo vs exit country | Europe/London with coordinates in Berlin |
free, and almost never accidental |
maxTouchPoints vs UA |
touch points on a non-touch desktop | cheap |
| WebGL vendor vs renderer | Intel vendor with an Apple renderer string | patched-string mismatch |
| GPU renderer vs claimed device class | an M1 string on an M2 profile | fleet hygiene |
| font list vs platform | MS Gothic on a macOS profile |
font enumeration is unforgiving |
| canvas hash vs the stock hash | the unpatched browser value | an unpatched canvas is worse than none |
hardwareConcurrency vs device class |
2 cores on an M2 Pro | implausible pairing |
| TLS version vs UA hints | TLS 1.2 with a Chrome 131 UA | see tls session resumption |
| header order vs expected order | accept-encoding before referer |
see http2 and header forging |
Fingerprint cross validation is the same exercise applied to cross-request consistency, and this is the cross-field version of it: one profile, checked against itself.
# Two regression checks that need no live target:
# 1. cross-field invariants over a stored device profile
# 2. a response fingerprint diffed against a stored baseline
PROFILE = {
"name": "nw-mac-07",
"user_agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) "
"AppleWebKit/537.36 (KHTML, like Gecko) "
"Chrome/131.0.0.0 Safari/537.36",
"navigator_platform": "MacIntel",
"sec_ch_ua_platform": '"macOS"',
"sec_ch_ua": '"Chromium";v="131", "Not_A Brand";v="24"',
"ua_platform_token": "macOS",
"screen": [1512, 982],
"viewport": [1440, 900],
"device_pixel_ratio": 2.0,
"languages": ["en-GB", "en"],
"timezone": "Europe/London",
"geo": {"lat": 51.51, "lon": -0.13, "country": "GB"},
"exit_country": "GB",
"hardware_concurrency": 8,
"device_memory": 8,
"max_touch_points": 0,
"webgl_vendor": "Google Inc. (Intel)",
"canvas_hash": ("6f1a9c0d4b8e27315fa4c9d0e7b2613485cf9a0e"
"1d2b3c4d5e6f70819a2b3c4d5e6f708192a3b4c5"),
"webgl_renderer": "ANGLE (Apple, ANGLE Metal Renderer: Apple M2 Pro, "
"Unspecified Version)",
"gpu_class": "apple-m2-pro",
"fonts": ["Helvetica Neue", "Menlo", "MS Gothic", "Times New Roman"],
"header_order": ["sec-ch-ua", "sec-ch-ua-mobile", "sec-ch-ua-platform",
"upgrade-insecure-requests", "user-agent", "accept",
"sec-fetch-site", "sec-fetch-mode", "sec-fetch-user",
"sec-fetch-dest", "referer", "accept-encoding"],
"tls_ja3": "771,4865-4866-4867-49195-49199,0-23-65281-...",
"tls_version": "1.3",
}
BASELINE = {
"taken": "2026-01-04T09:12Z",
"status": 200,
"body_bytes": 41230,
"body_digest": "9f2c41ab",
"header_names": ["cache-control", "content-encoding", "content-type",
"server", "set-cookie", "strict-transport-security",
"vary"],
"set_cookie_names": ["_cf_bm", "session"],
"challenge": False,
"hops": 1,
"ms_median": 84,
}
CURRENT = {
"status": 403,
"body_bytes": 640,
"body_digest": "0000beef",
"header_names": ["cache-control", "content-type", "server", "set-cookie"],
"set_cookie_names": ["session"],
"challenge": True,
"hops": 2,
"ms_median": 1310,
}
MS_TOLERANCE = 150
STOCK_CANVAS = ("6f1a9c0d4b8e27315fa4c9d0e7b2613485cf9a0e"
"1d2b3c4d5e6f70819a2b3c4d5e6f708192a3b4c5")
def invariants(p):
# Each check is a contradiction a server could act on, not a style rule.
checks = []
def check(name, ok, detail):
checks.append((name, bool(ok), detail))
mac = "Macintosh" in p["user_agent"]
check("ua platform token matches navigator.platform",
mac == (p["navigator_platform"] == "MacIntel"),
f"ua says {p['ua_platform_token']}, platform says "
f"{p['navigator_platform']}")
check("Sec-CH-UA-Platform matches the ua",
(p["sec_ch_ua_platform"].strip('"').lower()
== p["ua_platform_token"].lower()),
p["sec_ch_ua_platform"])
check("Sec-CH-UA major matches the ua major",
p["sec_ch_ua"].split('v="')[1].split('"')[0]
== p["user_agent"].split("Chrome/")[1].split(".")[0],
p["sec_ch_ua"])
check("screen is at least the viewport in both axes",
p["screen"][0] >= p["viewport"][0]
and p["screen"][1] >= p["viewport"][1],
f"screen {p['screen']} vs viewport {p['viewport']}")
check("devicePixelRatio is one of the values macOS ships",
p["device_pixel_ratio"] in (1.0, 2.0),
str(p["device_pixel_ratio"]))
check("primary language matches the exit country",
p["languages"][0].split("-")[1].lower() == p["geo"]["country"].lower(),
f"{p['languages'][0]} behind a {p['exit_country']} exit")
check("timezone country matches the exit country",
p["timezone"].split("/")[-1] in ("London", "Dublin", "Douglas"),
p["timezone"])
check("Mac profile advertises no touch points",
(p["max_touch_points"] == 0) == mac,
f"maxTouchPoints {p['max_touch_points']}")
vendor_token = p["webgl_vendor"].split("(")[-1].split(")")[0]
makers = [m for m in ("Apple", "Intel", "AMD", "Qualcomm")
if m in p["webgl_renderer"]]
check("UNMASKED_VENDOR_WEBGL names the same maker as the renderer",
vendor_token in makers,
f"vendor says {vendor_token}, renderer says "
f"{makers[0] if makers else 'no known maker'}")
check("the renderer names the GPU class the profile claims",
p["gpu_class"].split("-")[1].upper() in p["webgl_renderer"].upper(),
f"profile claims {p['gpu_class']}")
check("font set carries no family from another platform",
not ({"MS Gothic", "MS UI Gothic"} & set(p["fonts"])),
"MS Gothic is Windows-only and cannot enumerate on macOS")
missing = sorted({"Helvetica Neue", "Menlo"} - set(p["fonts"]))
check("font set carries the platform defaults",
not missing,
"all present" if not missing else "missing " + ", ".join(missing))
if p["canvas_hash"] == STOCK_CANVAS:
why = "equals the stock Chromium hash, so it is a real fingerprint"
elif len(p["canvas_hash"]) != 64:
why = f"{len(p['canvas_hash'])} chars, not a 64-char digest"
else:
why = "patched, and not the stock hash"
check("canvas hash is patched, not the stock hash",
p["canvas_hash"] != STOCK_CANVAS and len(p["canvas_hash"]) == 64,
why)
check("hardwareConcurrency is plausible for the GPU class",
p["hardware_concurrency"] in (4, 6, 8, 10, 12),
f"{p['hardware_concurrency']} cores")
check("TLS is 1.3 as the ua implies",
p["tls_version"] == "1.3", p["tls_version"])
check("header order starts at the client hints",
p["header_order"][0] == "sec-ch-ua", p["header_order"][0])
check("referer precedes accept-encoding, as Chrome sends it",
p["header_order"].index("referer")
< p["header_order"].index("accept-encoding"), "ordering")
return checks
def fingerprint_diff(base, cur, ms_tolerance):
drift = []
if base["status"] != cur["status"]:
drift.append(f"status {base['status']} -> {cur['status']}")
if base["challenge"] != cur["challenge"]:
drift.append(f"challenge {base['challenge']} -> {cur['challenge']}")
if base["hops"] != cur["hops"]:
drift.append(f"redirect hops {base['hops']} -> {cur['hops']}")
delta = cur["ms_median"] - base["ms_median"]
if abs(delta) > ms_tolerance:
drift.append(f"median latency {delta:+d} ms, tolerance "
f"{ms_tolerance}")
gone = sorted(set(base["header_names"]) - set(cur["header_names"]))
new = sorted(set(cur["header_names"]) - set(base["header_names"]))
if gone:
drift.append(f"response headers dropped: {', '.join(gone)}")
if new:
drift.append(f"response headers added: {', '.join(new)}")
lost = sorted(set(base["set_cookie_names"])
- set(cur["set_cookie_names"]))
if lost:
drift.append(f"cookies no longer set: {', '.join(lost)}")
if base["body_digest"] != cur["body_digest"]:
drift.append(f"body digest {base['body_digest']} -> "
f"{cur['body_digest']}, {base['body_bytes']} -> "
f"{cur['body_bytes']} bytes")
return drift
def main():
checks = invariants(PROFILE)
failed = [c for c in checks if not c[1]]
print(f"profile {PROFILE['name']}: {len(checks)} cross-field invariants, "
f"{len(failed)} contradiction(s)")
for name, ok, detail in checks:
print(f" {'ok ' if ok else 'FAIL'} {name:<52} {detail}")
print()
drift = fingerprint_diff(BASELINE, CURRENT, MS_TOLERANCE)
print(f"response fingerprint vs baseline {BASELINE['taken']} "
f"(tolerance {MS_TOLERANCE} ms)")
if drift:
for line in drift:
print(f" DRIFT {line}")
else:
print(" no drift")
blocking = [d for d in drift if not d.startswith("median latency")]
print(f" {len(drift)} difference(s), {len(blocking)} of them blocking; "
f"rerun this profile against the canary before touching a live job")
main()
profile nw-mac-07: 17 cross-field invariants, 3 contradiction(s)
ok ua platform token matches navigator.platform ua says macOS, platform says MacIntel
ok Sec-CH-UA-Platform matches the ua "macOS"
ok Sec-CH-UA major matches the ua major "Chromium";v="131", "Not_A Brand";v="24"
ok screen is at least the viewport in both axes screen [1512, 982] vs viewport [1440, 900]
ok devicePixelRatio is one of the values macOS ships 2.0
ok primary language matches the exit country en-GB behind a GB exit
ok timezone country matches the exit country Europe/London
ok Mac profile advertises no touch points maxTouchPoints 0
FAIL UNMASKED_VENDOR_WEBGL names the same maker as the renderer vendor says Intel, renderer says Apple
ok the renderer names the GPU class the profile claims profile claims apple-m2-pro
FAIL font set carries no family from another platform MS Gothic is Windows-only and cannot enumerate on macOS
ok font set carries the platform defaults all present
FAIL canvas hash is patched, not the stock hash equals the stock Chromium hash, so it is a real fingerprint
ok hardwareConcurrency is plausible for the GPU class 8 cores
ok TLS is 1.3 as the ua implies 1.3
ok header order starts at the client hints sec-ch-ua
ok referer precedes accept-encoding, as Chrome sends it ordering
response fingerprint vs baseline 2026-01-04T09:12Z (tolerance 150 ms)
DRIFT status 200 -> 403
DRIFT challenge False -> True
DRIFT redirect hops 1 -> 2
DRIFT median latency +1226 ms, tolerance 150
DRIFT response headers dropped: content-encoding, strict-transport-security, vary
DRIFT cookies no longer set: _cf_bm
DRIFT body digest 9f2c41ab -> 0000beef, 41230 -> 640 bytes
7 difference(s), 6 of them blocking; rerun this profile against the canary before touching a live job
Three failures, and only one of them is the kind you would have found by hand. The font list is visible if you read it. The WebGL vendor-versus-renderer split is not, because both strings look plausible in isolation. The canvas hash equalling the stock Chromium value is the worst of the three: it is not a contradiction at all, it is the absence of a patch, and it is invisible in every field-level check you would write by inspection. Only a stored expected value catches it.
Golden Comparison Against Recorded Traffic
A recorded capture is a fixture. The test is a diff, and the discipline is what you choose to store, because the temptation is to store everything and the result is a diff that reports forty changes and gets ignored.
Store per request index: status, latency, body size, body digest, and the names of the response headers and Set-Cookie values. Store the branch URL and the build identifier beside it, so two baselines are never compared across a deploy. Store the capture, not the response, for anything authenticated.
Then report three classes of finding separately, and this separation is what makes the diff actionable:
- Changed. Status moved, latency moved outside tolerance, body digest moved, a header appeared or disappeared.
- Missing. An entry in the baseline that did not happen. This is the one people forget, and it is the signature of a challenge adding or removing a step.
- New. An entry that was never reviewed. Treated as a review item, not as noise.
Tolerance has to be explicit and stated in the output. The snippet uses 150 ms on a median, which on an 84 ms baseline is generous enough for a bad transit day and tight enough to catch the 1.3 s that a challenge step adds. A latency drift you accept silently teaches you that latency drift is normal.
In the run above, seven differences, six blocking, and the header diff is the most informative line: content-encoding, strict-transport-security and vary all disappeared while a challenge appeared and _cf_bm stopped being set. That combination reads as "the request stopped reaching the origin", not as "the extractor broke", and it is diagnosable in a minute from the header names alone.
Canaries on Hosts You Own
A canary is a request to infrastructure you operate, through the identical code path: same profile, same headers, same TLS configuration, same proxy, same client. What it cannot tell you is whether the target's bot manager is happy. What it tells you is everything up to and including your own exit.
Three canaries, in increasing cost:
- A first-party endpoint you serve, hit through the scraper's own HTTP client. Catches egress, DNS, certificate and header-order regressions in about a second.
- A third-party echo service you have authorised, which reflects the headers and the observed client hints. Catches what the target's first-party script would have observed.
- A real vendor sandbox, where one exists and your account permits it. The only one that exercises a real bot manager, and it belongs in a manual, weekly, single-request check rather than in CI.
Run canaries on a schedule rather than on deploy, because some of the failures are not in your code at all: a proxy provider silently changed the ASN, a resolver started adding a search suffix, a TLS library upgraded and changed the cipher order. The canary is the only thing that sees those, and the check that catches them must be outside the deployment pipeline or nobody thinks to run it.
Drift Detection When the Target Moves
The target changes on a schedule you do not control, and the changes arrive in a fixed order. Learn the order.
- A patch. New version, new TLS fingerprint, new client hints. Cheap to detect with a fingerprint probe against a static endpoint.
- A new signal added. A canvas read, a WebGL extension query, a permission probe. Detectable by diffing the set of APIs the page touches, which is a capture-to-capture comparison rather than a field comparison.
- A challenge step added. The interstitial gains a phase, so the request sequence gains an entry. This is the missing-entry finding in a golden diff.
- A scoring change with no surface change. Same requests, different verdict. Only a live canary sees this, which is why the live smoke test exists and why its rate is the one number you argue about.
Measure drift on a schedule and store the series. A vendor whose patch cadence moved from weekly to monthly is telling you something, and a target whose challenge rate on your mature profiles rose from 2 to 8 per thousand over six weeks is telling you that your score distribution moved and your operating point should move with it.
The anti-pattern to avoid is the self-updating baseline. A diff that accepts itself teaches you nothing and leaves no record of when the flow moved, which is the one question you will need answered six months later.
The Suite, and How Often It Runs
A weekly stealth regression suite, five stages, in increasing risk. The whole thing fits in about twenty minutes and costs a handful of requests against the real target.
- Every commit. Static invariants over every stored profile in the fleet, and the device profile consistency rules: no two profiles may share a value that is supposed to be unique, and no profile may share a fingerprint with another. Zero network.
- Every commit. The header-order and TLS-configuration golden diff against the recorded fixture. Zero network.
- Nightly. Canary against the first-party endpoint you serve, plus a full context launch per profile with the health assertions from the pool. Catches browser upgrades, which are the single most common cause of a silent stealth regression.
- Nightly. A recorded capture replayed through the extractor. Catches changes in your parsing, which is a distinct failure from changes in the target's markup.
- Weekly, manually. One live request per vendor from a sacrificial profile, classified and diffed against the last baseline, with findings written down. Deliberately manual, deliberately low rate, and the only stage that is allowed to touch the target.
Add a monthly stage: rotate one sacrificial identity through the full lifecycle, cold to mature, so that a change in how the fleet is built shows up as a change in the warmup behaviour rather than as a ban three weeks later.
Why the Live Job Is Not a Test
The argument, because it is the one most teams lose.
A job that runs against the live target and reports 200s is testing three things at once: the target's tolerance, your pipeline's correctness, and the network in between. When it fails you cannot tell which, and the debugging starts by guessing. When it succeeds it tells you almost nothing, because success is compatible with the target having silently stopped collecting the signal you are spoofing.
A test suite separates those. Static invariants isolate your pipeline. Recorded fixtures isolate your parsing from the target's current mood. Canaries isolate the network from both. The live smoke test -- one request, one classification, one diff, once a week -- is the only place the target's opinion enters the system, and because it is the only place, its finding is unambiguous.
The second reason is that a job is a bad test because it is not repeatable. You cannot bisect a three-day production job. You cannot run it before the deploy. You cannot run the old version and the new version side by side. You cannot assert on it. A pipeline whose behaviour is only observable in production is a pipeline that is only debuggable in production, and that is a much larger cost than the ban rate you are trying to reduce.
Turning a Finding Into a Test
Every finding becomes a fixture, or it gets forgotten.
- Reproduce it offline. Take the profile or the capture that produced the finding and reduce it to the smallest thing that still shows it. If it cannot be reduced offline, it is a live-test problem, not a fixture problem.
- Write the assertion that would have caught it, at the level where it belongs: a field contradiction belongs in the static pass, a request-sequence change belongs in the golden diff, a network change belongs in the canary.
- Verify the assertion fails on the bad fixture and passes on the good one. An assertion that passes on both is worse than no assertion, because it looks like coverage.
- Commit both fixtures with the finding's date, vendor, and the build identifier, so that the next person to see it can tell what happened to what.
- Add it to the suite at the stage it belongs to, not at the stage that is easiest to run.
The pipeline that does this accumulates a library of specific, dated regressions, and each one is a real bug someone spent a day on. The pipeline that does not has the same bugs forever and rediscovers them on a schedule, which is what challenge debugging and observability ends up documenting. Pair the suite with observability and ban SLOS so that a finding raised in production becomes a fixture within a day, and the loop closes.
The Checklist
- Assert cross-field invariants over every stored profile, on every commit, before a context is created.
- Assert fleet-wide uniqueness: no shared fingerprint, no shared device id, no combination that should be unique.
- Diff recorded captures per request index; report changed, missing and new separately, with a stated tolerance.
- Store status, latency, size, digest and header names. Never cookie values, never authenticated bodies.
- Canary the first-party path nightly so that egress, DNS, proxy and TLS changes fail a test instead of a job.
- Run one live classified request per vendor weekly, by hand, and write the diff down.
- Measure drift on a schedule and store the series; the pace of a vendor's changes is itself a signal.
- Never let a baseline update itself. Reviewing the diff is the only thing that makes it a test.
- Turn every finding into a fixture within a day, and verify the new assertion fails on the bad input before it ships.
The last item is the whole lesson. A stealth pipeline fails quietly, on a schedule, with a green build and a 200 status code, and the only thing that catches it is a test that asserts something specific about the thing that broke.