"It Worked Yesterday" Is Not a Test

Every stealth pipeline has a half-life, and nothing in it announces the day it expires. A UA that was current last month is a curiosity. A patch that closed one leak silently unpatched another in the same init script. A challenge added a step and your extractor now returns a shell. All of these ship to production as a green build, because no test asserts the thing that broke. A stealth pipeline needs its own regression suite, and most of it has to run without touching the target.

The Four Kinds of Test

Split by what they can catch, because they have different costs and different tolerances.

  • Static invariants over a profile. Zero cost, runs in CI, catches the majority of regressions. Fields that contradict each other in a stored bundle: a UA that says Windows against a font list only macOS has.
  • Golden comparison against recorded traffic. No live target, catches target-side changes you have absorbed and then broken: a header set that changed shape, a cookie that stopped being set, a redirect that appeared.
  • Canary against a host you control. Small cost, catches what static checks cannot see, because it exercises the real network path with real timing: your egress, your DNS, your TLS stack, your proxy.
  • Live smoke against the real target, sampled. Expensive, risky, rate-limited to a handful of requests per hour, and the only one of the four that can tell you the truth about the target's current state. Everything else exists so you can keep this small.

The discipline is in the ratio: the cheap three should catch ninety percent, so the fourth can run at a rate that does not matter. A pipeline whose only real test is the live job has no test, it has a production environment you are watching.

Cross-Field Invariants, and Why They Are the Highest Value

A contradiction is a signal a server can act on even when every individual field is plausible. navigator.userAgent claiming macOS while the WebGL renderer string names an Intel GPU is not detectable by looking at either value; it is detectable by comparing them, and it is exactly the comparison a defender runs.

Write the checks as invariants over the stored bundle, not as assertions in the runtime code, so they run before a context is ever created. Seventeen of them is a realistic number for a desktop profile:

Field pair The contradiction Why it matters
UA string vs navigator.platform Macintosh UA, Win32 platform trivially cross-checked
Sec-CH-UA-Platform vs UA token "Windows" against a Mac UA the header is sent before any script runs
Sec-CH-UA major vs UA Chrome major 131 vs 130 version-rotation drift in a stored profile
screen vs viewport viewport wider than screen a headless default
devicePixelRatio vs platform 3.0 on a desktop macOS build impossible on that hardware
primary language vs exit country en-US behind a German residential exit one of the cheapest signals there is
timezone vs geo vs exit country Europe/London with coordinates in Berlin free, and almost never accidental
maxTouchPoints vs UA touch points on a non-touch desktop cheap
WebGL vendor vs renderer Intel vendor with an Apple renderer string patched-string mismatch
GPU renderer vs claimed device class an M1 string on an M2 profile fleet hygiene
font list vs platform MS Gothic on a macOS profile font enumeration is unforgiving
canvas hash vs the stock hash the unpatched browser value an unpatched canvas is worse than none
hardwareConcurrency vs device class 2 cores on an M2 Pro implausible pairing
TLS version vs UA hints TLS 1.2 with a Chrome 131 UA see tls session resumption
header order vs expected order accept-encoding before referer see http2 and header forging

Fingerprint cross validation is the same exercise applied to cross-request consistency, and this is the cross-field version of it: one profile, checked against itself.

# Two regression checks that need no live target:
#   1. cross-field invariants over a stored device profile
#   2. a response fingerprint diffed against a stored baseline

PROFILE = {
    "name": "nw-mac-07",
    "user_agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) "
                  "AppleWebKit/537.36 (KHTML, like Gecko) "
                  "Chrome/131.0.0.0 Safari/537.36",
    "navigator_platform": "MacIntel",
    "sec_ch_ua_platform": '"macOS"',
    "sec_ch_ua": '"Chromium";v="131", "Not_A Brand";v="24"',
    "ua_platform_token": "macOS",
    "screen": [1512, 982],
    "viewport": [1440, 900],
    "device_pixel_ratio": 2.0,
    "languages": ["en-GB", "en"],
    "timezone": "Europe/London",
    "geo": {"lat": 51.51, "lon": -0.13, "country": "GB"},
    "exit_country": "GB",
    "hardware_concurrency": 8,
    "device_memory": 8,
    "max_touch_points": 0,
    "webgl_vendor": "Google Inc. (Intel)",
    "canvas_hash": ("6f1a9c0d4b8e27315fa4c9d0e7b2613485cf9a0e"
                    "1d2b3c4d5e6f70819a2b3c4d5e6f708192a3b4c5"),
    "webgl_renderer": "ANGLE (Apple, ANGLE Metal Renderer: Apple M2 Pro, "
                      "Unspecified Version)",
    "gpu_class": "apple-m2-pro",
    "fonts": ["Helvetica Neue", "Menlo", "MS Gothic", "Times New Roman"],
    "header_order": ["sec-ch-ua", "sec-ch-ua-mobile", "sec-ch-ua-platform",
                     "upgrade-insecure-requests", "user-agent", "accept",
                     "sec-fetch-site", "sec-fetch-mode", "sec-fetch-user",
                     "sec-fetch-dest", "referer", "accept-encoding"],
    "tls_ja3": "771,4865-4866-4867-49195-49199,0-23-65281-...",
    "tls_version": "1.3",
}

BASELINE = {
    "taken": "2026-01-04T09:12Z",
    "status": 200,
    "body_bytes": 41230,
    "body_digest": "9f2c41ab",
    "header_names": ["cache-control", "content-encoding", "content-type",
                     "server", "set-cookie", "strict-transport-security",
                     "vary"],
    "set_cookie_names": ["_cf_bm", "session"],
    "challenge": False,
    "hops": 1,
    "ms_median": 84,
}

CURRENT = {
    "status": 403,
    "body_bytes": 640,
    "body_digest": "0000beef",
    "header_names": ["cache-control", "content-type", "server", "set-cookie"],
    "set_cookie_names": ["session"],
    "challenge": True,
    "hops": 2,
    "ms_median": 1310,
}

MS_TOLERANCE = 150
STOCK_CANVAS = ("6f1a9c0d4b8e27315fa4c9d0e7b2613485cf9a0e"
                "1d2b3c4d5e6f70819a2b3c4d5e6f708192a3b4c5")


def invariants(p):
    # Each check is a contradiction a server could act on, not a style rule.
    checks = []

    def check(name, ok, detail):
        checks.append((name, bool(ok), detail))

    mac = "Macintosh" in p["user_agent"]
    check("ua platform token matches navigator.platform",
          mac == (p["navigator_platform"] == "MacIntel"),
          f"ua says {p['ua_platform_token']}, platform says "
          f"{p['navigator_platform']}")
    check("Sec-CH-UA-Platform matches the ua",
          (p["sec_ch_ua_platform"].strip('"').lower()
           == p["ua_platform_token"].lower()),
          p["sec_ch_ua_platform"])
    check("Sec-CH-UA major matches the ua major",
          p["sec_ch_ua"].split('v="')[1].split('"')[0]
          == p["user_agent"].split("Chrome/")[1].split(".")[0],
          p["sec_ch_ua"])
    check("screen is at least the viewport in both axes",
          p["screen"][0] >= p["viewport"][0]
          and p["screen"][1] >= p["viewport"][1],
          f"screen {p['screen']} vs viewport {p['viewport']}")
    check("devicePixelRatio is one of the values macOS ships",
          p["device_pixel_ratio"] in (1.0, 2.0),
          str(p["device_pixel_ratio"]))
    check("primary language matches the exit country",
          p["languages"][0].split("-")[1].lower() == p["geo"]["country"].lower(),
          f"{p['languages'][0]} behind a {p['exit_country']} exit")
    check("timezone country matches the exit country",
          p["timezone"].split("/")[-1] in ("London", "Dublin", "Douglas"),
          p["timezone"])
    check("Mac profile advertises no touch points",
          (p["max_touch_points"] == 0) == mac,
          f"maxTouchPoints {p['max_touch_points']}")
    vendor_token = p["webgl_vendor"].split("(")[-1].split(")")[0]
    makers = [m for m in ("Apple", "Intel", "AMD", "Qualcomm")
              if m in p["webgl_renderer"]]
    check("UNMASKED_VENDOR_WEBGL names the same maker as the renderer",
          vendor_token in makers,
          f"vendor says {vendor_token}, renderer says "
          f"{makers[0] if makers else 'no known maker'}")
    check("the renderer names the GPU class the profile claims",
          p["gpu_class"].split("-")[1].upper() in p["webgl_renderer"].upper(),
          f"profile claims {p['gpu_class']}")
    check("font set carries no family from another platform",
          not ({"MS Gothic", "MS UI Gothic"} & set(p["fonts"])),
          "MS Gothic is Windows-only and cannot enumerate on macOS")
    missing = sorted({"Helvetica Neue", "Menlo"} - set(p["fonts"]))
    check("font set carries the platform defaults",
          not missing,
          "all present" if not missing else "missing " + ", ".join(missing))
    if p["canvas_hash"] == STOCK_CANVAS:
        why = "equals the stock Chromium hash, so it is a real fingerprint"
    elif len(p["canvas_hash"]) != 64:
        why = f"{len(p['canvas_hash'])} chars, not a 64-char digest"
    else:
        why = "patched, and not the stock hash"
    check("canvas hash is patched, not the stock hash",
          p["canvas_hash"] != STOCK_CANVAS and len(p["canvas_hash"]) == 64,
          why)
    check("hardwareConcurrency is plausible for the GPU class",
          p["hardware_concurrency"] in (4, 6, 8, 10, 12),
          f"{p['hardware_concurrency']} cores")
    check("TLS is 1.3 as the ua implies",
          p["tls_version"] == "1.3", p["tls_version"])
    check("header order starts at the client hints",
          p["header_order"][0] == "sec-ch-ua", p["header_order"][0])
    check("referer precedes accept-encoding, as Chrome sends it",
          p["header_order"].index("referer")
          < p["header_order"].index("accept-encoding"), "ordering")
    return checks


def fingerprint_diff(base, cur, ms_tolerance):
    drift = []
    if base["status"] != cur["status"]:
        drift.append(f"status {base['status']} -> {cur['status']}")
    if base["challenge"] != cur["challenge"]:
        drift.append(f"challenge {base['challenge']} -> {cur['challenge']}")
    if base["hops"] != cur["hops"]:
        drift.append(f"redirect hops {base['hops']} -> {cur['hops']}")
    delta = cur["ms_median"] - base["ms_median"]
    if abs(delta) > ms_tolerance:
        drift.append(f"median latency {delta:+d} ms, tolerance "
                     f"{ms_tolerance}")
    gone = sorted(set(base["header_names"]) - set(cur["header_names"]))
    new = sorted(set(cur["header_names"]) - set(base["header_names"]))
    if gone:
        drift.append(f"response headers dropped: {', '.join(gone)}")
    if new:
        drift.append(f"response headers added: {', '.join(new)}")
    lost = sorted(set(base["set_cookie_names"])
                  - set(cur["set_cookie_names"]))
    if lost:
        drift.append(f"cookies no longer set: {', '.join(lost)}")
    if base["body_digest"] != cur["body_digest"]:
        drift.append(f"body digest {base['body_digest']} -> "
                     f"{cur['body_digest']}, {base['body_bytes']} -> "
                     f"{cur['body_bytes']} bytes")
    return drift


def main():
    checks = invariants(PROFILE)
    failed = [c for c in checks if not c[1]]
    print(f"profile {PROFILE['name']}: {len(checks)} cross-field invariants, "
          f"{len(failed)} contradiction(s)")
    for name, ok, detail in checks:
        print(f"  {'ok  ' if ok else 'FAIL'}  {name:<52} {detail}")

    print()
    drift = fingerprint_diff(BASELINE, CURRENT, MS_TOLERANCE)
    print(f"response fingerprint vs baseline {BASELINE['taken']} "
          f"(tolerance {MS_TOLERANCE} ms)")
    if drift:
        for line in drift:
            print(f"  DRIFT  {line}")
    else:
        print("  no drift")
    blocking = [d for d in drift if not d.startswith("median latency")]
    print(f"  {len(drift)} difference(s), {len(blocking)} of them blocking; "
          f"rerun this profile against the canary before touching a live job")


main()
profile nw-mac-07: 17 cross-field invariants, 3 contradiction(s)
  ok    ua platform token matches navigator.platform         ua says macOS, platform says MacIntel
  ok    Sec-CH-UA-Platform matches the ua                    "macOS"
  ok    Sec-CH-UA major matches the ua major                 "Chromium";v="131", "Not_A Brand";v="24"
  ok    screen is at least the viewport in both axes         screen [1512, 982] vs viewport [1440, 900]
  ok    devicePixelRatio is one of the values macOS ships    2.0
  ok    primary language matches the exit country            en-GB behind a GB exit
  ok    timezone country matches the exit country            Europe/London
  ok    Mac profile advertises no touch points               maxTouchPoints 0
  FAIL  UNMASKED_VENDOR_WEBGL names the same maker as the renderer vendor says Intel, renderer says Apple
  ok    the renderer names the GPU class the profile claims  profile claims apple-m2-pro
  FAIL  font set carries no family from another platform     MS Gothic is Windows-only and cannot enumerate on macOS
  ok    font set carries the platform defaults               all present
  FAIL  canvas hash is patched, not the stock hash           equals the stock Chromium hash, so it is a real fingerprint
  ok    hardwareConcurrency is plausible for the GPU class   8 cores
  ok    TLS is 1.3 as the ua implies                         1.3
  ok    header order starts at the client hints              sec-ch-ua
  ok    referer precedes accept-encoding, as Chrome sends it ordering

response fingerprint vs baseline 2026-01-04T09:12Z (tolerance 150 ms)
  DRIFT  status 200 -> 403
  DRIFT  challenge False -> True
  DRIFT  redirect hops 1 -> 2
  DRIFT  median latency +1226 ms, tolerance 150
  DRIFT  response headers dropped: content-encoding, strict-transport-security, vary
  DRIFT  cookies no longer set: _cf_bm
  DRIFT  body digest 9f2c41ab -> 0000beef, 41230 -> 640 bytes
  7 difference(s), 6 of them blocking; rerun this profile against the canary before touching a live job

Three failures, and only one of them is the kind you would have found by hand. The font list is visible if you read it. The WebGL vendor-versus-renderer split is not, because both strings look plausible in isolation. The canvas hash equalling the stock Chromium value is the worst of the three: it is not a contradiction at all, it is the absence of a patch, and it is invisible in every field-level check you would write by inspection. Only a stored expected value catches it.

Golden Comparison Against Recorded Traffic

A recorded capture is a fixture. The test is a diff, and the discipline is what you choose to store, because the temptation is to store everything and the result is a diff that reports forty changes and gets ignored.

Store per request index: status, latency, body size, body digest, and the names of the response headers and Set-Cookie values. Store the branch URL and the build identifier beside it, so two baselines are never compared across a deploy. Store the capture, not the response, for anything authenticated.

Then report three classes of finding separately, and this separation is what makes the diff actionable:

  • Changed. Status moved, latency moved outside tolerance, body digest moved, a header appeared or disappeared.
  • Missing. An entry in the baseline that did not happen. This is the one people forget, and it is the signature of a challenge adding or removing a step.
  • New. An entry that was never reviewed. Treated as a review item, not as noise.

Tolerance has to be explicit and stated in the output. The snippet uses 150 ms on a median, which on an 84 ms baseline is generous enough for a bad transit day and tight enough to catch the 1.3 s that a challenge step adds. A latency drift you accept silently teaches you that latency drift is normal.

In the run above, seven differences, six blocking, and the header diff is the most informative line: content-encoding, strict-transport-security and vary all disappeared while a challenge appeared and _cf_bm stopped being set. That combination reads as "the request stopped reaching the origin", not as "the extractor broke", and it is diagnosable in a minute from the header names alone.

Canaries on Hosts You Own

A canary is a request to infrastructure you operate, through the identical code path: same profile, same headers, same TLS configuration, same proxy, same client. What it cannot tell you is whether the target's bot manager is happy. What it tells you is everything up to and including your own exit.

Three canaries, in increasing cost:

  • A first-party endpoint you serve, hit through the scraper's own HTTP client. Catches egress, DNS, certificate and header-order regressions in about a second.
  • A third-party echo service you have authorised, which reflects the headers and the observed client hints. Catches what the target's first-party script would have observed.
  • A real vendor sandbox, where one exists and your account permits it. The only one that exercises a real bot manager, and it belongs in a manual, weekly, single-request check rather than in CI.

Run canaries on a schedule rather than on deploy, because some of the failures are not in your code at all: a proxy provider silently changed the ASN, a resolver started adding a search suffix, a TLS library upgraded and changed the cipher order. The canary is the only thing that sees those, and the check that catches them must be outside the deployment pipeline or nobody thinks to run it.

Drift Detection When the Target Moves

The target changes on a schedule you do not control, and the changes arrive in a fixed order. Learn the order.

  1. A patch. New version, new TLS fingerprint, new client hints. Cheap to detect with a fingerprint probe against a static endpoint.
  2. A new signal added. A canvas read, a WebGL extension query, a permission probe. Detectable by diffing the set of APIs the page touches, which is a capture-to-capture comparison rather than a field comparison.
  3. A challenge step added. The interstitial gains a phase, so the request sequence gains an entry. This is the missing-entry finding in a golden diff.
  4. A scoring change with no surface change. Same requests, different verdict. Only a live canary sees this, which is why the live smoke test exists and why its rate is the one number you argue about.

Measure drift on a schedule and store the series. A vendor whose patch cadence moved from weekly to monthly is telling you something, and a target whose challenge rate on your mature profiles rose from 2 to 8 per thousand over six weeks is telling you that your score distribution moved and your operating point should move with it.

The anti-pattern to avoid is the self-updating baseline. A diff that accepts itself teaches you nothing and leaves no record of when the flow moved, which is the one question you will need answered six months later.

The Suite, and How Often It Runs

A weekly stealth regression suite, five stages, in increasing risk. The whole thing fits in about twenty minutes and costs a handful of requests against the real target.

  • Every commit. Static invariants over every stored profile in the fleet, and the device profile consistency rules: no two profiles may share a value that is supposed to be unique, and no profile may share a fingerprint with another. Zero network.
  • Every commit. The header-order and TLS-configuration golden diff against the recorded fixture. Zero network.
  • Nightly. Canary against the first-party endpoint you serve, plus a full context launch per profile with the health assertions from the pool. Catches browser upgrades, which are the single most common cause of a silent stealth regression.
  • Nightly. A recorded capture replayed through the extractor. Catches changes in your parsing, which is a distinct failure from changes in the target's markup.
  • Weekly, manually. One live request per vendor from a sacrificial profile, classified and diffed against the last baseline, with findings written down. Deliberately manual, deliberately low rate, and the only stage that is allowed to touch the target.

Add a monthly stage: rotate one sacrificial identity through the full lifecycle, cold to mature, so that a change in how the fleet is built shows up as a change in the warmup behaviour rather than as a ban three weeks later.

Why the Live Job Is Not a Test

The argument, because it is the one most teams lose.

A job that runs against the live target and reports 200s is testing three things at once: the target's tolerance, your pipeline's correctness, and the network in between. When it fails you cannot tell which, and the debugging starts by guessing. When it succeeds it tells you almost nothing, because success is compatible with the target having silently stopped collecting the signal you are spoofing.

A test suite separates those. Static invariants isolate your pipeline. Recorded fixtures isolate your parsing from the target's current mood. Canaries isolate the network from both. The live smoke test -- one request, one classification, one diff, once a week -- is the only place the target's opinion enters the system, and because it is the only place, its finding is unambiguous.

The second reason is that a job is a bad test because it is not repeatable. You cannot bisect a three-day production job. You cannot run it before the deploy. You cannot run the old version and the new version side by side. You cannot assert on it. A pipeline whose behaviour is only observable in production is a pipeline that is only debuggable in production, and that is a much larger cost than the ban rate you are trying to reduce.

Turning a Finding Into a Test

Every finding becomes a fixture, or it gets forgotten.

  1. Reproduce it offline. Take the profile or the capture that produced the finding and reduce it to the smallest thing that still shows it. If it cannot be reduced offline, it is a live-test problem, not a fixture problem.
  2. Write the assertion that would have caught it, at the level where it belongs: a field contradiction belongs in the static pass, a request-sequence change belongs in the golden diff, a network change belongs in the canary.
  3. Verify the assertion fails on the bad fixture and passes on the good one. An assertion that passes on both is worse than no assertion, because it looks like coverage.
  4. Commit both fixtures with the finding's date, vendor, and the build identifier, so that the next person to see it can tell what happened to what.
  5. Add it to the suite at the stage it belongs to, not at the stage that is easiest to run.

The pipeline that does this accumulates a library of specific, dated regressions, and each one is a real bug someone spent a day on. The pipeline that does not has the same bugs forever and rediscovers them on a schedule, which is what challenge debugging and observability ends up documenting. Pair the suite with observability and ban SLOS so that a finding raised in production becomes a fixture within a day, and the loop closes.

The Checklist

  • Assert cross-field invariants over every stored profile, on every commit, before a context is created.
  • Assert fleet-wide uniqueness: no shared fingerprint, no shared device id, no combination that should be unique.
  • Diff recorded captures per request index; report changed, missing and new separately, with a stated tolerance.
  • Store status, latency, size, digest and header names. Never cookie values, never authenticated bodies.
  • Canary the first-party path nightly so that egress, DNS, proxy and TLS changes fail a test instead of a job.
  • Run one live classified request per vendor weekly, by hand, and write the diff down.
  • Measure drift on a schedule and store the series; the pace of a vendor's changes is itself a signal.
  • Never let a baseline update itself. Reviewing the diff is the only thing that makes it a test.
  • Turn every finding into a fixture within a day, and verify the new assertion fails on the bad input before it ships.

The last item is the whole lesson. A stealth pipeline fails quietly, on a schedule, with a green build and a 200 status code, and the only thing that catches it is a test that asserts something specific about the thing that broke.