Capstone: A Stealth Architecture

Synthesis across the course. Link to signal inventory, fingerprint cross validation, browser pool architecture, account warmup and identity lifecycle, crawl budget and request scheduling, observability and ban SLOS, testing stealth pipelines, challenge debugging and observability.

Identity Planning

Plan identities with warmup and lifecycle states.

Topology

Maintain consistent DNS, proxy, TLS and header configuration.

Browser Management

Isolate profiles and prevent state leakage.

Scheduling

Enforce budgets with jitter and backoff.

Challenge Handling

Resolve challenges at appropriate layers.

Observability

Track key metrics and SLOs.

Testing

Run invariants and drift detection continuously.

Failure Modes at Scale

Address clock skew, config drift, leakage and cross-worker effects.

from datetime import date

# The assembled plan for one collection target. Every worker starts from this
# one structure; the invariant checks below run against it before any of it
# is handed to a fleet.
TODAY = date(2026, 10, 1)

TARGET = {
    "host": "shop.example.com",
    "edge": "cloudflare-managed-challenge",
    "detection": "turnstile, datadome secondary validation",
    "slo": {"request_success": 0.97, "ban_rate_per_day": 0.010,
            "p95_first_byte_ms": 4000},
}

BUDGET = {
    "max_requests_per_hour": 24000,
    "max_concurrency": 24,
    "daily_budget_usd": 380.0,
    "cost_per_request_usd": 0.00085,
    "min_warmup_requests": 20,
    "min_warmup_days": 7,
}

# id, platform, ua_platform string, ua_ch_platform, screen, cores,
# timezone, egress country, egress address, residential, resolver country,
# session cookie, credential, created, warmup log [(date, requests)],
# hourly cap
PROFILES = [
    ("p-a1", "windows", "Windows NT 10.0; Win64; x64", "Windows",
     "1920x1080", 8, "Europe/Berlin", "DE", "203.0.113.41", True, "DE",
     "sess_4a91", "acct_017", date(2026, 8, 14),
     [(date(2026, 8, 15), 4), (date(2026, 8, 17), 9),
      (date(2026, 8, 21), 16), (date(2026, 8, 28), 34),
      (date(2026, 9, 11), 71), (date(2026, 9, 25), 140)], 1200),
    ("p-a2", "windows", "Windows NT 10.0; Win64; x64", "Windows",
     "1680x1050", 12, "Europe/Paris", "FR", "198.51.100.17", True, "FR",
     "sess_4a92", "acct_018", date(2026, 8, 20),
     [(date(2026, 8, 22), 3), (date(2026, 8, 26), 11),
      (date(2026, 9, 2), 28), (date(2026, 9, 16), 63),
      (date(2026, 9, 29), 128)], 1400),
    ("p-b1", "macos", "Macintosh; Intel Mac OS X 10_15_7", "macOS",
     "2560x1440", 10, "America/New_York", "US", "192.0.2.88", True, "US",
     "sess_4a91", "acct_041", date(2026, 7, 30),
     [(date(2026, 8, 1), 5), (date(2026, 8, 9), 13),
      (date(2026, 8, 23), 30), (date(2026, 9, 6), 58),
      (date(2026, 9, 20), 96)], 1000),
    ("p-b2", "windows", "Windows NT 10.0; Win64; x64", "Windows",
     "1536x864", 8, "America/Chicago", "US", "192.0.2.91", True, "US",
     "sess_4b22", "acct_042", date(2026, 9, 1),
     [(date(2026, 9, 3), 6), (date(2026, 9, 8), 12),
      (date(2026, 9, 15), 24), (date(2026, 9, 24), 48),
      (date(2026, 9, 30), 70)], 900),
    ("p-c1", "macos", "Macintosh; Intel Mac OS X 10_15_7", "macOS",
     "1440x900", 8, "Asia/Tokyo", "US", "198.51.100.92", True, "JP",
     "sess_4b23", "acct_043", date(2026, 8, 11),
     [(date(2026, 8, 13), 4), (date(2026, 8, 19), 12),
      (date(2026, 8, 27), 25), (date(2026, 9, 9), 51),
      (date(2026, 9, 22), 88)], 800),
    ("p-c2", "android", "Linux; Android 14; Pixel 8", "Android",
     "412x915", 8, "Europe/Berlin", "DE", "203.0.113.55", True, "DE",
     "sess_4b24", "acct_044", date(2026, 8, 25),
     [(date(2026, 8, 27), 3), (date(2026, 9, 1), 10),
      (date(2026, 9, 10), 22), (date(2026, 9, 21), 47),
      (date(2026, 9, 30), 66)], 700),
    ("p-d1", "android", "Linux; Android 14; Pixel 8", "Android",
     "1920x1080", 8, "Europe/Berlin", "DE", "203.0.113.56", True, "DE",
     "sess_4b25", "acct_045", date(2026, 9, 2),
     [(date(2026, 9, 4), 4), (date(2026, 9, 9), 11),
      (date(2026, 9, 17), 24), (date(2026, 9, 26), 41),
      (date(2026, 9, 30), 55)], 600),
    ("p-d2", "windows", "Windows NT 10.0; Win64; x64", "Windows",
     "1920x1200", 6, "Europe/London", "GB", "203.0.113.61", True, "GB",
     "sess_4b26", "acct_046", date(2026, 9, 7),
     [(date(2026, 9, 2), 2), (date(2026, 9, 11), 5),
      (date(2026, 9, 19), 9)], 500),
    ("p-e1", "macos", "Macintosh; Intel Mac OS X 10_15_7", "macOS",
     "1512x982", 10, "Europe/Amsterdam", "NL", "198.51.100.73", True,
     "NL", "sess_4b27", "acct_047", date(2026, 8, 5),
     [(date(2026, 8, 7), 5), (date(2026, 8, 15), 13),
      (date(2026, 8, 24), 29), (date(2026, 9, 5), 55),
      (date(2026, 9, 18), 92), (date(2026, 9, 28), 130)], 1100),
]

# (worker, concurrency, declared impersonate target, declared egress country,
#  share of the global hourly budget)
WORKERS = [
    ("w1-fra", 4, "chrome131", "FR", 0.20),
    ("w2-iad", 4, "chrome131", "US", 0.20),
    ("w3-fra", 4, "chrome128", "FR", 0.20),
    ("w4-ams", 4, "chrome131", "NL", 0.20),
    ("w5-lhr", 4, "chrome131", "GB", 0.20),
]

BROWSER_BUILD = "131"
PLATFORM_SCREEN_MIN = {"windows": 1280, "macos": 1280, "android": 480}
PLATFORM_SCREEN_MAX = {"windows": 4096, "macos": 4096, "android": 480}
UA_PLATFORM = {"Windows": "Windows NT 10.0", "macOS": "Macintosh",
               "Android": "Linux"}

# An IANA zone identifier usually ends in a city, not a country code, so
# cross-checking it against the egress country needs an explicit map. Do not
# infer it from the suffix.
TIMEZONE_COUNTRY = {
    "America/Chicago": "US", "America/New_York": "US", "Asia/Tokyo": "JP",
    "Europe/Amsterdam": "NL", "Europe/Berlin": "DE", "Europe/London": "GB",
    "Europe/Paris": "FR",
}


def out(line=""):
    print(line.rstrip())


def screen_width(text):
    return int(text.split("x")[0])


def warmup_total(log):
    return sum(n for _d, n in log)


def monotonic(log):
    return all(log[i][0] < log[i + 1][0] for i in range(len(log) - 1))


def out_of_window(log):
    return [d for d, _n in log if d > TODAY]


def duplicated(values):
    seen, dupes = set(), set()
    for v in values:
        if v in seen:
            dupes.add(v)
        seen.add(v)
    return dupes


P = {}
for row in PROFILES:
    P[row[0]] = dict(zip(
        ("id", "platform", "ua_platform", "ua_ch_platform", "screen",
         "cores", "timezone", "geo", "egress", "residential", "resolver",
         "session", "credential", "created", "warmup", "rph_cap"), row))

CHECKS = []


def check(code, title):
    def wrap(fn):
        CHECKS.append((code, title, fn))
        return fn
    return wrap


@check("I1", "no two profiles share a session cookie")
def _i1():
    dupes = duplicated([p["session"] for p in P.values()])
    out_lines = []
    for s in sorted(duplicated([p["session"] for p in P.values()])):
        holders = sorted(p["id"] for p in P.values() if p["session"] == s)
        for pid in holders:
            others = [h for h in holders if h != pid]
            out_lines.append("%s shares session %s with %s"
                             % (pid, s, " ".join(others)))
    return out_lines


@check("I2", "no credential is shared across a pool")
def _i2():
    dupes = duplicated([p["credential"] for p in P.values()])
    return ["credential %s is used by %s"
            % (c, " ".join(sorted(p["id"] for p in P.values()
                                  if p["credential"] == c)))
            for c in sorted(dupes)]


@check("I3", "no two profiles share an egress address")
def _i3():
    return ["address %s is leased by %s"
            % (ip, " ".join(sorted(p["id"] for p in P.values()
                                  if p["egress"] == ip)))
            for ip in sorted(duplicated([p["egress"] for p in P.values()]))]


@check("I4", "timezone agrees with the egress country")
def _i4():
    return ["%s: timezone %s with egress %s"
            % (p["id"], p["timezone"], p["geo"])
            for p in sorted(P.values(), key=lambda p: p["id"])
            if TIMEZONE_COUNTRY.get(p["timezone"]) != p["geo"]]


@check("I5", "resolver country agrees with the egress country")
def _i5():
    return ["%s: resolver in %s, egress in %s"
            % (p["id"], p["resolver"], p["geo"])
            for p in sorted(P.values(), key=lambda p: p["id"])
            if p["resolver"] != p["geo"]]


@check("I6", "screen geometry agrees with the declared platform")
def _i6():
    bad = []
    for p in sorted(P.values(), key=lambda p: p["id"]):
        w = screen_width(p["screen"])
        if not (PLATFORM_SCREEN_MIN[p["platform"]] <= w
                <= PLATFORM_SCREEN_MAX[p["platform"]]):
            bad.append("%s: %s declares %s"
                       % (p["id"], p["platform"], p["screen"]))
        if not p["ua_platform"].startswith(UA_PLATFORM[p["ua_ch_platform"]]):
            bad.append("%s: user agent platform and ua-ch platform disagree"
                       % p["id"])
    return bad


@check("I7", "every profile has a complete, monotonic warmup history")
def _i7():
    bad = []
    for p in sorted(P.values(), key=lambda p: p["id"]):
        total = warmup_total(p["warmup"])
        age = (TODAY - p["created"]).days
        if total < BUDGET["min_warmup_requests"]:
            bad.append("%s: %d warmup requests, floor is %d"
                       % (p["id"], total, BUDGET["min_warmup_requests"]))
        if age < BUDGET["min_warmup_days"]:
            bad.append("%s: %d days old, floor is %d"
                       % (p["id"], age, BUDGET["min_warmup_days"]))
        if not monotonic(p["warmup"]):
            bad.append("%s: warmup log is not strictly increasing in time"
                       % p["id"])
    return bad


@check("I8", "no clock skew: created before first use, nothing dated ahead")
def _i8():
    bad = []
    for p in sorted(P.values(), key=lambda p: p["id"]):
        if p["warmup"] and p["warmup"][0][0] < p["created"]:
            bad.append("%s: first warmup %s precedes creation %s"
                       % (p["id"], p["warmup"][0][0], p["created"]))
        for d in out_of_window(p["warmup"]):
            bad.append("%s: warmup entry dated %s is in the future"
                       % (p["id"], d))
    return bad


@check("I9", "planned volume fits the global rate cap")
def _i9():
    planned = sum(p["rph_cap"] for p in P.values())
    shares = sum(w[4] for w in WORKERS)
    bad = []
    if planned > BUDGET["max_requests_per_hour"]:
        bad.append("profile caps sum to %d/hr, cap is %d/hr"
                   % (planned, BUDGET["max_requests_per_hour"]))
    if abs(shares - 1.0) > 1e-9:
        bad.append("worker budget shares sum to %.2f, not 1.00" % shares)
    if sum(w[1] for w in WORKERS) > BUDGET["max_concurrency"]:
        bad.append("worker concurrency %d exceeds %d"
                   % (sum(w[1] for w in WORKERS), BUDGET["max_concurrency"]))
    return bad


@check("I10", "projected egress spend fits the daily budget")
def _i10():
    planned = sum(p["rph_cap"] for p in P.values())
    daily = planned * 60 * BUDGET["cost_per_request_usd"]
    if daily > BUDGET["daily_budget_usd"]:
        return ["$%.2f/day projected against a $%.2f budget"
                % (daily, BUDGET["daily_budget_usd"])]
    return []


@check("I11", "no worker runs a stale impersonate target")
def _i11():
    want = "chrome" + BROWSER_BUILD
    return ["%s declares %s, plan is %s" % (w[0], w[2], want)
            for w in WORKERS if w[2] != want]


@check("I12", "every worker's egress geography is represented in the pool")
def _i12():
    have = set(p["geo"] for p in P.values())
    return ["%s claims %s but the pool has no profile there"
            % (w[0], w[3]) for w in WORKERS if w[3] not in have]


results = [(code, title, fn()) for code, title, fn in CHECKS]
failed = [r for r in results if r[2]]

out("readiness report: %s via %s"
    % (TARGET["host"], TARGET["edge"]))
out("detection: %s" % TARGET["detection"])
out("plan: %d profiles, %d workers, %d requests/hour ceiling, $%.0f/day"
    % (len(PROFILES), len(WORKERS), BUDGET["max_requests_per_hour"],
       BUDGET["daily_budget_usd"]))
out("slo: %.0f%% request success, %.1f%% daily ban rate, %d ms p95 first byte"
    % (100 * TARGET["slo"]["request_success"],
       100 * TARGET["slo"]["ban_rate_per_day"],
       TARGET["slo"]["p95_first_byte_ms"]))
out()
out("%-5s %-56s %s" % ("id", "invariant", "result"))
for code, title, bad in results:
    out("%-5s %-56s %s" % (code, title,
                           "pass" if not bad else "FAIL (%d)" % len(bad)))

out()
out("findings")
for code, title, bad in results:
    for line in bad:
        out("  %-5s %s" % (code, line))

quarantine = sorted(set(
    p["id"] for p in P.values()
    for code, _t, bad in results
    for line in bad
    if line.startswith(p["id"] + ":")
    or line.startswith(p["id"] + " ")))

out()
out("plan summary")
out("  invariants passing   %d of %d" % (len(results) - len(failed),
                                          len(results)))
out("  profiles quarantined %d of %d  [%s]"
    % (len(quarantine), len(PROFILES), " ".join(quarantine) or "none"))
planned_rph = sum(p["rph_cap"] for p in P.values()
                  if p["id"] not in quarantine)
out("  releasable volume    %d of %d requests/hour"
    % (planned_rph, sum(p["rph_cap"] for p in P.values())))
out("  projected spend      $%.2f of $%.2f per day"
    % (planned_rph * 60 * BUDGET["cost_per_request_usd"],
       BUDGET["daily_budget_usd"]))
out("  verdict              %s"
    % ("RELEASE" if not failed else "HOLD: fix %s first"
       % ", ".join(c for c, _t, b in failed if b)))
readiness report: shop.example.com via cloudflare-managed-challenge
detection: turnstile, datadome secondary validation
plan: 9 profiles, 5 workers, 24000 requests/hour ceiling, $380/day
slo: 97% request success, 1.0% daily ban rate, 4000 ms p95 first byte

id    invariant                                                result
I1    no two profiles share a session cookie                   FAIL (2)
I2    no credential is shared across a pool                    pass
I3    no two profiles share an egress address                  pass
I4    timezone agrees with the egress country                  FAIL (1)
I5    resolver country agrees with the egress country          FAIL (1)
I6    screen geometry agrees with the declared platform        FAIL (2)
I7    every profile has a complete, monotonic warmup history   FAIL (1)
I8    no clock skew: created before first use, nothing dated ahead FAIL (1)
I9    planned volume fits the global rate cap                  pass
I10   projected egress spend fits the daily budget             FAIL (1)
I11   no worker runs a stale impersonate target                FAIL (1)
I12   every worker's egress geography is represented in the pool pass

findings
  I1    p-a1 shares session sess_4a91 with p-b1
  I1    p-b1 shares session sess_4a91 with p-a1
  I4    p-c1: timezone Asia/Tokyo with egress US
  I5    p-c1: resolver in JP, egress in US
  I6    p-c2: android declares 412x915
  I6    p-d1: android declares 1920x1080
  I7    p-d2: 16 warmup requests, floor is 20
  I8    p-d2: first warmup 2026-09-02 precedes creation 2026-09-07
  I10   $418.20/day projected against a $380.00 budget
  I11   w3-fra declares chrome128, plan is chrome131

plan summary
  invariants passing   4 of 12
  profiles quarantined 6 of 9  [p-a1 p-b1 p-c1 p-c2 p-d1 p-d2]
  releasable volume    3400 of 8200 requests/hour
  projected spend      $173.40 of $380.00 per day
  verdict              HOLD: fix I1, I4, I5, I6, I7, I8, I10, I11 first