Observability and Ban SLOS
Your Failure Mode Is Invisible From the Outside
A conventional error rate tells you whether the requests you made succeeded. It says nothing about the requests you should have stopped making, the ones that returned a perfectly good 200 and an empty shell, or the ones that were slow enough that your timeout killed them before the server answered. When the failure mode is "your IP is flagged", the dashboard is the thing that has to change.
Ban Detection Is an Inference Problem
You are given a status code, a body and a latency. Everything else is a distribution you have to reason about. Five signals, in descending order of reliability.
Status code. The easy case. A 403 or 451 is a verdict. A 401 on a session you hold is a verdict. A 429 is a limiter, not a ban, and treating the two the same is the most common classification error in scrapers. A 503 is usually your target's problem and must be excluded from your ban rate, or your SLO is measuring someone else's uptime.
Body shape. More reliable than the status code in the cases that matter. A soft block is usually a 200 carrying an interstitial, a challenge page, or an application shell that renders to nothing. The tells are structural rather than textual: a body that dropped from 41 KB to 640 bytes, a Set-Cookie for a clearance value you never saw before, the presence of a vendor's challenge script, or the absence of the landmarks your extractor depends on. Validate content, not codes; that is the whole argument.
Latency shift. A throttle that returns 200 has to give itself away somewhere, and the place is the latency distribution. Watch the median against a per-endpoint baseline and watch a dispersion measure alongside it, because the difference between "slow" and "throttled" is that a slow origin has a wide spread and a throttle has a narrow one at a much higher floor.
Challenge rate. The challenge-to-request ratio is a graded signal where the status code is binary, which makes it the best early warning you have. A run at 2 challenges per thousand against a 30-day-old account, rising to 20 per thousand, is a score crossing a soft threshold, and you can act on it while the account is still recoverable.
Correlation with identity. A ban that appears simultaneously on six profiles sharing one exit is a network-level verdict. A ban on one profile out of six sharing that exit is an identity-level verdict, and the recovery is completely different.
Separating a Ban From a Broken Scraper
This is the diagnostic that saves the most time, and it is worth writing down as a decision order rather than guessing.
| Symptom | Likely verdict | Distinguishing test |
|---|---|---|
| 403 on every host from one exit | IP ban | another exit from the same pool works |
| 403 on one profile, others fine | identity ban | new profile on the same exit works |
| 200, body 640 bytes | soft block or ban | a fresh session on the same exit also gets 640 |
| 200, body 0 bytes, landmarks gone | banned, null-rounded | compare against a known-good capture |
| 200, latency 20x baseline | silent throttle | does latency recover after a rate cut |
| 5xx across all profiles and hosts | your target or your egress | does it reproduce outside your pipeline |
| 200, correct body, parser error | your scraper | the HTML parses in a browser and not in your code |
| empty results only on one endpoint | null-rounding | same endpoint at a lower rate |
The asymmetry to remember: a ban is a state, a breakage is an event. A break has a cause you can find in a deploy, a config change or a dependency; a ban has a cause you can only infer from the pattern. If you changed something in your own pipeline in the last 24 hours, suspect the pipeline first, and prove it by running the last known-good build against the same target with a canary profile.
Two specific traps. First, a 200 whose body is a correct page of the wrong thing -- an empty category, a consent wall, a region notice -- is not detectable by size or landmark alone and needs a content assertion specific to the page type. Second, a ban that your own retry logic is causing will look intermittent, because every retry succeeds after a cooldown, and the cooldown is long enough to hide the pattern from a dashboard sampled every five minutes.
The Metrics That Matter
Six, and the last two are the ones that are usually missing.
- Success ratio. Delivered pages over judgeable requests. Exclude transport failures from the denominator, because a 503 from the target's edge is not your success rate. But count it in the bill, which the snippet below does.
- Challenge rate. Challenges per thousand requests, per identity, tracked as a trend. This is the early-warning metric and it should alert on the slope, not the value.
- Ban incidence. Banned outcomes over total outcomes, per vendor, as an SLO. See below.
- Time to detect. The gap between the target's first signalling response and the first action your pipeline takes about it. This is not a property of the response stream; it is a property of your alerting and your runbook, and it is measured by how long the first 403 sat in a queue before a worker was told to stop. Most operations have no idea what this number is, and it is usually minutes to hours.
- Time to recover. The gap between the ban and the next good page on that identity, including the cooldown you chose rather than the minimum the server would have allowed.
- Cost per delivered page. Bytes fetched, requests issued, browser context-seconds and proxy bytes, all divided by pages that actually parsed. This is the metric that tells you whether a rate cut was worth it, and it is the only one that survives a conversation with whoever owns the bill.
Add the two that describe the fleet rather than the traffic: profiles retired per week, and accounts lost per week. Both are leading indicators of the two above and both are cheap to track.
Ban Rate as an SLO
An SLO is a number with a time window and a consequence, and a ban rate without both is just a metric you looked at once.
Set it per vendor, over a rolling window, with an error budget:
- Objective, usually ban rate at or below 1 percent of outcomes over a 24-hour window. For an authenticated target with a finite pool of accounts, use absolute counts instead -- zero bans per week -- because 1 percent of 200 requests is 2 bans and that is not a rate you can hold.
- Error budget, the number of bans you can absorb before the window is blown. If the objective is 5 bans per week and you have 4, you are allowed one more. A blown budget changes your behaviour: reduce rate, retire the worst identity, stop the job that is causing it. That consequence is the point of writing it down.
- Window, long enough to be stable. A daily ban rate on a low-volume vendor is 0 or 33 percent; use a weekly window there.
- Exclusions, declared in advance. A ban that arrived while you were running a load test you knew was out of policy is excluded if you said so first. Exclusions decided after the fact are how SLOs get abandoned.
The chart that makes this work is a stacked classification per vendor over time, with the objective drawn as a line. It immediately shows the thing a global error rate hides: one vendor at 8 percent and two at zero averages to 2.7, which reads as "within tolerance" while a third of your targets is on fire.
Why One Global Number Fails
Aggregation is a liability here, and there are four distinct reasons.
- Different vendors fail differently. In the run below, all three vendors are at zero challenges on two of them and non-zero on one, and the ban rates are 7.1, 8.3 and 10.0 percent. An average of 8.5 tells you nothing about which one to stop working on.
- Different denominators. A vendor that gives you 40 requests an hour and one that gives you 4,000 must not be averaged per-request. Weight by opportunity, and if the volumes differ by two orders of magnitude consider a per-vendor objective.
- Different consequences. A ban on an anonymous target costs you an address. A ban on an authenticated target costs you an account, and the account is the expensive asset. Same class of event, different SLO.
- Different recovery. One vendor recovers from a cooldown, another needs a new profile, another needs a new exit region and a new account. A single "we are banned" alert routes everyone to the wrong runbook.
Per-vendor dashboards are not a reporting preference. They are the minimum granularity at which an alert can carry an action.
# Classify recorded HTTP outcomes, then hold a ban SLO against the result.
# One hour of traffic, inlined. 'at' is a logical second offset, not a clock.
SLO = {"ban_rate": 0.01, "delivery": 0.95, "challenge_rate": 0.15}
# vendor, at, status, ms, bytes, markers present, landmarks missing
OUTCOMES = [
("northwind", 0, 200, 84, 41230, [], []),
("contoso", 3, 200, 142, 28400, [], []),
("fabrikam", 7, 200, 61, 12400, [], []),
("northwind", 9, 429, 12, 190, [], []),
("contoso", 13, 200, 151, 29100, [], []),
("fabrikam", 17, 200, 58, 12380, [], []),
("northwind", 19, 200, 260, 5140, ["cf-browser-verification"], []),
("fabrikam", 21, 403, 22, 480, ["account-restricted"], []),
("contoso", 24, 403, 27, 512, ["unusual-traffic"], []),
("fabrikam", 26, 200, 63, 12390, [], []),
("northwind", 31, 200, 88, 41100, [], []),
("fabrikam", 35, 200, 57, 12420, [], []),
("contoso", 38, 200, 147, 28600, [], []),
("northwind", 43, 200, 240, 5090, ["cf-browser-verification"], []),
("fabrikam", 46, 200, 1890, 12400, [], []),
("northwind", 47, 200, 84, 0, [], ["price", "sku"]),
("contoso", 49, 200, 139, 28800, [], []),
("fabrikam", 51, 200, 1840, 12400, [], []),
("northwind", 52, 403, 31, 640, ["access-denied"], []),
("fabrikam", 57, 200, 64, 12370, [], []),
("contoso", 58, 200, 144, 28300, [], []),
("northwind", 62, 502, 8, 210, [], []),
("fabrikam", 66, 200, 55, 12410, [], []),
("contoso", 71, 503, 6, 180, [], []),
("northwind", 76, 200, 91, 40980, [], []),
("fabrikam", 80, 200, 1890, 12400, [], []),
("contoso", 85, 429, 9, 140, [], []),
("northwind", 90, 200, 95, 40850, [], []),
("fabrikam", 94, 200, 63, 12390, [], []),
("contoso", 99, 200, 610, 27600, [], []),
("northwind", 104, 200, 86, 41200, [], []),
("fabrikam", 108, 200, 58, 12380, [], []),
("contoso", 113, 200, 142, 28400, [], []),
]
CHALLENGE_MARKERS = ("cf-browser-verification", "px-captcha", "datadome-captcha")
BLOCK_MARKERS = ("access-denied", "unusual-traffic", "account-restricted")
SILENT_THROTTLE_MS = 800
def classify(status, ms, size, markers, missing):
# Order matters: a transport failure is not a ban, and a ban is not a 429.
if status == 429:
return "rate-limited", f"429 at {ms} ms, page withheld"
if status >= 500:
return "broken", f"{status} from the edge, not the origin"
if status in (401, 403, 451):
return "banned", markers[0] if markers else "bare status"
hit = [m for m in markers if m in CHALLENGE_MARKERS]
if hit:
return "soft-challenge", f"{hit[0]} at {ms} ms"
if any(m in BLOCK_MARKERS for m in markers):
return "banned", markers[0]
if missing:
return "broken", f"{len(missing)} landmarks gone, {size} bytes"
if ms >= SILENT_THROTTLE_MS:
return "rate-limited", f"{ms} ms with a 200, silent throttle"
return "ok", f"{size} bytes in {ms} ms"
def kb(n):
return f"{n / 1024:.1f} KB"
def main():
classes = ("ok", "soft-challenge", "rate-limited", "banned", "broken")
tally = dict.fromkeys(classes, 0)
notes, vendors, sizes = {}, {}, []
for vendor, at, status, ms, size, markers, missing in OUTCOMES:
cls, why = classify(status, ms, size, markers, missing)
tally[cls] += 1
sizes.append((cls, size))
notes.setdefault(cls, []).append(f"t={at:>3}s {vendor:<10} {status} "
f"{why}")
row = vendors.setdefault(vendor, dict.fromkeys(classes, 0))
row[cls] += 1
print(f"classified {len(OUTCOMES)} outcomes from one hour")
print(f" {'class':<16}{'n':>4}{'share':>8}")
for cls in classes:
print(f" {cls:<16}{tally[cls]:>4}"
f"{100.0 * tally[cls] / len(OUTCOMES):>7.1f}%")
for cls in ("banned", "soft-challenge", "rate-limited", "broken"):
print(f" {cls}:")
for line in sorted(notes.get(cls, [])):
print(f" {line}")
print()
print("per vendor, because one global error rate hides all three")
print(f" {'vendor':<11}{'n':>4}{'ok':>4}{'chal':>6}{'rate':>6}"
f"{'ban':>5}{'brk':>5} ban rate outage wasted")
for vendor in sorted(vendors):
row = vendors[vendor]
n = sum(row.values())
rate = 100.0 * row["banned"] / n
verdict = "ok" if rate <= 100.0 * SLO["ban_rate"] else "BREACH"
mine = sorted(o for o in OUTCOMES if o[0] == vendor)
bans = [o[1] for o in mine
if classify(o[2], o[3], o[4], o[5], o[6])[0] == "banned"]
outage = "none"
wasted = 0
if bans:
after = [o for o in mine if o[1] > bans[0]]
dead = [o for o in after
if classify(o[2], o[3], o[4], o[5], o[6])[0] != "ok"]
back = [o[1] for o in after
if classify(o[2], o[3], o[4], o[5], o[6])[0] == "ok"]
outage = f"{back[0] - bans[0]:>3}s" if back else "open"
wasted = len(dead)
print(f" {vendor:<11}{n:>4}{row['ok']:>4}{row['soft-challenge']:>6}"
f"{row['rate-limited']:>6}{row['banned']:>5}{row['broken']:>5}"
f" {rate:5.1f}% {verdict:<6} {outage:>7} {wasted:>6}")
judged = len(OUTCOMES) - tally["broken"]
bans = tally["banned"]
delivery = 100.0 * tally["ok"] / judged
ban_rate = 100.0 * bans / len(OUTCOMES)
chal = 100.0 * tally["soft-challenge"] / len(OUTCOMES)
print()
print(f"SLO ban <= {100.0 * SLO['ban_rate']:.1f}% delivery >= "
f"{100.0 * SLO['delivery']:.1f}% of non-broken "
f"challenge <= {100.0 * SLO['challenge_rate']:.1f}%")
print(f"got ban {ban_rate:.1f}% "
f"{'PASS' if ban_rate <= 100 * SLO['ban_rate'] else 'FAIL'}"
f" delivery {delivery:.1f}% "
f"{'PASS' if delivery >= 100 * SLO['delivery'] else 'FAIL'}"
f" challenge {chal:.1f}% "
f"{'PASS' if chal <= 100 * SLO['challenge_rate'] else 'FAIL'}")
print(f" {judged} of {len(OUTCOMES)} requests were judgeable; "
f"{tally['broken']} were the target's fault and are excluded "
f"from delivery, not from the bill")
kept = sum(s for c, s in sizes if c == "ok")
fetched = sum(s for _c, s in sizes)
wasted_reqs = len(OUTCOMES) - tally["ok"]
print(f" fetched {kb(fetched)}, kept {kb(kept)} "
f"({100.0 * kept / fetched:.1f}%); {wasted_reqs} requests bought "
f"nothing at {kb(fetched / len(OUTCOMES))} each")
main()
classified 33 outcomes from one hour
class n share
ok 20 60.6%
soft-challenge 2 6.1%
rate-limited 5 15.2%
banned 3 9.1%
broken 3 9.1%
banned:
t= 21s fabrikam 403 account-restricted
t= 24s contoso 403 unusual-traffic
t= 52s northwind 403 access-denied
soft-challenge:
t= 19s northwind 200 cf-browser-verification at 260 ms
t= 43s northwind 200 cf-browser-verification at 240 ms
rate-limited:
t= 9s northwind 429 429 at 12 ms, page withheld
t= 46s fabrikam 200 1890 ms with a 200, silent throttle
t= 51s fabrikam 200 1840 ms with a 200, silent throttle
t= 80s fabrikam 200 1890 ms with a 200, silent throttle
t= 85s contoso 429 429 at 9 ms, page withheld
broken:
t= 47s northwind 200 2 landmarks gone, 0 bytes
t= 62s northwind 502 502 from the edge, not the origin
t= 71s contoso 503 503 from the edge, not the origin
per vendor, because one global error rate hides all three
vendor n ok chal rate ban brk ban rate outage wasted
contoso 10 7 0 1 1 1 10.0% BREACH 14s 2
fabrikam 12 8 0 3 1 0 8.3% BREACH 5s 3
northwind 11 5 2 1 1 2 9.1% BREACH 24s 1
SLO ban <= 1.0% delivery >= 95.0% of non-broken challenge <= 15.0%
got ban 9.1% FAIL delivery 66.7% FAIL challenge 6.1% PASS
30 of 33 requests were judgeable; 3 were the target's fault and are excluded from delivery, not from the bill
fetched 540.5 KB, kept 491.9 KB (91.0%); 13 requests bought nothing at 16.4 KB each
The last two sections are the argument. Twenty of 33 requests were judged ok; delivery against the judgeable 30 is 66.7 percent, well under the 95 percent objective, while the challenge rate passes comfortably at 6.1 percent. All three vendors breach the ban objective and all three by a similar margin -- which is itself the finding, because it means the problem is your behaviour and not one broken target. And the outage column shows something the ban rate does not: northwind was down for 24 seconds against fabrikam's 5, so the vendor you would have deprioritised first is the one that actually cost you the most time.
Time to Detect, Measured
Time to detect is a property of your system and it has to be measured, not assumed. Three instruments, cheapest first.
- An alert on the first banned classification, per vendor. Not on a rate, not on a threshold over an hour -- on the first occurrence, because one 403 on an authenticated identity is an event worth waking someone for. Then a second alert on the rolling rate for the pattern.
- A synthetic canary. One request per vendor per minute from a sacrificial identity you are willing to lose. When the canary bans, you know within a minute, and you know it is the target and not your pipeline because the canary runs the same code. This is the single highest-value instrument on the list, and testing stealth pipelines covers the rest of the suite that has to exist around it.
- A queue-depth assertion. When the first 403 lands in the dead-letter bucket, the queue's requeue rate jumps. Alerting on requeue rate detects the ban even when the classification is wrong, which is precisely when you need a second opinion.
Report the gap, not just the alert. "Alerted in 40 seconds, worker cooled down in 95 seconds" is a measurement; "we got an alert" is not.
What to Log, Per Request
Enough to reconstruct the decision later, which is a small fixed set.
- Identity: profile id, exit address, ASN, credential id. Without these three you cannot attribute anything.
- Classified outcome: the label and the reason, not just the status code.
403 access-deniedand403 bare statusmean different things about how much to trust the verdict. - Latency, body size, and body digest. The digest is what lets you compare this response to a known-good one without storing the response.
- Response header names and
Set-Cookienames. Not values. The names tell you the challenge machinery changed; the values are credentials and do not belong in a log. - Challenge machinery seen on this request, and whether it resolved.
- Attempt number and the classification of the previous attempt on this item.
Never log cookie values, authorization headers, or full response bodies from authenticated targets. The first two are credentials and the third is someone else's data; keeping them buys nothing that a digest and a size do not.
A Ban Runbook
Per vendor, written down, because nobody remembers the recovery ladder under pressure.
- Detect. Canary or alert. Record identity, exit, credential, and the classification with its reason.
- Freeze. The identity and its exit stop immediately. Continued requests escalate the block and extend it, and this is the step that people skip because it feels like giving up.
- Classify the scope. IP, identity or credential. One request from a different exit on the same profile answers IP. One request from the same exit on a fresh profile answers identity.
- Cool down. Hours for a 403, minutes for a 429. Reputation decays on a curve you do not control, so choose the conservative end.
- Return warm. Homepage, one detail page, then a tenth of the normal rate for five minutes. A re-rejection in that window means it has not lifted.
- Record. Which step resolved it, how long it took, and what preceded it. This is the only reason a second incident on the same vendor is cheaper than the first.
Encode it as a machine-readable playbook keyed by classification, not by prose, because a runbook that requires reading is a runbook that will be read wrong at three in the morning. And keep ban recovery and backoff for the algebra while this covers the decision.
Closing the Loop
An SLO you do not act on is documentation. Three habits turn the measurement into behaviour.
- The objective is a rate limit on your own operations. When the ban-rate budget is half spent, cut the crawl rate for the rest of the window. That is a real, pre-agreed consequence, and pre-agreed is what makes it survive a deadline.
- Every breach produces one change. Not seven. An incident review that produces a list of improvements produces none of them, and the ban rate does not move because you read the postmortem.
- Track the trend against the model, not against yourself. If your measured challenge rate is drifting up on young profiles while flat on mature ones, the model changed and your warmup policy should change with it. Keeping current is where the drift detection lives, and challenge debugging and observability is where the per-request instrumentation to make it work comes from.
The honest summary is that most scrapers are not short of instrumentation; they are short of classification. The difference between the two dashboards above is a twenty-line function and a decision order, and that function is the highest-leverage thing in this lesson.