The Bot You Blocked Was a Researcher With a Question

Every bot-management console has one chart that the team looks at far more often than the others: requests blocked per hour. Raise it and the chart goes up, so the team concludes the defence is working, and the assumption baked into that conclusion is that every blocked client was an attacker. That assumption is wrong most of the time. The traffic you block is dominated by search engines you failed to allowlist, price monitors, academic replication, SEO tooling, ticket aggregators, archiving projects, security researchers filing reports, and people saving their own reading history. Detecting automation is the easy half of the problem and it is the half everyone builds; deciding what to do about a client you have detected is the half that decides whether your site is usable and whether you are on the right side of a contract.

Who is actually crawling you

Before you tune anything, sort your traffic by purpose rather than by score. In practice six populations account for nearly all the volume, and only two of them are hostile.

Population Share of blocked volume Wants Your cheapest correct answer
Search engine indexer high Pages, ranked Allowlist by verified identity, not by pattern
Academic or journalist research low A corpus over time Named rate limit, an API key, a contact address
SEO and audit tooling medium Your own site's structure A sitemap, a crawl budget, a published API
Price aggregator or arbitrage monitor medium Fresh prices and stock A licensed feed or a rate-limited public endpoint
Archiver (web archive, personal save) low A durable copy Allow it, and set an honest crawl rate
Hostile: republisher, bottener, credential attacker medium in volume, high in cost Whatever you protect Authenticate, quota, challenge, ban, escalate

The column that matters is the last one, and the reason it is not the biggest is economic: an honest aggregator can resell your data, a hostile republisher can only resell it while you keep changing things so their copy goes stale. The population that should worry you is the one where your defence has no feedback loop - you cannot tell from a blocked request whether the block worked, so the defence gets no credit and the false positives get all the blame.

The taxonomy matters because the two hostile cases need different controls. A republisher is a volume and freshness problem, solved by authentication, per-account quotas and making the value of a fresh row decay. A credential attacker is an account problem, solved by multi-factor authentication, breached-credential checks, per-account lockout and device bans. Blend them into one bot score and you will rank both badly, which is why the classifier below routes before it scores.

robots.txt is a contract, not a control

robots.txt is the most-read and least-enforced document on the web. Three facts about it shape how you should use it.

First, it is advisory by construction. A well-behaved client fetches it, parses Disallow directives, and adjusts. Nothing stops a client from ignoring it, and nothing you write in it can. The robots.txt file is not an access control list and it is not a security boundary; treating it as one produces the specific failure of believing your expensive pages are protected.

Second, it is cheap, legible and mutually intelligible. It is the one mechanism by which a crawler you have never met can learn your preference without a negotiation. It costs you one cached file on your own domain, it is human-readable, and it can be linked to from a real page with a real email address. That combination - free, standard, legible - is why it is worth more than most of the paid controls for the population of well-behaved clients.

Third, its most valuable directives are the ones that are not Disallow. Crawl-delay (still honoured by some, ignored by Google, ignored by most new tooling) and Sitemap: are the two that change behaviour. A Sitemap: line plus a sane rate beats a wall of Disallow lines, because it converts an expensive discovery crawl into a cheap fetch of a manifest you wrote.

The honest framing: robots.txt is how you publish a contract. It is where you say what a cooperative client may do, what rate you will tolerate, and how to reach a human. Defence in this course runs the other way - it is a signal the site reads - but the same file is the only piece of written policy in the whole protocol.

Rate limits that actually hold

A rate limit is the control with the best ratio of cost to effect, and the one most often deployed wrong. The failure modes are specific.

Per-IP limits are the wrong unit. One abusive client with a residential proxy pool is a thousand IPs; one legitimate office is a thousand people behind a single NAT egress. A per-IP limit punishes the office and does nothing to the pool. Key on a session identifier when you have one, on a hashed (IP, user-agent, accept-language) triple when you do not, and keep a separate coarser IP bucket purely as abuse backstop.

Burst allowances decide whether you hurt readers. A bucket of 60 requests with a refill of 1 per second lets a legitimate page load with a large asset graph through untouched and stops a flat-rate scraper after the first few seconds of throughput. Set burst from your own page-load profile, not from a round number.

Sliding windows beat fixed windows. A fixed per-minute counter can be gamed exactly at the boundary and produces a periodic false-positive spike that looks like a bug to support. A sliding window over the last 60 seconds costs more memory and removes the artifact.

Concurrency caps beat rate caps for the expensive paths. A search endpoint that costs the site 40 ms of database time should have a global concurrency limit, not just a per-client rate. A client at 5 requests per second is polite; 200 clients at 5 requests per second is not, and only a fleet-level cap sees it.

Spend your engineering here before you spend it on fingerprinting. The behaviour you want to bound is cost per request, and rate and concurrency limits are the only controls that bound it directly.

Authentication, API tiers and per-account quotas

If you would rather not serve anonymous traffic at all, the clean answer is to make the authenticated path obviously better, and this is where Official APIs and Alternative Data meets the defensive side. An API key with a documented quota is worth more than a WAF rule because it identifies the client, it meters it exactly, it can be revoked, and it gives you somewhere to send an invoice or a takedown notice.

The shape that works in practice:

Tier Who gets it Quota What you give up
Anonymous Anyone 60 req/min burst 60, HTML only Nothing, and it is your sensor
Registered Free key, email verified 2,000 req/hour, JSON Identity, and a revocation handle
Partner Signed agreement 50,000 req/hour, bulk endpoints A contract you can point at in a dispute
Internal Your own services Unmetered A separate service identity, never a browser profile

Two rules make this work. Quotas are per account, not per key, so a client cannot mint ten keys to get ten quotas; and every account has a hard daily row cap on the fields that matter commercially, because the moment the data stops being fresh the scraper's business model changes. Per-account quotas also give you the one signal you cannot fabricate: how many authenticated requests a given identity made. A genuine client has a normal curve; a burner-account operation has a spike at exactly the quota and then a new account.

Graduated friction and what each rung costs

The escalation ladder in How Anti-Bots Work is described from the client side. From the server side, each rung is a cost you are choosing to impose on somebody, and the table below prices it on a 40 ms connection.

  • L0 allow. Zero cost, and you learn nothing beyond what the log already tells you.
  • L1 allow and count. Same cost, but you now have a metered identity: an ID, a volume, a repeat-visit rate.
  • L2 rate limit. Free to the reader, unless your token is shared across a household or an office.
  • L3 soft challenge. About 80 ms and three extra round trips. Invisible, and roughly 15 to 40 percent of scripted clients fail it.
  • L4 interstitial. About 900 ms and four round trips, plus an accessibility problem for anyone relying on the page loading.
  • L5 CAPTCHA. 8 to 20 seconds of a human's attention, and an accessibility catastrophe for screen-reader users, dyslexia, and anyone on a slow connection. Treat it as a last resort and instrument how often it fires on first-time human sessions.
  • L6 block. Free technically, and it is the rung that generates the complaints, the chargebacks and the bad press.

The discipline is to publish which rungs you serve to which population and to measure conversion and support-ticket rate by rung. A defence whose false-positive rate you cannot attribute is a defence you cannot tune, and the Signal Inventory lesson makes the same point from the collection side: log the score contributions, never just the verdict, or you can never delete a signal that is only adding friction.

A proportional classifier, with its own error bars

Here is the whole analysis, on a synthetic 24-hour edge view. Every input is a field the request log already contains; none of it requires serving a challenge, which is exactly why it is cheap to run on every request. The classifier routes authentication-shaped traffic to a different control family, scores the rest on six passive sub-scores, and places each client on the ladder.

import math

# One row per client, aggregated by the edge over a rolling 24-hour window.
# This is the whole input a proportional defence needs: every field is
# already in the request log, and none of it required running JavaScript.
#
#   requests      total requests in the window
#   unique_paths  distinct URL paths touched
#   top_paths     (path, hits) for the paths with the most hits
#   auth_frac     share of requests hitting login, session-create or account
#                 endpoints. The edge knows this from its own route table;
#                 never infer it from behaviour.
#   egress_ips    distinct source addresses
#   agents        (user agent family, percent of requests)
#   sessions      (session cookie value, requests, distinct ips it was seen on)
#   honoured      declared rules this client demonstrably obeys
#   declared      purpose, when the client proved its identity out of band
OBSERVED = [
    {
        "name": "partner-price-feed",
        "requests": 8640,
        "unique_paths": 14,
        "top_paths": [("/api/price", 4200), ("/api/stock", 3100),
                      ("/api/catalog", 1340)],
        "auth_frac": 0.0,
        "egress_ips": 4,
        "agents": [("partner-sdk/2.1", 100)],
        "sessions": [("sess_9f21", 8640, 1)],
        "honoured": ["declared-ua", "robots.txt", "api-key", "contact"],
        "declared": "licensed feed",
    },
    {
        "name": "verified-search-crawler",
        "requests": 41200,
        "unique_paths": 41200,
        "top_paths": [("/p/%d" % n, 1) for n in range(1, 13)],
        "auth_frac": 0.0,
        "egress_ips": 180,
        "agents": [("chrome/131", 100)],
        "sessions": [("sess_a41b", 2600, 1), ("sess_c7e0", 2600, 1),
                     ("sess_d913", 2600, 1), ("sess_fb20", 2600, 1)],
        "honoured": ["declared-ua", "robots.txt", "crawl-delay", "sitemap"],
        "declared": "search index",
    },
    {
        "name": "academic-longitudinal-study",
        "requests": 2600,
        "unique_paths": 310,
        "top_paths": [("/archive/2026", 900), ("/archive/2025", 880),
                      ("/search", 620), ("/docs/113", 200)],
        "auth_frac": 0.0,
        "egress_ips": 1,
        "agents": [("firefox/133", 100)],
        "sessions": [("sess_5c10", 2600, 1)],
        "honoured": ["robots.txt", "crawl-delay", "contact"],
        "declared": "research",
    },
    {
        "name": "seo-audit-crawler",
        "requests": 18400,
        "unique_paths": 18100,
        "top_paths": [("/p/%d" % n, 1) for n in range(1, 13)],
        "auth_frac": 0.0,
        "egress_ips": 6,
        "agents": [("headless-chrome/131", 34), ("chrome/130", 33),
                   ("chrome/131", 33)],
        "sessions": [("sess_11aa", 4600, 6), ("sess_22bb", 4600, 6),
                     ("sess_33cc", 4600, 6)],
        "honoured": ["crawl-delay", "sitemap"],
        "declared": "seo tool",
    },
    {
        "name": "personal-archiver",
        "requests": 740,
        "unique_paths": 740,
        "top_paths": [("/blog/%d" % n, 1) for n in range(1, 13)],
        "auth_frac": 0.0,
        "egress_ips": 1,
        "agents": [("firefox/133", 100)],
        "sessions": [("sess_7e44", 740, 1)],
        "honoured": ["robots.txt", "crawl-delay"],
        "declared": "personal archive",
    },
    {
        "name": "flash-sale-monitor",
        "requests": 612000,
        "unique_paths": 9400,
        "top_paths": [("/checkout/session", 210000), ("/cart", 180000)]
        + [("/p/%d" % n, 9000) for n in range(1, 10)],
        "auth_frac": 0.02,
        "egress_ips": 310,
        "agents": [("chrome/131", 61), ("chrome/130", 20),
                   ("safari/18", 11), ("firefox/133", 8)],
        "sessions": [("sess_a1", 204000, 74), ("sess_b2", 204000, 74),
                     ("sess_c3", 204000, 74)],
        "honoured": [],
        "declared": None,
    },
    {
        "name": "resale-listing-republisher",
        "requests": 94000,
        "unique_paths": 91200,
        "top_paths": [("/event/%d" % n, 1) for n in range(1, 13)],
        "auth_frac": 0.0,
        "egress_ips": 88,
        "agents": [("chrome/131", 44), ("chrome/130", 25),
                   ("headless-chrome/131", 20), ("safari/18", 11)],
        "sessions": [("sess_01x", 23500, 22), ("sess_02y", 23500, 22),
                     ("sess_03z", 23500, 22), ("sess_04w", 23500, 22)],
        "honoured": ["crawl-delay"],
        "declared": None,
    },
    {
        "name": "credential-stuffing",
        "requests": 148000,
        "unique_paths": 26,
        "top_paths": [("/login", 121000), ("/api/session", 21000),
                      ("/account/reset", 6000)],
        "auth_frac": 1.0,
        "egress_ips": 4200,
        "agents": [("python-requests/2.32", 100)],
        "sessions": [("sess_s%02d" % i, 3700, 210) for i in range(1, 41)],
        "honoured": [],
        "declared": None,
    },
]

MINUTES = 24 * 60
ABUSIVE = {"flash-sale-monitor", "resale-listing-republisher",
           "credential-stuffing"}
BENIGN = {"partner-price-feed", "verified-search-crawler",
          "academic-longitudinal-study", "seo-audit-crawler",
          "personal-archiver"}
AUTH_FLOOR = 0.50

# Six sub-scores in [0, 1], each normalised against the value at which the
# signal stops being ordinary and starts being a population outlier. The cut
# points are policy, not physics, and they are the thing to argue about.
SUBSCORES = [
    ("rate", 0.28, lambda f: clamp((f["rpm"] - 4.0) / 26.0)),
    ("novelty", 0.22, lambda f: clamp((f["novelty"] - 0.25) / 0.65)),
    ("concentration", 0.10, lambda f: clamp((f["conc"] - 1.40) / 2.00)),
    ("ua-spread", 0.16, lambda f: clamp((f["agent_spread"] - 0.02) / 0.45)),
    ("session", 0.14, lambda f: 0.5 * clamp((f["reuse"] - 0.90) / 0.09)
     + 0.5 * clamp(f["cross_ip"])),
    ("rules", 0.10, lambda f: 1.0 - clamp(f["rules_ok"] / 0.75)),
]

LADDER = [
    (0.16, "L0 allow", "no extra work, log only", 0, "0 ms"),
    (0.30, "L1 allow + count", "meter it, publish the API", 0, "0 ms"),
    (0.45, "L2 rate limit", "token bucket per session id", 0, "0 ms"),
    (0.60, "L3 soft challenge", "invisible script, no clearance needed", 3,
     "80 ms"),
    (0.75, "L4 interstitial", "visible check, issues clearance cookie", 4,
     "900 ms"),
    (0.88, "L5 captcha", "hard gate, the reader pays to pass", 5, "12000 ms"),
    (1.01, "L6 block", "403, retain logs for legal review", 6, "0 ms"),
]
RUNG_INDEX = dict((name, i) for i, (_c, name, _a, _h, _l) in enumerate(LADDER))


def out(line=""):
    print(line.rstrip())


def clamp(x):
    return 0.0 if x < 0.0 else (1.0 if x > 1.0 else x)


def concentration(top_paths):
    # Shannon entropy in bits over the busiest paths. 0.0 is one path only,
    # 3.58 is the twelve busiest paths hit exactly equally often.
    total = float(sum(hits for _p, hits in top_paths))
    return -sum((h / total) * math.log2(h / total) for _p, h in top_paths)


def features(row):
    req = float(row["requests"])
    return {
        "rpm": req / MINUTES,
        "novelty": row["unique_paths"] / req,
        "conc": concentration(row["top_paths"]),
        "auth": row["auth_frac"],
        "agent_spread": 1.0 - max(sh for _n, sh in row["agents"]) / 100.0,
        "reuse": max(h for _s, h, _i in row["sessions"]) / req,
        "cross_ip": sum(1 for _s, _h, ips in row["sessions"] if ips > 1)
        / float(len(row["sessions"])),
        "rules_ok": len(row["honoured"]) / 4.0,
    }


def score(f):
    return sum(w * fn(f) for _n, w, fn in SUBSCORES)


def rung(s):
    for _ceiling, name, action, _hops, _lat in LADDER:
        if s < _ceiling:
            return name, action
    return LADDER[-1][1], LADDER[-1][2]


out("24-hour edge view: %d clients, %s requests, passive signals only"
    % (len(OBSERVED), format(sum(r["requests"] for r in OBSERVED), ",")))
out("weights " + " ".join("%s=%.2f" % (n, w) for n, w, _f in SUBSCORES))
out()
out("%-30s %7s %6s %6s %7s %6s %6s %s"
    % ("client", "req/min", "novel", "conc", "ua-spr", "sess", "score",
       "rung"))

verdicts, routed = {}, []
for row in OBSERVED:
    f = features(row)
    s = score(f)
    if f["auth"] >= AUTH_FLOOR:
        routed.append((row["name"], row["requests"], f["auth"], s))
        verdicts[row["name"]] = (s, "routed: account controls")
    else:
        verdicts[row["name"]] = (s, rung(s)[0])
    out("%-30s %7.1f %6.3f %6.2f %7.2f %6.2f %6.2f %s"
        % (row["name"], f["rpm"], f["novelty"], f["conc"], f["agent_spread"],
           f["reuse"], s, verdicts[row["name"]][1]))

out()
out("what each rung costs the reader on a 40 ms connection")
for _c, name, action, hops, lat in LADDER:
    out("  %-19s %-42s %8s  %d extra round trips"
        % (name, action, lat, hops))

out()
out("one-signal screen: sub-score at least half saturated, no corroboration")
for name, _weight, fn in SUBSCORES:
    flagged = sorted(r["name"] for r in OBSERVED if fn(features(r)) >= 0.50)
    bad = [n for n in flagged if n in ABUSIVE]
    fp = [n for n in flagged if n in BENIGN]
    out("  %-13s flagged %d/%d  abusive %d  legitimate %d  %s"
        % (name, len(flagged), len(OBSERVED), len(bad), len(fp),
           " ".join(fp) or "-"))

out()
out("route first, rank second: endpoint semantics beat behavioural inference")
for n, req, auth, s in sorted(routed):
    out("  %-30s %8s requests, %.0f%% to auth endpoints, score %.2f"
        % (n, format(req, ","), 100 * auth, s))
    out("    off the crawler ladder entirely: MFA, per-account lockout, "
       "device and ASN bans")

in_scope = [r["name"] for r in OBSERVED if r["auth_frac"] < AUTH_FLOOR]
challenged = [n for n in in_scope if RUNG_INDEX[verdicts[n][1]] >= 3]
caught_bad = sorted(n for n in challenged if n in ABUSIVE)
false_pos = sorted(n for n in challenged if n in BENIGN)
out()
out("crawler ladder: %d clients in scope, %d held at L3 or above"
    % (len(in_scope), len(challenged)))
out("  abusive and caught: %d of %d  [%s]"
    % (len(caught_bad), len(ABUSIVE), ", ".join(caught_bad)))
out("  legitimate and challenged: %d of %d  [%s]"
    % (len(false_pos), len(BENIGN), ", ".join(false_pos)))
out("  precision %.2f  false-positive rate %.2f"
    % (len(caught_bad) / float(len(challenged)),
       len(false_pos) / float(len(BENIGN))))
out()
out("allowlist verified identity, keep the ladder for everyone else:")
for n in false_pos:
    out("  %-30s %.2f -> %-19s becomes L0"
        % (n, verdicts[n][0], verdicts[n][1]))
out()
out("detection tells you who to look at. the ladder is a cost decision,")
out("and every rung above L2 is paid for by someone who did nothing wrong.")
24-hour edge view: 8 clients, 925,580 requests, passive signals only
weights rate=0.28 novelty=0.22 concentration=0.10 ua-spread=0.16 session=0.14 rules=0.10

client                         req/min  novel   conc  ua-spr   sess  score rung
partner-price-feed                 6.0  0.002   1.45    0.00   1.00   0.09 L0 allow
verified-search-crawler           28.6  1.000   3.58    0.00   0.06   0.59 L3 soft challenge
academic-longitudinal-study        1.8  0.119   1.84    0.00   1.00   0.09 L0 allow
seo-audit-crawler                 12.8  0.984   3.58    0.66   0.25   0.68 L4 interstitial
personal-archiver                  0.5  1.000   3.58    0.00   1.00   0.42 L2 rate limit
flash-sale-monitor               425.0  0.015   2.03    0.39   0.33   0.61 L4 interstitial
resale-listing-republisher        65.3  0.970   3.58    0.56   0.25   0.90 L6 block
credential-stuffing              102.8  0.000   0.82    0.00   0.03   0.45 routed: account controls

what each rung costs the reader on a 40 ms connection
  L0 allow            no extra work, log only                        0 ms  0 extra round trips
  L1 allow + count    meter it, publish the API                      0 ms  0 extra round trips
  L2 rate limit       token bucket per session id                    0 ms  0 extra round trips
  L3 soft challenge   invisible script, no clearance needed         80 ms  3 extra round trips
  L4 interstitial     visible check, issues clearance cookie       900 ms  4 extra round trips
  L5 captcha          hard gate, the reader pays to pass         12000 ms  5 extra round trips
  L6 block            403, retain logs for legal review              0 ms  6 extra round trips

one-signal screen: sub-score at least half saturated, no corroboration
  rate          flagged 4/8  abusive 3  legitimate 1  verified-search-crawler
  novelty       flagged 4/8  abusive 1  legitimate 3  personal-archiver seo-audit-crawler verified-search-crawler
  concentration flagged 4/8  abusive 1  legitimate 3  personal-archiver seo-audit-crawler verified-search-crawler
  ua-spread     flagged 3/8  abusive 2  legitimate 1  seo-audit-crawler
  session       flagged 7/8  abusive 3  legitimate 4  academic-longitudinal-study partner-price-feed personal-archiver seo-audit-crawler
  rules         flagged 3/8  abusive 3  legitimate 0  -

route first, rank second: endpoint semantics beat behavioural inference
  credential-stuffing             148,000 requests, 100% to auth endpoints, score 0.45
    off the crawler ladder entirely: MFA, per-account lockout, device and ASN bans

crawler ladder: 7 clients in scope, 4 held at L3 or above
  abusive and caught: 2 of 3  [flash-sale-monitor, resale-listing-republisher]
  legitimate and challenged: 2 of 5  [seo-audit-crawler, verified-search-crawler]
  precision 0.50  false-positive rate 0.40

allowlist verified identity, keep the ladder for everyone else:
  seo-audit-crawler              0.68 -> L4 interstitial     becomes L0
  verified-search-crawler        0.59 -> L3 soft challenge   becomes L0

detection tells you who to look at. the ladder is a cost decision,
and every rung above L2 is paid for by someone who did nothing wrong.

Closing the Loop

You maintain the defence by tuning the ladder against observed outcomes, tracking time to detect and time to recover, and reviewing the effectiveness of each rung against real traffic patterns over time.