The List the Defender Already Keeps About You

Before a bot-management system decides anything about a request, it has already written a row about that request into a table. The row is assembled from dozens of independent observations, taken at different layers of the stack, by code that runs before, during and after your HTTP request. Most scrapers spend all their effort on the one layer they can see - the request itself - and never ask what else was recorded while it was in flight.

A signal is any observation the site can make about a request that is not the request's own content: who asked, from where, with what client, in what order, and what happened to the last ten thousand requests from the same neighbourhood. No single signal is damning. A Chrome user on a Comcast IP in Ohio who has visited the site before and clicked two links is unremarkable. The power is in the ledger: a modern WAF keeps a per-session risk score, and each signal is a weighted line item in it (see Risk Scoring and Session Profiles). This lesson is the catalogue that goes behind that score, one category at a time.

The Catalogue

Category Example signal Collected by Survives cookie clearing? Cost to the site
Transport TLS JA3/JA4 hash, ALPN list, GREASE values TLS terminator, no JS Yes, per connection CPU, ~1 ms per handshake
Transport QUIC transport parameters, HTTP/3 frame order Edge server Yes CPU, connection state
HTTP header order, sec-ch-ua, Sec-Fetch-*, Accept grammar HTTP parser Yes Effectively free
Network/IP ASN, whether the ASN is hosting, rDNS, geolocation Edge, GeoIP DB, DNS Yes Licence fee per lookup
Network/IP requests per IP per hour, distinct paths per IP Edge state store Yes Memory, millions of keys
Resolver which recursive resolver asked for the protected hostname Resolver-side logs, or the site if it runs one Yes Only if the site is the resolver
JS environment canvas and WebGL hashes, font metrics, audio hash, timezone, navigator surface Challenge script Yes, if stored server-side 20-80 ms of page time
JS environment webdriver, CDP artefacts, plugin lists, permissions state Challenge script Yes Included above
Behavioral pointer paths, key intervals, scroll velocity, dwell time Event listeners, post-hoc Yes Storage; the most expensive row
Persistent first-party cookie, localStorage probe, ETag, supercookie Site JS, server No, if cookie-based Storage, plus consent work
Account account age, login history, devices per account, MFA state Application No, if the account is recreated Application database
Historical past verdicts, challenge outcomes, captcha solves, past bans Vendor risk store Yes, indefinitely The vendor's core cost centre

Read the last column as an engineering budget, not a moral score. A challenge script that probes the environment costs the site 20 to 80 milliseconds of page latency per visitor; behavioural telemetry costs storage and an analytics pipeline; GeoIP and ASN data costs a licence. That asymmetry is why vendors converge on the same design: cheap passive signals by default, expensive active signals only for the minority of sessions a cheap signal already flagged. It is also why the answer to "which signal should I fix first" is almost never the behavioural one. Grouping by layer, as the first column does, tells you which subsystem a signal lives in and therefore what you have to change to move it. A TLS ClientHello is fixed before your application code runs. A navigator.languages value is whatever your environment says, so you own it outright.

Why Sites Collect More Than They Need

A WAF's failure mode is not the bot that gets through. It is the false positive: a real customer, on a corporate VPN, using a screen reader, gets a CAPTCHA or a block, and either fails to buy or files a ticket. That customer costs far more than a scraped row. Blocked legitimate traffic is measured directly - support tickets, conversion drop-off, complaint volume - while a bot that slips through costs nothing until someone acts on the data. So scoring is tuned to be permissive by default and strict under accumulation. A session clean on the first twenty cheap signals skips the expensive ones. A session that trips one expensive signal early often gets more of them, because the models are built to spend their budget where a human is watching. The other pressure is competitive: bot management is bought to beat the last vendor's scorecard, so features that measurably raise the score get instrumented even when their marginal detection value is small. If you are on the defending side, measure the cost of each signal in false positives rather than in detection percentage, and delete the rows that only add friction (see Defending Against Scrapers).

Passive and Active Signals

Passive signals are computed from data that already exists in the request path: the bytes of the ClientHello, the header block, the source IP, the cookie the site just received. They cost nothing per request beyond the edge lookup, they cannot be blocked by the client, and they are available even when the client never runs a line of JavaScript. Active signals require the site to serve something the client must execute: an obfuscated challenge script, an interstitial, a WebAssembly blob, a proof-of-work a script has to solve. This is the expensive class, and it buys the entire JavaScript environment - canvas, WebGL, fonts, Intl, storage probes, capability checks - along with the site's own latency, a second round trip and an uncacheable response path. A WAF that serves a static page from a CDN edge cannot serve a challenge from the same path, which is why the escalation ladder moves through several layers before it reaches a challenge (see JavaScript Challenges). The distinction matters operationally. A plain HTTP client produces passive signals only, and no amount of header polishing changes that. A real browser can choose whether to execute the active ones, and refusing to is itself a strong signal: a client that parses HTML but refuses to run script is not a human.

Retention: How Long a Fingerprint Lives

A signal's storage location decides its lifetime, and the three options behave very differently. A first-party cookie is visible to the user and deletable by them, but it expires with the browser and dies with the profile. A server-side record keyed by a fingerprint hash survives cookie clearing, survives a new browser profile, and can be joined to every other request from that hash for as long as the site keeps it - 30 days to 3 years is typical. Local storage dies with the profile unless the application deliberately restores it.

Cookie-based collection is the weakest and the most privacy-legible. Server-side retention is the one that makes fingerprinting powerful: a vendor that hashes a canvas value to a 32-bit identifier can still recognise the same device a year later, in a different session, with no cookie present, by matching the hash. The Browser Fingerprinting lesson covers how the identifier is built; the operational consequence here is that clearing cookies does not reset your identity, and neither does a fresh profile with a randomised fingerprint unless the whole profile changes at once. Sites that want durability without cross-site joinability build a first-party server-side record instead: an opaque random ID in a cookie, resolved to a fingerprint hash in their own database, deleted on a schedule. That is weaker for the defender and far better for the user. When you audit a site, retention is the question to ask, and it is the one most sites will not answer publicly.

The Cardinality Problem

How many distinct values does a signal actually have? Screen resolution has a few hundred plausible values, a timezone identifier about 420, a user agent string a few thousand, a truncated canvas hash on the order of millions. Multiply them and the joint space is astronomically larger than the human population, which is the entire basis of fingerprinting: any individual value is shared by millions of people, any combination is shared by nobody.

import math

# Order-of-magnitude counts of the values each signal actually takes across
# the installed base of browsers, ordered from weakest to strongest.
SIGNALS = [
    ("screen.width x height", 240),
    ("hardwareConcurrency", 16),
    ("deviceMemory (GiB)", 8),
    ("timezone id", 420),
    ("navigator.languages", 900),
    ("user agent string", 3_000),
    ("webgl vendor string", 240),
    ("webgl renderer string", 1_800),
    ("font metric signature", 12_000),
    ("audio context hash", 900_000),
    ("canvas 2d hash", 6_000_000),
]

HUMANS = 8.1e9          # world population, the only 'real' identity pool
TOLERATED = 0.01        # you want at least ~100 people sharing your tuple

print(f"{'signal':24} {'values':>11} {'bits':>7} {'joint space':>15} {'share/val':>10} {'hit?':>5}")
joint, bits, crowded_bits, crowded_at = 1, 0.0, 0.0, 0
for i, (name, values) in enumerate(SIGNALS, 1):
    joint *= values
    bits += math.log2(values)
    if HUMANS / joint >= TOLERATED:
        crowded_bits, crowded_at = bits, i
    print(f"{name:24} {values:>11,} {bits:>7.1f} {joint:>15.3e} "
          f"{HUMANS / joint:>10.2e} {'yes' if HUMANS / joint >= TOLERATED else 'NO':>5}")

print(f"\nall {len(SIGNALS)} signals: {bits:.0f} bits = {joint:.3e} combinations")
print(f"a random draw lands on a real device with p = {HUMANS / joint:.2e}")
print(f"the last crowded value is signal {crowded_at} at {crowded_bits:.0f} bits; "
      f"the remaining {len(SIGNALS) - crowded_at} buy uniqueness nobody has")
signal                        values    bits     joint space  share/val  hit?
screen.width x height            240     7.9       2.400e+02   3.38e+07   yes
hardwareConcurrency               16    11.9       3.840e+03   2.11e+06   yes
deviceMemory (GiB)                 8    14.9       3.072e+04   2.64e+05   yes
timezone id                      420    23.6       1.290e+07   6.28e+02   yes
navigator.languages              900    33.4       1.161e+10   6.98e-01   yes
user agent string              3,000    45.0       3.484e+13   2.33e-04    NO
webgl vendor string              240    52.9       8.361e+15   9.69e-07    NO
webgl renderer string          1,800    63.7       1.505e+19   5.38e-10    NO
font metric signature         12,000    77.3       1.806e+23   4.49e-14    NO
audio context hash           900,000    97.0       1.625e+29   4.98e-20    NO
canvas 2d hash             6,000,000   119.6       9.752e+35   8.31e-27    NO

all 11 signals: 120 bits = 9.752e+35 combinations
a random draw lands on a real device with p = 8.31e-27
the last crowded value is signal 5 at 33 bits; the remaining 6 buy uniqueness nobody has

The share/val column is the number of humans expected to share your exact tuple after each signal is added, and it is the number to watch. Four signals in, you are down to fewer than one expected person per tuple; every signal after that buys uniqueness that nobody has, and the value is now provably synthetic. A trained model flags the rarity, not the uniqueness. The practical rule for anyone building a profile is to spend a few cheap signals to reach a common cell and then stop - adding entropy past saturation does not make you look like a person, it makes you look like nothing at all, and it is exactly the combination that Cross-Validating Fingerprint Signals is built to catch.

What a Scraper Can and Cannot Influence

Signal Can you influence it? Cost to change it
ASN / hosting flag Only by changing egress network An entire proxy pool decision
rDNS record No, it belongs to the address owner Not yours to change
TLS JA3/JA4 Yes, by impersonating a browser stack Pin a new target on every vendor update
HTTP header grammar Yes, by replaying a real browser's bytes Re-capture on every browser release
JS environment values Yes, but only inside a real engine Deterministic noise, per profile
Behavioural timing Partially; synthetic motion is still synthetic Fitted distributions, per platform
Historical reputation Barely; reputation follows the IP and the ASN Days or weeks of clean traffic
Resolver identity Yes, by routing DNS through the proxy See DNS Resolver Signals

The pattern is that every cheap signal is also the one you can change, and the signals you cannot change are attached to the address you are borrowing. That is why buying a residential IP with a clean history is such a large part of the cost of this work: you are renting someone else's past, and the parts of their past you cannot edit are the parts a WAF weighs most heavily.

Auditing a Target: One Page Load, Every Request

The most useful procedure in this lesson takes two minutes. Load the target once in a real browser with request recording on, dump every request, and classify them. What comes back is the site's actual collection surface for that page rather than its marketing claims.

  1. Launch a browser with recording (Chrome DevTools' network panel, or a Playwright page with page.on('response')), load the page, and wait for the network to go quiet for three seconds.
  2. Export the list as (type, url, bytes).
  3. Sort the hosts by request count and bytes; anything unrecognised is a third party you did not know was present.
  4. Group by purpose, and note which groups only fire because JavaScript ran - those are the active signals.

The classifier below is the whole analysis, and it is deliberately crude: hostname and path patterns are enough to tell you which vendor is on the page and roughly what it is being told.

import re
from collections import Counter, defaultdict
from urllib.parse import urlsplit

# Every request one page load produced, recorded with Playwright's response
# event: (resource type, url, transferred bytes).
OBSERVED = [
    ("document", "https://shop.example.com/p/12345", 48210),
    ("script", "https://shop.example.com/assets/app.a91f.js", 241100),
    ("stylesheet", "https://shop.example.com/assets/app.a91f.css", 38120),
    ("image", "https://cdn.example-static.com/img/p/12345/main-1x.jpg", 88400),
    ("font", "https://cdn.example-static.com/font/inter-var.woff2", 51200),
    ("script", "https://shop.example.com/cdn-cgi/challenge-platform/h/g/flow/1.0/x.js", 21400),
    ("xhr", "https://shop.example.com/cdn-cgi/challenge-platform/h/b/orchestrate/chl_1", 900),
    ("xhr", "https://shop.example.com/cdn-cgi/bm/collect", 380),
    ("script", "https://cdn.mouseflow.com/2s/tracking.min.js", 94200),
    ("xhr", "https://api.mouseflow.com/2/track", 1240),
    ("script", "https://www.googletagmanager.com/gtm.js?id=GTM-K4Q2X", 61100),
    ("xhr", "https://www.google-analytics.com/g/collect?v=2", 640),
    ("script", "https://pagead2.googlesyndication.com/pagead/js/adsbygoogle.js", 105300),
    ("script", "https://cdn.cookielaw.org/1.18.0/consent.js", 47800),
    ("xhr", "https://px.ads.linkedin.com/collect", 340),
    ("xhr", "https://shop.example.com/api/v1/reviews?productId=12345", 14200),
    ("image", "https://cdn.example-static.com/img/hero@2x.avif", 142300),
]

# category, URL matcher, does the request only ever fire inside a JS engine?
RULES = [
    ("bot challenge", r"cdn-cgi/(challenge-platform|bm)/", True),
    ("session replay", r"(mouseflow|clarity\.ms|logrocket|fullstory|hotjar)", True),
    ("advertising", r"(googlesyndication|doubleclick|adsrvr|taboola|criteo|openx)", True),
    ("consent manager", r"(cookielaw|onetrust|didomi|trustarc|iubenda)", True),
    ("analytics", r"(google-analytics|gtag/js|googletagmanager\.com/gtm|segment|mixpanel)", True),
    ("marketing pixel", r"(connect\.facebook\.net|ads\.linkedin\.com|analytics\.tiktok|bat\.bing)", True),
    ("first party api", r"^https://shop\.example\.com/api/", False),
]
NEEDS_JS = {name: js for name, _, js in RULES}

def classify(url):
    for name, pattern, _ in RULES:
        if re.search(pattern, url):
            return name
    return "first party asset"

counts, bytes_by, js_requests, js_bytes = Counter(), defaultdict(int), 0, 0
for _rtype, url, size in OBSERVED:
    name = classify(url)
    counts[name] += 1
    bytes_by[name] += size
    if NEEDS_JS.get(name):
        js_requests, js_bytes = js_requests + 1, js_bytes + size

total = sum(bytes_by.values())
print(f"one page load: {len(OBSERVED)} requests, {total:,} bytes\n")
print(f"{'category':18} {'reqs':>4} {'bytes':>10} {'share':>7}  needs JS")
for name, n in counts.most_common():
    print(f"{name:18} {n:>4} {bytes_by[name]:>10,} {100 * bytes_by[name] / total:>6.1f}%  "
          f"{'yes' if NEEDS_JS.get(name) else 'no':>7}")
print(f"\nrequests a JS engine must execute: {js_requests} of {len(OBSERVED)} "
      f"({100 * js_bytes / total:.0f}% of the bytes)")
one page load: 17 requests, 956,830 bytes

category           reqs      bytes   share  needs JS
first party asset     6    609,330   63.7%       no
bot challenge         3     22,680    2.4%      yes
session replay        2     95,440   10.0%      yes
analytics             2     61,740    6.5%      yes
advertising           1    105,300   11.0%      yes
consent manager       1     47,800    5.0%      yes
marketing pixel       1        340    0.0%      yes
first party api       1     14,200    1.5%       no

requests a JS engine must execute: 10 of 17 (35% of the bytes)

Run this against three sites you are allowed to crawl and the results are sobering. A typical retail page load above makes 17 requests, and 10 of them only happen because a script executed. The challenge vendor is 2.4 percent of the bytes but 100 percent of the decision; the advertising and analytics stack is nearly a third of the bytes and tells the site nothing about whether you are a bot. If the goal is to be treated as a browser, the challenge requests are the ones that have to work and the rest can stay blocked.

Instrument Your Own Stack, and Collect Less

For a site owner the useful version of this inventory is a budget: collect only what changes a decision. Log the score contributions rather than the verdict, because a block you cannot attribute to a signal is a block you cannot tune or delete. Publish the false-positive cost per signal - conversion rate and support volume attributed to each trigger is the number that keeps collection honest. Prefer signals that cannot be faked cheaply: transport and IP history are expensive to fake, a client-reported string is not. Set retention explicitly, because a 30-day window with a rotating salt is nearly as good as a permanent one and far easier to justify. Make deletion real - if you promise to delete a fingerprint on request, the delete path has to be as reliable as the collect path. And re-audit quarterly, because third-party tags are added by marketing teams and the collection surface grows without anyone deciding to grow it. The privacy framing and the security framing point the same way here: the cheapest signal to defend against is the one never collected. A site with no session replay, no third-party ad tags and no persistent device ID has a smaller attack surface, less data to breach, and a challenge that only has to prove a person is present rather than prove a person is a returning, paying, identifiable person. From the other side, the legitimate routes - an API key, a partner feed, a rate-limit header, a permission - give you stable, cheap access, and Official APIs and Alternative Data is where to look before spending a week on a stealth pipeline. Related reading: How Anti-Bots Work for the scoring model, Browser Fingerprinting for the identifier itself, ASN, Cloud Ranges and Datacenter Detection for the network layer, and Challenge Vendor Deep Dive for four real implementations of this catalogue.