The Defender's Signal Inventory
The List the Defender Already Keeps About You
Before a bot-management system decides anything about a request, it has already written a row about that request into a table. The row is assembled from dozens of independent observations, taken at different layers of the stack, by code that runs before, during and after your HTTP request. Most scrapers spend all their effort on the one layer they can see - the request itself - and never ask what else was recorded while it was in flight.
A signal is any observation the site can make about a request that is not the request's own content: who asked, from where, with what client, in what order, and what happened to the last ten thousand requests from the same neighbourhood. No single signal is damning. A Chrome user on a Comcast IP in Ohio who has visited the site before and clicked two links is unremarkable. The power is in the ledger: a modern WAF keeps a per-session risk score, and each signal is a weighted line item in it (see Risk Scoring and Session Profiles). This lesson is the catalogue that goes behind that score, one category at a time.
The Catalogue
| Category | Example signal | Collected by | Survives cookie clearing? | Cost to the site |
|---|---|---|---|---|
| Transport | TLS JA3/JA4 hash, ALPN list, GREASE values | TLS terminator, no JS | Yes, per connection | CPU, ~1 ms per handshake |
| Transport | QUIC transport parameters, HTTP/3 frame order | Edge server | Yes | CPU, connection state |
| HTTP | header order, sec-ch-ua, Sec-Fetch-*, Accept grammar |
HTTP parser | Yes | Effectively free |
| Network/IP | ASN, whether the ASN is hosting, rDNS, geolocation | Edge, GeoIP DB, DNS | Yes | Licence fee per lookup |
| Network/IP | requests per IP per hour, distinct paths per IP | Edge state store | Yes | Memory, millions of keys |
| Resolver | which recursive resolver asked for the protected hostname | Resolver-side logs, or the site if it runs one | Yes | Only if the site is the resolver |
| JS environment | canvas and WebGL hashes, font metrics, audio hash, timezone, navigator surface |
Challenge script | Yes, if stored server-side | 20-80 ms of page time |
| JS environment | webdriver, CDP artefacts, plugin lists, permissions state |
Challenge script | Yes | Included above |
| Behavioral | pointer paths, key intervals, scroll velocity, dwell time | Event listeners, post-hoc | Yes | Storage; the most expensive row |
| Persistent | first-party cookie, localStorage probe, ETag, supercookie | Site JS, server | No, if cookie-based | Storage, plus consent work |
| Account | account age, login history, devices per account, MFA state | Application | No, if the account is recreated | Application database |
| Historical | past verdicts, challenge outcomes, captcha solves, past bans | Vendor risk store | Yes, indefinitely | The vendor's core cost centre |
Read the last column as an engineering budget, not a moral score. A challenge script that probes the environment costs the site 20 to 80 milliseconds of page latency per visitor; behavioural telemetry costs storage and an analytics pipeline; GeoIP and ASN data costs a licence. That asymmetry is why vendors converge on the same design: cheap passive signals by default, expensive active signals only for the minority of sessions a cheap signal already flagged. It is also why the answer to "which signal should I fix first" is almost never the behavioural one. Grouping by layer, as the first column does, tells you which subsystem a signal lives in and therefore what you have to change to move it. A TLS ClientHello is fixed before your application code runs. A navigator.languages value is whatever your environment says, so you own it outright.
Why Sites Collect More Than They Need
A WAF's failure mode is not the bot that gets through. It is the false positive: a real customer, on a corporate VPN, using a screen reader, gets a CAPTCHA or a block, and either fails to buy or files a ticket. That customer costs far more than a scraped row. Blocked legitimate traffic is measured directly - support tickets, conversion drop-off, complaint volume - while a bot that slips through costs nothing until someone acts on the data. So scoring is tuned to be permissive by default and strict under accumulation. A session clean on the first twenty cheap signals skips the expensive ones. A session that trips one expensive signal early often gets more of them, because the models are built to spend their budget where a human is watching. The other pressure is competitive: bot management is bought to beat the last vendor's scorecard, so features that measurably raise the score get instrumented even when their marginal detection value is small. If you are on the defending side, measure the cost of each signal in false positives rather than in detection percentage, and delete the rows that only add friction (see Defending Against Scrapers).
Passive and Active Signals
Passive signals are computed from data that already exists in the request path: the bytes of the ClientHello, the header block, the source IP, the cookie the site just received. They cost nothing per request beyond the edge lookup, they cannot be blocked by the client, and they are available even when the client never runs a line of JavaScript. Active signals require the site to serve something the client must execute: an obfuscated challenge script, an interstitial, a WebAssembly blob, a proof-of-work a script has to solve. This is the expensive class, and it buys the entire JavaScript environment - canvas, WebGL, fonts, Intl, storage probes, capability checks - along with the site's own latency, a second round trip and an uncacheable response path. A WAF that serves a static page from a CDN edge cannot serve a challenge from the same path, which is why the escalation ladder moves through several layers before it reaches a challenge (see JavaScript Challenges). The distinction matters operationally. A plain HTTP client produces passive signals only, and no amount of header polishing changes that. A real browser can choose whether to execute the active ones, and refusing to is itself a strong signal: a client that parses HTML but refuses to run script is not a human.
Retention: How Long a Fingerprint Lives
A signal's storage location decides its lifetime, and the three options behave very differently. A first-party cookie is visible to the user and deletable by them, but it expires with the browser and dies with the profile. A server-side record keyed by a fingerprint hash survives cookie clearing, survives a new browser profile, and can be joined to every other request from that hash for as long as the site keeps it - 30 days to 3 years is typical. Local storage dies with the profile unless the application deliberately restores it.
Cookie-based collection is the weakest and the most privacy-legible. Server-side retention is the one that makes fingerprinting powerful: a vendor that hashes a canvas value to a 32-bit identifier can still recognise the same device a year later, in a different session, with no cookie present, by matching the hash. The Browser Fingerprinting lesson covers how the identifier is built; the operational consequence here is that clearing cookies does not reset your identity, and neither does a fresh profile with a randomised fingerprint unless the whole profile changes at once. Sites that want durability without cross-site joinability build a first-party server-side record instead: an opaque random ID in a cookie, resolved to a fingerprint hash in their own database, deleted on a schedule. That is weaker for the defender and far better for the user. When you audit a site, retention is the question to ask, and it is the one most sites will not answer publicly.
The Cardinality Problem
How many distinct values does a signal actually have? Screen resolution has a few hundred plausible values, a timezone identifier about 420, a user agent string a few thousand, a truncated canvas hash on the order of millions. Multiply them and the joint space is astronomically larger than the human population, which is the entire basis of fingerprinting: any individual value is shared by millions of people, any combination is shared by nobody.
import math
# Order-of-magnitude counts of the values each signal actually takes across
# the installed base of browsers, ordered from weakest to strongest.
SIGNALS = [
("screen.width x height", 240),
("hardwareConcurrency", 16),
("deviceMemory (GiB)", 8),
("timezone id", 420),
("navigator.languages", 900),
("user agent string", 3_000),
("webgl vendor string", 240),
("webgl renderer string", 1_800),
("font metric signature", 12_000),
("audio context hash", 900_000),
("canvas 2d hash", 6_000_000),
]
HUMANS = 8.1e9 # world population, the only 'real' identity pool
TOLERATED = 0.01 # you want at least ~100 people sharing your tuple
print(f"{'signal':24} {'values':>11} {'bits':>7} {'joint space':>15} {'share/val':>10} {'hit?':>5}")
joint, bits, crowded_bits, crowded_at = 1, 0.0, 0.0, 0
for i, (name, values) in enumerate(SIGNALS, 1):
joint *= values
bits += math.log2(values)
if HUMANS / joint >= TOLERATED:
crowded_bits, crowded_at = bits, i
print(f"{name:24} {values:>11,} {bits:>7.1f} {joint:>15.3e} "
f"{HUMANS / joint:>10.2e} {'yes' if HUMANS / joint >= TOLERATED else 'NO':>5}")
print(f"\nall {len(SIGNALS)} signals: {bits:.0f} bits = {joint:.3e} combinations")
print(f"a random draw lands on a real device with p = {HUMANS / joint:.2e}")
print(f"the last crowded value is signal {crowded_at} at {crowded_bits:.0f} bits; "
f"the remaining {len(SIGNALS) - crowded_at} buy uniqueness nobody has")
signal values bits joint space share/val hit?
screen.width x height 240 7.9 2.400e+02 3.38e+07 yes
hardwareConcurrency 16 11.9 3.840e+03 2.11e+06 yes
deviceMemory (GiB) 8 14.9 3.072e+04 2.64e+05 yes
timezone id 420 23.6 1.290e+07 6.28e+02 yes
navigator.languages 900 33.4 1.161e+10 6.98e-01 yes
user agent string 3,000 45.0 3.484e+13 2.33e-04 NO
webgl vendor string 240 52.9 8.361e+15 9.69e-07 NO
webgl renderer string 1,800 63.7 1.505e+19 5.38e-10 NO
font metric signature 12,000 77.3 1.806e+23 4.49e-14 NO
audio context hash 900,000 97.0 1.625e+29 4.98e-20 NO
canvas 2d hash 6,000,000 119.6 9.752e+35 8.31e-27 NO
all 11 signals: 120 bits = 9.752e+35 combinations
a random draw lands on a real device with p = 8.31e-27
the last crowded value is signal 5 at 33 bits; the remaining 6 buy uniqueness nobody has
The share/val column is the number of humans expected to share your exact tuple after each signal is added, and it is the number to watch. Four signals in, you are down to fewer than one expected person per tuple; every signal after that buys uniqueness that nobody has, and the value is now provably synthetic. A trained model flags the rarity, not the uniqueness. The practical rule for anyone building a profile is to spend a few cheap signals to reach a common cell and then stop - adding entropy past saturation does not make you look like a person, it makes you look like nothing at all, and it is exactly the combination that Cross-Validating Fingerprint Signals is built to catch.
What a Scraper Can and Cannot Influence
| Signal | Can you influence it? | Cost to change it |
|---|---|---|
| ASN / hosting flag | Only by changing egress network | An entire proxy pool decision |
| rDNS record | No, it belongs to the address owner | Not yours to change |
| TLS JA3/JA4 | Yes, by impersonating a browser stack | Pin a new target on every vendor update |
| HTTP header grammar | Yes, by replaying a real browser's bytes | Re-capture on every browser release |
| JS environment values | Yes, but only inside a real engine | Deterministic noise, per profile |
| Behavioural timing | Partially; synthetic motion is still synthetic | Fitted distributions, per platform |
| Historical reputation | Barely; reputation follows the IP and the ASN | Days or weeks of clean traffic |
| Resolver identity | Yes, by routing DNS through the proxy | See DNS Resolver Signals |
The pattern is that every cheap signal is also the one you can change, and the signals you cannot change are attached to the address you are borrowing. That is why buying a residential IP with a clean history is such a large part of the cost of this work: you are renting someone else's past, and the parts of their past you cannot edit are the parts a WAF weighs most heavily.
Auditing a Target: One Page Load, Every Request
The most useful procedure in this lesson takes two minutes. Load the target once in a real browser with request recording on, dump every request, and classify them. What comes back is the site's actual collection surface for that page rather than its marketing claims.
- Launch a browser with recording (Chrome DevTools' network panel, or a Playwright page with
page.on('response')), load the page, and wait for the network to go quiet for three seconds. - Export the list as
(type, url, bytes). - Sort the hosts by request count and bytes; anything unrecognised is a third party you did not know was present.
- Group by purpose, and note which groups only fire because JavaScript ran - those are the active signals.
The classifier below is the whole analysis, and it is deliberately crude: hostname and path patterns are enough to tell you which vendor is on the page and roughly what it is being told.
import re
from collections import Counter, defaultdict
from urllib.parse import urlsplit
# Every request one page load produced, recorded with Playwright's response
# event: (resource type, url, transferred bytes).
OBSERVED = [
("document", "https://shop.example.com/p/12345", 48210),
("script", "https://shop.example.com/assets/app.a91f.js", 241100),
("stylesheet", "https://shop.example.com/assets/app.a91f.css", 38120),
("image", "https://cdn.example-static.com/img/p/12345/main-1x.jpg", 88400),
("font", "https://cdn.example-static.com/font/inter-var.woff2", 51200),
("script", "https://shop.example.com/cdn-cgi/challenge-platform/h/g/flow/1.0/x.js", 21400),
("xhr", "https://shop.example.com/cdn-cgi/challenge-platform/h/b/orchestrate/chl_1", 900),
("xhr", "https://shop.example.com/cdn-cgi/bm/collect", 380),
("script", "https://cdn.mouseflow.com/2s/tracking.min.js", 94200),
("xhr", "https://api.mouseflow.com/2/track", 1240),
("script", "https://www.googletagmanager.com/gtm.js?id=GTM-K4Q2X", 61100),
("xhr", "https://www.google-analytics.com/g/collect?v=2", 640),
("script", "https://pagead2.googlesyndication.com/pagead/js/adsbygoogle.js", 105300),
("script", "https://cdn.cookielaw.org/1.18.0/consent.js", 47800),
("xhr", "https://px.ads.linkedin.com/collect", 340),
("xhr", "https://shop.example.com/api/v1/reviews?productId=12345", 14200),
("image", "https://cdn.example-static.com/img/hero@2x.avif", 142300),
]
# category, URL matcher, does the request only ever fire inside a JS engine?
RULES = [
("bot challenge", r"cdn-cgi/(challenge-platform|bm)/", True),
("session replay", r"(mouseflow|clarity\.ms|logrocket|fullstory|hotjar)", True),
("advertising", r"(googlesyndication|doubleclick|adsrvr|taboola|criteo|openx)", True),
("consent manager", r"(cookielaw|onetrust|didomi|trustarc|iubenda)", True),
("analytics", r"(google-analytics|gtag/js|googletagmanager\.com/gtm|segment|mixpanel)", True),
("marketing pixel", r"(connect\.facebook\.net|ads\.linkedin\.com|analytics\.tiktok|bat\.bing)", True),
("first party api", r"^https://shop\.example\.com/api/", False),
]
NEEDS_JS = {name: js for name, _, js in RULES}
def classify(url):
for name, pattern, _ in RULES:
if re.search(pattern, url):
return name
return "first party asset"
counts, bytes_by, js_requests, js_bytes = Counter(), defaultdict(int), 0, 0
for _rtype, url, size in OBSERVED:
name = classify(url)
counts[name] += 1
bytes_by[name] += size
if NEEDS_JS.get(name):
js_requests, js_bytes = js_requests + 1, js_bytes + size
total = sum(bytes_by.values())
print(f"one page load: {len(OBSERVED)} requests, {total:,} bytes\n")
print(f"{'category':18} {'reqs':>4} {'bytes':>10} {'share':>7} needs JS")
for name, n in counts.most_common():
print(f"{name:18} {n:>4} {bytes_by[name]:>10,} {100 * bytes_by[name] / total:>6.1f}% "
f"{'yes' if NEEDS_JS.get(name) else 'no':>7}")
print(f"\nrequests a JS engine must execute: {js_requests} of {len(OBSERVED)} "
f"({100 * js_bytes / total:.0f}% of the bytes)")
one page load: 17 requests, 956,830 bytes
category reqs bytes share needs JS
first party asset 6 609,330 63.7% no
bot challenge 3 22,680 2.4% yes
session replay 2 95,440 10.0% yes
analytics 2 61,740 6.5% yes
advertising 1 105,300 11.0% yes
consent manager 1 47,800 5.0% yes
marketing pixel 1 340 0.0% yes
first party api 1 14,200 1.5% no
requests a JS engine must execute: 10 of 17 (35% of the bytes)
Run this against three sites you are allowed to crawl and the results are sobering. A typical retail page load above makes 17 requests, and 10 of them only happen because a script executed. The challenge vendor is 2.4 percent of the bytes but 100 percent of the decision; the advertising and analytics stack is nearly a third of the bytes and tells the site nothing about whether you are a bot. If the goal is to be treated as a browser, the challenge requests are the ones that have to work and the rest can stay blocked.
Instrument Your Own Stack, and Collect Less
For a site owner the useful version of this inventory is a budget: collect only what changes a decision. Log the score contributions rather than the verdict, because a block you cannot attribute to a signal is a block you cannot tune or delete. Publish the false-positive cost per signal - conversion rate and support volume attributed to each trigger is the number that keeps collection honest. Prefer signals that cannot be faked cheaply: transport and IP history are expensive to fake, a client-reported string is not. Set retention explicitly, because a 30-day window with a rotating salt is nearly as good as a permanent one and far easier to justify. Make deletion real - if you promise to delete a fingerprint on request, the delete path has to be as reliable as the collect path. And re-audit quarterly, because third-party tags are added by marketing teams and the collection surface grows without anyone deciding to grow it. The privacy framing and the security framing point the same way here: the cheapest signal to defend against is the one never collected. A site with no session replay, no third-party ad tags and no persistent device ID has a smaller attack surface, less data to breach, and a challenge that only has to prove a person is present rather than prove a person is a returning, paying, identifiable person. From the other side, the legitimate routes - an API key, a partner feed, a rate-limit header, a permission - give you stable, cheap access, and Official APIs and Alternative Data is where to look before spending a week on a stealth pipeline. Related reading: How Anti-Bots Work for the scoring model, Browser Fingerprinting for the identifier itself, ASN, Cloud Ranges and Datacenter Detection for the network layer, and Challenge Vendor Deep Dive for four real implementations of this catalogue.