Cross-Validating Fingerprint Signals
Every Signal Can Be Plausible and the Session Still Dies
The most common reason a hardened session gets banned is not one broken value. It is fourteen individually defensible values that have never appeared together on a real machine. A WAF does not ask "is this fingerprint correct?" It asks "how many real users look like this?", and the answer for a randomly assembled profile is close to zero.
That is the whole argument of this lesson. Consistency is scored, not correctness, and the joint distribution of real traffic is extremely clustered: almost nobody uses a Firefox on Linux with a HiDPI Apple GPU and a Windows screen resolution. Randomising each component independently maximises your chance of landing outside the cluster.
The Contradiction Table Is the Real Detector
A contradiction is any pair of signals where at least one of the two values would have to be wrong. The pairs below cover the axes a WAF can see from a single request plus a single script run.
| Pair | Contradiction | Why no real user shows it |
|---|---|---|
User-Agent platform vs screen |
iPhone OS string with 1920x1080, or Windows with 393x852 | the OS string comes from the device the screen belongs to |
User-Agent platform vs maxTouchPoints |
Windows with 5 touch points, iOS with 0 | a touchscreen Windows laptop exists, but not a non-touch iPhone |
User-Agent platform vs WebGL renderer |
Apple GPU on a Windows string |
Apple GPUs only exist in Apple hardware |
User-Agent platform vs Accept-Encoding |
Windows string without zstd |
Chromium has sent zstd since 123, WebKit never does |
sec-ch-ua-platform vs User-Agent |
hint says Windows, UA says macOS |
the hint is generated by the browser, not by your script |
Accept-Language vs Intl locale |
de-DE list, en-US formatter |
both come from one OS language preference |
Intl timezone vs IP geolocation |
Asia/Tokyo behind a German IP |
one person is in one place |
navigator.languages vs Accept-Language |
header en-GB,en, JS ["de-DE","de"] |
the header is derived from the JS list |
| Canvas hash vs WebGL renderer | an Apple-GPU canvas hash on an Intel renderer string | both come from the same rasteriser |
| Canvas hash vs font set | Segoe UI metrics on a machine with no Segoe UI | text metrics follow installed fonts |
WebGL renderer vs webgl_max_texture |
an integrated Intel UHD with a 32768 limit | the limit is a property of the driver generation |
devicePixelRatio vs screen |
3840x2160 at ratio 1 with a HiDPI-class UA | a 4K panel is always above ratio 1 |
hardwareConcurrency vs screen |
64 logical threads on a 1366x768 laptop | a VM, not a machine with a display |
TLS JA3 vs User-Agent |
a Safari JA3 hash on a Chrome UA string | the TLS stack is chosen by the browser binary |
TLS JA3 vs sec-ch-ua version list |
a 2023 JA3 with a Chrome 140 brand list | JA3 changes when the TLS stack changes |
| HTTP/2 SETTINGS vs ALPN | an h2 fingerprint with http/1.1 only in ALPN |
the same build produced both |
Sec-Fetch-* vs Sec-CH-UA-Mobile |
Sec-Fetch-Dest: document on ?0 mobile hints |
same browser, same request |
| Battery API presence vs platform | navigator.getBattery on an iPhone UA |
never existed in Safari |
The second column is the check. The third column is the payload: a detector does not need to know what the right answer is, only that the two cannot both be right.
Impossible Magnitudes Are a Separate Failure Mode
The pairs above are categorical. Magnitudes fail differently: each value is individually legal for its class, and the combination of sizes is not. A 4K viewport at devicePixelRatio: 1 on a profile claiming a modern HiDPI laptop is a real tell, because a 3840-pixel-wide CSS viewport at ratio 1 implies a physical panel nobody buys. hardwareConcurrency: 2 next to a WebGL extension set reporting limits appropriate to eight discrete GPUs is the same shape of problem. A claimed iPhone at 1920x1080 with zero touch points is a third.
The rule is arithmetic before it is statistical: multiply the plausible screen size by the plausible pixel ratio, and check the product against the plausible panel. Do it for cores against rendered work, for deviceMemory against screen size, for the number of media devices against the device class. Magnitude checks are cheap, they need no training data, and they catch the profiles that pass every categorical rule.
Version Contradictions Are What Date Your Spoof
Every component of a browser identity carries a version, and the versions come from different codebases with different release rhythms. A profile that says Chrome 140 in the user agent, Chrome 122 in the client hints, a 2023 TLS fingerprint, an ICU version from 2022 and an Accept-Encoding list without zstd is claiming to be five browsers at once.
The specific anchors to watch:
| Signal | Anchor | What it dates |
|---|---|---|
Accept-Encoding containing zstd |
Chrome 123, March 2024 | the HTTP stack |
Sec-CH-UA-Platform-Version |
updated with every OS release | the OS build |
sec-ch-ua-full-version-list |
Chromium major version | the browser build |
TLS supported_versions plus key_share GREASE |
Chrome 115+ | the TLS stack generation |
Intl.DateTimeFormat().resolvedOptions() timeZone list |
IANA 2024 releases | the ICU data |
RTCRtpSender.getCapabilities header extensions |
per-Chromium-version | the WebRTC stack |
navigator.plugins PDF plugin name |
Chrome 100+ unified PDF | the plugin surface |
The practical rule for a scraper: derive the version from one place and let everything else follow. Set the browser version, then read the client hints, Accept-Encoding, TLS fingerprint and codec table that the actual binary produces, and only then decide whether you need to change any of them. Hardcoding a version string in one place and hoping the rest follows is how a profile ends up internally inconsistent.
Independent Randomisation Is the Worst Possible Strategy
If each component is drawn from its own realistic distribution, the profile is still wrong, and it is wrong in a specific way. Real devices are correlated: the screen size, the GPU, the RAM bucket, the core count, the codec set and the font list move together because they were sold as one machine. Independent sampling destroys every one of those correlations at once and lands you in the empty middle of the space, which is exactly where a classifier's confidence is highest.
The correct move is the opposite of diversification. Pick a device, not a set of options:
- A 2023 mid-market Windows 11 laptop on Chrome 140. Intel Iris Xe or Ryzen 7 5800U, 16 GiB RAM reported as
8, 16 threads, 1920x1080 at ratio 1,hardwareConcurrencyabove 8,maxTouchPoints: 0, a Segoe UI and Calibri font set,zstdinAccept-Encoding, a 2024-dated JA3,Europe/Londonwithen-GB. - An iPhone 15 on iOS 18 with Safari 18. 393x852 at ratio 3, six threads,
deviceMemory: 8, five touch points,Apple GPU, nozstd, nozstd-era extensions, a Safari JA3,navigator.getBatteryabsent,deviceMemoryandhardwareConcurrencyboth from the A17 Pro's six cores.
Then change nothing you cannot justify. If you need a different exit IP, change the IP and the geo it implies, and change nothing else. Device Profiles covers the locale and timezone half of the same bundle, and Timezone, Locale and Device Profiles is the earlier treatment of the same idea.
Scoring Coherence Yourself
You do not need the vendor's model to catch your own mistakes. Enumerate the pairs, count how many are contradictory, and treat a non-zero count as a build failure. The code below holds two complete real bundles and a rule per pair, then reports a deliberately scrambled profile against both of them.
PROFILES = {
"win11-xeon-laptop": {
"ua_platform": "Windows NT 10.0; Win64; x64", "sec_ch_ua_platform": "Windows",
"ua_engine": "Chrome/140.0.0.0", "accept_encoding": "gzip, deflate, br, zstd",
"accept_language": "en-GB,en;q=0.9", "intl_locale": "en-GB",
"timezone": "Europe/London", "country": "GB", "screen": "1920x1080", "dpr": 1,
"max_touch_points": 0, "hardware_concurrency": 16, "webgl_max_texture": 16384,
"webgl_renderer": "ANGLE (Intel, Iris(R) Xe Graphics, Direct3D11)",
"tls_ja3": "e7d705a3286e19ea42f587b344ee6865",
"font_hashi": "arial=0.42;calibri=0.55;cambria=0.51;segoe=0.44;times=0.46",
},
"iphone15-ios18": {
"ua_platform": "iPhone; CPU iPhone OS 18_1 like Mac OS X",
"sec_ch_ua_platform": "iOS", "ua_engine": "Version/18.1 Mobile/15E148 Safari/604.1",
"accept_encoding": "gzip, deflate, br", "accept_language": "en-GB,en;q=0.9",
"intl_locale": "en-GB", "timezone": "Europe/London", "country": "GB",
"screen": "393x852", "dpr": 3, "max_touch_points": 5,
"hardware_concurrency": 6, "webgl_max_texture": 16384,
"webgl_renderer": "Apple GPU", "tls_ja3": "b32309a26951912be7dba376398abc3b",
"font_hashi": "helvetica=0.41;menlo=0.52;monaco=0.50;times=0.46",
},
}
def pair_is_sane(profile, a, b):
va, vb = profile[a], profile[b]
if a == "timezone":
return va.split("/")[0] == {"GB": "Europe"}[vb]
if a == "country":
return {"GB": "Europe"}[va] == vb.split("/")[0]
if (a, b) == ("ua_platform", "screen"):
return va.startswith("iPhone") == (int(vb.split("x")[0]) < 500)
if (a, b) == ("ua_platform", "max_touch_points"):
return va.startswith("iPhone") == (vb > 0)
if (a, b) == ("ua_platform", "webgl_renderer"):
return ("Apple GPU" in vb) == va.startswith("iPhone")
if (a, b) == ("ua_platform", "ua_engine"):
return va.startswith("iPhone") == ("Safari" in vb and "Chrome" not in vb)
if (a, b) == ("ua_platform", "accept_encoding"):
return va.startswith("iPhone") == ("zstd" not in vb)
if (a, b) == ("sec_ch_ua_platform", "ua_platform"):
return vb.startswith({"Windows": "Windows NT", "iOS": "iPhone"}[va])
if (a, b) == ("ua_engine", "tls_ja3"):
safari = "Safari" in va and "Chrome" not in va
return safari == (vb == PROFILES["iphone15-ios18"]["tls_ja3"])
if (a, b) == ("dpr", "screen"):
return profile["ua_platform"].startswith("iPhone") == (va >= 2)
if (a, b) == ("screen", "dpr"):
return not (int(va.split("x")[0]) >= 2560 and vb == 1)
if (a, b) == ("hardware_concurrency", "webgl_max_texture"):
return not (va <= 2 and vb >= 16384)
if (a, b) == ("hardware_concurrency", "webgl_renderer"):
discrete = "Intel" not in vb and "Apple" not in vb
return not (va >= 32 and not discrete)
if (a, b) == ("intl_locale", "accept_language"):
return vb.lower().startswith(va.lower())
if (a, b) == ("font_hashi", "ua_platform"):
fonts = {f.split("=")[0] for f in va.split(";")}
iphone = vb.startswith("iPhone")
return ("menlo" in fonts) == iphone and ("segoe" in fonts) == (not iphone)
return True
PAIR_RULES = [
("ua_platform", "screen", "a desktop OS with a phone screen, or an iPhone at 1920x1080"),
("ua_platform", "max_touch_points", "Windows claims 0 touch points, iOS always reports 5"),
("ua_platform", "webgl_renderer", "only Apple GPUs ship an 'Apple GPU' renderer"),
("ua_platform", "ua_engine", "a Chrome UA on an iPhone OS string is a spoofed desktop"),
("ua_platform", "accept_encoding", "Chromium sends zstd, WebKit does not"),
("sec_ch_ua_platform", "ua_platform", "the client hint is fetched, not spoofed"),
("ua_engine", "tls_ja3", "a Safari JA3 cannot belong to a Chrome build"),
("dpr", "screen", "an iPhone-class device at device pixel ratio 1 is a resized desktop"),
("screen", "dpr", "a 2560-wide desktop at ratio 1 is an RDP session or a stitched profile"),
("hardware_concurrency", "webgl_max_texture", "2 vCPUs never ship a 16384 px texture limit"),
("hardware_concurrency", "webgl_renderer", "64 threads with an integrated GPU is a VM"),
("intl_locale", "accept_language", "both are the same language preference list"),
("timezone", "country", "timezone and IP country come from the same trip"),
("font_hashi", "ua_platform", "Segoe UI is Windows-only, Menlo is macOS and iOS only"),
]
def check(profile):
rows = []
for a, b, why in PAIR_RULES:
if a in profile and b in profile:
rows.append((a, profile[a], b, profile[b], pair_is_sane(profile, a, b), why))
return rows
def report(name, profile):
rows = check(profile)
bad = [r for r in rows if not r[4]]
print("=" * 96)
print("%-20s %d signal pairs checked, %d contradictions" % (name, len(rows), len(bad)))
print("=" * 96)
for a, va, b, vb, sane, why in bad:
print(" %-20s x %-20s" % (a + "=" + str(va)[:24], b + "=" + str(vb)[:24]))
print(" %s" % why)
if not bad:
print(" no contradictions: every checked pair is one real device does")
print()
report("win11-xeon-laptop", PROFILES["win11-xeon-laptop"])
report("iphone15-ios18", PROFILES["iphone15-ios18"])
scrambled = dict(PROFILES["iphone15-ios18"], ua_platform="Windows NT 10.0; Win64; x64",
sec_ch_ua_platform="Windows", ua_engine="Chrome/140.0.0.0",
accept_encoding="gzip, deflate, br", screen="1920x1080", dpr=1,
hardware_concurrency=64)
report("scrambled-iphone", scrambled)
report("win11-4k-at-ratio-1", dict(PROFILES["win11-xeon-laptop"], screen="3840x2160"))
================================================================================================
win11-xeon-laptop 14 signal pairs checked, 0 contradictions
================================================================================================
no contradictions: every checked pair is one real device does
================================================================================================
iphone15-ios18 14 signal pairs checked, 0 contradictions
================================================================================================
no contradictions: every checked pair is one real device does
================================================================================================
scrambled-iphone 14 signal pairs checked, 6 contradictions
================================================================================================
ua_platform=Windows NT 10.0; Win64; x max_touch_points=5
Windows claims 0 touch points, iOS always reports 5
ua_platform=Windows NT 10.0; Win64; x webgl_renderer=Apple GPU
only Apple GPUs ship an 'Apple GPU' renderer
ua_platform=Windows NT 10.0; Win64; x accept_encoding=gzip, deflate, br
Chromium sends zstd, WebKit does not
ua_engine=Chrome/140.0.0.0 x tls_ja3=b32309a26951912be7dba376
a Safari JA3 cannot belong to a Chrome build
hardware_concurrency=64 x webgl_renderer=Apple GPU
64 threads with an integrated GPU is a VM
font_hashi=helvetica=0.41;menlo=0.5 x ua_platform=Windows NT 10.0; Win64;
Segoe UI is Windows-only, Menlo is macOS and iOS only
================================================================================================
win11-4k-at-ratio-1 14 signal pairs checked, 1 contradictions
================================================================================================
screen=3840x2160 x dpr=1
a 2560-wide desktop at ratio 1 is an RDP session or a stitched profile
The scrambled profile is the shape almost every hand-rolled spoof produces: an iPhone identity with a desktop operating system, a desktop screen, no HiDPI, no zstd and a server's core count. Six of fourteen pairs fail, and the six failures point at four different subsystems, which is what "not a real device" looks like from the inside. Media Codecs and Capability Probing extends the same table to hardware capabilities, and TLS and JA3 Spoofing covers the transport side of the JA3 rows.
What Defenders Actually Score
A production scorer is usually a mixture of three things: a small set of hard rules (automation flags, a HeadlessChrome token, a cdc_ global, an ASN in a cloud range), a joint-density model over the vector, and a session-level model over behaviour and history. The hard rules are cheap and catch unsophisticated traffic. The density model is what catches a well-built fake, and it works because real traffic occupies a thin, well-populated manifold and a synthetic profile almost never lands on it.
The session model is the third layer and the reason a correct profile still gets banned eventually: a profile that passes every request check but visits the same three URLs in the same order at the same second every hour is scored on the trajectory, not the fingerprint. That is the argument in Evading ML Anomaly Detection, and the mirror image is Behavioral Biometrics read from the server side.
Failure Modes
- Fixing one contradiction and creating another. Setting
sec-ch-ua-platformto match the spoofed user agent is a good move until you realise the hint is generated by the browser and now the hint disagrees with the real platform. - Mixing two reference profiles. Taking the timezone and language from the Windows bundle and the GPU and fonts from the iPhone bundle is worse than either, because you have destroyed the correlation that made each bundle plausible.
- Treating a passing score as a pass. Vendors weight signals dynamically, and a profile that scores 0.02 today scores 0.4 after their next model release. Re-run the check on a schedule.
- Auditing only the pairs you know about. The value of the table is the pairs you have not thought of. Keep the rules in a versioned file and add a row every time you read a new detection write-up.
- Not re-auditing after a browser update. A new Chromium release changes the client hints, the TLS stack and the extension lists, and every hardcoded string you own is now a version contradiction.
- Assuming magnitude checks need training data. They do not. Screen times pixel ratio, cores times GPU class, memory times screen size: all pure arithmetic, all worth asserting in CI.
Coherence Checklist
- One named reference device, documented in a single file, with every signal written down.
- Every pair in the contradiction table above has a rule, and the rule set runs in CI against the live profile.
- Magnitude assertions exist for screen times ratio, cores times GPU class, and memory times screen size.
- The browser version appears in exactly one place, and every version-derived value is read from the running binary rather than hardcoded.
- A change to any single signal is reviewed as a change to the whole profile, not to that field.
- The reference profile is current: a device class that shipped two years ago still exists, but a WebKit build from three years ago does not.
- The profile is stable for the lifetime of a browser identity, and different identities do not share signals. See Distributed Worker Fleets at Scale.
The Legitimate Route
Every rule in this lesson is a rule you would write against your own traffic, and the version above is strictly a self-audit. What does not fit is a fleet of profiles tuned to defeat a specific vendor's scorer on a site that has refused automated access, where the "contradiction" being optimised is the one that keeps a block in place: see Legal and Compliance Boundaries. Where a site offers an API, a partner feed or a documented crawler token, that is the route that survives a model update, and the coherence work above is still worth doing because it makes your own regression suite meaningful.