Inside Open-Source Fingerprint Collectors
Read the Detector's Source
Every commercial fingerprint product is a variation on the same open-source designs. Browser Fingerprinting explains what a fingerprint is and why spoofing one component in isolation fails. This lesson goes the other way: how the collectors that everyone studies are actually built, what they gather, how the entropy arithmetic works, and what to read in a site's bundle when you want to know what it collects.
Knowing the code is the difference between guessing at a signal and confirming one. It is also how you tell a real device from a pool of clones without guessing at thresholds.
The Math: A Vector, a Hash, an Estimate
A fingerprint is three things stacked:
- A vector of component values, each collected from a JavaScript API or an HTTP header. A typical collector ends up with 20 to 60 of them.
- A hash over the normalised vector, so the value can be stored, compared and looked up cheaply. Component values are joined with a separator, length-prefixed or delimited so that
["ab","c"]and["a","bc"]cannot collide, then hashed with SHA-256 or xxHash. - An entropy estimate, a number of bits, computed from how many distinct values each component takes across the population that vendor has measured. This is what a marketing figure quotes, and it is the least useful of the three for detection.
The entropy estimate is the part to distrust, for a reason the arithmetic makes obvious: it is capped by the population. If a vendor has measured a million browsers, no fingerprint can carry more than about 19.9 bits of information about them, whatever the components sum to. Published figures cluster at 18 to 20 bits for exactly this reason. The sum of raw per-component claims is always far larger, and quoting it is a way of describing the collector's ambitions rather than its result.
Two Ways to Generate a Fingerprint
Open-source generation hashes the values in the page. fingerprintjs is the reference implementation: it builds a component tree, each leaf returning a raw value, a hashed value, and optionally a per-component bit estimate, then walks the tree in a fixed key order and concatenates. The result is reproducible from the source, which is why it is the thing everyone studies and the thing vendors fork.
Browser-computed generation asks the browser for a number. A WebGL extension, a canvas-based one, or increasingly a server-side component computes a hash on real hardware and hands back a single identifier. The bits per component are much higher, and the implementation is much harder to reproduce. Where a collector cannot compute something natively, it falls back to the open-source path for that component alone.
A collector that mixes the two is the norm, and the mix is where the interesting behaviour is: a component can return a value, a hash, or an error, and the shape of that return is itself observed.
The Component Catalogue
| Component | Source | Typical value | Why it is stable |
|---|---|---|---|
| canvas2d | canvas.toDataURL() on a drawn scene |
64-char hex | driver, rasteriser, font rendering, GPU |
| webgl | WEBGL_debug_renderer_info |
ANGLE (NVIDIA, ...) |
the GPU string itself |
| webglParams | ~20 gl.getParameter() calls |
20 numeric values | limits and extensions per driver |
| audio | AudioContext through an oscillator and compressor |
float in a narrow band | the audio backend's maths |
| fonts | measured width of a test string in each candidate font | bitmask | which fonts are installed |
| fontMetrics | measureText widths |
several floats | hinting, subpixel layout, driver |
| screen | screen.width/height/avail* |
e.g. 1920x1080 | the display |
| pixelRatio | devicePixelRatio |
1, 1.25, 2 | the display's scaling |
| colorDepth | screen.colorDepth |
24 | almost always 24 |
| hardwareConcurrency | navigator.hardwareConcurrency |
4, 8, 12 | the CPU's core count |
| deviceMemory | navigator.deviceMemory |
0.25 to 8, quantised | the RAM bucket Chrome reports |
| platform | navigator.platform |
Win32, MacIntel |
legacy, but still read |
| languages | navigator.languages |
["en-GB","en"] |
the user's language list |
| timezone | Intl.DateTimeFormat().resolvedOptions() |
Europe/Berlin |
the OS time zone |
| intlLocales | Intl.supportedLocalesOf over a probe list |
a list of tags | the ICU build's contents |
| mediaCodecs | canPlayType per MIME string |
probably/maybe/"" |
the codec build |
| touchPoints | navigator.maxTouchPoints |
0, 5, 10 | the input hardware |
| cookiesEnabled | navigator.cookieEnabled |
true | almost always true |
| doNotTrack | navigator.doNotTrack |
null, "1", "unspecified" |
the user's setting |
| storage | localStorage, IndexedDB, cookie write |
throws or not | the policy, per browser |
| adBlock | a bait element, or extension globals | detected / not | the user's extensions |
| permissions | Notification.permission, query() results |
default |
the browser's state |
| error behaviour | what happens when the probe throws | value or error | the browser's failure mode |
The last row is the one people forget, and it is covered in its own section below.
Multiplying the Bits
The entropy arithmetic is a sum of logarithms, capped by the population, and the entropy code below prints it for five device classes:
# Entropy arithmetic for a browser fingerprint.
#
# Three things are worth computing: what each component is worth, how the bits
# multiply into a uniqueness estimate, and why that estimate cannot tell a real
# device from a spoofed one on its own.
import math
WORLD = 1_000_000 # browsers in the wild a collector tracks
CEILING = math.log2(WORLD) # no fingerprint can beat the population
# name, raw bits, source, varies within a device family?, stable across visits?
COMPONENTS = [
("userAgent", 10.0, "request header", False, True),
("platform", 3.5, "navigator", False, True),
("vendor", 2.0, "navigator", False, True),
("languages", 4.5, "navigator", False, True),
("timezone", 8.0, "Intl", False, True),
("screen", 9.0, "navigator.screen", False, True),
("intlLocales", 6.5, "Intl.supportedLocalesOf", False, True),
("mediaCodecs", 4.0, "canPlayType", False, True),
("colorDepth", 1.5, "navigator.screen", False, True),
("cookiesEnabled", 0.6, "navigator", False, True),
("doNotTrack", 0.9, "navigator", False, True),
("hardwareConcurrency", 3.0, "navigator", True, True),
("deviceMemory", 2.5, "navigator", True, True),
("pixelRatio", 3.0, "devicePixelRatio", True, True),
("maxTouchPoints", 2.0, "navigator", True, True),
("canvas2d", 12.0, "canvas.toDataURL", True, True),
("webgl", 11.0, "WEBGL_debug_renderer_info", True, True),
("webglParams", 9.5, "gl.getParameter()", True, True),
("audio", 10.5, "AudioContext", True, True),
("fonts", 9.0, "offsetWidth probe", True, True),
("fontMetrics", 7.5, "canvas measureText", True, True),
("adBlock", 1.2, "heuristic probe", True, False),
("perfTiming", 5.0, "performance.now()", True, False),
("storageProbe", 3.0, "localStorage / IDB", True, False),
("webrtc", 4.0, "RTCPeerConnection", True, False),
]
# One row per device class. `size` is how many browsers of that class exist;
# the remaining values are the *modal share* of each component: how many of
# them report the same value as you do. A component shared by 200 browsers
# carries log2(200) = 7.6 bits and no more, however many bits it could in
# principle have carried.
POPULATIONS = {
"Windows 11 + Chrome desktop": {
"size": 300_000, "hardwareConcurrency": 10, "deviceMemory": 6,
"pixelRatio": 8, "maxTouchPoints": 40, "colorDepth": 300_000,
"canvas2d": 400, "webgl": 300, "webglParams": 250,
"audio": 500, "fonts": 200, "fontMetrics": 300,
},
"macOS 14 + Safari": {
"size": 90_000, "hardwareConcurrency": 8, "deviceMemory": 5,
"pixelRatio": 12, "maxTouchPoints": 90_000, "colorDepth": 90_000,
"canvas2d": 260, "webgl": 220, "webglParams": 180,
"audio": 90, "fonts": 150, "fontMetrics": 240,
},
"Android 13 + Chrome mobile": {
"size": 400_000, "hardwareConcurrency": 24, "deviceMemory": 14,
"pixelRatio": 16, "maxTouchPoints": 30, "colorDepth": 400_000,
"canvas2d": 700, "webgl": 500, "webglParams": 600,
"audio": 120, "fonts": 90, "fontMetrics": 200,
},
"Headless Linux container": {
"size": 4_000, "hardwareConcurrency": 1_100, "deviceMemory": 1_500,
"pixelRatio": 1_900, "maxTouchPoints": 3_900, "colorDepth": 4_000,
"canvas2d": 2_600, "webgl": 2_900, "webglParams": 3_000,
"audio": 800, "fonts": 1_800, "fontMetrics": 2_400,
},
"Spoofed Win11 on headless host": {
"size": 400_000, "hardwareConcurrency": 12, "deviceMemory": 9,
"pixelRatio": 5, "maxTouchPoints": 3, "colorDepth": 400_000,
"canvas2d": 1, "webgl": 1, "webglParams": 2,
"audio": 1, "fonts": 3, "fontMetrics": 2,
},
}
INFORMATIVE_SHARE = 0.01 # a component needs >1% spread to count
REALISTIC_FLOOR = 5 # no real value is shared by fewer than this
def collide(bits, world=WORLD):
# P(at least one other browser among `world` matches on every component).
if bits <= 0:
return 1.0
return -math.expm1(world * math.log1p(-(2.0 ** -bits)))
def census(profile):
# Score one device class the way a plausibility check would.
# Summing component bits always overshoots, because the components are
# correlated and because a class of N browsers only has log2(N) bits to give.
# So uniqueness is the *normal* case, and the checks worth running are the two
# ends of the distribution: components so common inside the class that they
# carry nothing, and values so rare that no real population produces them
size = profile["size"]
bits = coarse = 0
rarest_name, rarest = None, size
for name, raw, _, varies, _ in COMPONENTS:
if not varies:
continue # selects a device family, not a person
share = profile.get(name, size)
if share < rarest:
rarest_name, rarest = name, share
if share > size * INFORMATIVE_SHARE:
coarse += 1 # shared by most of the class: no bits
continue
bits += min(raw, math.log2(share))
return min(bits, math.log2(size)), max(1, size / 2 ** bits), coarse, rarest_name, rarest
def main():
stable = [c for c in COMPONENTS if c[4]]
lost = sum(b for _, b, _, _, ok in COMPONENTS if not ok)
total = sum(b for _, b, _, _, _ in COMPONENTS)
within_family = sum(1 for c in COMPONENTS if c[3])
print(f"{len(COMPONENTS)} components, {len(stable)} stable across visits, "
f"{within_family} that vary inside a device family\n")
print(f"world population {WORLD:,} -> uniqueness ceiling {CEILING:.1f} bits\n")
print("bits as they accumulate; nothing above the ceiling adds anything:")
running = 0.0
for index, (name, bits, source, varies, stable_flag) in enumerate(COMPONENTS, 1):
before = running
running = min(running + bits, CEILING)
note = "" if running - before == bits else f" only +{running - before:.1f} useful"
if not varies:
note += " (selects a device family, not an individual)"
if not stable_flag:
note += " [unstable]"
print(f" {index:2}. {name:22} {source:26} +{bits:5.1f} -> {running:5.1f}{note}")
print(f" {'total':22} {'':26} {total:5.1f} raw bits, {running:.1f} usable\n")
print(f"raw total {total:6.1f} bits")
print(f"raw total, stable components {total - lost:6.1f} bits "
f"({lost:.1f} lost to components that change per visit)")
print(f"at the population ceiling {min(total, CEILING):6.1f} bits, "
f"collision odds {collide(CEILING):.4f}")
print(" which is why every vendor publishes a number near 18 to 20 bits:\n")
print(f"{'device class':30} {'bits':>5} {'crowd':>7} {'coarse':>7} rarest value reading")
for label, profile in POPULATIONS.items():
bits, crowd, coarse, rarest_name, rarest = census(profile)
if rarest < REALISTIC_FLOOR:
reading = f"{rarest_name} shared by only {rarest}: not a real population"
elif coarse >= 7:
reading = "a pool of clones, one value repeated everywhere"
elif crowd > 20:
reading = "unusually rare, cross-check before trusting it"
else:
reading = "an ordinary device in an ordinary crowd"
print(f"{label:30} {bits:5.1f} {crowd:7.0f} {coarse:7} "
f"{rarest_name}={rarest:<7} {reading}")
main()
25 components, 21 stable across visits, 14 that vary inside a device family
world population 1,000,000 -> uniqueness ceiling 19.9 bits
bits as they accumulate; nothing above the ceiling adds anything:
1. userAgent request header + 10.0 -> 10.0 (selects a device family, not an individual)
2. platform navigator + 3.5 -> 13.5 (selects a device family, not an individual)
3. vendor navigator + 2.0 -> 15.5 (selects a device family, not an individual)
4. languages navigator + 4.5 -> 19.9 only +4.4 useful (selects a device family, not an individual)
5. timezone Intl + 8.0 -> 19.9 only +0.0 useful (selects a device family, not an individual)
6. screen navigator.screen + 9.0 -> 19.9 only +0.0 useful (selects a device family, not an individual)
7. intlLocales Intl.supportedLocalesOf + 6.5 -> 19.9 only +0.0 useful (selects a device family, not an individual)
8. mediaCodecs canPlayType + 4.0 -> 19.9 only +0.0 useful (selects a device family, not an individual)
9. colorDepth navigator.screen + 1.5 -> 19.9 only +0.0 useful (selects a device family, not an individual)
10. cookiesEnabled navigator + 0.6 -> 19.9 only +0.0 useful (selects a device family, not an individual)
11. doNotTrack navigator + 0.9 -> 19.9 only +0.0 useful (selects a device family, not an individual)
12. hardwareConcurrency navigator + 3.0 -> 19.9 only +0.0 useful
13. deviceMemory navigator + 2.5 -> 19.9 only +0.0 useful
14. pixelRatio devicePixelRatio + 3.0 -> 19.9 only +0.0 useful
15. maxTouchPoints navigator + 2.0 -> 19.9 only +0.0 useful
16. canvas2d canvas.toDataURL + 12.0 -> 19.9 only +0.0 useful
17. webgl WEBGL_debug_renderer_info + 11.0 -> 19.9 only +0.0 useful
18. webglParams gl.getParameter() + 9.5 -> 19.9 only +0.0 useful
19. audio AudioContext + 10.5 -> 19.9 only +0.0 useful
20. fonts offsetWidth probe + 9.0 -> 19.9 only +0.0 useful
21. fontMetrics canvas measureText + 7.5 -> 19.9 only +0.0 useful
22. adBlock heuristic probe + 1.2 -> 19.9 only +0.0 useful [unstable]
23. perfTiming performance.now() + 5.0 -> 19.9 only +0.0 useful [unstable]
24. storageProbe localStorage / IDB + 3.0 -> 19.9 only +0.0 useful [unstable]
25. webrtc RTCPeerConnection + 4.0 -> 19.9 only +0.0 useful [unstable]
total 133.7 raw bits, 19.9 usable
raw total 133.7 bits
raw total, stable components 120.5 bits (13.2 lost to components that change per visit)
at the population ceiling 19.9 bits, collision odds 0.6321
which is why every vendor publishes a number near 18 to 20 bits:
device class bits crowd coarse rarest value reading
Windows 11 + Chrome desktop 18.2 1 4 deviceMemory=6 an ordinary device in an ordinary crowd
macOS 14 + Safari 16.5 1 5 deviceMemory=5 an ordinary device in an ordinary crowd
Android 13 + Chrome mobile 18.6 1 4 deviceMemory=14 an ordinary device in an ordinary crowd
Headless Linux container 0.0 4000 14 audio=800 a pool of clones, one value repeated everywhere
Spoofed Win11 on headless host 13.0 49 4 canvas2d=1 canvas2d shared by only 1: not a real population
What the table actually shows is the reason a uniqueness score is the wrong tool for detection. The three real device classes all saturate the ceiling and all have a crowd of one, because a real desktop browser is unique among a million of them. Uniqueness is the normal case, not the signal.
The two non-real rows are the informative ones, and they are informative for opposite reasons:
- The headless container has 14 coarse components. Every hardware value is shared by a large fraction of a 4,000-browser pool, so the profile carries almost no information and the crowd is the whole pool. This is the signature of automation: not a rare device but a clone farm, where a thousand visitors report an identical vector.
- The spoofed profile claims Windows but has a
canvas2dvalue shared by exactly one browser in the class. Real canvas hashes cluster, because millions of machines run the same handful of driver versions; a value that unique inside a 400,000-browser Windows population does not exist. This is the signature of a copy-pasted fingerprint: too good to be true.
So the checks worth running are at the two ends of the distribution. Components so common inside the class that they carry nothing, and values so rare that no real population produces them. A WAF doing cross-validation is doing exactly this, and Fingerprint Cross-Validation covers the contradiction matrix.
Visitor ID Versus Fingerprint
A fingerprint is the hash of the component vector. A visitor ID is a server-side bucket: the vendor assigns your fingerprint to one of a number of groups, and a visit counts as unique if the group is new. FingerprintingJS reports a visitor ID as a two-part identifier (a surrogate ID and a random site-specific ID) precisely because the raw hash should not be exposed.
The distinction matters because of instability. A component that changes between two visits makes the vector change, and a collector that hashes the whole vector produces a different fingerprint each time. A few components are known unstable and are handled specially:
| Component | Instability | Handling |
|---|---|---|
screen.availWidth/Height |
changes when the taskbar moves or a display is unplugged | dropped, or normalised |
performance.now() resolution |
varies with hardware and power state | binned, or excluded |
| ad-block detection | changes when the user installs an extension | a separate low-weight key |
| storage state | changes when the user clears site data | excluded from the stable key |
| WebRTC candidates | change per network, and per session | excluded, or a separate key |
When a collector cannot get a component to be stable it has two bad options: include it and lose stability, or drop it and lose coverage. Vendors differ here, which is the main reason their published bit counts differ for the same browser. The consistency lesson makes the same point from the other direction: a profile that is unique but changes every visit is a worse identity than one that is slightly coarser and stable, because stability across visits is what a WAF correlates on.
When a Component Throws
Probes fail, and how they fail is data. The common cases:
- Safari and
Intl: aDateTimeFormatconstructed with an invalid calendar throws aRangeError, and which inputs throw varies by ICU version. - WebGL in a restricted context:
getContext('webgl')returnsnullwhen GPU access is blocked, in a headless build without a GPU, or under a privacy extension. - Audio without an output device: an
AudioContextin a machine with no sound card can start but produce silence, and a collector that expects a non-zero sample gets a suspiciously round number. - Storage under a strict policy:
localStoragethrows aSecurityErrorin some privacy modes rather than returning null. - Permissions API:
query()resolves with a state the user never chose, or rejects outright.
A well-built collector catches the error, records a sentinel, and continues. A naive one lets the exception propagate and the whole probe returns undefined for that key. Both are observable: the sentinel value is a different value from the real one, and a key that is missing entirely is different again. A collector that reports canvas2d: undefined on a desktop Chrome profile is telling you it did not run a real probe.
This is the same lesson as Fixing CDP & Headless Leaks, from the collector's side: the failure mode is the signal, and both parties know it.
The Public Test Pages
The open-source collectors are public, and the sites that run them are the best regression suite available:
| Page | What it uniquely probes |
|---|---|
| browserleaks | per-API depth: separate pages for canvas, WebGL, fonts, audio, screen, Java, Flash-era plugins, navigator in full |
| bot.sannysoft | the automation surface: navigator.webdriver, plugin counts, window.chrome, permissions, WebGL renderer, error-stack artefacts |
| creepjs | lying detectors: it plants deliberate falsehoods (a fake navigator.plugins entry, a spoofed screen) and reports which ones you did not correct, on the theory that a bot fixes the values it knows about and misses the rest |
| fingerprintjs demo | the open-source collector itself, component by component, with each component's raw value and hash shown |
| bot-detection (a) | leak tests run one per line, easy to diff between runs |
creepjs deserves the emphasis because it inverts the usual assumption. A naive scraper patches navigator.webdriver and stops; a slightly better one patches the documented list; creepjs succeeds by checking whether the profile is internally consistent, so a partial patch is worse than none. The lesson it teaches is covered in Device Profile Consistency.
What to Read in a Site's Bundle
When you want to know what a specific site collects, four checks in order of cost:
- Grep the bundle for API names. The minified bundle still contains property names, because they are not mangled. Search for
toDataURL,getParameter,WEBGL_debug_renderer_info,hardwareConcurrency,AudioContext,maxTouchPoints,Intl.,measureText,fonts. A hit tells you the component is collected; the surrounding code tells you how. - Look for the hashing call.
crypto.subtle.digest, or a bundled xxHash or murmur, tells you the values are being combined rather than just sent. The separator used is worth reading: it is what makes the hash reproducible. - Watch the network calls. A collector either posts the components to your own server, or sends a hash to a third party. The endpoint name identifies the vendor, and the request body shape tells you whether you are looking at raw components or an already-hashed value.
- Read what it does with errors. Search for
try,catch, and anyundefineddefault near a component. The error path is where the vendor's assumptions are visible.
Wireshark or a proxy that records the collector's own request is often faster than reading minified code, and the MITM and HAR lesson covers the capture. Read the code on a system you are authorised to test, or your own site.
Why Vendors' Numbers Differ
When two products report different entropy for the same browser, the reasons are usually these, in order of how much they matter:
| Difference | Effect on the number |
|---|---|
| different component set | a missing canvas drops it by 10+ bits |
| different population measured | a vendor with 10M browsers can claim more bits than one with 1M |
| different stability filtering | components that change per visit get dropped, costing bits |
| different error handling | a sentinel vs a missing key changes the hash, not the count |
| different normalisation | is 1920x1080 one value or three? rounding rules differ |
| different entropy method | Shannon over measured frequencies vs an assumed uniform split |
| a browser-computed component | one native hash can be worth 20 bits of the total |
The practical consequence: a published bit count is a claim about a vendor's measurement, not a property of the browser. Treat it as a hint about which components matter and ignore the number itself.
The Legitimate Route
Running a collector against your own browser to confirm that your profile is internally self-consistent, and reading the source to understand which components a site requests before you decide what to send, are both normal engineering. Using that same source knowledge to enumerate exactly what a particular site harvests in order to defeat its controls is a different act and belongs behind written authorisation. On the defensive side, the same reading is simply how you audit your own site: if you cannot justify each component your own collector requests, it is a signal you are collecting more than you use. The defending against scrapers lesson covers the other side of that decision.
Cross-links: Browser Fingerprinting, Inside Open-Source Fingerprint Collectors, Spoofing Canvas, WebGL & Audio, Device Profile Consistency, Fingerprint Cross-Validation, Fixing CDP & Headless Leaks, Font Metrics and Intl Fingerprinting, Defending Against Scrapers.