Read the Detector's Source

Every commercial fingerprint product is a variation on the same open-source designs. Browser Fingerprinting explains what a fingerprint is and why spoofing one component in isolation fails. This lesson goes the other way: how the collectors that everyone studies are actually built, what they gather, how the entropy arithmetic works, and what to read in a site's bundle when you want to know what it collects.

Knowing the code is the difference between guessing at a signal and confirming one. It is also how you tell a real device from a pool of clones without guessing at thresholds.

The Math: A Vector, a Hash, an Estimate

A fingerprint is three things stacked:

  1. A vector of component values, each collected from a JavaScript API or an HTTP header. A typical collector ends up with 20 to 60 of them.
  2. A hash over the normalised vector, so the value can be stored, compared and looked up cheaply. Component values are joined with a separator, length-prefixed or delimited so that ["ab","c"] and ["a","bc"] cannot collide, then hashed with SHA-256 or xxHash.
  3. An entropy estimate, a number of bits, computed from how many distinct values each component takes across the population that vendor has measured. This is what a marketing figure quotes, and it is the least useful of the three for detection.

The entropy estimate is the part to distrust, for a reason the arithmetic makes obvious: it is capped by the population. If a vendor has measured a million browsers, no fingerprint can carry more than about 19.9 bits of information about them, whatever the components sum to. Published figures cluster at 18 to 20 bits for exactly this reason. The sum of raw per-component claims is always far larger, and quoting it is a way of describing the collector's ambitions rather than its result.

Two Ways to Generate a Fingerprint

Open-source generation hashes the values in the page. fingerprintjs is the reference implementation: it builds a component tree, each leaf returning a raw value, a hashed value, and optionally a per-component bit estimate, then walks the tree in a fixed key order and concatenates. The result is reproducible from the source, which is why it is the thing everyone studies and the thing vendors fork.

Browser-computed generation asks the browser for a number. A WebGL extension, a canvas-based one, or increasingly a server-side component computes a hash on real hardware and hands back a single identifier. The bits per component are much higher, and the implementation is much harder to reproduce. Where a collector cannot compute something natively, it falls back to the open-source path for that component alone.

A collector that mixes the two is the norm, and the mix is where the interesting behaviour is: a component can return a value, a hash, or an error, and the shape of that return is itself observed.

The Component Catalogue

Component Source Typical value Why it is stable
canvas2d canvas.toDataURL() on a drawn scene 64-char hex driver, rasteriser, font rendering, GPU
webgl WEBGL_debug_renderer_info ANGLE (NVIDIA, ...) the GPU string itself
webglParams ~20 gl.getParameter() calls 20 numeric values limits and extensions per driver
audio AudioContext through an oscillator and compressor float in a narrow band the audio backend's maths
fonts measured width of a test string in each candidate font bitmask which fonts are installed
fontMetrics measureText widths several floats hinting, subpixel layout, driver
screen screen.width/height/avail* e.g. 1920x1080 the display
pixelRatio devicePixelRatio 1, 1.25, 2 the display's scaling
colorDepth screen.colorDepth 24 almost always 24
hardwareConcurrency navigator.hardwareConcurrency 4, 8, 12 the CPU's core count
deviceMemory navigator.deviceMemory 0.25 to 8, quantised the RAM bucket Chrome reports
platform navigator.platform Win32, MacIntel legacy, but still read
languages navigator.languages ["en-GB","en"] the user's language list
timezone Intl.DateTimeFormat().resolvedOptions() Europe/Berlin the OS time zone
intlLocales Intl.supportedLocalesOf over a probe list a list of tags the ICU build's contents
mediaCodecs canPlayType per MIME string probably/maybe/"" the codec build
touchPoints navigator.maxTouchPoints 0, 5, 10 the input hardware
cookiesEnabled navigator.cookieEnabled true almost always true
doNotTrack navigator.doNotTrack null, "1", "unspecified" the user's setting
storage localStorage, IndexedDB, cookie write throws or not the policy, per browser
adBlock a bait element, or extension globals detected / not the user's extensions
permissions Notification.permission, query() results default the browser's state
error behaviour what happens when the probe throws value or error the browser's failure mode

The last row is the one people forget, and it is covered in its own section below.

Multiplying the Bits

The entropy arithmetic is a sum of logarithms, capped by the population, and the entropy code below prints it for five device classes:

# Entropy arithmetic for a browser fingerprint.
#
# Three things are worth computing: what each component is worth, how the bits
# multiply into a uniqueness estimate, and why that estimate cannot tell a real
# device from a spoofed one on its own.

import math

WORLD = 1_000_000                     # browsers in the wild a collector tracks
CEILING = math.log2(WORLD)            # no fingerprint can beat the population

# name, raw bits, source, varies within a device family?, stable across visits?
COMPONENTS = [
    ("userAgent", 10.0, "request header", False, True),
    ("platform", 3.5, "navigator", False, True),
    ("vendor", 2.0, "navigator", False, True),
    ("languages", 4.5, "navigator", False, True),
    ("timezone", 8.0, "Intl", False, True),
    ("screen", 9.0, "navigator.screen", False, True),
    ("intlLocales", 6.5, "Intl.supportedLocalesOf", False, True),
    ("mediaCodecs", 4.0, "canPlayType", False, True),
    ("colorDepth", 1.5, "navigator.screen", False, True),
    ("cookiesEnabled", 0.6, "navigator", False, True),
    ("doNotTrack", 0.9, "navigator", False, True),
    ("hardwareConcurrency", 3.0, "navigator", True, True),
    ("deviceMemory", 2.5, "navigator", True, True),
    ("pixelRatio", 3.0, "devicePixelRatio", True, True),
    ("maxTouchPoints", 2.0, "navigator", True, True),
    ("canvas2d", 12.0, "canvas.toDataURL", True, True),
    ("webgl", 11.0, "WEBGL_debug_renderer_info", True, True),
    ("webglParams", 9.5, "gl.getParameter()", True, True),
    ("audio", 10.5, "AudioContext", True, True),
    ("fonts", 9.0, "offsetWidth probe", True, True),
    ("fontMetrics", 7.5, "canvas measureText", True, True),
    ("adBlock", 1.2, "heuristic probe", True, False),
    ("perfTiming", 5.0, "performance.now()", True, False),
    ("storageProbe", 3.0, "localStorage / IDB", True, False),
    ("webrtc", 4.0, "RTCPeerConnection", True, False),
]

# One row per device class. `size` is how many browsers of that class exist;
# the remaining values are the *modal share* of each component: how many of
# them report the same value as you do. A component shared by 200 browsers
# carries log2(200) = 7.6 bits and no more, however many bits it could in
# principle have carried.
POPULATIONS = {
    "Windows 11 + Chrome desktop": {
        "size": 300_000, "hardwareConcurrency": 10, "deviceMemory": 6,
        "pixelRatio": 8, "maxTouchPoints": 40, "colorDepth": 300_000,
        "canvas2d": 400, "webgl": 300, "webglParams": 250,
        "audio": 500, "fonts": 200, "fontMetrics": 300,
    },
    "macOS 14 + Safari": {
        "size": 90_000, "hardwareConcurrency": 8, "deviceMemory": 5,
        "pixelRatio": 12, "maxTouchPoints": 90_000, "colorDepth": 90_000,
        "canvas2d": 260, "webgl": 220, "webglParams": 180,
        "audio": 90, "fonts": 150, "fontMetrics": 240,
    },
    "Android 13 + Chrome mobile": {
        "size": 400_000, "hardwareConcurrency": 24, "deviceMemory": 14,
        "pixelRatio": 16, "maxTouchPoints": 30, "colorDepth": 400_000,
        "canvas2d": 700, "webgl": 500, "webglParams": 600,
        "audio": 120, "fonts": 90, "fontMetrics": 200,
    },
    "Headless Linux container": {
        "size": 4_000, "hardwareConcurrency": 1_100, "deviceMemory": 1_500,
        "pixelRatio": 1_900, "maxTouchPoints": 3_900, "colorDepth": 4_000,
        "canvas2d": 2_600, "webgl": 2_900, "webglParams": 3_000,
        "audio": 800, "fonts": 1_800, "fontMetrics": 2_400,
    },
    "Spoofed Win11 on headless host": {
        "size": 400_000, "hardwareConcurrency": 12, "deviceMemory": 9,
        "pixelRatio": 5, "maxTouchPoints": 3, "colorDepth": 400_000,
        "canvas2d": 1, "webgl": 1, "webglParams": 2,
        "audio": 1, "fonts": 3, "fontMetrics": 2,
    },
}

INFORMATIVE_SHARE = 0.01                # a component needs >1% spread to count
REALISTIC_FLOOR = 5                     # no real value is shared by fewer than this


def collide(bits, world=WORLD):
    # P(at least one other browser among `world` matches on every component).
    if bits <= 0:
        return 1.0
    return -math.expm1(world * math.log1p(-(2.0 ** -bits)))


def census(profile):
    # Score one device class the way a plausibility check would.
    # Summing component bits always overshoots, because the components are
    # correlated and because a class of N browsers only has log2(N) bits to give.
    # So uniqueness is the *normal* case, and the checks worth running are the two
    # ends of the distribution: components so common inside the class that they
    # carry nothing, and values so rare that no real population produces them
    size = profile["size"]
    bits = coarse = 0
    rarest_name, rarest = None, size
    for name, raw, _, varies, _ in COMPONENTS:
        if not varies:
            continue                            # selects a device family, not a person
        share = profile.get(name, size)
        if share < rarest:
            rarest_name, rarest = name, share
        if share > size * INFORMATIVE_SHARE:
            coarse += 1                           # shared by most of the class: no bits
            continue
        bits += min(raw, math.log2(share))
    return min(bits, math.log2(size)), max(1, size / 2 ** bits), coarse, rarest_name, rarest


def main():
    stable = [c for c in COMPONENTS if c[4]]
    lost = sum(b for _, b, _, _, ok in COMPONENTS if not ok)
    total = sum(b for _, b, _, _, _ in COMPONENTS)
    within_family = sum(1 for c in COMPONENTS if c[3])

    print(f"{len(COMPONENTS)} components, {len(stable)} stable across visits, "
          f"{within_family} that vary inside a device family\n")
    print(f"world population {WORLD:,} -> uniqueness ceiling {CEILING:.1f} bits\n")

    print("bits as they accumulate; nothing above the ceiling adds anything:")
    running = 0.0
    for index, (name, bits, source, varies, stable_flag) in enumerate(COMPONENTS, 1):
        before = running
        running = min(running + bits, CEILING)
        note = "" if running - before == bits else f"  only +{running - before:.1f} useful"
        if not varies:
            note += "  (selects a device family, not an individual)"
        if not stable_flag:
            note += "  [unstable]"
        print(f"  {index:2}. {name:22} {source:26} +{bits:5.1f} -> {running:5.1f}{note}")
    print(f"      {'total':22} {'':26} {total:5.1f} raw bits, {running:.1f} usable\n")

    print(f"raw total                      {total:6.1f} bits")
    print(f"raw total, stable components   {total - lost:6.1f} bits "
          f"({lost:.1f} lost to components that change per visit)")
    print(f"at the population ceiling      {min(total, CEILING):6.1f} bits, "
          f"collision odds {collide(CEILING):.4f}")
    print("  which is why every vendor publishes a number near 18 to 20 bits:\n")

    print(f"{'device class':30} {'bits':>5} {'crowd':>7} {'coarse':>7}  rarest value      reading")
    for label, profile in POPULATIONS.items():
        bits, crowd, coarse, rarest_name, rarest = census(profile)
        if rarest < REALISTIC_FLOOR:
            reading = f"{rarest_name} shared by only {rarest}: not a real population"
        elif coarse >= 7:
            reading = "a pool of clones, one value repeated everywhere"
        elif crowd > 20:
            reading = "unusually rare, cross-check before trusting it"
        else:
            reading = "an ordinary device in an ordinary crowd"
        print(f"{label:30} {bits:5.1f} {crowd:7.0f} {coarse:7}  "
              f"{rarest_name}={rarest:<7} {reading}")


main()
25 components, 21 stable across visits, 14 that vary inside a device family

world population 1,000,000 -> uniqueness ceiling 19.9 bits

bits as they accumulate; nothing above the ceiling adds anything:
   1. userAgent              request header             + 10.0 ->  10.0  (selects a device family, not an individual)
   2. platform               navigator                  +  3.5 ->  13.5  (selects a device family, not an individual)
   3. vendor                 navigator                  +  2.0 ->  15.5  (selects a device family, not an individual)
   4. languages              navigator                  +  4.5 ->  19.9  only +4.4 useful  (selects a device family, not an individual)
   5. timezone               Intl                       +  8.0 ->  19.9  only +0.0 useful  (selects a device family, not an individual)
   6. screen                 navigator.screen           +  9.0 ->  19.9  only +0.0 useful  (selects a device family, not an individual)
   7. intlLocales            Intl.supportedLocalesOf    +  6.5 ->  19.9  only +0.0 useful  (selects a device family, not an individual)
   8. mediaCodecs            canPlayType                +  4.0 ->  19.9  only +0.0 useful  (selects a device family, not an individual)
   9. colorDepth             navigator.screen           +  1.5 ->  19.9  only +0.0 useful  (selects a device family, not an individual)
  10. cookiesEnabled         navigator                  +  0.6 ->  19.9  only +0.0 useful  (selects a device family, not an individual)
  11. doNotTrack             navigator                  +  0.9 ->  19.9  only +0.0 useful  (selects a device family, not an individual)
  12. hardwareConcurrency    navigator                  +  3.0 ->  19.9  only +0.0 useful
  13. deviceMemory           navigator                  +  2.5 ->  19.9  only +0.0 useful
  14. pixelRatio             devicePixelRatio           +  3.0 ->  19.9  only +0.0 useful
  15. maxTouchPoints         navigator                  +  2.0 ->  19.9  only +0.0 useful
  16. canvas2d               canvas.toDataURL           + 12.0 ->  19.9  only +0.0 useful
  17. webgl                  WEBGL_debug_renderer_info  + 11.0 ->  19.9  only +0.0 useful
  18. webglParams            gl.getParameter()          +  9.5 ->  19.9  only +0.0 useful
  19. audio                  AudioContext               + 10.5 ->  19.9  only +0.0 useful
  20. fonts                  offsetWidth probe          +  9.0 ->  19.9  only +0.0 useful
  21. fontMetrics            canvas measureText         +  7.5 ->  19.9  only +0.0 useful
  22. adBlock                heuristic probe            +  1.2 ->  19.9  only +0.0 useful  [unstable]
  23. perfTiming             performance.now()          +  5.0 ->  19.9  only +0.0 useful  [unstable]
  24. storageProbe           localStorage / IDB         +  3.0 ->  19.9  only +0.0 useful  [unstable]
  25. webrtc                 RTCPeerConnection          +  4.0 ->  19.9  only +0.0 useful  [unstable]
      total                                             133.7 raw bits, 19.9 usable

raw total                       133.7 bits
raw total, stable components    120.5 bits (13.2 lost to components that change per visit)
at the population ceiling        19.9 bits, collision odds 0.6321
  which is why every vendor publishes a number near 18 to 20 bits:

device class                    bits   crowd  coarse  rarest value      reading
Windows 11 + Chrome desktop     18.2       1       4  deviceMemory=6       an ordinary device in an ordinary crowd
macOS 14 + Safari               16.5       1       5  deviceMemory=5       an ordinary device in an ordinary crowd
Android 13 + Chrome mobile      18.6       1       4  deviceMemory=14      an ordinary device in an ordinary crowd
Headless Linux container         0.0    4000      14  audio=800     a pool of clones, one value repeated everywhere
Spoofed Win11 on headless host  13.0      49       4  canvas2d=1       canvas2d shared by only 1: not a real population

What the table actually shows is the reason a uniqueness score is the wrong tool for detection. The three real device classes all saturate the ceiling and all have a crowd of one, because a real desktop browser is unique among a million of them. Uniqueness is the normal case, not the signal.

The two non-real rows are the informative ones, and they are informative for opposite reasons:

  • The headless container has 14 coarse components. Every hardware value is shared by a large fraction of a 4,000-browser pool, so the profile carries almost no information and the crowd is the whole pool. This is the signature of automation: not a rare device but a clone farm, where a thousand visitors report an identical vector.
  • The spoofed profile claims Windows but has a canvas2d value shared by exactly one browser in the class. Real canvas hashes cluster, because millions of machines run the same handful of driver versions; a value that unique inside a 400,000-browser Windows population does not exist. This is the signature of a copy-pasted fingerprint: too good to be true.

So the checks worth running are at the two ends of the distribution. Components so common inside the class that they carry nothing, and values so rare that no real population produces them. A WAF doing cross-validation is doing exactly this, and Fingerprint Cross-Validation covers the contradiction matrix.

Visitor ID Versus Fingerprint

A fingerprint is the hash of the component vector. A visitor ID is a server-side bucket: the vendor assigns your fingerprint to one of a number of groups, and a visit counts as unique if the group is new. FingerprintingJS reports a visitor ID as a two-part identifier (a surrogate ID and a random site-specific ID) precisely because the raw hash should not be exposed.

The distinction matters because of instability. A component that changes between two visits makes the vector change, and a collector that hashes the whole vector produces a different fingerprint each time. A few components are known unstable and are handled specially:

Component Instability Handling
screen.availWidth/Height changes when the taskbar moves or a display is unplugged dropped, or normalised
performance.now() resolution varies with hardware and power state binned, or excluded
ad-block detection changes when the user installs an extension a separate low-weight key
storage state changes when the user clears site data excluded from the stable key
WebRTC candidates change per network, and per session excluded, or a separate key

When a collector cannot get a component to be stable it has two bad options: include it and lose stability, or drop it and lose coverage. Vendors differ here, which is the main reason their published bit counts differ for the same browser. The consistency lesson makes the same point from the other direction: a profile that is unique but changes every visit is a worse identity than one that is slightly coarser and stable, because stability across visits is what a WAF correlates on.

When a Component Throws

Probes fail, and how they fail is data. The common cases:

  • Safari and Intl: a DateTimeFormat constructed with an invalid calendar throws a RangeError, and which inputs throw varies by ICU version.
  • WebGL in a restricted context: getContext('webgl') returns null when GPU access is blocked, in a headless build without a GPU, or under a privacy extension.
  • Audio without an output device: an AudioContext in a machine with no sound card can start but produce silence, and a collector that expects a non-zero sample gets a suspiciously round number.
  • Storage under a strict policy: localStorage throws a SecurityError in some privacy modes rather than returning null.
  • Permissions API: query() resolves with a state the user never chose, or rejects outright.

A well-built collector catches the error, records a sentinel, and continues. A naive one lets the exception propagate and the whole probe returns undefined for that key. Both are observable: the sentinel value is a different value from the real one, and a key that is missing entirely is different again. A collector that reports canvas2d: undefined on a desktop Chrome profile is telling you it did not run a real probe.

This is the same lesson as Fixing CDP & Headless Leaks, from the collector's side: the failure mode is the signal, and both parties know it.

The Public Test Pages

The open-source collectors are public, and the sites that run them are the best regression suite available:

Page What it uniquely probes
browserleaks per-API depth: separate pages for canvas, WebGL, fonts, audio, screen, Java, Flash-era plugins, navigator in full
bot.sannysoft the automation surface: navigator.webdriver, plugin counts, window.chrome, permissions, WebGL renderer, error-stack artefacts
creepjs lying detectors: it plants deliberate falsehoods (a fake navigator.plugins entry, a spoofed screen) and reports which ones you did not correct, on the theory that a bot fixes the values it knows about and misses the rest
fingerprintjs demo the open-source collector itself, component by component, with each component's raw value and hash shown
bot-detection (a) leak tests run one per line, easy to diff between runs

creepjs deserves the emphasis because it inverts the usual assumption. A naive scraper patches navigator.webdriver and stops; a slightly better one patches the documented list; creepjs succeeds by checking whether the profile is internally consistent, so a partial patch is worse than none. The lesson it teaches is covered in Device Profile Consistency.

What to Read in a Site's Bundle

When you want to know what a specific site collects, four checks in order of cost:

  1. Grep the bundle for API names. The minified bundle still contains property names, because they are not mangled. Search for toDataURL, getParameter, WEBGL_debug_renderer_info, hardwareConcurrency, AudioContext, maxTouchPoints, Intl., measureText, fonts. A hit tells you the component is collected; the surrounding code tells you how.
  2. Look for the hashing call. crypto.subtle.digest, or a bundled xxHash or murmur, tells you the values are being combined rather than just sent. The separator used is worth reading: it is what makes the hash reproducible.
  3. Watch the network calls. A collector either posts the components to your own server, or sends a hash to a third party. The endpoint name identifies the vendor, and the request body shape tells you whether you are looking at raw components or an already-hashed value.
  4. Read what it does with errors. Search for try, catch, and any undefined default near a component. The error path is where the vendor's assumptions are visible.

Wireshark or a proxy that records the collector's own request is often faster than reading minified code, and the MITM and HAR lesson covers the capture. Read the code on a system you are authorised to test, or your own site.

Why Vendors' Numbers Differ

When two products report different entropy for the same browser, the reasons are usually these, in order of how much they matter:

Difference Effect on the number
different component set a missing canvas drops it by 10+ bits
different population measured a vendor with 10M browsers can claim more bits than one with 1M
different stability filtering components that change per visit get dropped, costing bits
different error handling a sentinel vs a missing key changes the hash, not the count
different normalisation is 1920x1080 one value or three? rounding rules differ
different entropy method Shannon over measured frequencies vs an assumed uniform split
a browser-computed component one native hash can be worth 20 bits of the total

The practical consequence: a published bit count is a claim about a vendor's measurement, not a property of the browser. Treat it as a hint about which components matter and ignore the number itself.

The Legitimate Route

Running a collector against your own browser to confirm that your profile is internally self-consistent, and reading the source to understand which components a site requests before you decide what to send, are both normal engineering. Using that same source knowledge to enumerate exactly what a particular site harvests in order to defeat its controls is a different act and belongs behind written authorisation. On the defensive side, the same reading is simply how you audit your own site: if you cannot justify each component your own collector requests, it is a signal you are collecting more than you use. The defending against scrapers lesson covers the other side of that decision.

Cross-links: Browser Fingerprinting, Inside Open-Source Fingerprint Collectors, Spoofing Canvas, WebGL & Audio, Device Profile Consistency, Fingerprint Cross-Validation, Fixing CDP & Headless Leaks, Font Metrics and Intl Fingerprinting, Defending Against Scrapers.