The Driver Is Part of the Fingerprint

Choosing between Playwright, Puppeteer and Selenium is usually made on ergonomics: language binding, API shape, debugging story. That is the wrong criterion, because each framework launches the browser with a different set of command-line switches, a different debugging transport, and a different set of injected globals. A WAF cannot see your Python process, but it can see all three.

The practical consequence is that the framework is a decision you make once and then cannot change per request, and a framework that is fine for a functional test can be an instant ban for a scraper. Pick the one whose launch surface you can control, and audit the switches it adds.

What Playwright Does at Launch

Playwright launches Chromium through CDP over its own transport. It does not use --remote-debugging-port; it uses --remote-debugging-pipe, passing a file descriptor 3 and 4 to the child process, which means there is no listening TCP port for a page or a local process to discover. It sets --enable-automation, which is what makes navigator.webdriver return true unless you pass --disable-blink-features=AutomationControlled or install a context init script that redefines the property.

Its default switches are a fixed list that has changed across releases, and it injects a __playwright__binding__ global plus a __pwInitScripts registry when you use add_init_script. On a persistent context it also writes a playwright section into the profile's Preferences file. Playwright's advantage over the other two is that the context model is clean: browser.new_context() gives an isolated BrowserContext with its own cookie jar, storage and proxy, which makes session isolation a one-liner rather than a profile directory. Its disadvantage is that the default argument list is not documented as a contract, so it drifts with releases.

What Puppeteer Does Differently

Puppeteer also speaks CDP, and it also defaults to a pipe rather than a port, but its default argument list is longer and older: --disable-background-networking, --disable-background-timer-throttling, --disable-backgrounding-occluded-windows, --disable-breakpad, --disable-client-side-phishing-detection, --disable-component-extensions-with-background-pages, --disable-default-apps, --disable-dev-shm-usage, --disable-extensions, --disable-features=Translate,AcceptCHFrame,MediaRouter,OptimizationHints, --disable-hang-monitor, --disable-ipc-flooding-protection, --disable-popup-blocking, --disable-prompt-on-repost, --disable-renderer-backgrounding, --disable-sync, --enable-automation, --force-color-profile=srgb, --metrics-recording-only, --no-first-run, --password-store=basic, --use-mock-keychain, --use-gl=angle, --font-render-hinting=none.

Two of those are fingerprinting surfaces on their own. --enable-automation is the source of navigator.webdriver. --force-color-profile=srgb and --font-render-hinting=none change the pixel output of canvas and the measured advance width of text, so a Puppeteer canvas hash is not the canvas hash of the same machine in Chrome. The puppeteer-extra-plugin-stealth package exists precisely to undo this list, and it works by removing arguments as well as patching JavaScript.

Selenium 4 and the WebDriver Contract

Selenium is a specification implemented by a driver, and the specification itself is the tell. The W3C WebDriver algorithm requires the driver to set navigator.webdriver to true when it is running under automation, so this is not a Chromium bug you can patch around: it is the contract. ChromeDriver also injects the cdc_adoQpoasnfa76pfcZLmcfl_* globals when it executes scripts, and it sets document.webdriver-evaluate and document.selenium-evaluate properties on every execute_script call. The full enumeration of those names is in CDP and Instrumentation Leaks.

Selenium 4.6 and later added Selenium Manager, which downloads the correct chromedriver, geckodriver or msedgedriver binary at run time by reading the browser's version from the filesystem. That is convenient and it is a fingerprint: a network request to a driver CDN from a process that has never visited Google, a chromedriver process tree, and a --user-data-dir that Selenium creates under a temporary path with a name that encodes the session id.

The 2022 to 2024 shift from JSON Wire Protocol to W3C also changed what is on the wire. The old protocol sent a body of unprefixed command objects over HTTP; W3C sends a POST /session/{id}/... with a JSON body carrying a capitalized command name, and the response uses a value envelope with a numeric status inside it. The getCapabilities response shape is what tells you the browser version, and the platformName and browserVersion inside it are what a WAF-visible fingerprint would be derived from on the client side.

Headless Shell, Old Headless and New Headless

Chromium has three distinct modes and conflating them is the most common version error in this area. The headless_shell binary is a separate download, a stripped build used by Puppeteer's headless: 'shell' and by older tooling; it has no GPU process, no extensions, no printing and a different media stack. --headless in a full Chrome build is "old headless", a mode that keeps the full browser but swaps out the windowing, and it leaves HeadlessChrome in the user agent string. --headless=new is the modern mode, which runs the real browser with a virtual window, and since Chrome 132 it is simply --headless.

The user agent token is the observable difference and it is trivially checked, so the practical rule is: never launch a full Chrome build in any headless mode for a long-lived identity. Launch a real browser, windowed, under a virtual display, and pay the memory cost. The comparison is in Stealth Browsers and the patching options are in Patching and Building Chromium for Stealth.

WebDriver BiDi and What Standardisation Removes

WebDriver BiDi replaces the HTTP-and-JSON command wire with a WebSocket that carries a single command envelope, {"id": 1, "method": "script.evaluate", "params": {...}}, in both directions, with events pushed over the same socket. It is the standard the other drivers are converging on: Selenium 4 has shipped BiDi support, Playwright has experimental BiDi, and geckodriver has implemented it for Firefox.

The reason it matters for fingerprinting is subtle and worth stating precisely. BiDi does not remove the automation marks: navigator.webdriver is still true under a conformant implementation, and a conformant implementation must set it. What BiDi removes is a class of implementation-specific artifacts, because the spec requires the same session shape regardless of which browser is underneath, so the driver-specific globals and the vendor-specific script bridges have fewer reasons to exist. What it adds is a new observable: a page can see that a WebSocket connection exists, though it cannot see the protocol, and Firefox's BiDi implementation has its own set of about: and preference side effects that differ from classic Marionette. Standardisation reduces one fingerprint family and introduces another, and it does not make the automation question go away.

Connecting to a Real Browser Versus Letting a Driver Launch One

This is the most important distinction in the lesson. When you launch the browser, you choose the switches, the user data directory, the extensions, the window size, the proxy and the timing. When a driver launches it, the driver chooses all of that and you get whatever defaults it has, and those defaults are part of the fingerprint whether or not you asked for them.

Aspect You launch it Driver launches it
Switch list exactly what you wrote driver defaults plus yours
Debug transport pipe you control pipe or port the driver picked
Profile directory yours, persistent driver temp dir with a session-encoded name
navigator.webdriver controllable via init script or a patched build forced to true by the WebDriver contract
Extra globals none unless you add them cdc_*, __webdriver_evaluate, __selenium_unwrapped
chrome.loadTimes, chrome.csi present in a full build absent in headless_shell
Version control you choose the binary driver resolves a matching version, sometimes at run time

For a scraping fleet the answer is consistent: launch the browser yourself, then attach over a pipe with connect_over_cdp or the Playwright equivalent. You pay one extra process boundary and you get control of every switch.

Switches Are a Fingerprint Too

The page cannot read the command line directly, but the consequences leak through chrome://version-adjacent surfaces, through the GPU and font subsystems the switches disable, and through the WAF's own correlation of the resulting signals. The switches that matter most in practice:

Switch Effect the page can observe
--headless, --headless=new HeadlessChrome token, window.chrome shape, no window manager
--enable-automation navigator.webdriver === true
--disable-gpu UNMASKED_RENDERER_WEBGL becomes SwiftShader or llvmpipe
--single-process no renderer crash isolation, different task timing
--mute-audio AudioContext runs, output silent, MediaElement.volume behaviour odd
--font-render-hinting=none measured glyph advance widths change, font hash changes
--force-color-profile=srgb canvas pixel values differ from the display profile
--no-sandbox no direct page read, but the sandbox status is inferable from process and crash behaviour
--disable-dev-shm-usage no direct read; changes renderer memory pressure and crash timing
--remote-debugging-port=N the port is discoverable, and devtools:// endpoints are reachable
--proxy-server=... leaks through WebRTC candidate gathering and DNS resolution order

Individually most of these are weak. Jointly they are a signature: no human launches Chrome with --no-sandbox --disable-dev-shm-usage --disable-gpu --single-process --mute-audio --font-render-hinting=none --headless=new --remote-debugging-port=9222, and the combination is what a scorer keys on. Hardening Playwright with CDP covers removing them, and the audit script below is the checklist form of the same idea.

Auditing a Launch Configuration

Keep your argument list in one place, review it like code, and score it before it reaches production. The script classifies each switch and then applies joint rules, because a set can be suspicious even when every member is individually defensible.

FLAG_CLASS = {
    "--no-sandbox":                  ("automation-default", "Docker/CI image default"),
    "--disable-setuid-sandbox":      ("automation-default", "pairs with --no-sandbox"),
    "--disable-dev-shm-usage":       ("automation-default", "Playwright default"),
    "--disable-gpu":                 ("developer",          "software rasteriser"),
    "--headless=new":                ("suspicious",         "new headless token"),
    "--headless":                    ("suspicious",         "old headless, UA leaks"),
    "--remote-debugging-port=0":     ("suspicious",         "leaves a CDP port open"),
    "--disable-blink-features=AutomationControlled": ("suspicious", "deletes webdriver"),
    "--single-process":              ("suspicious",         "one process for all"),
    "--mute-audio":                  ("developer",          "silent output"),
    "--font-render-hinting=none":    ("developer",          "changes glyph widths"),
    "--window-size=1920,1080":       ("real user",          "ordinary windowed launch"),
    "--lang=en-GB":                  ("real user",          "ordinary launch"),
    "--disable-background-timer-throttling": ("real user", "real Chrome on a bg tab"),
    "--enable-automation":           ("suspicious",         "sets the Blink auto bit"),
    "--password-store=basic":        ("developer",          "no keychain prompt"),
    "--disable-extensions":          ("developer",          "empty extension list"),
    "--use-mock-keychain":           ("automation-default", "Puppeteer headless default"),
    "--proxy-server=http://127.0.0.1:8080": ("real user", "a user's manual proxy"),
    "--host-resolver-rules=MAP * ~NOTFOUND": ("developer", "blocks DNS unless paired"),
    "--safebrowsing-disable-auto-update": ("developer",  "stable in CI only"),
}

PATTERN_HINTS = (
    ("automation", "suspicious", "name alone is a tell"),
    ("headless", "suspicious", "name alone is a tell"),
    ("remote-debug", "suspicious", "CDP is reachable from the page"),
    ("disable-", "automation-default", "looks like a CI default"),
    ("no-", "automation-default", "looks like a CI default"),
    ("mute-", "developer", "silent output"),
)


def classify(flag):
    if flag in FLAG_CLASS:
        return FLAG_CLASS[flag]
    if "=" in flag and flag.split("=", 1)[0] in FLAG_CLASS:
        return FLAG_CLASS[flag.split("=", 1)[0]]
    for needle, klass, why in PATTERN_HINTS:
        if needle in flag:
            return klass, why
    return "real user", "no automation signal in the name"


CANDIDATE = [
    "--no-sandbox",
    "--disable-setuid-sandbox",
    "--disable-dev-shm-usage",
    "--disable-gpu",
    "--single-process",
    "--mute-audio",
    "--font-render-hinting=none",
    "--headless=new",
    "--remote-debugging-port=9222",
    "--disable-blink-features=AutomationControlled",
    "--window-size=1920,1080",
    "--lang=en-GB",
    "--proxy-server=http://127.0.0.1:8080",
]

tally = {"real user": 0, "developer": 0, "automation-default": 0, "suspicious": 0}
rows = []
for flag in CANDIDATE:
    klass, why = classify(flag)
    tally[klass] += 1
    rows.append((flag, klass, why))

print("launch config audit: %d flags" % len(rows))
print("=" * 98)
print("%-49s %-19s %s" % ("flag", "class", "why"))
print("-" * 98)
for flag, klass, why in rows:
    print("%-49s %-19s %s" % (flag, klass, why))
print("=" * 98)
for klass in ("suspicious", "automation-default", "developer", "real user"):
    print("%-19s %d" % (klass, tally[klass]))
print("=" * 98)

joint = []
if tally["suspicious"] >= 2:
    joint.append("two or more single-flag tells: the set is itself the signal")
if tally["automation-default"] >= 3:
    joint.append("three or more CI defaults: a stock image, not a user")
if {"--single-process", "--disable-gpu"} <= set(CANDIDATE):
    joint.append("--single-process with --disable-gpu: no GPU process at all")
if tally["suspicious"] == 0:
    joint.append("no flag-level tell; remaining risk is inside the binary")
print("joint verdict:")
for line in joint:
    print("  - " + line)
print("  risk score: %d  (3 x suspicious + 1 x automation-default)"
      % (tally["suspicious"] * 3 + tally["automation-default"]))
launch config audit: 13 flags
==================================================================================================
flag                                              class               why
--------------------------------------------------------------------------------------------------
--no-sandbox                                      automation-default  Docker/CI image default
--disable-setuid-sandbox                          automation-default  pairs with --no-sandbox
--disable-dev-shm-usage                           automation-default  Playwright default
--disable-gpu                                     developer           software rasteriser
--single-process                                  suspicious          one process for all
--mute-audio                                      developer           silent output
--font-render-hinting=none                        developer           changes glyph widths
--headless=new                                    suspicious          new headless token
--remote-debugging-port=9222                      suspicious          CDP is reachable from the page
--disable-blink-features=AutomationControlled     suspicious          deletes webdriver
--window-size=1920,1080                           real user           ordinary windowed launch
--lang=en-GB                                      real user           ordinary launch
--proxy-server=http://127.0.0.1:8080              real user           a user's manual proxy
==================================================================================================
suspicious          4
automation-default  3
developer           3
real user           3
==================================================================================================
joint verdict:
  - two or more single-flag tells: the set is itself the signal
  - three or more CI defaults: a stock image, not a user
  - --single-process with --disable-gpu: no GPU process at all
  risk score: 15  (3 x suspicious + 1 x automation-default)

Four suspicious switches and three CI defaults. The joint verdict lines are the ones a per-flag review would miss: --single-process with --disable-gpu reads as two unrelated performance tweaks and is in fact the description of a machine with no graphics hardware at all.

The Decision Matrix

Job Reasonable choice Why
Functional testing your own site Selenium or Playwright, headless detection is not the concern
Scraping a site with no bot management httpx or curl_cffi, no browser cheaper and faster than any browser
Scraping with JavaScript challenges Playwright, launched by you, attached over a pipe full control of the switch list
Firefox-only target Playwright's Firefox, or a patched geckodriver over BiDi avoids the cdc_ family entirely
Safari-only target WebKit via Playwright, with the honest limits stated see Hardening Firefox and WebKit Profiles
Long-lived identities at volume anti-detect browser or a patched build, driven by Playwright the JS surface is not the binding constraint
A queue of small jobs one browser, many contexts see Browser Pool Architecture

The general rule: use the framework with the smallest launch surface you can get away with, launch the browser yourself, and treat the switch list as a reviewed artefact.

Failure Modes

  • Patching navigator.webdriver and stopping. --enable-automation also affects other Blink behaviour, and Function.prototype.toString on your own override is native-looking only if you patch that too.
  • Passing --headless=new for a profile that is meant to persist. The mode changes the window manager, and a headless browser has no tab lifecycle, so the visibilitychange and blur signals never fire.
  • Using chromium channel with headless: 'shell'. You get the stripped build: no extensions, no chrome.loadTimes, a different media stack, and a canvas hash that matches nobody.
  • Relying on Puppeteer's default args. They change between major versions, so a working configuration can become a fingerprint after an upgrade with no code change on your side.
  • Connecting over --remote-debugging-port=9222. The port is a well-known tell, and any local process can read the page. Use a pipe.
  • Comparing framework fingerprints across releases. The correct baseline is a real user on the same browser version, not last month's driver defaults.
  • Ignoring the driver download. Selenium Manager fetching chromedriver at run time is a network and process artefact that has nothing to do with your script.

Framework Checklist

  • The browser is launched by your code, not by a driver, and attached over a pipe.
  • The switch list lives in one versioned file and is audited by the script above before every release.
  • No --headless in any form for a long-lived identity; a virtual display is used instead.
  • No --remote-debugging-port; the transport is --remote-debugging-pipe.
  • --enable-automation is either removed or paired with an init script, and the init script also handles Function.prototype.toString.
  • --disable-gpu and --font-render-hinting are absent, so canvas, WebGL and font metrics come from real subsystems.
  • The framework version is pinned, and the default argument list is diffed on upgrade.
  • The choice is documented next to the profile it serves, so a future change to the driver is reviewed as a fingerprint change.

The Legitimate Route

A driver is a tool; the framework choice is an engineering decision, and the audit script belongs in your own regression suite. The line is about what you point it at: a site you operate, a site whose terms allow automated access, or a site that has given you a crawler token or an API. Automating access to a site that has blocked you, or rotating identities specifically to defeat a ban you were given, is not a fingerprinting problem and none of this changes it: see Ban Recovery and Backoff for what a legitimate client does after a 403.