Playwright, Puppeteer, Selenium and BiDi
The Driver Is Part of the Fingerprint
Choosing between Playwright, Puppeteer and Selenium is usually made on ergonomics: language binding, API shape, debugging story. That is the wrong criterion, because each framework launches the browser with a different set of command-line switches, a different debugging transport, and a different set of injected globals. A WAF cannot see your Python process, but it can see all three.
The practical consequence is that the framework is a decision you make once and then cannot change per request, and a framework that is fine for a functional test can be an instant ban for a scraper. Pick the one whose launch surface you can control, and audit the switches it adds.
What Playwright Does at Launch
Playwright launches Chromium through CDP over its own transport. It does not use --remote-debugging-port; it uses --remote-debugging-pipe, passing a file descriptor 3 and 4 to the child process, which means there is no listening TCP port for a page or a local process to discover. It sets --enable-automation, which is what makes navigator.webdriver return true unless you pass --disable-blink-features=AutomationControlled or install a context init script that redefines the property.
Its default switches are a fixed list that has changed across releases, and it injects a __playwright__binding__ global plus a __pwInitScripts registry when you use add_init_script. On a persistent context it also writes a playwright section into the profile's Preferences file. Playwright's advantage over the other two is that the context model is clean: browser.new_context() gives an isolated BrowserContext with its own cookie jar, storage and proxy, which makes session isolation a one-liner rather than a profile directory. Its disadvantage is that the default argument list is not documented as a contract, so it drifts with releases.
What Puppeteer Does Differently
Puppeteer also speaks CDP, and it also defaults to a pipe rather than a port, but its default argument list is longer and older: --disable-background-networking, --disable-background-timer-throttling, --disable-backgrounding-occluded-windows, --disable-breakpad, --disable-client-side-phishing-detection, --disable-component-extensions-with-background-pages, --disable-default-apps, --disable-dev-shm-usage, --disable-extensions, --disable-features=Translate,AcceptCHFrame,MediaRouter,OptimizationHints, --disable-hang-monitor, --disable-ipc-flooding-protection, --disable-popup-blocking, --disable-prompt-on-repost, --disable-renderer-backgrounding, --disable-sync, --enable-automation, --force-color-profile=srgb, --metrics-recording-only, --no-first-run, --password-store=basic, --use-mock-keychain, --use-gl=angle, --font-render-hinting=none.
Two of those are fingerprinting surfaces on their own. --enable-automation is the source of navigator.webdriver. --force-color-profile=srgb and --font-render-hinting=none change the pixel output of canvas and the measured advance width of text, so a Puppeteer canvas hash is not the canvas hash of the same machine in Chrome. The puppeteer-extra-plugin-stealth package exists precisely to undo this list, and it works by removing arguments as well as patching JavaScript.
Selenium 4 and the WebDriver Contract
Selenium is a specification implemented by a driver, and the specification itself is the tell. The W3C WebDriver algorithm requires the driver to set navigator.webdriver to true when it is running under automation, so this is not a Chromium bug you can patch around: it is the contract. ChromeDriver also injects the cdc_adoQpoasnfa76pfcZLmcfl_* globals when it executes scripts, and it sets document.webdriver-evaluate and document.selenium-evaluate properties on every execute_script call. The full enumeration of those names is in CDP and Instrumentation Leaks.
Selenium 4.6 and later added Selenium Manager, which downloads the correct chromedriver, geckodriver or msedgedriver binary at run time by reading the browser's version from the filesystem. That is convenient and it is a fingerprint: a network request to a driver CDN from a process that has never visited Google, a chromedriver process tree, and a --user-data-dir that Selenium creates under a temporary path with a name that encodes the session id.
The 2022 to 2024 shift from JSON Wire Protocol to W3C also changed what is on the wire. The old protocol sent a body of unprefixed command objects over HTTP; W3C sends a POST /session/{id}/... with a JSON body carrying a capitalized command name, and the response uses a value envelope with a numeric status inside it. The getCapabilities response shape is what tells you the browser version, and the platformName and browserVersion inside it are what a WAF-visible fingerprint would be derived from on the client side.
Headless Shell, Old Headless and New Headless
Chromium has three distinct modes and conflating them is the most common version error in this area. The headless_shell binary is a separate download, a stripped build used by Puppeteer's headless: 'shell' and by older tooling; it has no GPU process, no extensions, no printing and a different media stack. --headless in a full Chrome build is "old headless", a mode that keeps the full browser but swaps out the windowing, and it leaves HeadlessChrome in the user agent string. --headless=new is the modern mode, which runs the real browser with a virtual window, and since Chrome 132 it is simply --headless.
The user agent token is the observable difference and it is trivially checked, so the practical rule is: never launch a full Chrome build in any headless mode for a long-lived identity. Launch a real browser, windowed, under a virtual display, and pay the memory cost. The comparison is in Stealth Browsers and the patching options are in Patching and Building Chromium for Stealth.
WebDriver BiDi and What Standardisation Removes
WebDriver BiDi replaces the HTTP-and-JSON command wire with a WebSocket that carries a single command envelope, {"id": 1, "method": "script.evaluate", "params": {...}}, in both directions, with events pushed over the same socket. It is the standard the other drivers are converging on: Selenium 4 has shipped BiDi support, Playwright has experimental BiDi, and geckodriver has implemented it for Firefox.
The reason it matters for fingerprinting is subtle and worth stating precisely. BiDi does not remove the automation marks: navigator.webdriver is still true under a conformant implementation, and a conformant implementation must set it. What BiDi removes is a class of implementation-specific artifacts, because the spec requires the same session shape regardless of which browser is underneath, so the driver-specific globals and the vendor-specific script bridges have fewer reasons to exist. What it adds is a new observable: a page can see that a WebSocket connection exists, though it cannot see the protocol, and Firefox's BiDi implementation has its own set of about: and preference side effects that differ from classic Marionette. Standardisation reduces one fingerprint family and introduces another, and it does not make the automation question go away.
Connecting to a Real Browser Versus Letting a Driver Launch One
This is the most important distinction in the lesson. When you launch the browser, you choose the switches, the user data directory, the extensions, the window size, the proxy and the timing. When a driver launches it, the driver chooses all of that and you get whatever defaults it has, and those defaults are part of the fingerprint whether or not you asked for them.
| Aspect | You launch it | Driver launches it |
|---|---|---|
| Switch list | exactly what you wrote | driver defaults plus yours |
| Debug transport | pipe you control | pipe or port the driver picked |
| Profile directory | yours, persistent | driver temp dir with a session-encoded name |
navigator.webdriver |
controllable via init script or a patched build | forced to true by the WebDriver contract |
| Extra globals | none unless you add them | cdc_*, __webdriver_evaluate, __selenium_unwrapped |
chrome.loadTimes, chrome.csi |
present in a full build | absent in headless_shell |
| Version control | you choose the binary | driver resolves a matching version, sometimes at run time |
For a scraping fleet the answer is consistent: launch the browser yourself, then attach over a pipe with connect_over_cdp or the Playwright equivalent. You pay one extra process boundary and you get control of every switch.
Switches Are a Fingerprint Too
The page cannot read the command line directly, but the consequences leak through chrome://version-adjacent surfaces, through the GPU and font subsystems the switches disable, and through the WAF's own correlation of the resulting signals. The switches that matter most in practice:
| Switch | Effect the page can observe |
|---|---|
--headless, --headless=new |
HeadlessChrome token, window.chrome shape, no window manager |
--enable-automation |
navigator.webdriver === true |
--disable-gpu |
UNMASKED_RENDERER_WEBGL becomes SwiftShader or llvmpipe |
--single-process |
no renderer crash isolation, different task timing |
--mute-audio |
AudioContext runs, output silent, MediaElement.volume behaviour odd |
--font-render-hinting=none |
measured glyph advance widths change, font hash changes |
--force-color-profile=srgb |
canvas pixel values differ from the display profile |
--no-sandbox |
no direct page read, but the sandbox status is inferable from process and crash behaviour |
--disable-dev-shm-usage |
no direct read; changes renderer memory pressure and crash timing |
--remote-debugging-port=N |
the port is discoverable, and devtools:// endpoints are reachable |
--proxy-server=... |
leaks through WebRTC candidate gathering and DNS resolution order |
Individually most of these are weak. Jointly they are a signature: no human launches Chrome with --no-sandbox --disable-dev-shm-usage --disable-gpu --single-process --mute-audio --font-render-hinting=none --headless=new --remote-debugging-port=9222, and the combination is what a scorer keys on. Hardening Playwright with CDP covers removing them, and the audit script below is the checklist form of the same idea.
Auditing a Launch Configuration
Keep your argument list in one place, review it like code, and score it before it reaches production. The script classifies each switch and then applies joint rules, because a set can be suspicious even when every member is individually defensible.
FLAG_CLASS = {
"--no-sandbox": ("automation-default", "Docker/CI image default"),
"--disable-setuid-sandbox": ("automation-default", "pairs with --no-sandbox"),
"--disable-dev-shm-usage": ("automation-default", "Playwright default"),
"--disable-gpu": ("developer", "software rasteriser"),
"--headless=new": ("suspicious", "new headless token"),
"--headless": ("suspicious", "old headless, UA leaks"),
"--remote-debugging-port=0": ("suspicious", "leaves a CDP port open"),
"--disable-blink-features=AutomationControlled": ("suspicious", "deletes webdriver"),
"--single-process": ("suspicious", "one process for all"),
"--mute-audio": ("developer", "silent output"),
"--font-render-hinting=none": ("developer", "changes glyph widths"),
"--window-size=1920,1080": ("real user", "ordinary windowed launch"),
"--lang=en-GB": ("real user", "ordinary launch"),
"--disable-background-timer-throttling": ("real user", "real Chrome on a bg tab"),
"--enable-automation": ("suspicious", "sets the Blink auto bit"),
"--password-store=basic": ("developer", "no keychain prompt"),
"--disable-extensions": ("developer", "empty extension list"),
"--use-mock-keychain": ("automation-default", "Puppeteer headless default"),
"--proxy-server=http://127.0.0.1:8080": ("real user", "a user's manual proxy"),
"--host-resolver-rules=MAP * ~NOTFOUND": ("developer", "blocks DNS unless paired"),
"--safebrowsing-disable-auto-update": ("developer", "stable in CI only"),
}
PATTERN_HINTS = (
("automation", "suspicious", "name alone is a tell"),
("headless", "suspicious", "name alone is a tell"),
("remote-debug", "suspicious", "CDP is reachable from the page"),
("disable-", "automation-default", "looks like a CI default"),
("no-", "automation-default", "looks like a CI default"),
("mute-", "developer", "silent output"),
)
def classify(flag):
if flag in FLAG_CLASS:
return FLAG_CLASS[flag]
if "=" in flag and flag.split("=", 1)[0] in FLAG_CLASS:
return FLAG_CLASS[flag.split("=", 1)[0]]
for needle, klass, why in PATTERN_HINTS:
if needle in flag:
return klass, why
return "real user", "no automation signal in the name"
CANDIDATE = [
"--no-sandbox",
"--disable-setuid-sandbox",
"--disable-dev-shm-usage",
"--disable-gpu",
"--single-process",
"--mute-audio",
"--font-render-hinting=none",
"--headless=new",
"--remote-debugging-port=9222",
"--disable-blink-features=AutomationControlled",
"--window-size=1920,1080",
"--lang=en-GB",
"--proxy-server=http://127.0.0.1:8080",
]
tally = {"real user": 0, "developer": 0, "automation-default": 0, "suspicious": 0}
rows = []
for flag in CANDIDATE:
klass, why = classify(flag)
tally[klass] += 1
rows.append((flag, klass, why))
print("launch config audit: %d flags" % len(rows))
print("=" * 98)
print("%-49s %-19s %s" % ("flag", "class", "why"))
print("-" * 98)
for flag, klass, why in rows:
print("%-49s %-19s %s" % (flag, klass, why))
print("=" * 98)
for klass in ("suspicious", "automation-default", "developer", "real user"):
print("%-19s %d" % (klass, tally[klass]))
print("=" * 98)
joint = []
if tally["suspicious"] >= 2:
joint.append("two or more single-flag tells: the set is itself the signal")
if tally["automation-default"] >= 3:
joint.append("three or more CI defaults: a stock image, not a user")
if {"--single-process", "--disable-gpu"} <= set(CANDIDATE):
joint.append("--single-process with --disable-gpu: no GPU process at all")
if tally["suspicious"] == 0:
joint.append("no flag-level tell; remaining risk is inside the binary")
print("joint verdict:")
for line in joint:
print(" - " + line)
print(" risk score: %d (3 x suspicious + 1 x automation-default)"
% (tally["suspicious"] * 3 + tally["automation-default"]))
launch config audit: 13 flags
==================================================================================================
flag class why
--------------------------------------------------------------------------------------------------
--no-sandbox automation-default Docker/CI image default
--disable-setuid-sandbox automation-default pairs with --no-sandbox
--disable-dev-shm-usage automation-default Playwright default
--disable-gpu developer software rasteriser
--single-process suspicious one process for all
--mute-audio developer silent output
--font-render-hinting=none developer changes glyph widths
--headless=new suspicious new headless token
--remote-debugging-port=9222 suspicious CDP is reachable from the page
--disable-blink-features=AutomationControlled suspicious deletes webdriver
--window-size=1920,1080 real user ordinary windowed launch
--lang=en-GB real user ordinary launch
--proxy-server=http://127.0.0.1:8080 real user a user's manual proxy
==================================================================================================
suspicious 4
automation-default 3
developer 3
real user 3
==================================================================================================
joint verdict:
- two or more single-flag tells: the set is itself the signal
- three or more CI defaults: a stock image, not a user
- --single-process with --disable-gpu: no GPU process at all
risk score: 15 (3 x suspicious + 1 x automation-default)
Four suspicious switches and three CI defaults. The joint verdict lines are the ones a per-flag review would miss: --single-process with --disable-gpu reads as two unrelated performance tweaks and is in fact the description of a machine with no graphics hardware at all.
The Decision Matrix
| Job | Reasonable choice | Why |
|---|---|---|
| Functional testing your own site | Selenium or Playwright, headless | detection is not the concern |
| Scraping a site with no bot management | httpx or curl_cffi, no browser |
cheaper and faster than any browser |
| Scraping with JavaScript challenges | Playwright, launched by you, attached over a pipe | full control of the switch list |
| Firefox-only target | Playwright's Firefox, or a patched geckodriver over BiDi | avoids the cdc_ family entirely |
| Safari-only target | WebKit via Playwright, with the honest limits stated | see Hardening Firefox and WebKit Profiles |
| Long-lived identities at volume | anti-detect browser or a patched build, driven by Playwright | the JS surface is not the binding constraint |
| A queue of small jobs | one browser, many contexts | see Browser Pool Architecture |
The general rule: use the framework with the smallest launch surface you can get away with, launch the browser yourself, and treat the switch list as a reviewed artefact.
Failure Modes
- Patching
navigator.webdriverand stopping.--enable-automationalso affects other Blink behaviour, andFunction.prototype.toStringon your own override is native-looking only if you patch that too. - Passing
--headless=newfor a profile that is meant to persist. The mode changes the window manager, and a headless browser has no tab lifecycle, so thevisibilitychangeandblursignals never fire. - Using
chromiumchannel withheadless: 'shell'. You get the stripped build: no extensions, nochrome.loadTimes, a different media stack, and a canvas hash that matches nobody. - Relying on Puppeteer's default
args. They change between major versions, so a working configuration can become a fingerprint after an upgrade with no code change on your side. - Connecting over
--remote-debugging-port=9222. The port is a well-known tell, and any local process can read the page. Use a pipe. - Comparing framework fingerprints across releases. The correct baseline is a real user on the same browser version, not last month's driver defaults.
- Ignoring the driver download. Selenium Manager fetching chromedriver at run time is a network and process artefact that has nothing to do with your script.
Framework Checklist
- The browser is launched by your code, not by a driver, and attached over a pipe.
- The switch list lives in one versioned file and is audited by the script above before every release.
- No
--headlessin any form for a long-lived identity; a virtual display is used instead. - No
--remote-debugging-port; the transport is--remote-debugging-pipe. --enable-automationis either removed or paired with an init script, and the init script also handlesFunction.prototype.toString.--disable-gpuand--font-render-hintingare absent, so canvas, WebGL and font metrics come from real subsystems.- The framework version is pinned, and the default argument list is diffed on upgrade.
- The choice is documented next to the profile it serves, so a future change to the driver is reviewed as a fingerprint change.
The Legitimate Route
A driver is a tool; the framework choice is an engineering decision, and the audit script belongs in your own regression suite. The line is about what you point it at: a site you operate, a site whose terms allow automated access, or a site that has given you a crawler token or an API. Automating access to a site that has blocked you, or rotating identities specifically to defeat a ban you were given, is not a fingerprinting problem and none of this changes it: see Ban Recovery and Backoff for what a legitimate client does after a 403.