A Proxy Is Real Only If Everything Goes Through It

Setting --proxy-server and seeing your exit IP change proves almost nothing. The browser has at least three independent network paths -- the HTTP stack, the DNS resolver, and the UDP/QUIC and WebRTC paths -- and only the first one is obviously the proxy. A configuration that hides your origin on requests while still asking your own resolver for every hostname has moved the leak, not closed it.

This lesson is the plumbing: the three ways to attach a proxy, what each one does to DNS, and how to prove from inside the browser that the wiring is real.

Three Ways to Attach a Proxy

Method Scope DNS behaviour Detectability
--proxy-server whole browser process still uses the system resolver for names a chrome://net-export capture shows it plainly
PAC file (--proxy-pac-url) whole process, decided per URL depends on what the PAC returns a PAC fetch at startup is itself an odd request
browser.new_context(proxy=...) one context only resolved at the proxy invisible to the page

The command-line switch is the blunt one. It applies to every socket the process opens, and it is the right tool when one browser serves one identity through one exit. The PAC file is more expressive -- you can route different hosts through different exits -- and correspondingly more suspicious, because a real user's browser does not fetch a proxy auto-config script from a proxy vendor. Fetch the PAC from your own infrastructure, and expect it in the net-export log.

The per-context proxy is the interesting one, and the reason to use Playwright rather than driving a raw binary over CDP yourself. The automation layer opens a separate connection per context, applies the proxy to that connection only, and keeps the browser process itself off the target's network entirely. Two contexts can therefore sit on two different exits with two different profiles inside one browser process -- powerful, and the correlation risk pool architecture warns about.

DNS Is a Separate Path

This is the leak everyone forgets. For an HTTP proxy, Chromium opens a TCP connection to the proxy and sends CONNECT target.example:443. The hostname travels inside that request, so the proxy resolves it. But Chromium separately needs to know things for its own reasons -- preconnect hints, DNS prefetch, HSTS and DNS-based pinning checks, the SVCB and HTTPS records some sites now publish -- and those lookups go to whatever resolver the OS is configured with. Your ISP's resolver, or your cloud provider's, sees a query for a hostname you are about to visit while your requests come from a residential IP in another country.

The switch that closes it is --host-resolver-rules, and the canonical value maps every name to a failure so the browser can never resolve anything itself:

--host-resolver-rules="MAP * ~NOTFOUND, EXCLUDE localhost"

MAP * ~NOTFOUND makes every resolution attempt fail locally, which pushes the browser to ask the proxy for the address; EXCLUDE localhost keeps loopback names resolvable so that localhost development and internal admin pages still work. If you would rather keep DNS working locally for non-target hosts, map the target's wildcard explicitly (MAP *.target.example ~NOTFOUND) and leave the rest alone -- but be aware that the unmapped traffic is still visible to your resolver, so it is only honest if you do not care about it. The DNS resolver signals lesson covers what the resolver itself gives away.

SOCKS5 vs SOCKS5h

SOCKS5 has a hostname-addressing command that lets the client send a name to the proxy instead of an address. Almost nothing uses it by default, and the difference decides where your DNS runs.

  • socks5://host:port -- the client resolves the name locally and sends an IP address. Your resolver sees every hostname, and if the exit IP is in a different country the geo mismatch is a signal on its own.
  • socks5h://host:port -- the client sends the hostname and the proxy resolves it. The local resolver sees nothing, and the exit IP matches the DNS geography the site observes.

Most proxy vendors document socks5h on their own dashboards and almost nobody enables it, which is why socks5 configurations leak at scale. Check it explicitly: look at the gateway's logs, or resolve a name from inside the browser and compare the answer with your own resolver's. The leak is easy to confirm and easy to miss, because everything looks fine from inside the page.

Authenticated Proxies and the 407

A proxy that requires credentials answers 407 Proxy Authentication Required with a Proxy-Authenticate challenge, and the client resends with a Proxy-Authorization: Basic base64(user:pass) header. Two ways to supply it:

  • In the URL, http://user:pass@gate.example:8000. Simple, and it works in Playwright's proxy= option and in most CLI tools. The cost is that the credential now lives in your process arguments, your config file and your log lines; percent-encode any special characters, and treat the gateway URL as a secret.
  • At the browser level, the Proxy-Authorization header on a navigation. This does not work for CONNECT tunnels in most implementations, because the challenge-and-retry happens inside the network stack where your header has already been sent and discarded. Do not build a scheme on it.

Two operational notes. Rotating credentials per profile means a 407 storm mid-job, so put a 407-aware retry on the gateway rather than in the scrape loop. And an authentication failure looks like a connection error, not a 407, at the page.goto call site -- a run of "navigation failed" with zero HTTP status is a credentials problem until proven otherwise. The proxy pools and session stickiness lesson covers keeping a sticky exit for the whole flow; the credential has to be part of that stickiness, not re-injected per request.

The Other Two Escape Routes

QUIC. Chromium negotiates HTTP/3 over UDP/443 when the server advertises it, and UDP does not traverse an HTTP proxy at all. A browser behind --proxy-server can still open a direct UDP socket to the origin's IP, and QUIC's connection migration is built around exactly that. --disable-quic is the blunt fix and the right one; the finer-grained approach is to strip the alt=svc-style advertisement path, but you cannot control the server's record, so disabling is the honest answer. The HTTP/3 and transport fingerprints lesson covers what a QUIC fingerprint looks like from the server's side.

WebRTC. RTCPeerConnection gathers ICE candidates, and with the default policy it will happily produce host candidates carrying your real local IP, plus a UDP path to a STUN server. --force-webrtc-ip-handling-policy=disable_non_proxied_udp forces the network layer to relay everything through the configured proxy; default_public_interface_only is weaker because it still enumerates interfaces. Note the force- prefix matters: the unforced --webrtc-ip-handling-policy is a per-page preference a page can change for itself, and a page that changes it is a page that is testing you. The WebRTC and network information leaks lesson covers the API side.

Verifying the Wiring

Three checks, in increasing order of rigour:

  1. Exit IP through the browser. Navigate the page to a public "what is my IP" endpoint and read the body. This proves the HTTP path, and nothing else. It is the check everyone runs and it catches the least.
  2. Resolver check. A DNS-leak test page reports the resolvers it observed while loading, and the answer to look for is your gateway's or your proxy's resolver, not your ISP's or your cloud provider's. Run it through the same context and the same proxy as the real job.
  3. chrome://net-export. The authoritative source. Start a capture, load the page, stop the capture, and open the .json file: it lists every request, the proxy each connection used, and the resolution path. A connection with a remote_endpoint that is not your proxy, or a host resolved by a local resolver, is the leak, sitting in a file you can attach to a bug.

Check one, two and three after any change to launch flags, and check three whenever a new target is added to the fleet. It is the only check that sees paths nobody asked about.

Per-Context Scope and the Correlation Risk

A per-context proxy is scoped correctly, and it creates a new problem. If one browser process holds contexts on eight different residential exits, each with a different profile, the site sees eight unrelated devices that all share a TLS session cache, a shared font configuration and a shared GPU renderer string within the same process, arriving with connection-level similarities. Worse, some of those contexts will carry identical fingerprints while sitting on different IPs, which is a device hopping countries.

The safe shapes are: one browser process per identity, with the proxy fixed at launch; or one process with contexts that share a profile and an exit, used sequentially rather than concurrently. What to avoid is a shared process fanning out across many exits at once, which is cheap and which the fingerprint cross-validation rules will eventually catch. Decide the scope deliberately, and record it in the config, because it is invisible in the code that uses the context.

Auditing a Launch Command

Argument lists grow by accretion, and a flag added for one target quietly applies to every target. Auditing them as data -- rather than by reading the code -- is cheap and catches the whole class of mistake. The analyser below classifies each argument by which network path it touches, then renders a verdict for the three channels that can leak: DNS resolution, UDP/QUIC, and WebRTC ICE.

CONFIGS = {
    "A  --proxy-server only": [
        "--headless=new",
        "--proxy-server=http://gate.example.net:8000",
        "--disable-blink-features=AutomationControlled",
    ],
    "C  fully wired": [
        "--headless=new",
        "--proxy-server=http://user:pass@gate.example.net:8000",
        "--host-resolver-rules=MAP * ~NOTFOUND, EXCLUDE localhost",
        "--disable-quic",
        "--force-webrtc-ip-handling-policy=disable_non_proxied_udp",
        "--proxy-bypass-list=<-loopback>",
    ],
}


def classify(args):
    rows = []
    for arg in args:
        if arg.startswith("--proxy-server="):
            scheme = arg.split("=", 1)[1].split(":", 1)[0]
            if scheme == "socks5":
                note = "SOCKS5: resolved locally, not at the proxy"
            elif scheme == "socks5h":
                note = "SOCKS5h: the proxy resolves the hostname"
            else:
                note = "HTTP CONNECT sends hostname, 407 on auth"
            rows.append((arg, "proxy", note))
        elif arg.startswith("--host-resolver-rules="):
            rows.append((arg, "dns", "MAP * ~NOTFOUND means no local lookups"))
        elif arg == "--no-proxy-server":
            rows.append((arg, "proxy", "proxies disabled for the whole process"))
        elif arg.startswith("--proxy-pac-url="):
            rows.append((arg, "proxy", "PAC fetched at startup, net-export shows it"))
        elif arg.startswith("--proxy-bypass-list="):
            rows.append((arg, "proxy", "bypass rules: " + arg.split("=", 1)[1]))
        elif arg == "--disable-quic":
            rows.append((arg, "udp", "UDP/443 HTTP/3 off, no QUIC bypass"))
        elif arg.startswith("--force-webrtc-ip-handling-policy="):
            rows.append((arg, "webrtc", "forced " + arg.split("=", 1)[1]))
        elif arg.startswith("--webrtc-ip-handling-policy="):
            rows.append((arg, "webrtc", "not forced, a page can override it"))
        elif arg == "--enable-automation":
            rows.append((arg, "automation", "turns on navigator.webdriver"))
        elif arg.startswith("--headless"):
            rows.append((arg, "automation", "HeadlessChrome token, no plugins"))
        elif arg.startswith("--disable-blink-features="):
            rows.append((arg, "automation", "JS-level fix, no transport effect"))
        else:
            rows.append((arg, "-", "no effect on the network path"))
    return rows


def verdicts(args):
    def has(prefix):
        return any(a == prefix or a.startswith(prefix) for a in args)

    def value(prefix):
        return next((a.split("=", 1)[1] for a in args
                     if a.startswith(prefix) and "=" in a), None)

    out = []
    if has("--no-proxy-server"):
        out.append(("DNS", "LEAK", "the proxy is off, every lookup is local"))
    elif has("--proxy-pac-url"):
        out.append(("DNS", "REVIEW", "a PAC decides routing; read its FindProxyForURL"))
    elif "~NOTFOUND" in (value("--host-resolver-rules") or ""):
        out.append(("DNS", "OK", "all names fail locally, the proxy resolves them"))
    elif has("--proxy-server="):
        out.append(("DNS", "LEAK", "hostnames still go to the system resolver"))
    else:
        out.append(("DNS", "OK", "direct connection, nothing to leak"))

    if has("--disable-quic"):
        out.append(("UDP", "OK", "QUIC disabled, no UDP/443 path out"))
    elif (value("--proxy-server") or "").startswith("socks5://"):
        out.append(("UDP", "LEAK", "SOCKS5 resolves locally, QUIC still on"))
    else:
        out.append(("UDP", "REVIEW", "QUIC on; does the proxy carry UDP at all?"))

    policy = value("--force-webrtc-ip-handling-policy")
    if policy == "disable_non_proxied_udp":
        out.append(("WEBRTC", "OK", "non-proxied UDP candidates suppressed"))
    elif policy:
        out.append(("WEBRTC", "REVIEW", "forced policy is " + policy))
    else:
        out.append(("WEBRTC", "LEAK", "ICE can expose the local interface address"))

    return out


for name, args in CONFIGS.items():
    print("config " + name)
    print("-" * 78)
    for arg, touches, note in classify(args):
        print("  " + arg)
        print("      {:<10} {}".format(touches, note))
    print()
    print("  {:<8} {:<7} {}".format("CHANNEL", "VERDICT", "REASON"))
    print("  " + "-" * 74)
    rows = verdicts(args)
    for channel, verdict, reason in rows:
        print("  {:<8} {:<7} {}".format(channel, verdict, reason))
    leaking = [c for c, v, _ in rows if v == "LEAK"]
    print()
    print("  overall: {}".format(
        "CLEAN" if not leaking else "LEAKING through " + ", ".join(leaking)))
    print()
config A  --proxy-server only
------------------------------------------------------------------------------
  --headless=new
      automation HeadlessChrome token, no plugins
  --proxy-server=http://gate.example.net:8000
      proxy      HTTP CONNECT sends hostname, 407 on auth
  --disable-blink-features=AutomationControlled
      automation JS-level fix, no transport effect

  CHANNEL  VERDICT REASON
  --------------------------------------------------------------------------
  DNS      LEAK    hostnames still go to the system resolver
  UDP      REVIEW  QUIC on; does the proxy carry UDP at all?
  WEBRTC   LEAK    ICE can expose the local interface address

  overall: LEAKING through DNS, WEBRTC

config C  fully wired
------------------------------------------------------------------------------
  --headless=new
      automation HeadlessChrome token, no plugins
  --proxy-server=http://user:pass@gate.example.net:8000
      proxy      HTTP CONNECT sends hostname, 407 on auth
  --host-resolver-rules=MAP * ~NOTFOUND, EXCLUDE localhost
      dns        MAP * ~NOTFOUND means no local lookups
  --disable-quic
      udp        UDP/443 HTTP/3 off, no QUIC bypass
  --force-webrtc-ip-handling-policy=disable_non_proxied_udp
      webrtc     forced disable_non_proxied_udp
  --proxy-bypass-list=<-loopback>
      proxy      bypass rules: <-loopback>

  CHANNEL  VERDICT REASON
  --------------------------------------------------------------------------
  DNS      OK      all names fail locally, the proxy resolves them
  UDP      OK      QUIC disabled, no UDP/443 path out
  WEBRTC   OK      non-proxied UDP candidates suppressed

  overall: CLEAN

Config A is the shape almost every first attempt has: a proxy, a JS-level automation fix, and two live leaks. The REVIEW on UDP is not a pass -- it means the answer depends on whether your proxy carries UDP at all, and you should find out rather than guess. Config C differs only in the three flags that close the other paths, which is the entire argument for keeping the launch configuration in one reviewable place.

Checklist

  • One proxy scope per identity, recorded explicitly in config: process-level switch, PAC URL, or per context. Never all three at once.
  • --host-resolver-rules set whenever a proxy is set, with the target's own names covered by the MAP rule and loopback excluded.
  • socks5h:// rather than socks5:// for any SOCKS proxy, confirmed in the gateway's own resolution logs.
  • Credentials supplied through the URL or the gateway, never as a Proxy-Authorization header on a navigation, and a 407 retry handled at the gateway.
  • --disable-quic set, and --force-webrtc-ip-handling-policy=disable_non_proxied_udp rather than the unforced per-page preference.
  • --proxy-bypass-list reviewed for entries that would silently send a target request direct, such as a bare * or an over-broad suffix.
  • Exit IP, observed resolver and chrome://net-export re-checked after every flag change, and net-export captured for every new target.

Staying Inside the Line

Everything here is about making your own authorised traffic look like the traffic you are entitled to send, and about not accidentally identifying your own infrastructure to a third party. It is not a licence to rotate through a residential pool to strip a rate limit that a site set deliberately, to re-add traffic from an address you have been blocked from, or to make a paywalled endpoint look like an ordinary page load. Respect Retry-After, keep a rate the site could plausibly have intended for a human, and prefer the official API, the partner feed or a negotiated agreement when one exists. The IP reputation lesson is where the ethics of proxy use are treated properly; this page is only the wiring.