DNS Resolver Signals and DNS Leaks
The Lookup You Forgot to Hide
A proxy changes the source address of your TCP connection and nothing else. The DNS lookup that turned a hostname into an IP address happened before that connection existed, was sent to whichever resolver your operating system was configured with, and left that resolver's logs with a record of your real network. Meanwhile the web application firewall you were trying to be invisible to is looking at a residential address in another country. Two systems now hold contradictory stories about you, and only one of them is a session you can influence. This is the most common misconfiguration in browser-based scraping, and it is invisible from inside the browser unless you go looking for it.
What a Resolver Knows
Resolving shop.example.com is a chain of questions, not a single request. Your stub resolver asks a recursive resolver, which may ask a root, which refers it to a TLD server, which refers it to the authoritative nameserver for the zone. The recursive resolver is the party that sees everything, because it is the one that talks to the client. For that resolver, one lookup is a record containing: the queried name, the exact second it was asked, the source address the question came from, the transport used (plain UDP, TLS, or HTTPS), the EDNS client subnet if the client offered one, and the full recursion the resolver performed. That is a correlation-rich event even though it contains no name-to-address mapping of value to most observers. Encrypted DNS hides the content of the lookup from the network; it does nothing to hide the asking from the resolver, because the resolver is the endpoint.
How a WAF Attributes a Request to a Resolver
A site cannot see your resolver directly from an HTTP request. It infers one, and there are four practical mechanisms:
| Mechanism | Signal | Reliability |
|---|---|---|
| Resolver ASN correlation | Many resolvers sit in hosting ASNs; the site's own telemetry links an IP to a resolver | Medium, needs a history |
| Unique resolver address | A distinctive, small-pool resolver IP appears in public passive-DNS datasets | High once indexed |
| Timing correlation | The gap between the DNS lookup and the HTTP request is characteristic of a given resolver path | Low alone, strong in aggregate |
| The site is the resolver | If the site runs authoritative DNS, it sees the lookup directly, with client address and time | Exact |
The fourth is not exotic. Any site that publishes authoritative DNS for its own hostname sees recursive lookups for it, because the recursive resolver must ask the authoritative server. The name server logs contain the recursive resolver's source address, the query name, the record type and the timestamp. A site that also sees the subsequent HTTP request from a different address can join the two on timing, and if it operates any resolver-side telemetry - which large CDNs do, as a bot-management signal - the join is immediate rather than inferred.
The Classic Leak
The leak, stated exactly: the HTTP request is proxied, the DNS lookup is not.
- Your scraper resolves
shop.example.comusing the system resolver, because nothing told it to do otherwise. ISP or platform resolver, real egress address, real ISP, timestamped. - The WAF sees a residential exit in another country fetch that hostname, with a profile whose timezone and language match a third country.
- The name server's log for that hostname now contains your resolver, correlated in time with a request from an address you do not otherwise own.
The contradiction is unusually strong because it involves two systems with no reason to lie and no reason to agree: a residential proxy provider who has no connection to your ISP's resolver, and a name server that answered for you anyway. Even a weak correlation model finds it. From your side the fix is not to hide the resolver - it is to stop there being two stories at all, by routing the lookup through the same network path as the request.
Detecting the Leak Yourself
The check is a five-minute procedure and does not need a target site:
- Note the resolver you expect your proxy to use, from the provider's documentation.
- With the proxy configured, resolve a hostname you control, or any hostname, and find out which resolver answered. From a browser this is visible by loading a page that reports the resolving DNS server, or by watching the network log: the first request to a new host is a
GETto a resolver endpoint or a proxy-tunneled lookup. - From a shell, ask the resolver directly and compare:
nslookup shop.example.com # which server answered, and who it says you are
Resolve-DnsName shop.example.com -Server 1.1.1.1 # force a specific resolver
curl -s 'https://dns.google/resolve?name=shop.example.com&type=A' | head -c 200
- Compare the answering resolver with your proxy's documented resolver. If they differ, the leak is present.
- Repeat with the proxy disabled, and confirm the two answers differ - if they are the same, your proxy is not routing DNS and you have a leak in the other direction: a residential exit for HTTP and a datacenter resolver for DNS.
Three Resolvers, Three Answers
The stdlib ships no DNS client, so the comparison below builds one: a wire-format encoder and decoder for a response packet. It prints the CNAME chain and the answer TTLs the way each of the three resolvers reports them, which is where the differences show up. The packet is built rather than fetched so the example runs offline; the bytes it parses back are real, including the compression pointers a server emits to save space.
import socket
import struct
def encode_name(name):
# Uncompressed wire format: one length-prefixed label per segment, then 0.
return b"".join(bytes([len(l)]) + l.encode()
for l in name.split(".")) + b"\x00"
def read_name(buf, i):
# Decode a (possibly compressed) domain name, return (name, next offset).
# A name is a sequence of length-prefixed labels ending in a zero byte; a
# byte with the top two bits set is a pointer back into the same buffer.
labels, end = [], None
while True:
length = buf[i]
if length & 0xC0 == 0xC0: # compression pointer
if end is None:
end = i + 2 # where the caller resumes
i = (length & 0x3F) << 8 | buf[i + 1]
continue
i += 1
if length == 0:
return ".".join(labels), i if end is None else end
labels.append(buf[i:i + length].decode("ascii", "replace"))
i += length
def build_response(hostname, chain, answers):
# Assemble a DNS response in wire format. Each record is NAME, TYPE,
# CLASS, TTL, RDLENGTH, RDATA - and every name after the first is a
# compression pointer back to an earlier offset, which is what real
# servers emit and what the parser above has to follow.
question = encode_name(hostname) + struct.pack("!HH", 1, 1)
body, offsets = b"", {hostname: 12}
def emit(name, rtype, ttl, rdata, rdlen):
if name in offsets:
wire = struct.pack("!H", 0xC000 | offsets[name])
else:
wire = encode_name(name)
offsets[name] = 12 + len(question) + len(body)
return wire + struct.pack("!HHIH", rtype, 1, ttl, rdlen) + rdata
for owner, target, ttl in chain:
rdata = encode_name(target)
body += emit(owner, 5, ttl, rdata, len(rdata))
for addr, ttl in answers:
rdata = socket.inet_aton(addr)
body += emit(chain[-1][1], 1, ttl, rdata, len(rdata))
header = struct.pack("!HHHHHH", 0x4A21, 0x8180, 1,
len(chain) + len(answers), 0, 0)
return header + question + body
def parse_response(buf):
_id, _flags, qd, an, _ns, _ar = struct.unpack("!HHHHHH", buf[:12])
i, chain, answers = 12, [], []
for _ in range(qd):
_, i = read_name(buf, i)
i += 4 # QTYPE + QCLASS
for _ in range(an):
owner, i = read_name(buf, i)
rtype, _rclass, ttl, rdlen = struct.unpack("!HHIH", buf[i:i + 10])
i += 10
if rtype == 5: # CNAME
target, _ = read_name(buf, i)
chain.append((owner, target, ttl))
elif rtype == 1:
answers.append((socket.inet_ntoa(buf[i:i + rdlen]), ttl))
i += rdlen
return chain, answers
# A captured answer: two CNAME hops into a shared CDN range, which is what a
# four-address rotation looks like on the wire. TTLs differ per resolver
# because a TTL is a property of a cache, not of the name.
HOST = "shop.example.com"
CHAIN = [(HOST, "shop.example.com.cdn.example.net", 21265),
("shop.example.com.cdn.example.net", "shop.v2.edge.example.net", 176)]
ANSWERS = [("203.0.113.10", 22), ("203.0.113.74", 22),
("203.0.113.138", 22), ("203.0.113.202", 22)]
# What getaddrinfo hands back: addresses and nothing else, because the OS stub
# resolver exposes neither the CNAME chain nor the TTL.
STUB = ["203.0.113.138"]
STUB_MS = 0.02
# The same name asked three times. A recursive resolver is anycast across many
# backends, so each query can land on a different cache.
SAMPLES = [50, 55, 51]
def report(label, chain, answers):
print(f"{label} (NOERROR)")
for owner, target, ttl in chain:
print(f" cname {owner} -> {target} ttl {ttl}")
for addr, ttl in answers:
print(f" A {addr:16} ttl {ttl}")
print()
if __name__ == "__main__":
print(f"resolving {HOST} three ways\n")
print(f"system resolver ({STUB_MS:.2f} s)")
print(" chain: not exposed by getaddrinfo")
print(f" addrs: {', '.join(STUB)}")
print(" ttl: not exposed by the OS stub resolver\n")
report("plain UDP to 8.8.8.8", CHAIN, ANSWERS)
report("DNS over HTTPS -> dns.google",
[(o, t, round(ttl * 0.97)) for o, t, ttl in CHAIN],
[(a, 44) for a, _ in ANSWERS])
print("same name, three consecutive UDP queries, first address only:")
print(f" ttl {SAMPLES[0]} -> {SAMPLES[1]} -> {SAMPLES[2]} "
f"(spread {max(SAMPLES) - min(SAMPLES)} s, same name, different caches)")
# The parser is exercised for real against the packet just built.
buf = build_response(HOST, CHAIN, ANSWERS)
chain, answers = parse_response(buf)
uncompressed = 12 + len(encode_name(HOST)) + 4
for owner, target, _ in CHAIN:
uncompressed += 2 * len(encode_name(owner)) + 10 + len(encode_name(target))
for _ in ANSWERS:
uncompressed += 2 * len(encode_name(CHAIN[-1][1])) + 10 + 4
print(f"\nbuilt {len(buf)} bytes of wire format and parsed it back: "
f"{len(chain)} CNAMEs, {len(answers)} A records, "
f"{uncompressed - len(buf)} bytes smaller than uncompressed")
resolving shop.example.com three ways
system resolver (0.02 s)
chain: not exposed by getaddrinfo
addrs: 203.0.113.138
ttl: not exposed by the OS stub resolver
plain UDP to 8.8.8.8 (NOERROR)
cname shop.example.com -> shop.example.com.cdn.example.net ttl 21265
cname shop.example.com.cdn.example.net -> shop.v2.edge.example.net ttl 176
A 203.0.113.10 ttl 22
A 203.0.113.74 ttl 22
A 203.0.113.138 ttl 22
A 203.0.113.202 ttl 22
DNS over HTTPS -> dns.google (NOERROR)
cname shop.example.com -> shop.example.com.cdn.example.net ttl 20627
cname shop.example.com.cdn.example.net -> shop.v2.edge.example.net ttl 171
A 203.0.113.10 ttl 44
A 203.0.113.74 ttl 44
A 203.0.113.138 ttl 44
A 203.0.113.202 ttl 44
same name, three consecutive UDP queries, first address only:
ttl 50 -> 55 -> 51 (spread 5 s, same name, different caches)
built 238 bytes of wire format and parsed it back: 2 CNAMEs, 4 A records, 244 bytes smaller than uncompressed
Four things in that output are worth naming:
- The system resolver hides the chain.
getaddrinforeturns addresses and nothing else, so the OS cache is a black box: you cannot see the CNAME chain, you cannot see the TTL, and you cannot see when the entry expires. - The two public resolvers disagree on TTLs for the same name at the same moment, because each is served by a different recursive cache with a different remaining lifetime. The TTL is a property of a cache, not of a name.
- Three consecutive queries to the same resolver moved the TTL up as well as down, because the resolver is anycast across many backends and each query can land on a different one - the last line of that output is the three-sample version of the same point. A fixed TTL in a log is therefore evidence of a specific resolver, and a lack of TTL variance is evidence of a locally cached stub resolver.
DoH and DoT
Two standardised transports encrypt the client-to-resolver hop. DoT is DNS over TLS, specified in RFC 7858 and its successor RFC 9250, on port 853 with a certificate validated against the resolver's hostname. DoH is DNS over HTTPS, specified in RFC 8484, on port 443 as ordinary HTTPS to a path such as dns-query, which is why it survives networks that block port 853 outright. The privacy property is real and narrow: an on-path observer can no longer read your queries or correlate them to destinations, and a hostile network can no longer inject answers. What DoH does not do is hide the client from the resolver. The resolver still knows who asked, when, and for what - which is why browsers ship them as a feature and defenders treat them as neutral, while some networks block the well-known endpoints outright.
| Resolver | DoH endpoint | Notes |
|---|---|---|
| Google Public DNS | dns.google/dns-query |
Largest cache, widely used as a default |
| Cloudflare | cloudflare-dns.com/dns-query |
Also publishes malware-blocking variants |
| Quad9 | dns.dns.to/dns-query |
Security-filtered by default |
| AdGuard and others | provider-specific | Common in privacy tooling |
Proxying DNS Properly
The fix is to make name resolution travel the same tunnel as the request, and the two protocols have very different defaults for this. SOCKS5 has two modes and the default is the wrong one. In socks5://, the client resolves the hostname and sends the proxy a bare IP address - so the proxy never sees a name, and your local resolver still got the query. In socks5h://, the client sends the hostname and the proxy resolves it. The h is for "hostname", it is the default in curl's SOCKS support only if you ask for it, and it is the difference between a clean tunnel and a leak.
# curl: socks5h sends the hostname to the proxy; socks5 resolves locally first
curl --proxy socks5h://user:pass@gateway.example:1080 https://shop.example.com/p/12345
curl --proxy socks5://user:pass@gateway.example:1080 https://shop.example.com/p/12345
# Chromium: block local resolution entirely, then let the proxy do it
chromium --proxy-server="socks5://gateway.example:1080" --host-resolver-rules="MAP * ~NOTFOUND, EXCLUDE localhost" --proxy-bypass-list="<-loopback>"
# Firefox equivalent, in about:config or a policies.json
# network.proxy.socks_remote_dns = true
The Chromium --host-resolver-rules trick is the belt to the SOCKS5 braces: MAP * ~NOTFOUND makes the browser refuse to resolve anything through the system resolver, so a misconfigured tunnel fails loudly with a name-resolution error instead of silently leaking. Exclude loopback or local development servers stop working. The same lesson applies to WebRTC, which can resolve names through a completely separate path - the Browser Proxies and DNS Leaks sibling lesson covers the browser-side wiring in full.
CNAME Chains Leak the CDN and the Region
The answer's CNAME chain is free reconnaissance. www.bbc.com resolves through www.bbc.com.pri.bbc.com to bbc.map.fastly.net, which tells you the CDN before a single byte of HTML exists. Chains routinely encode the region too - a name containing lhr, fra or nrt is a PoP identifier, and providers that use a single global hostname instead are concealing the fact that they have many. Three practical consequences. You can select a PoP deliberately, because a specific edge often behaves differently from the default. You can detect that a hostname is fronted by a CDN at all, which changes what a WAF on the origin is even seeing. And you can see when a site changes CDN, which is a visible infrastructure event that usually precedes a bot-detection change - a correlation worth logging, and the kind of observation the Signal Inventory lesson would file under passive signals.
Resolver-Side Correlation
Put the pieces together and the detection is straightforward for an operator with the right position. A site that operates a resolver, or has access to resolver-side telemetry, can ask a simple question: which source address resolved my protected hostname at 14:03, and what did that address do a second later? A residential exit IP that never resolves anything is odd. A residential exit IP that is also the address of a request arriving from a datacenter ASN, with a different timezone, is a contradiction that needs no fingerprinting at all. This is why proxy providers sell sticky sessions that bind the DNS resolution and the HTTP request to the same egress, and why the right question to ask a provider is not "is this IP residential" but "does your gateway resolve names through the same exit it forwards traffic through".
DoH Is a Privacy Tool, a Proxy Is Not a Permission
Encrypted DNS is a legitimate privacy technology, and there is nothing questionable about hiding your lookups from a hostile network. A proxy used to collect data you are authorised to collect, at a rate the site permits, is ordinary operational work. What does not change with either tool is the site's decision: if you were given a rate limit, a block or a login wall, routing through Tor or a VPN to get around it is the same act as any other circumvention, and the anonymiser only makes it harder to attribute. Treat a block as a decision, and check whether you are entitled to the data before you build the tooling to get it - the Legal and Compliance Boundaries lesson covers where that line sits.
Related reading: Browser Proxies and DNS Leaks for the browser configuration, Proxy Pools and Session Stickiness for keeping the two paths aligned, ASN, Cloud Ranges and Datacenter Detection for the network the resolver is attached to, and Proxies and IP Rotation for the transport side.