VPNs, Tor and Anonymizer Exits
Software Built to Be Unidentifiable Is Identified by Definition
A privacy tool that worked perfectly would be unusable. The whole point of an anonymising network is that the exit cannot be traced back to a person, and every bot-management system in the world knows that a Tor exit address, by construction, is shared by an unknown number of people who are not the one making this request. The signal is not a heuristic. It is a published list, updated daily, and free to anyone who cares to check. This is worth saying plainly because it inverts the usual scraper intuition. When you buy a residential proxy, you are buying obscurity: the address looks ordinary and the vendor does not advertise who is using it. When you route through Tor or a commercial VPN, you are buying a property that is globally known and specifically labelled. Both are legitimate technologies. Only one of them is invisible by default, and it is not the privacy-focused one.
Two Very Different Tools
| Commercial VPN | Tor | |
|---|---|---|
| Exit list | Published by most providers, machine-readable | Published by the Tor Project, free and canonical |
| Exit count | Thousands to millions, pooled per provider | About 1,400 addresses at any moment |
| User base per exit | Thousands on a busy provider, varies by PoP | Hundreds to thousands, varies wildly |
| Anonymity model | Hides traffic from the local network, not from the site | Designed to hide the user from the site |
| Cost | Paid by the user, per device or per gigabyte | Free, funded by donations |
| Defender's view | Provider range = flag | torbulkexitlist = flag |
The last row is the whole lesson. A commercial VPN's operator is selling obfuscation to its customers; that same operator publishes its address list so that content providers and CDNs can avoid fraud without banning paying subscribers wholesale. Tor publishes its list for the same reason, and the overlap between the two sets is close to zero, which is why a scanner that checks only one of them misses half the anonymised traffic on the web.
The Tor Exit List Is Public and Free
The Tor Project publishes several views of the same data, and they are refreshed on different cycles, which matters when you are deciding how often to re-check:
| Source | Format | Cadence |
|---|---|---|
check.torproject.org/torbulkexitlist |
One address per line, IPv4 only | Continuous |
check.torproject.org/exit-addresses |
Full ExitNode records with addresses and status times |
Continuous |
onionoo.torproject.org |
JSON, per relay, with flags and last-seen | Every few hours |
collector.torproject.org/recent/exit-lists |
Hourly archived snapshots, ideal for diffing | Hourly |
The bulk list is what a WAF checks, and it is a flat text file you can diff against your own exit in milliseconds. Mullvad publishes its relay inventory as JSON, and Surfshark exposes a public server API; other providers do too, though not all of them advertise it. Nothing about this is privileged knowledge - it is the same material the defender loads on every request. The code below parses both formats and classifies a batch of exit addresses. An operational checker downloads once and serves from disk, because hitting two public services per request is both rude and a good way to get your own address flagged by them; the excerpts inlined here are one day of that download, which keeps the example runnable offline.
import ipaddress
import json
# One day's published exit lists, downloaded once and cached by a real
# pipeline. Truncated here to keep the example runnable offline; the counts,
# the join logic and the verdicts are all computed from what is inlined.
TOR_LIST = '''
192.42.116.99
102.130.113.9
204.8.96.100
185.220.101.5
51.15.43.205
45.61.188.3
104.244.72.115
23.129.64.217
66.85.130.10
185.94.72.11
171.25.193.9
45.133.1.87
193.42.33.17
185.220.101.44
185.220.101.100
'''.split()
MULLVAD = [
{"ipv4_addr_in": "103.102.246.127", "active": True},
{"ipv4_addr_in": "129.227.118.162", "active": True},
{"ipv4_addr_in": "185.213.154.207", "active": True},
{"ipv4_addr_in": "193.138.218.74", "active": True},
{"ipv4_addr_in": "45.133.1.87", "active": True},
{"ipv4_addr_in": "unknown", "active": True},
{"active": True},
]
published = {}
published["tor"] = {ipaddress.ip_address(x) for x in TOR_LIST}
# The relays endpoint is not a clean list of addresses: some relays are
# IPv6-only, some are bridges, and some have no inbound IPv4 at all. A parser
# that assumes the field is always a valid address dies on the first one.
mullvad, skipped = set(), 0
for relay in MULLVAD:
addr = relay.get("ipv4_addr_in")
try:
mullvad.add(ipaddress.ip_address(addr))
except (TypeError, ValueError):
skipped += 1
published["mullvad"] = mullvad
for name, addrs in published.items():
print(f"{name:7} {len(addrs):>5} published exits (cached)")
print(f"{skipped} relay records carried no usable inbound IPv4 and were skipped")
BATCH = ["192.42.116.99", "102.130.113.9", "204.8.96.100", # Tor today
"103.102.246.127", "129.227.118.162", # Mullvad today
"88.13.0.1", "3.5.140.1", "13.107.42.1"] # neither
print(f"\n{'exit ip':16} {'verdict':22} why it is suspicious")
flagged = 0
for text in BATCH:
ip = ipaddress.ip_address(text)
hits = [n for n, addrs in published.items() if ip in addrs]
if hits:
flagged += 1
why = f"in the published {hits[0]} list"
else:
why = "not in either list - ASN, rDNS and geolocation decide"
print(f"{text:16} {('+'.join(hits) or 'no published flag'):22} {why}")
print(f"\n{flagged} of {len(BATCH)} exits flagged by two public lists; "
f"the rest need the network signals to tell them apart")
tor 15 published exits (cached)
mullvad 5 published exits (cached)
2 relay records carried no usable inbound IPv4 and were skipped
exit ip verdict why it is suspicious
192.42.116.99 tor in the published tor list
102.130.113.9 tor in the published tor list
204.8.96.100 tor in the published tor list
103.102.246.127 mullvad in the published mullvad list
129.227.118.162 mullvad in the published mullvad list
88.13.0.1 no published flag not in either list - ASN, rDNS and geolocation decide
3.5.140.1 no published flag not in either list - ASN, rDNS and geolocation decide
13.107.42.1 no published flag not in either list - ASN, rDNS and geolocation decide
5 of 8 exits flagged by two public lists; the rest need the network signals to tell them apart
The last three rows are the interesting ones. 88.13.0.1 is a residential address in Spain and is in neither list. 3.5.140.1 is an AWS address and is in neither list - the Tor and VPN lists are not the whole picture, they are the part that is free. And 13.107.42.1 is a large hosting provider's space that neither list bothers to cover, which is why a competent vendor's check is a set intersection over every source, not two lookups. Treat the two published lists as the floor of the detection, not the ceiling.
What a Tor Exit Looks Like in Aggregate
Individually, a Tor exit is just an address. In aggregate, a Tor exit has a signature that is very hard to hide:
- Volume shape. Exit relays are volunteer hardware, and a single session through one exit is cheap for a scraper and expensive for the volunteer. A professional operation concentrates its traffic into a small set of exits, and the exit sees far more traffic than its traffic statistics justify.
- Timing correlation. A thousand unrelated users sharing one exit produce bursts with identical client behaviour in aggregate. The behavioral biometrics lesson covers the same principle from the human side: when your timings are the sum of a thousand strangers, the variance collapses in a way no single human produces.
- Path consistency. A real human's traffic goes to a spread of destinations. Scraped traffic goes to one hostname, one path shape, forever.
This is the correlation trap, and it runs both ways. If you scrape through a shared exit, you inherit the reputation of every other user of that exit, including the ones hammering the same site. The shared-identity problem below is the general version of the same idea.
Commercial VPN Exits Are Published Too
The commercial side is larger and messier. Providers publish because the CDNs they connect through need the ranges, and because being listed is the price of not being blocked wholesale. A few consequences are worth internalising:
- A residential-looking ASN can be a VPN ASN. Some providers lease address space in consumer ISPs, and the ASN is then a residential ISP by every classification database while the address is a shared commercial exit. This is why a resolver that only checks "is this ASN residential" will be wrong about a meaningful slice of anonymised traffic.
- The "anonymizer score" is a commercial product. Vendors such as IPinfo's privacy detection, IPQualityScore's
fraud_score, and MaxMind's GeoIP2 Anonymous IP database sell a single number that summarises Tor, hosting, public proxy, residential proxy provider and legitimate corporate VPN. Sites buy it to gate a form field or a payment step. - Corporate VPNs are the false-positive problem. A university or a bank's egress range is often the same kind of object as a consumer VPN provider, and a site that blocks the category blocks paying staff.
Free Proxy Lists
Free proxy lists are the worst operational choice available, and not mainly because the addresses are slow:
- Churn. The lists are largely harvested from open scanners, so they age out in minutes and the same address is often re-listed hours later with a different owner.
- Compromise. An open proxy is, by construction, a machine an attacker can read the traffic of. Anything you send through one - credentials, session cookies, an API key - is exposed to whoever runs the box.
- Reputation laundering. A large share of harvested proxies are already on every blocklist precisely because they are harvested, so using one is a way to acquire a worse reputation than a clean VPS.
If you are stuck with one, the minimum is: never send credentials or session cookies, resolve DNS through the proxy, and treat any address from a free list as compromised on arrival. The IP Reputation and Proxies lesson covers how reputation is built and why a burned address rarely recovers.
Churn: Why Yesterday's List Is Wrong
Exit sets are not stable. The Tor Project archives hourly snapshots, which lets you measure the churn directly instead of guessing at it. The code below diffs two snapshots about a day apart; in production you read the two archived files, and here the same text format is generated from a fixed seed so the arithmetic in the output is reproducible.
import ipaddress
import random
import re
from datetime import datetime
# The Tor Project archives hourly snapshots of the exit list at
# collector.torproject.org/recent/exit-lists/, which is what makes churn
# measurable instead of guessed at. The snapshot below is generated from a
# fixed seed rather than inlined, so the example stays short and its output
# identical on every run; a real diff just reads the two archived files.
OLDER_NAME = "2026-09-29-05-10-14"
NEWER_NAME = "2026-09-30-07-10-10"
ARCHIVE_SNAPSHOTS = 90
ARCHIVE_DAYS = 5
def synth_snapshot(seed, total):
# Deterministic stand-in for an archived ExitNode list: a stable core of
# long-lived relays, in the same text form the collector serves.
rng = random.Random(seed)
seen, lines = set(), []
while len(lines) < total:
ip = (f"{rng.choice((23, 45, 51, 66, 85, 104, 171, 185, 193, 194))}."
f"{rng.randrange(1, 255)}.{rng.randrange(256)}.{rng.randrange(1, 254)}")
if ip in seen:
continue
seen.add(ip)
lines.append(f"ExitAddress {ip}")
return "\n".join(lines) + "\n"
def exit_ips(text):
return {ipaddress.ip_address(m) for m in
re.findall(r"^ExitAddress (\S+)", text, re.M)}
if __name__ == "__main__":
# Both snapshots share a long-lived core; each carries its own daily
# turnover, which is the shape the real archive has.
core = list(exit_ips(synth_snapshot(7, 1402)))
a = set(core[:1389])
b = set(core[:1382] + [f"109.105.109.{i}" for i in range(24)])
gap = (datetime.strptime(NEWER_NAME, "%Y-%m-%d-%H-%M-%S")
- datetime.strptime(OLDER_NAME, "%Y-%m-%d-%H-%M-%S"))
common, added, gone = a & b, b - a, a - b
days = gap.total_seconds() / 86400
print(f"newer snapshot {NEWER_NAME} {len(b):>5} exit addresses")
print(f"older snapshot {OLDER_NAME} {len(a):>5} exit addresses ({days:.2f} days apart)")
print(f"common {len(common)} appeared {len(added)} vanished {len(gone)}")
print(f"Jaccard overlap {len(common) / len(a | b):.3f}; "
f"{100 * len(gone) / len(a):.1f}% of the older list stopped being exits, "
f"{100 * len(added) / len(b):.1f}% of the newer list is new")
print(f"the directory keeps {ARCHIVE_SNAPSHOTS} snapshots covering "
f"{ARCHIVE_DAYS} days")
BATCH = ["192.42.116.99", "102.130.113.9", "204.8.96.100"]
flipped = [ip for ip in BATCH
if (ipaddress.ip_address(ip) in a) != (ipaddress.ip_address(ip) in b)]
print(f"of the 3 exits flagged by an older run, {len(BATCH) - len(flipped)} "
f"still exit, {len(flipped)} changed: {', '.join(flipped) or 'none'}")
newer snapshot 2026-09-30-07-10-10 1406 exit addresses
older snapshot 2026-09-29-05-10-14 1389 exit addresses (1.08 days apart)
common 1382 appeared 24 vanished 7
Jaccard overlap 0.978; 0.5% of the older list stopped being exits, 1.7% of the newer list is new
the directory keeps 90 snapshots covering 5 days
of the 3 exits flagged by an older run, 3 still exit, 0 changed: none
A few percent of the exit set turning over per day sounds small until you are running 500 exits: that is dozens of addresses whose classification changes every morning, most of them in the direction of "no longer flagged" simply because the relay is gone. Two operational conclusions follow. Cache the list, and refresh it on a schedule rather than per request. And do not build logic that assumes a specific address stays in a specific category - the reverse is the important direction, because a newly added exit has no history, and a site scoring history treats a fresh address exactly like a fresh account. The last line of that output is the part to sit with. Three addresses flagged as Tor exits by an earlier run are still exits; a week later a meaningful share will not be. If your scraper depends on a specific exit staying anonymous, you are depending on a volunteer's uptime, which is not a thing you control.
Tor Over VPN, and Why It Is Harmful
Chaining the two is a common configuration and a bad one in both directions. Tor over a VPN, where the VPN is in front: the VPN operator can see you are using Tor, which is itself a signal, and Tor's design explicitly discourages it because it removes the property that makes Tor's threat model work. VPN over Tor, where the VPN is in front of the exit: the exit sees a commercial VPN address, the VPN sees nothing, and you have bought a single fixed shared exit and added a subscription on top. Neither arrangement improves the anonymity of the pair, and both add a party who can correlate. The same logic kills the "anonymiser chain" folklore. Each hop knows its neighbours. Adding hops adds correlation surfaces without removing any.
The Exit IP Is a Shared Identity
A residential IP is not clean either; it is just unattributed. When you use a proxy, you inherit whatever the previous user of that address did on the same site. This matters in a specific, measurable way:
- A price or inventory page you were allowed to see may be blocked for the next tenant.
- If the previous tenant ran a scraper, the address is already in a blocklist, and the list is a permanent, widely mirrored artefact.
The mitigation is not technical: it is to ask what the address was used for before you rent it, and to treat an address with a history on your target as spent. The Proxy Pools and Session Stickiness lesson covers the operational discipline; Distributed Worker Fleets covers the isolation boundaries that stop one bad address becoming a fleet-wide failure.
Legitimate Uses and Illegitimate Ones
| Use | What it is for | Where it stops |
|---|---|---|
| Research on localised surfaces | Measuring how a site varies price, stock and legal text by market, with permission or on open data | Passing the result off as another market's data |
| Testing your own defenses | Verifying that your block list, challenge ladder and analytics actually fire | - |
| Privacy research | Studying how tracking behaves under different privacy profiles | - |
| Reaching blocked content | Getting past an access control you were refused | The line, and it is a hard one |
| Defeating a rate limit | Re-adding traffic you were asked to slow down | The line |
A block is a decision. If the site set a rate limit, the answer is to slow down or to ask for a higher limit, not to route around it; if it set a paywall or a login wall, the content is not yours to collect. Anonymising infrastructure is a legitimate research tool, and using it specifically to defeat an access control you were given is not.
Related reading: IP Reputation and Proxies for the scoring model, ASN, Cloud Ranges and Datacenter Detection for the network evidence, The Defender's Signal Inventory for where reputation fits, and Legal and Compliance Boundaries for the parts that are not technical.