ASN, Cloud Ranges & Datacenter Detection
Every Server Has a Postcode Address
The first question a bot-management system asks about a connection is not about HTTP at all. It arrives at the edge with a source address, looks that address up in a table, and learns which organisation operates the network it came from. If the answer is a cloud provider, a hosting company or a VPN operator, the session starts life with a handicap no amount of header polishing will remove.
This is the cheapest signal in the entire stack to compute and one of the hardest to escape, because the answer lives in the routing system rather than in anything you control. The data is public, machine-readable, refreshed daily, and the exact same files the vendor uses are one curl away from you.
What an ASN Is
The internet is a graph of independently operated networks, and each one is assigned an Autonomous System Number (ASN) - a 32-bit integer, so 32-bit ASNs like 16509 (AWS) and 15169 (Google) coexist with 4-byte ASNs from 4200000000 upward. The announcements that bind an ASN to a set of IP prefixes travel between routers using BGP, and the authoritative record of who owns what is spread across the five Regional Internet Registries: ARIN, RIPE NCC, APNIC, LACNIC and AFRINIC.
Two artefacts matter for this work. Delegation files are the RIRs' machine-readable statements of who has been allocated which ranges - a daily delegated-*-extended-latest file per registry, and a corresponding RDAP query interface for live lookups. Route collector data comes from projects that passively observe the global BGP table, such as RouteViews and RIPE RIS, and it tells you which prefixes an ASN is actually announcing, which is a different and more useful question than who was allocated them. A reseller announces customer space under its own ASN; the delegation file and the routing table disagree in ways that are informative.
| Question | Source of the answer |
|---|---|
| Who was allocated this range? | RIR delegation file, RDAP, WHOIS |
| Who is announcing it right now? | BGP route collectors, reverse DNS on the router interface |
| What kind of company is that? | Commercial classification DB, your own ASN notes |
| How bad is its reputation? | The vendor's own history, which nobody publishes |
The Feeds Vendors Buy
Commercial vendors do not classify ASNs by hand at query time. They buy a GeoIP or IP intelligence database that is regenerated nightly, join it to a second table of ASN categories and per-ASN abuse history, and serve the result as a single lookup with a few hundred bytes of JSON. MaxMind, IP2Location, IPinfo, DB-IP and ipapi all sell some form of this, at prices spanning three orders of magnitude, and the differences between them are mostly in the parts they refresh nightly and how they handle mobile and IPv6.
Two biases matter more than the price. Population bias: an address is attributed to the registered location of the block, or to the location where the block was announced, and for dynamic or reassigned space that is often the operator's headquarters rather than the subscriber. Vendor bias: a database that is popular with the defender community optimises its own categories for the categories the defender community asks about, which means hosting and residential are well separated and everything in between is lumped together.
The Published Cloud Range Lists
The larger the provider, the more likely it publishes its own address list, and publishing is a hard requirement for joining most CDNs because the CDN's own fraud screening needs the list. These files are the reason this signal is not a secret.
| Provider | File | Format | Cadence |
|---|---|---|---|
| AWS | ip-ranges.json |
JSON, ip_prefix plus region and service |
Continuous, with a createDate |
| Google Cloud | cloud.json |
JSON, ipv4Prefix plus scope and service |
Daily, creationTime |
| Google front-end | goog.json |
Same schema, Google's own edge space | Daily |
| Azure | ServiceTags_Public_*.json |
JSON, arrays of addressPrefixes |
Weekly, needs a free Azure account |
| Cloudflare | ips-v4, ips-v6 |
One CIDR per line | On change |
| Oracle Cloud | publicip.json |
JSON, addressPrefixes per region |
On change |
| Hetzner | No official file | Read the RIPE delegated file or a route dump | - |
The key property of these lists is that they are complete. AWS publishes every address it routes, so the defender's lookup is a binary search through ten thousand prefixes and the answer is exact - not a guess, not a heuristic. For a scraper, that makes the signal entirely self-service: download the same file, classify your own exits, and you know exactly what the WAF knows about you.
The code below does the whole job with the standard library: parse, flatten, then classify a batch of exit addresses. A production version fetches these URLs once a day and caches them to disk so the network is never touched per lookup; the excerpts inlined here are a single day's download, which keeps the example runnable offline and its output identical on every run. The reverse DNS column is the bonus signal.
import ipaddress
import json
# Excerpts of the three published files, copied verbatim from a single day's
# download. Real deployments fetch these URLs once and cache the result; the
# excerpts keep this example runnable offline and its output reproducible.
AWS = {"createDate": "2026-09-30-04-57-06", "prefixes": [
{"ip_prefix": "3.5.140.0/22", "region": "ap-northeast-2", "service": "EC2"},
{"ip_prefix": "13.107.42.0/24", "region": "Global", "service": "WAF"},
]}
GCP = {"creationTime": "2026-09-30T03:45:18Z", "prefixes": [
{"ipv4Prefix": "34.80.0.0/15", "scope": "asia-east1", "service": "Cloud"},
]}
CLOUDFLARE = "1.0.0.0/24\n103.21.244.0/22\n104.16.0.0/13\n131.0.72.0/22\n"
def load():
# Every list reduced to one flat list of (network, source, detail).
nets, provenance = [], []
docs = [("aws", AWS, "json"), ("gcp", GCP, "json"),
("cloudflare", CLOUDFLARE, "lines")]
for name, doc, kind in docs:
if kind == "lines":
rows = [(line, "edge/anycast") for line in doc.split()]
published = "-"
else:
published = doc.get("createDate") or doc.get("creationTime", "")[:19]
rows = []
for p in doc["prefixes"]:
cidr = p.get("ip_prefix") or p.get("ipv4Prefix") # AWS / GCP
if cidr is None:
continue # IPv6 row
where = p.get("region") or p.get("scope") # region / scope
rows.append((cidr, f'{where}/{p["service"]}'))
for cidr, detail in rows:
nets.append((ipaddress.ip_network(cidr), name, detail))
provenance.append(f"{name:11} {len(rows):>6,} prefixes {published}")
return nets, provenance
def classify(addr, nets):
# First match wins, in file order - exactly what a naive vendor does.
ip = ipaddress.ip_address(addr)
for net, source, detail in nets:
if ip.version == net.version and ip in net:
return source, str(net), detail
return "unknown", "-", "-"
PTR = { # what reverse DNS returned for each address, sampled once
"3.5.140.1": None,
"34.80.0.1": "1.0.80.34.bc.googleusercontent.com",
"13.107.42.14": None,
"1.1.1.1": "one.one.one.one",
"104.16.123.96": None,
"88.13.0.1": "1.red-88-13-0.dynamicip.rima-tde.net",
"73.162.0.1": "c-73-162-0-1.hsd1.co.comcast.net",
}
KEYWORDS = ("cloud", "ec2", "googleusercontent", "akamaitechnologies",
"kunden", "dyn", "comcast")
def rdns_keyword(addr):
# rDNS is a network round trip per uncached address, so a real pipeline
# caches it hard and a miss looks identical to a hit in a log. Fetch once,
# keyword-scan the result, never decide on it alone.
name = PTR.get(addr)
if name is None:
return "(none)"
hits = [k for k in KEYWORDS if k in name]
return f"{name} [{','.join(hits)}]" if hits else name
if __name__ == "__main__":
nets, provenance = load()
print("\n".join(provenance))
print(f"{len(nets):,} IPv4 prefixes in memory\n")
EXITS = ["3.5.140.1", "34.80.0.1", "13.107.42.14", "1.1.1.1",
"104.16.123.96", "88.13.0.1", "73.162.0.1"]
print(f"{'exit ip':16} {'verdict':10} {'matched prefix':18} {'detail':20} rDNS")
for addr in EXITS:
source, net, detail = classify(addr, nets)
print(f"{addr:16} {source:10} {net:18} {detail:20} {rdns_keyword(addr)}")
aws 2 prefixes 2026-09-30-04-57-06
gcp 1 prefixes 2026-09-30T03:45:18
cloudflare 4 prefixes -
7 IPv4 prefixes in memory
exit ip verdict matched prefix detail rDNS
3.5.140.1 aws 3.5.140.0/22 ap-northeast-2/EC2 (none)
34.80.0.1 gcp 34.80.0.0/15 asia-east1/Cloud 1.0.80.34.bc.googleusercontent.com [googleusercontent]
13.107.42.14 aws 13.107.42.0/24 Global/WAF (none)
1.1.1.1 unknown - - one.one.one.one
104.16.123.96 cloudflare 104.16.0.0/13 edge/anycast (none)
88.13.0.1 unknown - - 1.red-88-13-0.dynamicip.rima-tde.net [dyn]
73.162.0.1 unknown - - c-73-162-0-1.hsd1.co.comcast.net [comcast]
Three things in that output are worth pausing on. 1.1.1.1 is Cloudflare's resolver and is not in Cloudflare's published CDN ranges, because a resolver is not an edge - "resolves from Cloudflare" and "is a Cloudflare server" are different claims, and a vendor that conflates them produces false positives on a very large population. 3.5.140.1 is AWS and has no PTR record at all, which is a reminder that the absence of reverse DNS is not evidence of residential. And rDNS is not free: a real pipeline pays a network round trip per uncached address, and a miss that takes several seconds is paid per address rather than per session. A vendor that reverse-resolves every request has learned to cache aggressively, which is exactly why the miss looks identical to a hit in a log. The bracketed keyword column is the part worth copying, because it extracts the operator name from the PTR record instead of depending on the name appearing in a particular shape.
rDNS and the PTR Record
Reverse DNS maps an address back to a name. It is the oldest, crudest signal in the family, and it is still worth checking because the names themselves leak the operator.
- A cloud instance often resolves to something like
ec2-3-5-140-1.ap-northeast-2.compute.amazonaws.comor1.0.80.34.bc.googleusercontent.com, which states the provider and often the region in plain text. - A residential address often resolves to the customer's router:
1.red-88-13-0.dynamicip.rima-tde.net, whererima-tde.netidentifies a Spanish residential ISP. - A CDN edge often has a generic name (
one.one.one.one,a23-1-2-3-4.deploy.akamaitechnologies.net) or none at all.
Use it as a weak signal and know its three failure modes. False negatives: plenty of hosting providers publish no PTR, and RFC 1918 space is not reverse-published at all, so a miss is common among legitimate servers. False positives: plenty of residential ISPs delegate PTR to the customer's own DNS, so a no-reverse.example.com style record says nothing. Latency: it is a network round trip per uncached address, and a batch of a thousand looks nothing like a batch of ten from a timing model. The right use is a keyword scan over the name - does it contain cloud, ec2, googleusercontent, akamaitechnologies, kunden, dyn - combined with the CIDR check, never as a decision on its own.
The Categories Vendors Maintain
| Category | Typical ASNs | Default posture | How it is usually built |
|---|---|---|---|
| Hosting / datacenter | AWS, GCP, Azure, Hetzner, OVH, DigitalOcean | Hostile by default | Published CIDR list plus manual ASN notes |
| Business / enterprise | corporate leased lines, some university ranges | Neutral | Allocation plus geofeed |
| Education | university and research ASNs | Neutral to hostile | RIR allocation |
| Mobile carrier | cellular ASNs | Trusted | Allocation plus carrier confirmation |
| Residential ISP | consumer ISP ASNs | Trusted | Allocation, sometimes split by prefix type |
| CDN / edge | Cloudflare, Akamai, Fastly | Neutral, and often not meaningful | Self-published list |
| Anonymizer | VPN, Tor, open proxy operators | Hostile | Published lists plus observed behaviour |
The important part is that "residential" is not one bucket. Residential space includes a long tail of dynamic reassignment, a small amount of space that hosting companies lease in bulk, and a great deal of space behind carrier-grade NAT. A subcategory that separates genuinely residential access from resold dynamic space removes a surprising amount of the false-positive rate, and a subcategory that recognises mobile carrier space is what stops a WAF from blocking five thousand real phones when it bans one CGNAT address.
Data Center, VPN, Proxy, Anonymizer
These four words are used interchangeably and mean different things with different consequences.
| Label | What it actually is | How it is identified | Blast radius of a ban |
|---|---|---|---|
| Data center | An address the provider knows it routes to servers | Published CIDR list, BGP, rDNS | Your whole fleet, if it is one cloud |
| VPN | A shared exit operated by a commercial provider | Provider list, or ASN plus port/protocol behaviour | The provider's whole pool |
| Proxy | An open relay forwarding traffic for anyone | Mass-scanning results, abuse reports | The one address, plus its neighbourhood |
| Anonymizer | A commercial score summarising all of the above | A vendor flag, e.g. an anonymous-IP database's categories | Whatever the score gates |
The reason this distinction is worth money is blast radius. Banning one open proxy is cheap and effective. Banning a /24 costs you the twenty-eight other customers behind it. Banning an ASN costs you everyone the provider has, and if that provider is a cloud you also rent, it can cost you your compute. That asymmetry is why vendors escalate to softer responses first - a challenge, a rate limit, a degraded response - and reserve address-level bans for the cases where the address is the problem.
Rotating Inside a Single ASN
Rotation inside one provider is worse than a fixed IP in a clean ASN, and the reason is granularity. Reputation is scored per IP, per prefix and per ASN, and a fleet that rotates produces thousands of distinct addresses inside a narrow block, all behaving identically and all arriving within a short window. From the edge that is a textbook coordinated pattern; from the reputation system it is a single noisy host with a suspiciously diverse address list.
The buyer-side rule follows directly: a provider whose pool is 90 percent one ASN is one incident away from a fleet-wide ban. Diversify by ASN first and by city second, because ASN diversity is what survives a ban and city diversity only helps if the ASNs are also different. This is the same failure mode described in IP Reputation and Proxies from the pool-buying side, and it is worth writing the ASN mix down as a procurement requirement rather than discovering it during an incident. The Proxies and IP Rotation lesson covers the transport mechanics once the network decision is made.
Staying Inside the Line
These lists are the site owner's primary tool and they are published in the open precisely so that infrastructure can be configured correctly - that is the stated purpose of most of them. Using them for authorised testing, or to check your own fleet's exposure before pointing a scraper at a third party, is straightforward. Using them to pick an exit that a target's terms forbid is the same act as using any other means to ignore an access control. Where a target has published a rate limit, an API or a crawl directive, that is the constraint to work within, and Official APIs and Alternative Data is where to look for a sanctioned route before you build one.
Related reading: IP Reputation and Proxies for the scoring model, IP Geolocation for what the address claims to be, VPNs, Tor and Anonymizer Exits for the shared-exit problem, and The Defender's Signal Inventory for the layers you cannot change from the network.