IP Geolocation: Grades, Anycast and Mismatch
The Database Is Confident and the Database Is Wrong
A geolocation lookup returns a country, a region, a city, a postal code, a latitude, a longitude and an accuracy radius, in under a millisecond, for two dollars per thousand queries. That is the most impressive and least reliable number in the entire anti-bot signal set, and every downstream decision that depends on it - the language you are served, the prices you see, the challenge you are shown - inherits the error.
The useful thing to understand is not that geolocation is inaccurate. It is how it is inaccurate, in a structured way, which means the error is predictable and therefore usable as a signal in both directions.
Four Questions, Four Accuracies
| Granularity | Typical error | Good enough for | Fails at |
|---|---|---|---|
| Country | 95-99.5% correct | Coarse routing, sanctions, fraud rules | Nothing much |
| Region / state | 70-90% | Currency defaults, tax region hints | Rural and mobile |
| City | 40-75% | Sanity checks, timezone defaults | Hosted and mobile IP |
| Metro / postcode | 10-30% | Store locators | Anything dynamic |
Country accuracy is genuinely good, because a country is a large contiguous block of addresses and databases only have to be right about boundaries. City accuracy is a different problem: the address has to have been geolocated to a specific street, and for dynamic residential space the correct street is wherever the subscriber's CPE was registered months ago. A database that says 100 km of confidence is telling you that it knows a circle, not a point.
Why the Error Radius Exists
Commercial databases publish a radius for a reason. When a block is delegated to an ISP, the geolocation vendor maps the allocated range to the ISP's registered service area, and the ISP's service area is often the size of a state. Rather than pretend, the honest vendors publish an accuracy_radius in kilometres, and good ones refuse to return a city outside that radius.
The convention that matters in practice: 20 to 50 km for a well-served urban block, 100 to 200 km for a rural or dynamically allocated block, and a country-only answer with no radius at all for a hosting address that has never been geolocated from an end-user perspective. A well-behaved site uses country for compliance decisions, city only for presentation, and never for anything load-bearing. If you see a site gating a purchase on city-level IP geolocation, that is a bug in their risk model, not a precision they have earned.
The Major Databases and Their Biases
| Provider | Refresh | Characteristic bias |
|---|---|---|
| MaxMind | Daily, GeoLite2 free, GeoIP2 paid | Ties city to ISP registration; large installed base |
| IP2Location | Weekly | More aggressive on mobile carrier inference |
| IPinfo | Continuous | Strong on ASN, privacy and hosted-network flags |
| DB-IP | Monthly | Cheap, country-level emphasis |
| IPwho.is / ipapi | Aggregated | Inherits whichever upstream it mirrors |
| Internal (ISP self-reported) | Real time | Only the operator knows |
Two of these biases matter to a scraper. First, aggregation: many free APIs resolve a city by asking a paid database and caching it, so two "independent" services frequently return the same upstream guess and their agreement is not evidence. Second, self-report: the most accurate location for a residential address is the ISP's own record, available only to the ISP, which is why a carrier's own lookup is the only one that can be right about a CGNAT pool.
The practical consequence is that agreement between databases proves nothing, and disagreement is a signal. Two major databases that place the same residential address in two different countries are telling you that the underlying block is dynamic or reassigned, which is a property of the address, not of the databases.
Anycast and the CDN Edge Problem
The worst case for geolocation is anycast, and every large CDN uses it. The same address is announced from hundreds of points of presence; a visitor in Frankfurt and a visitor in São Paulo may be served by the same address; and the address is registered to the CDN operator, not to either visitor.
The result is the familiar absurdity: a United States exit address, a site that only exists in one country, an Accept-Language of ar-SA, and a page rendered in a currency that matches neither. Every one of those values is individually plausible and the combination is impossible for a human, which is precisely why mismatch scoring works. The more interesting question is what a defender should do with anycast addresses, and the honest answer is to treat geolocation as unknown for them rather than to believe the registrant - the ASN and datacenter lesson explains why the address is in a CDN's range in the first place.
Mismatch Is the Signal
The signal is not the location. It is the disagreement between locations, computed from independent claims:
| Claim | Source | Independent of the IP? |
|---|---|---|
| Exit country | GeoIP database | No - it is the IP |
| Timezone | Intl.DateTimeFormat().resolvedOptions().timeZone |
Yes, and hard to spoof convincingly |
Accept-Language |
HTTP header, sent by the client | Yes |
navigator.language |
JS environment | Yes |
| Currency and price display | Server-side selection from the exit country | No, derived |
| Local time of the request | Edge clock minus the exit country's offset | Yes, and the edge controls it |
| Viewport and platform | JS environment and UA | Yes |
A single claim can be wrong. Four claims that agree on a country are very hard to be wrong about simultaneously by accident, which is why the matrix below is the useful unit of analysis rather than any individual lookup. = means the two claims can both be true of one person, x means they cannot, ? means one of them is unknown.
ZONE_COUNTRY = {
"Europe/Berlin": {"DE"}, "America/New_York": {"US"},
"America/Los_Angeles": {"US"}, "Asia/Tokyo": {"JP"},
"Europe/London": {"GB"}, "Africa/Lagos": {"NG"},
}
LANG_REGION = { # what a language tag says, not what it proves
"de": {"DE", "AT", "CH"}, "en": {"US", "GB", "IE", "NG", "AU", "ZA"},
"ar": {"SA", "EG", "AE", "JO"}, "ja": {"JP"}, "nl": {"NL", "BE"},
}
CURRENCY_COUNTRY = {
"EUR": {"DE", "FR", "NL", "IT", "ES", "BE", "AT", "IE", "PT", "FI", "IE"},
"USD": {"US", "EC", "SV", "PA"}, "GBP": {"GB"}, "JPY": {"JP"},
"SAR": {"SA"}, "NGN": {"NG"}, "EGP": {"EG"},
}
MOBILE_WIDTHS = range(320, 600) # CSS px, portrait phones
PROFILES = [
# name, exit country, timezone, Accept-Language, currency, viewport, UA platform
("de-shop", "DE", "Europe/Berlin", "de-DE,de;q=0.9,en;q=0.8", "EUR", 1920, "Windows"),
("us-shop", "US", "America/New_York", "en-US,en;q=0.9", "USD", 1920, "Windows"),
("ar-feed", "US", "America/Los_Angeles", "ar-SA,ar;q=0.9", "SAR", 412, "Android"),
("jp-mobile", "JP", "Asia/Tokyo", "ja-JP,ja;q=0.9", "JPY", 412, "Android"),
("nl-lab", "NL", "Africa/Lagos", "en-NG,en;q=0.9", "NGN", 1366, "Windows"),
("us-mixed", "DE", "Europe/Berlin", "ar-EG,en;q=0.8", "EGP", 1440, "macOS"),
("de-ua-drift", "DE", "Europe/Berlin", "de-DE,de;q=0.9", "EUR", 412, "Windows"),
("no-geo", None, "Asia/Tokyo", "ja-JP,ja;q=0.9", "JPY", 1920, "Windows"),
]
def countries(profile):
_name, ip_cc, zone, accept, currency, _w, _p = profile
out = {"exit IP": {ip_cc} if ip_cc else set(),
"timezone": ZONE_COUNTRY.get(zone, set()),
"language": LANG_REGION.get(accept.split(",")[0].split("-")[0], set()),
"currency": CURRENCY_COUNTRY.get(currency, set())}
return out
def verdict(a, b):
if not a or not b:
return "?"
return "x" if a.isdisjoint(b) else "="
SIGNALS = ["exit IP", "timezone", "language", "currency"]
PAIRS = [(a, b) for i, a in enumerate(SIGNALS) for b in SIGNALS[i + 1:]]
print(f"{'profile':12}" + "".join(f"{a[:2]}/{b[:2]:<2} " for a, b in PAIRS)
+ " device coherence contradictions")
for profile in PROFILES:
name = profile[0]
sets = countries(profile)
cells, bad = [], []
for a, b in PAIRS:
mark = verdict(sets[a], sets[b])
cells.append(mark)
if mark == "x":
bad.append(f"{a} vs {b}")
looks_mobile = profile[5] in MOBILE_WIDTHS
device = f"{profile[5]}px/{profile[6]:8} {'clash' if looks_mobile != (profile[6] == 'Android') else 'ok'}"
agree = sum(1 for c in cells if c == "=") / len(cells)
print(f"{name:12}" + "".join(f"{c:^7}" for c in cells)
+ f" {device:22} {agree:4.0%} "
+ (", ".join(bad) if bad else "-"))
profile ex/ti ex/la ex/cu ti/la ti/cu la/cu device coherence contradictions
de-shop = = = = = = 1920px/Windows ok 100% -
us-shop = = = = = = 1920px/Windows ok 100% -
ar-feed = x x x x = 412px/Android ok 33% exit IP vs language, exit IP vs currency, timezone vs language, timezone vs currency
jp-mobile = = = = = = 412px/Android ok 100% -
nl-lab x x x = = = 1366px/Windows ok 50% exit IP vs timezone, exit IP vs language, exit IP vs currency
us-mixed = x x x x = 1440px/macOS ok 33% exit IP vs language, exit IP vs currency, timezone vs language, timezone vs currency
de-ua-drift = = = = = = 412px/Windows clash 100% -
no-geo ? ? ? = = = 1920px/Windows ok 50% -
The de-ua-drift row is the one to notice: every geographic claim agrees, so a country-level check passes it, and it is still wrong, because a 412-pixel viewport on a Windows user agent is a contradiction with no geographic content. Cross-signal consistency has to include the device, not just the place, and the machine-readable form of that check is the contradiction matrix covered in Cross-Validating Fingerprint Signals.
Latency Beats the Database
Network round-trip time is the one geolocation signal that reflects the actual path rather than a registration record. Light in fibre travels at roughly 200,000 km per second, so one millisecond of RTT is about 100 km of fibre each way. Probe a handful of well-distributed anycast resolvers and solve for the point closest to all of them, and you frequently land within a few hundred kilometres - comparable to a commercial city database, at no licence cost, and using a measurement rather than an inference.
# RTT to anycast resolvers; the low, stable one is usually the nearest PoP
for h in 1.1.1.1 8.8.8.8 9.9.9.9 208.67.222.222; do
ping -n 6 "$h" | tail -2 | head -1
done
tracert -d -h 12 -w 200 example.com # path length is itself a region signal
Two cautions. First, the mapping is only as good as the anchor set, and a probe set concentrated in one continent will happily report the wrong continent. Second, transit routing is not distance: a submarine cable, a congested transit hop or a sponsored peering can add hundreds of kilometres of delay, so a latency-derived city is a prior, not a verdict. A defender combining a database lookup with latency and a declared timezone gets a much tighter bound than any of the three alone, and a scraper who wants to be believed has to satisfy all three.
What a Legitimate Localised Scrape Does
If a job genuinely needs a specific market - prices, availability, a localised landing page - the correct approach is to say so and to be consistent about it:
- Pick the market from the data requirement, not from the proxy pool you happen to have.
- Use an exit geolocated to that market, at city level if the data is city level, and accept that city-level is a best guess.
- Set the profile's timezone,
Accept-Language,navigator.languages,Intllocale and currency to the same market, as a bundle. - Check the rendered page, not the headers: a store that ignores your region and serves a currency chooser has told you the alignment did not work.
- Never mix markets inside a session. A session whose IP, timezone and language change together is a stronger anomaly than any single wrong value.
- Expect the challenge anyway. A residential proxy in the right city with a coherent profile is less suspicious, not invisible.
Geo-Targeting and the Law
Geo-targeting is where anti-bot work touches regulation most directly. If the collected data identifies or can identify a person, the GDPR applies to processing it outside the EU, and consent, purpose limitation and retention are not optional. Where a site displays prices, tax or availability by country, scraping the wrong market and presenting the results as if they were the right one can mislead buyers and breach consumer-protection rules, and several jurisdictions have specific pricing-display statutes. Treat the geography of a dataset as part of the dataset's meaning: record which market a row came from, and do not launder a price list through three relocations.
Related reading: ASN, Cloud Ranges and Datacenter Detection for the address before the location, Timezone, Locale and Device Profiles for the client-side half of the matrix, VPNs, Tor and Anonymizer Exits for what geolocation says about an anonymised exit, and Proxies and IP Rotation for the transport side of a regional exit.