Knowing Who Is Knocking

A large share of the requests hitting any public site are automated. Some of that automation you want (search engine crawlers, uptime monitors, feed readers, link previews in chat apps), some you may or may not want (AI training crawlers, SEO tools, price scrapers), and some is plainly hostile (vulnerability scanners, credential stuffing, fake signups, scalping). Blocking "bots" wholesale breaks search visibility and integrations; ignoring them wastes capacity and invites abuse. Bot management means identifying each class and giving it a deliberate response.

This lesson takes the site owner's side. Volumetric floods belong to DDoS defence, and how commercial detection systems score clients is covered in how anti-bot systems work.

A Working Taxonomy

Class Examples What it wants Usual stance
Search crawlers Googlebot, Bingbot Pages to index Allow, but verify identity
AI crawlers GPTBot, ClaudeBot, CCBot Content for model training or answers A policy decision; most honour robots.txt
SEO and marketing tools AhrefsBot, SemrushBot Link and ranking data Allow or rate-limit
Service bots Uptime monitors, feed readers, chat link previews One URL, occasionally Allow
Content and price scrapers Custom scripts, headless browsers Your catalogue, repeatedly Rate-limit, or block per your terms
Vulnerability scanners Mass internet scanners /.env, /wp-login.php, /.git/config Block; fix what they look for
Account and payment abuse Credential stuffing, card testing, fake signups, scalping Your login, checkout, signup Challenge or block on those flows

The last row does the most damage and hides best, because it only needs a few endpoints and often uses real browsers on residential IP addresses.

Reading Your Own Logs

Before buying anything, look at what your access logs already say. Useful per-client signals:

  • Asset ratio: browsers fetch CSS, JavaScript and images along with pages; simple scripts fetch only HTML.
  • Path choice: probes for files you do not have (/.env, /phpmyadmin/) are scanners by definition.
  • Behaviour on sensitive endpoints: many POST /login with mostly 401s is guessing.
  • robots.txt: well-behaved crawlers fetch it first; most abusive clients never do.
  • User-Agent claims: worth recording, but only as a claim to verify.

A compact classifier over a combined-format log:

import re
from collections import defaultdict

LINE = re.compile(r'^(\S+) \S+ \S+ \[[^\]]+\] "(\S+) (\S+) [^"]*" (\d{3}) \S+ "[^"]*" "([^"]*)"')
ASSET = re.compile(r"\.(css|js|png|jpe?g|webp|svg|woff2?|ico)$")
PROBES = ("/.env", "/.git/", "/wp-login.php", "/phpmyadmin", "/xmlrpc.php")
CRAWLER_UA = re.compile(r"googlebot|bingbot", re.I)

stats = defaultdict(lambda: {"n": 0, "assets": 0, "probes": 0, "logins": 0,
                             "fails": 0, "robots": False, "ua": ""})
with open("access.log", encoding="utf-8") as f:
    for line in f:
        if not (m := LINE.match(line)):
            continue
        ip, method, path, status, ua = m.groups()
        s = stats[ip]
        s["n"] += 1
        s["ua"] = ua
        s["assets"] += bool(ASSET.search(path))
        s["probes"] += path.startswith(PROBES)
        s["robots"] |= path == "/robots.txt"
        if method == "POST" and path == "/login":
            s["logins"] += 1
            s["fails"] += status == "401"

def classify(ip, s):
    if s["probes"] >= 2:
        return "vulnerability scanner"
    if s["logins"] >= 3 and s["fails"] / s["logins"] > 0.5:
        return "login abuse"
    if claim := CRAWLER_UA.search(s["ua"]):
        ok, why = verify_crawler(ip, claim.group(0).lower())   # next section
        return f"verified {why}" if ok else f"FAKE {claim.group(0)}: {why}"
    if s["n"] >= 3 and s["assets"] == 0:
        return "page-only client (script?)"
    return "browser-like"

for ip, s in sorted(stats.items(), key=lambda kv: -kv[1]["n"]):
    print(f"{ip:14} req={s['n']} assets={s['assets']} robots={s['robots']!s:5} -> {classify(ip, s)}")

Run against a small sample log:

192.0.2.150    req=5 assets=0 robots=False -> login abuse
203.0.113.10   req=4 assets=2 robots=False -> browser-like
198.51.100.23  req=4 assets=0 robots=False -> FAKE Googlebot: no PTR record
192.0.2.77     req=4 assets=0 robots=False -> vulnerability scanner
66.249.66.1    req=2 assets=0 robots=True  -> verified crawl-66-249-66-1.googlebot.com

Two clients send the identical Googlebot User-Agent; only one is Google. Heuristics like these are a starting point, not a verdict: a headless browser loads assets, and a privacy extension can block them. Treat each signal as evidence and look at clients that trip several.

Verifying Good Bots

The User-Agent header is free text, so "Googlebot" is the most common disguise for scrapers hoping to be let through. Major search engines document a verification method: a reverse DNS lookup on the IP must return a hostname in their domain, and a forward lookup of that hostname must return the same IP. The forward step matters because whoever controls an IP block controls its PTR records and can make them say anything.

import socket
from functools import lru_cache

CRAWLER_DOMAINS = {
    "googlebot": (".googlebot.com", ".google.com", ".googleusercontent.com"),
    "bingbot": (".search.msn.com",),
}

@lru_cache(maxsize=10_000)
def verify_crawler(ip: str, claimed: str) -> tuple[bool, str]:
    suffixes = CRAWLER_DOMAINS[claimed]
    try:
        host = socket.gethostbyaddr(ip)[0]                 # reverse (PTR) lookup
    except (socket.herror, socket.gaierror):
        return False, "no PTR record"
    if not host.endswith(suffixes):
        return False, f"PTR {host} not in {claimed} domains"
    try:
        forward = {ai[4][0] for ai in socket.getaddrinfo(host, None)}  # A and AAAA
    except socket.gaierror:
        return False, f"{host} does not resolve"
    if ip not in forward:
        return False, f"{host} resolves to {forward}, not {ip}"
    return True, host
66.249.66.1  googlebot -> (True, 'crawl-66-249-66-1.googlebot.com')
157.55.39.1  bingbot   -> (True, 'msnbot-157-55-39-1.search.msn.com')
8.8.8.8      googlebot -> (False, 'PTR dns.google not in googlebot domains')

Details that matter: the suffixes start with a dot, so evilgooglebot.com does not match; results are cached, because DNS lookups on every request would be slow; and this should run off the request path (in a log pipeline or background job), not inline. Google and several other operators, including some AI crawler operators, also publish their crawler IP ranges as JSON, which is simpler to enforce at the edge than DNS: download the lists on a schedule and match with the ipaddress module or a web server geo block.

robots.txt Is a Request, Not a Control

robots.txt tells cooperative crawlers what you would like. It is the right tool for managing search engines and the AI crawlers that honour it:

User-agent: GPTBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: *
Disallow: /cart/
Disallow: /search

Google-Extended is a control token, not a separate crawler: it tells Google not to use your content for its Gemini models, while Googlebot keeps crawling for search. Know the limits:

  • Nothing enforces it. Abusive clients ignore it, and some read it for a list of interesting paths. Never list a secret admin URL there.
  • Disallow is not noindex. A disallowed URL can still appear in search results if other pages link to it; to keep a page out of the index, allow crawling and send noindex.
  • It applies per host and per protocol, and it cannot express rate limits that major crawlers honour.

The same file, from the other side, is discussed in robots.txt and ethics.

A Ladder of Responses

Response Use for Cost
Allow Verified crawlers, service bots None
Serve from cache Scrapers of public pages Nearly none; origin untouched
Rate-limit with 429 and Retry-After Heavy but tolerable clients Well-behaved bots slow down
Challenge (JavaScript check, CAPTCHA) Suspicious clients on sensitive flows Friction and accessibility cost for humans
Block with 403 Fake crawlers, scanners, confirmed abuse Risk of false positives

Blocking by IP works for scanners on a single cloud host and fails for botnets and residential proxy networks, where each address sends a handful of requests. Blocking by User-Agent only stops honest bots. The durable controls sit on the endpoints that matter: login, signup, password reset, checkout and search get their own limits (rate limiting and brute-force protection), and business rules often beat detection outright: per-account purchase limits stop scalpers, and requiring email verification before an account can do anything stops most fake signups.

Enforcing at the Edge with nginx

This configuration rejects impostor crawlers and rate-limits everyone except verified ones. The ranges file is regenerated by the scheduled job described above:

geo $verified_crawler {
    default 0;
    include /etc/nginx/verified-crawlers.conf;   # lines like "192.0.2.0/24 1;"
}

map $http_user_agent $claims_crawler {
    default               0;
    ~*(googlebot|bingbot) 1;
}

map "$claims_crawler:$verified_crawler" $fake_crawler {
    default 0;
    "1:0"   1;                                   # claims to be one, is not
}

map $verified_crawler $limit_key {
    0 $binary_remote_addr;
    1 "";                                        # empty key = not limited
}

limit_req_zone $limit_key zone=pages:20m rate=5r/s;

server {
    if ($fake_crawler) { return 403; }

    location / {
        limit_req zone=pages burst=30 nodelay;
        limit_req_status 429;
        proxy_pass http://app;
    }
}

limit_req does not count requests whose key is empty, which is how the verified ranges are exempted. Behind a CDN, set up the real-IP module first, or every client shares the CDN's address.

Challenges and Managed Products

JavaScript challenges and CAPTCHAs stop simple scripts, but they add friction, can exclude users of assistive technology, and break legitimate API and feed clients, so apply them to specific flows rather than the whole site. Determined operators use real browsers and CAPTCHA-solving services (captchas and challenges shows the client side). Managed bot products from CDNs and WAF vendors combine signals you cannot collect yourself: TLS and HTTP fingerprints, cross-site IP reputation, and behavioural telemetry. Some also verify bots that cryptographically sign their requests using HTTP Message Signatures (RFC 9421), an approach that is still emerging but removes the guesswork from identifying well-behaved automation.

Measure What You Block

Every rule creates false positives somewhere. Log each decision with its reason, graph 403 and 429 rates per rule, and review a sample of blocked clients weekly. After any change, check search engine crawl statistics (Google Search Console reports crawl requests and errors) to make sure you did not block the crawler you wanted, and keep a documented allowlist for partners and monitoring services so they survive the next tightening.