The Limiter You Never Saw

You get a 429 and you know where you stand. The dangerous case is the run where you never receive one, the responses stay 200, and the throughput you have been quietly achieving drops to a third without a single status code changing. Rate limiting is mostly invisible, and the discipline is reading the limiter's shape out of whatever it does tell you.

The Headers Are the Specification

Four headers carry most of the information, and they are worth parsing rather than ignoring.

  • Retry-After is either delta-seconds or an HTTP-date. Delta-seconds is what you will see in practice. It is a minimum, and the strongest signal a limiter gives you, because a server that emits it is telling you exactly how long it wants you to wait. Ignoring it and applying your own exponential backoff is a good way to be banned by a system that was being polite.
  • X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset are the de-facto convention. Remaining at zero with a 200 is the clearest statement a server can make: the request succeeded and you are now out. Many implementations only decrement Remaining on success, so the header is a better rate signal than the status code.
  • X-RateLimit-Policy and vendor variants encode the algorithm in a string: "20;w=60" for twenty per sixty seconds, or the draft form "100;w=3600". If you parse it, you get the window size directly instead of inferring it.
  • X-Debug-..., X-Cache, X-Served-By, X-Request-Id are not rate limits, and they matter anyway: the presence and shape of X-Served-By tells you whether your request reached a cache, and a response served from cache skips the limiter entirely. A run that shifts from origin to cache mid-flight looks like a sudden massive increase in your allowance.

Absence is information too. A host that returns Retry-After on one path and not another has per-endpoint limiters, and the paths without one are usually the expensive ones you would rather not hammer.

Three Shapes of Bucket

Token bucket is a bucket that refills at a fixed rate and holds at most a capacity. You can spend the capacity as a burst at any moment; when it is empty, requests wait for a token. Its property is that a client sending steadily at the refill rate is never throttled no matter how long it has been running, and only bursts above the rate are. This is the friendliest shape for a scraper, and it is the one you should hope for.

Fixed window counts requests inside windows aligned to the server's clock, and resets at each boundary. It is trivial to implement with two integers and a cron job, which is why it is everywhere. Its property is the edge burst: the quota is granted twice in quick succession across a boundary, so a client can send 2 * limit requests inside a few seconds and be entirely within the rules while being nothing like the intended rate.

Sliding window enforces the limit over any contiguous window rather than over aligned buckets. An exact implementation keeps a log of acceptance timestamps and drops those older than the window; an approximate one uses two counters with weighted interpolation, which is what most production systems use because it is O(1) memory. Its property is that limit is a true ceiling on any window of that length, with no edge allowance, which makes it the strictest of the three and the least forgiving of a client that does not read Retry-After.

Leaky bucket is the queue-shaped variant: arrivals join a queue at a fixed service rate and the queue depth is the limit. You do not get rejected, you get delayed, which is why it appears as latency rather than as a status code. It is also the mechanism behind soft throttling, discussed below.

Same Pattern, Three Verdicts

The snippet implements all three and replays one request pattern through each. Then it moves the fixed window's boundary and replays the same pattern again, and finally runs forty requests inside a quarter of a second across a window edge.

# Three limiters, one request pattern. Same client, three different verdicts.
# Logical seconds only: t=0.0 is the start of the run, nothing sleeps.

ARRIVAL_GAP = 0.185      # 5.40 requests per second, a plausible "fast crawler"
N_REQUESTS = 378
WINDOW = 60.0


class TokenBucket:
    # Capacity C, refill R per second. Bursts allowed up to C.

    def __init__(self, cap, rate):
        self.cap, self.rate, self.tokens = float(cap), float(rate), float(cap)
        self.now = 0.0

    def take(self, now):
        self.tokens = min(self.cap,
                          self.tokens + (now - self.now) * self.rate)
        self.now = now
        if self.tokens >= 1.0:
            self.tokens -= 1.0
            return True, 0.0
        return False, (1.0 - self.tokens) / self.rate


class FixedWindow:
    # Quota Q per window W. The boundary is the server's clock, not yours.

    def __init__(self, quota, window, phase=0.0):
        self.quota, self.window = quota, float(window)
        self.start, self.used = phase, 0

    def take(self, now):
        while now - self.start >= self.window:
            self.start += self.window
            self.used = 0
        if self.used < self.quota:
            self.used += 1
            return True, 0.0
        return False, self.start + self.window - now


class SlidingLog:
    # Quota Q in any W seconds, enforced exactly. Phase-independent.

    def __init__(self, quota, window):
        self.quota, self.window = quota, float(window)
        self.log = []

    def take(self, now):
        while self.log and now - self.log[0] >= self.window:
            self.log.pop(0)
        if len(self.log) < self.quota:
            self.log.append(now)
            return True, 0.0
        return False, self.window - (now - self.log[0])


def build():
    return [
        ("token bucket 20 @5.0/s", TokenBucket(20, 5.0)),
        ("fixed window 30 / 60s", FixedWindow(30, WINDOW)),
        ("sliding log 30 / 60s", SlidingLog(30, WINDOW)),
    ]


def run(limiter, pattern):
    allowed_at, first, denied = [], None, 0
    for t in pattern:
        ok, _wait = limiter.take(t)
        if ok:
            allowed_at.append(t)
        else:
            denied += 1
            if first is None:
                first = t
    return allowed_at, denied, first


def replay(pattern):
    print(f"pattern: {N_REQUESTS} requests, gap {ARRIVAL_GAP}s "
          f"({1 / ARRIVAL_GAP:.2f}/s), run spans {pattern[-1]:.1f}s")
    print()
    print("  limiter                  200   429  first 429  implied  "
          "edge-bound")
    for name, limiter in build():
        allowed, denied, first = run(limiter, pattern)
        span = pattern[-1] - pattern[0]
        print(f"  {name:<22}{len(allowed):>4}{denied:>6}   "
              f"t={first:>5.2f}s  {len(allowed) / span:5.2f}/s   "
              f"{'yes' if isinstance(limiter, FixedWindow) else 'no'}")


def phase_test():
    print()
    print("the same pattern again, only the server's window boundary moved")
    for phase in (0.0, 3.0, 29.5, 47.0):
        limiter = FixedWindow(30, WINDOW, phase=phase)
        allowed, denied, first = run(limiter, [round(phase + i * ARRIVAL_GAP, 4)
                                              for i in range(N_REQUESTS)])
        print(f"  boundary at t=+{phase:>4.1f}s   200s {len(allowed):>3}  "
              f"429s {denied:>3}   first 429 at t={first:>5.2f}s")
    print()
    print("  a fixed window grants its quota twice inside 2s, once per edge")
    print("  40 requests inside 0.26s, straddling the boundary at t=60.0")
    burst = ([round(59.88 + i * 0.006, 4) for i in range(20)]
             + [round(60.02 + i * 0.006, 4) for i in range(20)])
    for name, limiter in build():
        allowed, _denied, _first = run(limiter, burst)
        print(f"  {name:<22}{len(allowed):>4}{len(burst) - len(allowed):>6}")


def adaptive(label, limiter, start_rate, seed, attempts=240):
    # A client that reads Retry-After, creeps on success, jitters its gaps.
    import random
    rng = random.Random(seed)
    rate = start_rate
    ok = denied = 0
    t = 0.0
    for _ in range(attempts):
        t += rng.uniform(0.55, 1.55) / rate
        good, wait = limiter.take(t)
        if good:
            ok += 1
            rate = min(start_rate * 1.6, rate * 1.02)
        else:
            denied += 1
            rate = max(0.05, rate * (0.5 if wait <= 0 else 0.9))
    print(f"  {label:<22}{rate:>6.2f}/s  200s {ok:>3}  429s {denied:>3}  "
          f"{100.0 * ok / attempts:5.1f}% accepted")


def main():
    pattern = [round(i * ARRIVAL_GAP, 4) for i in range(N_REQUESTS)]
    replay(pattern)
    phase_test()
    print()
    print("the same client behind a controller that reads Retry-After")
    for n, (label, _unused) in enumerate(build()):
        adaptive(label, build()[n][1], start_rate=6.0, seed=17 + n)


main()
pattern: 378 requests, gap 0.185s (5.41/s), run spans 69.7s

  limiter                  200   429  first 429  implied  edge-bound
  token bucket 20 @5.0/s 368    10   t=46.99s   5.28/s   no
  fixed window 30 / 60s   60   318   t= 5.55s   0.86/s   yes
  sliding log 30 / 60s    60   318   t= 5.55s   0.86/s   no

the same pattern again, only the server's window boundary moved
  boundary at t=+ 0.0s   200s  60  429s 318   first 429 at t= 5.55s
  boundary at t=+ 3.0s   200s  60  429s 318   first 429 at t= 8.55s
  boundary at t=+29.5s   200s  60  429s 318   first 429 at t=35.05s
  boundary at t=+47.0s   200s  60  429s 318   first 429 at t=52.55s

  a fixed window grants its quota twice inside 2s, once per edge
  40 requests inside 0.26s, straddling the boundary at t=60.0
  token bucket 20 @5.0/s  21    19
  fixed window 30 / 60s   40     0
  sliding log 30 / 60s    30    10

the same client behind a controller that reads Retry-After
  token bucket 20 @5.0/s  5.42/s  200s 206  429s  34   85.8% accepted
  fixed window 30 / 60s   0.71/s  200s 186  429s  54   77.5% accepted
  sliding log 30 / 60s    0.63/s  200s 185  429s  55   77.1% accepted

Three findings worth arguing with. First, the token bucket let 368 of 378 through and only started refusing at t=46.99s, because its capacity absorbs exactly the burst a fixed window would have refused in the first five seconds. Second, the fixed window's total barely moved when its boundary moved but the moment the first 429 arrived moved by 47 seconds -- which means a single probe against a fixed-window limiter tells you nearly nothing about your real allowance. Third, the edge burst: the fixed window accepted all forty requests that the sliding log refused ten of, and the token bucket refused nineteen.

Per Endpoint, Per Account, Per Address

The scope of the limiter decides which budget you are actually spending.

  • Per IP. The default. Your exit address is the unit, which means a residential exit shared by eight contexts has an eighth of the quota, and rotating addresses inside a burst is the fastest way to trip a per-account limit that sits behind it.
  • Per account or credential. Binds the quota to the logged-in identity regardless of where the request comes from. This is invisible until you log in, at which point your effective rate can drop by an order of magnitude with no header change. If your authenticated phase runs at a lower rate than your anonymous phase for no reason you can identify, this is why.
  • Per endpoint. Search is usually cheap and generously limited; the detail-page and inventory endpoints are where the real quota lives. A budget expressed per host is wrong for any target with a cost gradient.
  • Per session. Ties to the cookie or clearance value, so it survives an IP change and resets when you drop the jar. Combined with per-account, this is a two-dimensional quota and the practical consequence is that you cannot raise the rate by rotating addresses if your account is the binding constraint.
  • Per ASN or subnet. Rare, and aimed squarely at cloud ranges. It means a hundred residential exits routed through the same provider's aggregation prefix get one budget, and that budget is shared with every other customer of that provider.

Find out which one binds before you tune anything: run a fixed-rate probe against each scope separately and see which one's Remaining moves. Browser pool architecture treats per-exit limits as a hard concurrency ceiling for the same reason.

Soft Throttling Is Not a 429

The most expensive limiter is the one that returns 200 OK and takes four seconds. You cannot back off from what you were never told about.

How it works: the origin accepts the connection, runs your request, and applies a deliberate delay before responding, sized to bring aggregate throughput down to the target rate. Every status code is correct, every body is correct, every cookie is set as usual. The only evidence is the latency distribution.

What that does to a client is worth spelling out. A pipeline with a fixed per-request timeout will start shedding requests that used to succeed. A pipeline with a throughput SLA will miss it while reporting a 100 percent success rate. A retry layer keyed on status codes will never fire, so the latency climbs until a socket somewhere closes, and the resulting error looks like an infrastructure problem -- which is exactly how you will misdiagnose it for a day.

Detect it by tracking the median and a dispersion measure rather than the mean, and by comparing against a per-endpoint baseline. A median that moves from 84 ms to 1.9 s with a tight distribution is a throttle; the same median with a wide spread is a slow origin. In the snippet above, fabrikam shows three 200s at 1840-1890 ms against a 55-64 ms baseline, with no status change at all, and that is the shape to alert on.

The defensive response is to treat latency as a first-class rate-limit signal: when the median exceeds a multiple of the baseline, drop your rate by the same proportion, whether or not anything told you to.

Why Scaling Concurrency Trips a Limiter You Never Heard From

This is the common way to get banned without ever seeing a 429, and the mechanism is worth being precise about.

Most limiters are not concurrency limiters. A concurrency cap rejects the eleventh simultaneous request and tells you. A token bucket with capacity 20 does something else: it accepts the first twenty, and the twenty-first waits for a token. If your client blocks on that wait, nothing fails -- the request just takes longer. But if your client has a queue and a deadline, the deadline is what fails. And if your client has a pool of eight workers each holding a connection open while it waits, you have done the one thing a leaky-bucket limiter is designed to punish: you have converted rate into concurrency, which increases the server's queue depth without increasing its throughput.

Three specific ways this happens:

  • Per-worker rate instead of per-fleet rate. Eight workers each sending 5/s is 40/s. Every worker believes it is under the limit. This is the single most common version, and the fix is a shared token bucket rather than a per-worker sleep.
  • Burst on every reconnect. A pool that recycles contexts or proxies sends a burst of initial requests every time one comes up. Multiply that by the pool churn rate and the average looks fine while the peaks do the damage, and the churn itself should be damped, because a token bucket with a generous capacity is what absorbs those peaks.
  • Retry amplification. Every layer that retries independently doubles the load at the worst possible moment. See crawl budget and request scheduling for the retry budget, and note that a 429 with a Retry-After you ignore is not free politeness.

The general rule: if you cannot state, in one sentence, what your aggregate request rate to one host is at the worst moment of a minute, you do not control it.

Probing Without Burning Your Reputation

You have to measure, and measurement is not free. Four methods, cheapest first.

  • Respect the advertised numbers. If X-RateLimit-Limit and X-RateLimit-Reset are present, you have the answer without probing. This is the correct default and it costs nothing.
  • Binary search on the observed rate. Hold a low rate, then step up by 20 percent and hold for a full window. The first step where Remaining starts hitting zero is your ceiling. Budget roughly a logarithmic number of windows and it converges in three or four.
  • Measure the refill, then the capacity. Send one request, read Remaining, wait exactly one advertised window, send one request, read again. The difference is the refill rate. The peak Remaining you ever observe is the capacity. Two probes, one window, no risk.
  • Read the challenge rate instead. Where a limiter is invisible, the bot manager's intervention rate is not. A run at 2 challenges per thousand requests against a run at 20 is a rate signal that arrived as a status code of 200.

Do all probing from the identity you intend to keep. A ceiling measured from a throwaway exit tells you about the exit, and the account's ceiling is usually lower.

Turning Measurements Into a Setpoint

Once you know the limiter's shape you have four numbers: refill rate, burst capacity, scope, and whether it tells you. The controller is straightforward and it is the same one described in ban recovery and backoff: obey Retry-After absolutely, back off on 429, creep up on success by a small factor, and stop creeping the moment the challenge rate or the latency median rises.

The setpoint you hold is not the ceiling you found. It is the ceiling minus a margin, sized so that a 20 percent traffic spike from something else on your side does not take you over. A common operating point is 60 to 70 percent of measured capacity, because the alternative is discovering the real ceiling during an incident.

Hold the ban rate and the challenge rate as the objective and the rate as a knob. Waf evasion tips makes the same argument about behavioral limits; the mechanism differs but the arithmetic is identical, and observability and ban SLOS is where the objective function gets defined.