Crawl Budget and Request Scheduling
A Budget Is a Constraint, Not a Suggestion
Most rate limiting happens in the client's own code, minutes before any server gets a chance to complain. You decide what a worker does on its next iteration, and that decision is where throughput, politeness and survival are actually settled. A crawl with a budget is a crawler; a crawl without one is a flood that happens to be polite until it is not.
The Four Budgets, and Why One Is Not Enough
A single global rate limit is the wrong abstraction because the constraints that matter live at different scopes, and exceeding the wrong one is what gets you banned.
| Budget | Unit | What exceeding it costs you |
|---|---|---|
| Per host, per minute | requests/s to one origin | a 429, or a silent throttle, or an IP-scoped ban |
| Per credential, per hour | requests on one logged-in identity | account review, then account ban |
| Per identity, per day | requests on one device profile | the profile's trust stops accumulating |
| Global, per hour | everything you send | your own egress bill and your own saturation |
The per-host budget protects the address. The per-credential budget protects the account, and it is usually the lower of the two, which surprises people who tuned against a host-level Retry-After and never saw the account limit. The per-identity daily budget protects the thing you spent weeks building, and it is the one most often absent, because nothing complains until the account is reviewed.
Track all four as separate counters with separate reset windows. A run that respects three of them and blows the fourth still ends, and the diagnosis takes a day when the counters are not independent.
Politeness as Survival
The economics are simple enough to be worth stating plainly. Your crawl has a completion date. A ban costs you the identity, the accumulated trust on every profile sharing that exit, and the elapsed time. Cutting your rate by 30 percent to avoid that is almost always correct, because the data you lose by running slower is data you were going to get next week anyway.
Two mechanisms make politeness pay directly rather than merely avoid cost.
Server-side queueing. Under a leaky bucket your requests are not rejected, they are delayed. Being under the limit does not make the site faster, but being over it makes your own latency explode, and latency is what your pool sizing and your timeout budget are built on.
Reputation decay. IP and account reputation are decaying functions. A steady 4 requests per second for an hour leaves a very different ledger entry than 20 for a minute followed by a 30-minute gap. If your job can be spread over the window you have, spreading it is worth more than the throughput you gave up.
Which is not a licence to be slow. The right frame is that you have a fixed hourly budget and you want to spend it where it returns value, and politeness is what stops you from spending it all on one cheap endpoint in the first four minutes.
Spending the Budget Where It Returns Value
Every URL you fetch is not worth the same. A crawl that treats a sitemap index, a product page, an image manifest and a pagination endpoint identically is spending 60 percent of its budget on rows it does not need.
Score by value density, not value:
- Row value. What the page gives you. A product detail page with price and stock is worth thirty times an image URL.
- Prerequisite depth. A pagination page is worth almost nothing on its own, but it is the only route to forty product pages, so its density is forty times its row value. Decide whether a prerequisite gets funded out of a shared reserve, because funding it at its own density starves it.
- Freshness decay. If prices change hourly and reviews change monthly, a review page fetched twice in a day is 90 percent waste. Store a per-URL-type TTL and skip anything inside it.
- Marginal information. Two pages from the same listing page, fetched through different sort orders, share 95 percent of their content. The second one is worth a fraction of the first.
The practical shape is a priority queue over the frontier, drained under the budget, with a small reserve for the prerequisites you cannot skip. That is what the schedule below does: 502 rows queued, 190 of them affordable, and it spends them on the highest-value rows while respecting a per-host politeness ceiling that is tighter than the budget would otherwise allow.
Concurrency Caps That Survive a Slowdown
Here is the failure mode, and it is the reason naive pools die in exactly the wrong way.
The target slows from 84 ms to 1.9 s -- a soft throttle, an incident, a bad deploy, your provider's transit. With a fixed worker pool of eight, request throughput collapses from 95/s to 4.2/s. With a worker pool that grows to fill the queue, eight become forty become two hundred, because every connection is "slow" and the pool interprets slow as a reason to try more. You have turned a slowdown into a self-inflicted denial of service against your own egress and, on a shared block, against everyone behind it.
Four rules, and the pool in browser pool architecture enforces all four.
- Cap concurrency, never grow it on latency. Latency is a signal to stop, not to expand. If anything, reduce the cap.
- Cap per host independently. A slow host must not consume workers that another host could use.
- Apply a global in-flight ceiling derived from your egress and your proxy gateway's published concurrency allowance, not from your queue length.
- Trip a breaker on the latency median, not on a request count. Four consecutive responses over 2x the rolling median is a stop, and the reset is a cooldown rather than a retry.
The measurement that tells you the cap is right is the ratio of in-flight requests to completed ones. Healthy is under 2. If it is above 8 you are running an unbounded queue and the latency number is your queue depth, not the server's.
Jitter, and Why Uniform Spacing Is a Signal
Uniform spacing is worse than randomness, because it is not noise -- it is a constant, and constants are detectable with a much smaller sample.
A sleep(2.0) between requests produces deltas of 2.001, 2.000, 2.002 seconds. Any classifier that looks at the variance of inter-arrival times separates that from a human's within a few dozen requests, and the variance of a human's is large: 0.3 to 9 seconds, occasionally 40 seconds while they read. What identifies a person is not the mean but the spread, and any schedule with a small spread is a machine.
Two properties to include in a jitter function:
- A heavy right tail. Most gaps short, occasionally very long. A truncated exponential or a lognormal does it; a uniform distribution does not have the right shape.
- Session structure. Gaps are not independent. A burst of 40 requests at 300 ms during a listing crawl, then a 4-second gap while a detail page renders, then 300 ms again -- that is what a workflow looks like. Independent per-request jitter produces an aperiodic stream that is less human than a structured one.
Decorrelated jitter, from AWS's architecture blog, is the backoff form of this and it is the one worth using: each wait is drawn from uniform(base, previous * 3) and capped, which gives you a spread that grows under pressure instead of a fixed interval. Full-jitter (uniform(0, base * 2^n)) is a good and simpler alternative when you do not carry state between attempts.
There is a second reason jitter matters that has nothing to do with detection: it decorrelates your workers. Eight workers each on a fixed schedule hit in lockstep. Eight workers each on independent jittered schedules spread out, and the aggregate arrival process looks like a network of independent users rather than one scheduler.
Retry Budgets and the Amplification Bug
Retries are the only mechanism in a crawler that can increase load on a system that is already failing. Everything else reduces it.
The bug is this: a layer retries on 429, the layer above it retries on timeout, the layer above that retries on empty results, and each has its own backoff with its own randomness. During a throttling event all three fire simultaneously, each independently deciding to try again, and the resulting load on the target is three to eight times the intended rate at exactly the moment the target can least afford it. The amplification is not steady-state throughput; it is a spike shaped like the failure, which means it makes the failure longer.
The fix is a budget, enforced centrally:
- Cap retries as a fraction of successful requests. A common setting is 10 percent, so a run that gets 900 successes may spend 90 retries and then stop trying entirely for the window. A hard absolute cap per hour works too, and is easier to reason about.
- Never retry a 4xx other than 429. A 403 is a verdict, not a transient. Retrying a banned identity is what converts a soft block into a hard one, which is the escalation rule in ban recovery and backoff.
- One retry layer, not three. The retry decision belongs in the queue consumer that owns the request, and every other layer must pass failure up rather than retrying.
- Spend from the reserve, not from the plan. If the hourly budget is 260 and you reserve 24 for retries, the plan gets 236. Spending the reserve is what the schedule below does: the credential budget runs to zero before a single retry does.
- Requeue, do not retry, when the whole fleet is affected. If the ban is per-host rather than per-identity, the right move is to stop and cool down, not to try the same request from a different worker thirty seconds later.
Track retries as a first-class metric with the same alerting as success rate. A crawl where retries are 30 percent of traffic is a crawl you are about to lose.
A Schedule Under Budget
Five hundred and two rows, a global ceiling of 8 per minute, 5 per host per minute, 260 requests of total hourly budget, 190 on the credential, and 24 held back for retries. Twenty-five minutes of runway, then the credential budget is the binding constraint and 312 rows stay queued.
# A crawl scheduled under a token budget, with decorrelated-jitter backoff.
# Logical minutes. No clock, no sleeping, no real network.
MINUTES = 40
GLOBAL_CAP = 8 # requests/minute, all hosts, one credential
HOST_CAP = 5 # requests/minute to any single host
BUDGET_TOTAL = 260
BUDGET_CRED = 190
BUDGET_HOST = {"shop.example.com": 150, "docs.example.com": 70}
RETRY_RESERVE = 24 # retries kept back out of the plan
def make_items():
# Value falls off with id so the plan visibly walks down a priority list.
items = [("login", "shop.example.com", 10, "flow"),
("cart", "shop.example.com", 10, "flow")]
for n in range(1, 301):
items.append((f"p/{1000 + n}", "shop.example.com",
10 - (n % 7), "detail"))
for n in range(1, 201):
items.append((f"d/{300 + n}", "docs.example.com",
8 - (n % 6), "doc"))
return sorted(items, key=lambda it: (-it[2], it[0]))
def schedule():
pool = make_items()
left_host = dict(BUDGET_HOST)
left_total, left_cred = BUDGET_TOTAL, BUDGET_CRED
plan = []
for minute in range(MINUTES):
per_host, picked = {}, []
for item in list(pool):
if len(picked) >= GLOBAL_CAP or left_cred <= 0 or left_total <= 0:
break
_iid, host, _value, _kind = item
if per_host.get(host, 0) >= HOST_CAP:
continue
if left_host.get(host, 0) <= 0:
continue
per_host[host] = per_host.get(host, 0) + 1
picked.append(item)
pool.remove(item)
left_host[host] -= 1
left_total -= 1
left_cred -= 1
plan.append({
"minute": minute, "picked": picked,
"per_host": dict(sorted(per_host.items())),
"left_total": left_total, "left_cred": left_cred,
"left_host": dict(sorted(left_host.items())),
})
if not picked:
break
deferred = [i[0] for i in pool]
return plan, deferred
def decorrelated_jitter(outcomes, seed, base=0.5, cap=60.0):
# wait_n = min(cap, uniform(base, wait_{n-1} * 3)), wait_0 = base.
import random
rng = random.Random(seed)
wait = base
out = []
for i, ok in enumerate(outcomes):
if ok:
out.append((i, "200", 0.0))
continue
wait = min(cap, rng.uniform(base, wait * 3.0))
out.append((i, "429", wait))
return out
def main():
plan, deferred = schedule()
print(f"budget: {BUDGET_TOTAL}/hour total, {BUDGET_CRED} on the "
f"credential, {sum(BUDGET_HOST.values())} across hosts")
print(f"ceilings: {GLOBAL_CAP}/min global, {HOST_CAP}/min/host; "
f"{len(make_items())} rows queued; {RETRY_RESERVE} retries "
f"held back")
print()
print(" min spend shop docs total cred left tot left "
"lowest value")
for row in plan:
shop = row["per_host"].get("shop.example.com", 0)
docs = row["per_host"].get("docs.example.com", 0)
low = min((i[2] for i in row["picked"]), default="-")
print(f" {row['minute']:>3} {len(row['picked']):>5} {shop:>4} "
f"{docs:>4} {len(row['picked']):>5} "
f"{row['left_cred']:>9} {row['left_total']:>8} "
f"{str(low):>12}")
last = plan[-1]
blocked = [name for name, n in sorted(last["left_host"].items()) if n == 0]
print()
print(f" ran {len(plan)} minutes, spent {BUDGET_TOTAL - last['left_total']}"
f" of {BUDGET_TOTAL}; still queued {len(deferred)} rows")
print(f" binding constraint: credential budget at {last['left_cred']}, "
f"host budgets exhausted for {blocked or 'none'}")
print(f" retries reserved {RETRY_RESERVE}; six workers retrying a 429 "
f"spend that in {RETRY_RESERVE // 6} minutes")
print()
print("decorrelated jitter, seed 5, base 0.5s, cap 60s, against 429s")
outcomes = [False] * 6 + [True] + [False] * 5 + [True, True]
waits = []
for i, status, w in decorrelated_jitter(outcomes, seed=5):
waits.append(w)
print(f" attempt {i:>2} {status} wait {w:6.2f}s")
ordered = sorted(waits)
print(f" idle {sum(waits):.2f}s over {len(outcomes)} attempts; "
f"max {ordered[-1]:.2f}s, median {ordered[len(ordered) // 2]:.2f}s")
main()
budget: 260/hour total, 190 on the credential, 220 across hosts
ceilings: 8/min global, 5/min/host; 502 rows queued; 24 retries held back
min spend shop docs total cred left tot left lowest value
0 8 5 3 8 182 252 8
1 8 5 3 8 174 244 8
2 8 5 3 8 166 236 8
3 8 5 3 8 158 228 8
4 8 5 3 8 150 220 8
5 8 5 3 8 142 212 8
6 8 5 3 8 134 204 8
7 8 5 3 8 126 196 8
8 8 5 3 8 118 188 8
9 8 5 3 8 110 180 8
10 8 5 3 8 102 172 8
11 8 5 3 8 94 164 7
12 8 5 3 8 86 156 7
13 8 5 3 8 78 148 7
14 8 5 3 8 70 140 7
15 8 5 3 8 62 132 7
16 8 5 3 8 54 124 7
17 8 5 3 8 46 116 7
18 8 5 3 8 38 108 7
19 8 5 3 8 30 100 7
20 8 5 3 8 22 92 7
21 8 5 3 8 14 84 7
22 8 5 3 8 6 76 6
23 6 5 1 6 0 70 6
24 0 0 0 0 0 70 -
ran 25 minutes, spent 190 of 260; still queued 312 rows
binding constraint: credential budget at 0, host budgets exhausted for ['docs.example.com']
retries reserved 24; six workers retrying a 429 spend that in 4 minutes
decorrelated jitter, seed 5, base 0.5s, cap 60s, against 429s
attempt 0 429 wait 1.12s
attempt 1 429 wait 2.63s
attempt 2 429 wait 6.37s
attempt 3 429 wait 18.04s
attempt 4 429 wait 40.18s
attempt 5 429 wait 60.00s
attempt 6 200 wait 0.00s
attempt 7 429 wait 5.71s
attempt 8 429 wait 8.24s
attempt 9 429 wait 23.34s
attempt 10 429 wait 45.62s
attempt 11 429 wait 60.00s
attempt 12 200 wait 0.00s
attempt 13 200 wait 0.00s
idle 271.26s over 14 attempts; max 60.00s, median 8.24s
Read the plan top to bottom and the design is visible in it. The first minutes are full because high-value rows exist. From minute 22 the credential budget starts thinning the plan, and at minute 23 the docs host's own 70-request budget runs out mid-minute, so that minute spends 6 instead of 8. Minute 24 spends nothing, and the run stops with 70 tokens unspent on shop.example.com because spending them would have breached the credential budget. That last behavior -- holding tokens back rather than using them because a different scope is exhausted -- is the whole reason budgets are tracked per scope.
The lowest value column is the priority walk: the scheduler descends through value tiers as the hour fills, so the rows you drop are the ones you would have dropped anyway.
Queue Topology
Four things about the queue that matter more than the rate.
- Partition by identity, not by URL. One queue per profile means one worker always holds one identity, which is what makes per-identity rate limiting expressible and what keeps a ban from spreading across the fleet.
- Fair-share between hosts. Round-robin across host queues, or a deficit round robin weighted by what each host is worth. A single FIFO across all targets means one slow host starves every other, and starving a fast host to serve a slow one is a bad trade.
- Durable and idempotent. The queue outlives the workers, because the workers are the thing that crashes. An item carries its identity, its attempt count and its deadline, and a worker that dies mid-item leaves it requeueable rather than lost.
- Dead-letter with a reason. An item that failed three times leaves the queue into a dead-letter bucket with its classified failure. That bucket is your detector: when 200 items land there in a minute, you have a ban, and the classification tells you which kind.
Resolving the Egress Question
Two of the constraints in a budget are not really about the target: your own egress and your proxy gateway.
Your own egress. The DNS resolution question matters more than people expect. If a large share of your lookups go to a resolver you do not control, you have given away a correlation point, and dns resolver signals covers what a site can learn from it. Resolve through a resolver you operate, on a schedule, cache the answers for their TTL, and make the resolver's query pattern for your own domains look like any other resolver's.
Your gateway. Residential providers publish a per-gateway concurrency allowance, usually 20 to 50, and exceeding it produces a 429 from your supplier that is indistinguishable from the target's. Log which side produced every 429 -- the header is usually there, the connection's TLS session is not -- or you will spend a day blaming the target for your supplier's limit.
Bandwidth is the third: if each page is 41 KB and you budget 260 per hour, your ceiling is 10.4 MB per hour, which is not a number that needs optimising unless you have an expensive egress plan. Know it anyway.
A Working Budget
Defaults, in the order you should set them.
- Per-host rate at 60-70 percent of measured capacity, from rate limit discovery. Below that if the identity is young.
- Retry reserve of 8-12 percent of the plan, held out of the hourly budget, spent only on 429 and 5xx.
- Concurrency cap from the gateway's published allowance, per host and globally, never raised on latency.
- Per-credential daily cap set from the account's age: low for a 20-day account, and the thing you should be most conservative about because it protects the asset that took longest to build.
- Session rhythm, so the hourly spend has a human shape: active windows, quiet middles, one long idle block, matching the ledger model in risk score and session profiles.
- A hard stop when the budget is exhausted. Finishing early is a feature; overrunning is the failure mode the whole mechanism exists to prevent.
Then measure the outcome against the objective rather than against the rate: success ratio, challenge rate, and cost per delivered row, defined in observability and ban SLOS. A budget that meets its rate and misses its ban rate is not working, however well the plan looks.