Browser Pool Architecture
A Browser Context Is a Half-Gigabyte Lease
A Chromium context is not a function call. It is a renderer process, a GPU process share, a font and audio stack, a network context and a warm JIT, and it occupies 300-700 MB of resident memory whether it is navigating or idle. Multiply that by a fleet and the browser stops being part of your scraper and becomes the reason your scraper will not fit on the machine.
The answer is a pool: keep contexts warm, hand them out, take them back, and throw them away on a policy rather than on a crash. The design is unglamorous and it is where most of the operational reliability lives.
The Cost Model
Get these four numbers into your head, because every sizing decision follows from them:
| Quantity | Typical value | Why it matters |
|---|---|---|
| RSS per context | 300-700 MB | sets how many contexts a machine can hold |
| Cold start to first paint | 1-3 s | the latency you pay per job without a pool |
| Renderer processes per origin | 1 per site isolation bucket | a page fanning out to ten origins costs ten processes |
| CPU per active challenge | 0.2-1.0 core at peak | challenges are the real CPU cost, not navigation |
A 16 GB machine holds roughly 20-25 warm contexts with headroom, not 16, because the browser process, the GPU process and the OS page cache all need room. Cold start matters less than it looks, because a cold start is not just latency: it is a signature. A renderer that initialises, navigates and completes a challenge in 900 ms with no prior history is a behavioural outlier, and warm-up is as much about looking continuous as about speed.
Pool Anatomy
Five components, and it is worth naming them separately because they fail separately:
- The manager. A thread-safe structure holding a queue of idle contexts and a registry of everything alive.
acquire()andrelease()are the only two operations most code touches. - The profile factory. Turns a stored bundle -- fingerprint, locale, timezone, storage state, proxy lease -- into a context. This is where device profile consistency is enforced, and it is the only place a context's identity is decided.
- The warm-up routine.
about:blank, then a neutral page on the exit IP's own region, so fonts, GPU and audio initialise and one or two connections are already open when the target arrives. - The health check. A cheap assertion after each job: the context still answers, the expected media stack still initialises, the profile's key fields still read the values the bundle specified. Failures retire the context instead of returning it.
- The eviction policy. LRU plus a maximum age, so a context does not live forever and does not accumulate state that a detector can notice changing.
A pool that only holds contexts, with no health check, becomes a pool that quietly poisons itself. A context whose renderer wedged on a challenge will be handed to the next job and fail it, and you will spend a day blaming the target.
Warm-Up Is Not a No-Op
What "warm" means has to be defined per job, because a challenge that needs an AudioContext will still fail on a context where audio has never been instantiated. A defensible routine is: about:blank, one navigation to a first-party or neutral page, a call that forces font loading, a call that creates and closes an AudioContext, and a WebGL context creation and disposal. On a residential exit, that page should be a real site rather than example.com, so the connection you warm is to a host on the same network as the target.
The cost is 300-800 ms and a few hundred kilobytes per context, once. Skip it and every job pays 1-3 s plus a cold-renderer signature. Do it and the pool's value is both latency and plausibility, which is why it is worth the complexity even at modest scale.
State Reset and Its Fingerprint Cost
Between jobs you have a choice, and the choice is a fingerprint decision, not a performance one:
| Policy | Speed | What the site sees |
|---|---|---|
| Full reset (new context) | fastest | a brand-new device with no history, every time |
| Keep storage, drop cookies | medium | a device that keeps localStorage but loses its session |
| Keep everything | slowest | a returning device with accumulating history |
Full reset is what Playwright's new_context() gives you by default, and at scale it is the most detectable of the three. Thousands of sessions that each present a pristine profile with a plausible history elsewhere are not a population; they are one fingerprint repeated, and repetition at volume is the signal anomaly models find first. Keeping storage and history is slower and produces a device that ages, which is what a real device does.
The practical middle: keep a per-profile storage state on disk, reload it, and let a small share of jobs be first visits for a genuinely new profile. The cookie jar and storage persistence lesson covers the mechanics; the discipline is that a profile's history must be monotonic -- never cleared, never reset to empty, only added to.
Four Concurrency Limits
A single global limit is a mistake, because the constraint that matters differs by scope and exceeding the wrong one is what gets you banned:
- Per host. The target's own tolerance, learned from
Retry-Afterand 429 behaviour. This is the limit that protects you. - Per profile. How many jobs one identity may have in flight. Two concurrent jobs on one profile means two tabs with the same fingerprint hitting the site at once, which is a shared-device pattern some detectors model explicitly.
- Per exit IP. A residential exit shared by six contexts is six times the request rate on one address, and the IP reputation cost lands on the address, not the context.
- Global. Your own CPU and memory ceiling. This is the one that should be derived from
sizein the pool, and it is the one you enforce last, as a backstop.
Set per-host from measurement, per-profile to 1 or 2, per-exit from the provider's own concurrency allowance (ask; most residential gateways publish one), and global from the pool size you can afford to hold in RAM.
Memory Budgeting and the OOM Killer
The failure mode is not a clean exception. Linux's OOM killer chooses a process to kill, and on a container host it frequently chooses the process with the largest RSS, which is your browser. The job dies with no traceback, the pool loses contexts it thought were healthy, and the supervisor restarts the browser into a machine that is still out of memory, so it happens again.
Three defences, in order of value:
- A hard headroom line. Set the pool's
sizeso that browser RSS plus the scraper's own footprint plus 2 GB of slack fits in physical memory. Compute it from the measured 300-700 MB per context, not from the theoretical maximum. - A memory ceiling in the cgroup.
memory.maxon the container turns a host-wide OOM into a cgroup OOM you can observe, and often into a clean shutdown of the biggest child rather than the parent. - A restart policy that drains. When the supervisor sees the browser process gone, it must requeue in-flight jobs rather than assume the pool is still valid, and it must refuse to rebuild the pool until memory has actually been reclaimed.
Watch RSS as a first-class metric, not as a diagnostic. A pool that drifts upward over hours is leaking, and it will take the machine with it on the worst possible night.
Eviction and Shutdown
LRU alone is wrong for fingerprinting. A context that has served two hundred jobs is not a context you want to keep, and neither is one that has been idle for six hours while the target's TLS tickets have rotated. Use three rules together: retire on a job count (retire_after), retire on a wall-clock age (max_age), and evict the least recently used when the idle count exceeds max_idle. The idle floor (min_idle) is separate and is what absorbs a burst: you pre-warm N contexts so the first requests of a busy minute do not each pay 2 seconds of cold start.
Graceful shutdown is the same discipline applied to the end of a run. Drain in-flight jobs, persist the storage state you want to keep, release proxy leases, close contexts in an orderly way, then close the browser. The signal to watch for is a shutdown that takes longer than your orchestrator's grace period, which is usually a context that is mid-navigation and will not stop on request.
Failure Handling
A browser crash must be a local event. The contract is: every context belonging to a dead browser is dead, the jobs that were using it are requeued with their backoff, the queue itself is untouched, and the supervisor rebuilds only what it lost. The two mistakes that turn a crash into an outage are requeueing the whole queue (which loses the ordering and the rate discipline) and rebuilding the full pool immediately (which re-creates the memory pressure that may have caused the crash).
Per-profile isolation is the same idea one level up: a profile that has drawn three challenges in a row is moved to the cooling tier and its in-flight work is requeued to maturing profiles, not retried on the same identity. And the per-host circuit breaker belongs here too -- once a host has returned enough challenges in a window, the pool stops handing out contexts for that host until the window expires. Fighting a confirmed detection burns the IP and, more importantly, teaches the target nothing about your rate limits.
Sizing a Pool for a Target
Take a concrete job: two pages per second sustained, where each job means navigating a product page, waiting for a challenge to clear, and extracting.
A warmed context handles about 2.5 s per navigation including the challenge, and a job is four navigations plus a 1.5 s warm-up, so a job occupies a context for roughly 11.5 s. At two jobs per second you need 2 x 11.5 = 23 contexts resident at all times, not 23 average. At 450 MB each that is about 10 GB of browser, which fits one 16 GB machine with headroom.
Now the same job at six pages per second: 69 contexts, 31 GB, and no single machine that also has room for anything else. Three machines behind the shared queue from distributed worker fleets handles it, or you change the architecture. The better change is the one the production pipeline already uses: let the browser solve the challenge once, harvest the clearance cookie, and do the 99 percent of requests through curl_cffi on a fast HTTP client. That moves the browser's job from "two pages per second" to "one solve every twenty minutes", and the pool shrinks to a handful of contexts per IP.
The lesson from the arithmetic is that throughput targets should be expressed in context-seconds, not pages. Once you price a job in context-seconds, the browser's cost stops being a mystery and becomes a line in the capacity plan.
import contextlib
import queue
import random
import threading
import time
WARMUP = 0.002 # real seconds spent standing in for new_context()
JOB_TIME = 0.008 # real seconds spent standing in for a navigation
TICK = 0.5 # logical seconds one job advances the pool's clock
MAX_AGE = 4.0 # logical seconds a context may live
PROFILES = ["profile-0", "profile-1", "profile-2"]
class Clock:
'''A logical clock: pool decisions must not depend on real elapsed time.'''
def __init__(self):
self.now = 0.0
self.lock = threading.Lock()
def advance(self, delta):
with self.lock:
self.now += delta
return self.now
class Context:
'''Stands in for a warmed browser context: a profile bundle plus a live renderer.'''
def __init__(self, index, profile, born):
self.index = index
self.profile = profile
self.born = born
self.last_used = born
self.jobs = 0
class WarmPool:
'''One pool per fingerprint profile: a context never serves a different profile.'''
def __init__(self, profile, factory, stats,
size=4, min_idle=1, max_idle=2, retire_after=3):
self.profile = profile
self.factory = factory
self.stats = stats
self.clock = Clock() # one logical clock per pool
self.size = size
self.min_idle = min_idle
self.max_idle = max_idle
self.retire_after = retire_after
self._idle = queue.Queue()
self._live = []
self._lock = threading.Lock()
self._counter = 0
def _create(self):
with self._lock:
self._counter += 1
index = self._counter
born = self.clock.now
ctx = self.factory(index, self.profile, born)
with self._lock:
self._live.append(ctx)
self.stats["created"] += 1
return ctx
def _forget(self, ctx, reason):
with self._lock:
if ctx in self._live:
self._live.remove(ctx)
self.stats[reason] += 1
def _sweep(self):
'''Retire aged contexts, close idle ones over max_idle (LRU), then top up.'''
now = self.clock.now
idle = []
while True:
try:
idle.append(self._idle.get_nowait())
except queue.Empty:
break
fresh = []
for ctx in idle:
if now - ctx.born > MAX_AGE:
self._forget(ctx, "retired_on_age")
elif ctx.jobs >= self.retire_after:
self._forget(ctx, "retired_after_jobs")
else:
fresh.append(ctx)
fresh.sort(key=lambda c: c.last_used) # oldest first
while len(fresh) > self.max_idle:
self._forget(fresh.pop(0), "evicted_lru")
for ctx in fresh:
self._idle.put(ctx)
live = len(self._live)
while len(fresh) < self.min_idle and live < self.size:
self._idle.put(self._create())
fresh.append(True)
live += 1
@contextlib.contextmanager
def acquire(self, timeout=2.0):
try:
ctx = self._idle.get_nowait()
except queue.Empty:
self.stats["acquires_waited"] += 1
with self._lock:
room = len(self._live) < self.size
ctx = self._create() if room else self._idle.get(timeout=timeout)
try:
yield ctx
finally:
ctx.jobs += 1
ctx.last_used = self.clock.advance(TICK)
if ctx.index % 5 == 4: # failed health probe
self._forget(ctx, "health_failures")
else:
self._idle.put(ctx)
self._sweep()
def gauge(self):
idle = self._idle.qsize()
with self._lock:
live = len(self._live)
return idle, live - idle, live
def make_context(index, profile, born):
time.sleep(WARMUP) # new_context() plus the warm-up navigation
return Context(index, profile, born)
def worker(pool, tid, jobs, totals):
for n in jobs:
for attempt in range(3):
with pool.acquire() as ctx:
if ctx.profile != pool.profile:
totals["profile_mixups"] += 1
time.sleep(JOB_TIME) # one navigation
ok = random.Random(hash((tid, n, attempt))).random() > 0.10
if ok:
break
else:
totals["exhausted"] += 1
def main():
totals = {"profile_mixups": 0, "exhausted": 0}
stats = {"created": 0, "health_failures": 0, "retired_on_age": 0,
"retired_after_jobs": 0, "evicted_lru": 0, "acquires_waited": 0}
pools = {p: WarmPool(p, make_context, stats) for p in PROFILES}
for pool in pools.values():
for _ in range(pool.size): # pre-warm the whole pool
pool._idle.put(pool._create())
pool._sweep() # trim to max_idle by LRU
per_profile = 15
crew = [threading.Thread(target=worker, args=(pools[p], t, list(range(per_profile)), totals))
for t, p in enumerate(PROFILES)] # one thread per profile pool
for t in crew:
t.start()
for t in crew:
t.join()
for pool in pools.values():
pool._sweep()
for pool in pools.values(): # the fleet then idles for an hour
pool.clock.advance(MAX_AGE + TICK)
for pool in pools.values():
pool._sweep() # the age rule reclaims the warm contexts
first = pools["profile-0"]
print("warm pool: {} profiles x (size={} idle={}..{} retire_after={} jobs max_age={}s)".format(
len(PROFILES), first.size, first.min_idle, first.max_idle, first.retire_after, MAX_AGE))
print("workload: {} threads (one per profile) x {} jobs = {} jobs".format(
len(PROFILES), per_profile, len(PROFILES) * per_profile))
print()
print("{:<24} {}".format("METRIC", "VALUE"))
print("-" * 42)
for name, value in [
("jobs", len(PROFILES) * per_profile),
("contexts_created", stats["created"]),
("contexts_live", sum(p.gauge()[2] for p in pools.values())),
("contexts_idle", sum(p.gauge()[0] for p in pools.values())),
("contexts_in_use", sum(p.gauge()[1] for p in pools.values())),
("health_failures", stats["health_failures"]),
("retired_on_age", stats["retired_on_age"]),
("retired_after_jobs", stats["retired_after_jobs"]),
("evicted_lru", stats["evicted_lru"]),
("acquires_waited", stats["acquires_waited"]),
("profile_mixups", totals["profile_mixups"]),
("jobs_exhausted", totals["exhausted"])]:
print("{:<24} {}".format(name, value))
print()
print("contexts per job: {:.2f} (a cold browser per job would be 1.00)".format(
stats["created"] / float(len(PROFILES) * per_profile)))
print("idle floor held: {}".format(
all(p.gauge()[0] >= p.min_idle for p in pools.values())))
print("no context crossed profile: {}".format(totals["profile_mixups"] == 0))
if __name__ == "__main__":
main()
warm pool: 3 profiles x (size=4 idle=1..2 retire_after=3 jobs max_age=4.0s)
workload: 3 threads (one per profile) x 15 jobs = 45 jobs
METRIC VALUE
------------------------------------------
jobs 45
contexts_created 30
contexts_live 3
contexts_idle 3
contexts_in_use 0
health_failures 4
retired_on_age 3
retired_after_jobs 14
evicted_lru 6
acquires_waited 0
profile_mixups 0
jobs_exhausted 0
contexts per job: 0.67 (a cold browser per job would be 1.00)
idle floor held: True
no context crossed profile: True
Forty-five jobs created thirty contexts, so 0.67 contexts per job where a naive implementation creates one per job and pays a cold start for each. Every churn counter is doing visible work: six contexts were evicted by LRU when the pre-warmed pool was trimmed from size to max_idle, fourteen were retired by the job count, four failed their health probe and were dropped rather than returned, and three were reclaimed by the age rule after the fleet idled. The two lines that must stay at zero are profile_mixups and jobs_exhausted: a pool that hands a worker a context carrying another profile's fingerprint has created a device that changes identity mid-session, and an exhausted retry means the pool could not serve work it had already accepted.
Note how the pool decides anything at all: a logical clock per pool advances one tick per job, so age retirement is a function of jobs served rather than of wall-clock timing, which is what makes the table above reproducible instead of a different set of numbers on every run. In production the same rules read a real clock, the numbers move, and that is exactly why these counters belong in a metrics table and in a dashboard rather than in an assertion.
Checklist
- A pool manager with
acquire()andrelease(), a registry of every live context, and metrics for idle, in-use, created, retired, evicted and failed. - The factory is the only place a context's identity is built, and contexts are never shared across profiles.
- Warm-up runs per context before first use: blank page, neutral page, fonts, audio, WebGL.
- A health check runs after every job, and a failure retires the context instead of returning it.
- Three eviction rules together: job count, wall-clock age, and LRU over a maximum idle count, with a separate warm floor.
- Per-host, per-profile, per-exit and global concurrency limits, all enforced, with per-host derived from measurement.
- A memory ceiling and a supervisor that requeues in-flight jobs on a browser crash instead of rebuilding the pool blindly.
- Graceful shutdown that drains, persists, releases leases and closes inside the orchestrator's grace period.
The Legitimate Route
A pool is an efficiency structure, and building one is only worth it for work you are entitled to do. The capacity arithmetic above is a useful filter on its own: if the number of contexts a legitimate, rate-respecting job needs will not fit on one machine, the answer is usually to slow down, use an official API or a partner feed, or negotiate access -- not to add machines until it does. Deliberately sizing a fleet to out-run a site's stated rate limit, rotating through fresh addresses to shed a block you were given, or keeping a banned identity alive by rebuilding its context is the line this course stops at, and none of the pool mechanics above change that. The crawl budgets and request scheduling lesson is the honest version of the same arithmetic: measure the tolerance, work inside it, and let the pool absorb bursts rather than manufacture them.