Cookie Jars, Storage State and Session Persistence
A Cookie Jar Is Not a Device
You can persist a session perfectly and still show up as a stranger, because what you save is not what the server is holding. Cookies carry a scope, a lifetime and a partition; storage carries state the jar does not know about; and the server keeps its own view of who you are on the other side. Hand it a file called cookies.json from last Tuesday and it will correctly conclude that a brand-new machine has appeared with a stolen session token.
Three Layers, and Only One of Them Is the Jar
A persisted session is three separate things, and most confusion comes from treating them as one.
- The cookie store. Name/value pairs with scope and lifetime. This is what you can export, and it is the smallest layer by far.
- The storage state. Cookies plus
localStorage,sessionStorage, IndexedDB databases, and the service-worker registrations that Playwright'sstorage_statesnapshot serialises. One JSON file, several kilobytes to a couple of megabytes, and the only part a naivecookies.jsonkeeps. - Device state. The profile directory behind them: the HTTP cache, the HSTS table, the client certificate store, the permissions grants, the Service Worker caches, the font cache, the visited-links database. Not exportable, not portable, and the layer that decides whether you look like a machine that has been here before.
A jar replay restores the first layer perfectly and the second partially, and says nothing about the third. That is why a replayed session frequently works for two requests and then fails: the tokens are valid, but nothing behind them corroborates them. Browser pool architecture covers the third layer, because the pool is where device state lives and where it decides whether an identity ages.
A Cookie Is an Addressed Record, Not a Pair
RFC 6265 stores each cookie with a tuple of attributes that decide, independently, whether it is sent, when it expires, and whether script can see it. The two that people get wrong most often are Domain and Path.
Domain has two modes. A host-only cookie, written without a leading dot, goes back only to the exact host that set it. A domain cookie, written as .example.com, matches example.com and every subdomain. The match is a suffix test on a dot boundary, not a substring test, which is why notexample.com does not receive a cookie scoped to .example.com but a naive endswith("example.com") in your own code would happily send it there. Subdomain takeover is the other reason this matters: a host-only cookie on shop.example.com survives someone pointing shop.example.com at a different origin, and a domain cookie does not.
Path is a prefix match with a boundary rule. /cart matches /cart and /cart/items, and does not match /carthage. This is what lets one host carry two cookies with the same name for different sections, and it is why "the cookie is set, why is it not sent" is usually a path answer rather than an expiry answer.
The Attributes That Silently Remove a Cookie
Four more decide whether a record that is present ever reaches the server.
Secure is the one that breaks scrapers. Any Secure cookie is withheld from a plain http:// request. If your base URL is http:// and your client follows a redirect to https:// only to later replay against http:// again, the jar looks full and the wire is empty. One request in the matcher below returns exactly one cookie out of seven for precisely this reason.
HttpOnly hides the cookie from document.cookie. It has no effect on whether it is sent, which is a common misconception: HttpOnly is a defence against exfiltration, not against the server. If your own extraction code reads session tokens out of document.cookie and comes up empty, that is HttpOnly working.
SameSite controls whether a cookie rides along on a cross-site request, and its interaction with your flow topology is the reason captchas in session flows can break in a way that looks like a token problem. Lax sends the cookie on top-level GET navigations but not on subresource loads or cross-site POSTs. None requires Secure and is what a third-party context needs. Strict withholds the cookie on any cross-site arrival, which means a visitor who arrives from a link in an email or a chat app lands with no session at all and is re-authenticated. If you replay a session by loading a deep link first, Strict cookies are absent on that first request and the server decides you are new.
Expires and Max-Age decide the end. A cookie with neither is a session cookie that lives until the browser closes, and reloading the process to "start clean" is a fingerprint event in itself. Max-Age wins over Expires when both are present, which matters when you are rewriting a cookie's lifetime on the way through: modify the wrong one and the server still expires it on the original schedule.
In-Memory, Persisted, and the Snapshot
Three ways to hold a session, and they are not interchangeable.
| Approach | Survives a crash | Survives a restart | Sees a returning device | Cost |
|---|---|---|---|---|
| In-memory jar | no | no | no | free |
| Persisted jar file | yes | yes, if you reload it | no | one file read per worker |
storage_state snapshot |
yes | yes, if you reload it | no | a context launch per load |
| Persistent browser profile | yes | yes | yes | 100-400 MB of disk per identity |
The persistent profile is the only one that produces a device the target has met before. Everything else produces a device with a plausible history, which is a weaker claim. In Playwright, storage_state is the snapshot: cookies plus origins' local storage, written by context.storage_state() and replayed by browser.new_context(storage_state=...). It is the right tool for a stateless HTTP client that needs cookies and nothing else, and the wrong tool when the target's own script reads something you did not capture.
Two failure modes live here. Capturing at the wrong moment produces a snapshot taken during a challenge, which contains the pre-clearance cookies and no clearance cookie. Capturing too often produces a snapshot whose history never advances, so the identity is always exactly one step behind itself. Capture once per completed flow, at the point where the session is genuinely valid, and store it under the profile's own name.
Partitioned Cookies Will Break Your Reuse
CHIPS partitioned a cookie by the top-level site, so the same name can hold three different values in three different embedding contexts. A session cookie partitioned under https://news.example is not sent on a request whose top-level document is shop.example.com, even though the domain, path and expiry all match.
This breaks naive reuse in a specific and confusing way. If you captured a state inside an iframe, in a context you drive directly, then replay it top-level, every partitioned cookie silently disappears and your session is gone. The header looks short rather than wrong; nothing errors. The matcher below shows one cookie dropped for partition-mismatch on a request that matches on every other axis.
The same partitioning shows up as third-party cookie blocking, which is a browser policy rather than an attribute, and it is the reason storage-and-supercookie-probes matters: the sites that still identify you reliably after third-party cookies are gone are the ones fingerprinting through storage, and a jar that captures nothing else gives those sites a free pass. When you load a stored state, load it in the same context topology it was captured in, and record that topology alongside the snapshot.
Ordering Is Specified, and Servers Parse It
RFC 6265 §5.4 sorts the accepted cookies: longer Path first, then earlier creation time. Chromium honours that order by default and lets a Cookie-Priority attribute move entries around it, with Low demoted and High promoted.
Order matters because the first occurrence of a duplicated name wins in most server frameworks. Send /cart's sid before /'s sid and a server reading cookies.get('sid') gets the section-scoped value. Send them in the wrong order and it gets the root value, which may belong to a different session on the same host. The result is a request that is internally inconsistent rather than obviously wrong, which is the hardest class of bug to notice from a status code.
Create time is the tiebreaker, and it is the reason restoring a jar must preserve insertion order. http.cookiejar and most file formats do not round-trip creation time reliably, so a restored jar may reorder two same-path cookies that were genuinely distinct. Capture the order, and if your format cannot carry it, record the name list separately and re-set the cookies in the recorded order on load.
Clearing and Overwriting Are Different Signals
This is the part that separates a returning visitor from a fresh one, and it is a server-visible difference in how the request is shaped.
Overwriting means the jar already holds sid and you set it again. Most implementations replace the value in place, keeping the original creation time, and the server sees a normal request with a session it recognises.
Clearing means deleting the record and letting it be reissued, or setting it to an empty value with a past expiry. The server sees a request with no session cookie, then a request with a new session id issued seconds later. That pairing is a login event as far as most session layers are concerned: a cookie that vanishes and reappears is a session rollover. Session rollovers on an account that has never logged out are exactly the pattern that account warmup and identity lifecycle exists to avoid.
Third state: overwriting with a widened scope. If you take a host-only cookie captured on shop.example.com and write it back as .example.com, every subdomain now sends a session token that was never scoped for it. If any of those subdomains are third-party-embedded, you have just leaked a session cookie into contexts you do not control, and a SameSite=None cookie on a domain you do not fully own will do it silently.
What Replaying a Stale Jar Costs You
Four costs, in ascending order of how long they take to notice.
- A rejected replay, cheaply. The server notices the token is unknown or the signature does not verify, and returns 401. This is the good case.
- A silent re-authentication. The server issues a fresh anonymous session, keeps serving 200, and your extractor now writes an empty result set. Status-code monitoring sees nothing wrong. This is the expensive case, and it is why observability and ban SLOS validates content and not just codes.
- Session fixation, against you. If the app adopts the client-supplied session id, your stale jar is now a valid session that an attacker who logged it could use. This is the vulnerability a good session layer prevents, and you should not count on every target doing so.
- A consumed one-time token. Clearance cookies, CSRF tokens, and CAPTCHA responses are frequently single-use. Replaying a jar that carries a spent token gets you a token-reuse rejection, and if the vendor counts reuses against the site key's reputation, it degrades the clearance for every other session you hold.
None of these are visible from the status code alone, which is the whole argument for validating the body shape after every restore.
Matching a Jar Against a Request
Here is the rule set, implemented against the store and requests from the matcher discussion above. It is deliberately small: domain match, path match, expiry, Secure, partition key, then the §5.4 ordering. Everything it prints is derived from the inline store, and the last three lines are the failure mode in one line -- replaying the raw store sends eight cookies where the specification sends five.
# RFC 6265 selection: given a cookie store, what does each request send?
# Fixed clock, inline store, no network. Every rule below is quoted from the RFC.
NOW = 1767225600 # 2026-01-01T00:00:00Z. Nothing here reads the real clock.
COOKIES = [
{"label": "sid", "value": "a1b2", "domain": "shop.example.com", "path": "/",
"secure": True, "http_only": True, "same_site": "Lax",
"expires": NOW + 2592000, "created": 10, "partitioned": False, "key": None},
{"label": "sid@/cart", "value": "c3d4", "domain": "shop.example.com",
"path": "/cart", "secure": True, "http_only": True, "same_site": "Lax",
"expires": NOW + 2592000, "created": 40, "partitioned": False, "key": None},
{"label": "csrf", "value": "t9", "domain": "shop.example.com", "path": "/",
"secure": False, "http_only": False, "same_site": "Strict",
"expires": NOW + 3600, "created": 12, "partitioned": False, "key": None},
{"label": "consent", "value": "eu", "domain": ".example.com", "path": "/",
"secure": True, "http_only": True, "same_site": "None",
"expires": NOW + 15552000, "created": 3, "partitioned": False, "key": None},
{"label": "tier", "value": "platinum", "domain": "example.com", "path": "/",
"secure": True, "http_only": False, "same_site": "Lax",
"expires": NOW + 31536000, "created": 1, "partitioned": False, "key": None},
{"label": "psid", "value": "p1", "domain": "shop.example.com", "path": "/",
"secure": True, "http_only": True, "same_site": "None",
"expires": NOW + 2592000, "created": 60, "partitioned": True,
"key": "https://news.example"},
{"label": "legacy", "value": "stale", "domain": "shop.example.com", "path": "/",
"secure": False, "http_only": False, "same_site": "Lax",
"expires": NOW - 86400, "created": 5, "partitioned": False, "key": None},
{"label": "_ga", "value": "GA1.2.9", "domain": "partner.net", "path": "/",
"secure": False, "http_only": False, "same_site": "Lax",
"expires": NOW + 63072000, "created": 2, "partitioned": False, "key": None},
]
REQUESTS = [
("https://shop.example.com/cart", "https://shop.example.com"),
("http://shop.example.com/checkout/step2", "https://shop.example.com"),
("https://shop.example.com/checkout/step2", "https://news.example"),
("https://img.partner.net/i.png", "https://news.example"),
("https://tier.partner.net/t.js", "https://shop.example.com"),
]
def domain_match(host, domain):
# 5.1.3: identical, or a suffix match on a dot boundary
if host == domain:
return True
stripped = domain[1:] if domain.startswith(".") else domain
return host == stripped or host.endswith("." + stripped)
def path_match(request_path, cookie_path):
# 5.1.4: exact, prefix ending in "/", or prefix cut on a "/" boundary
if cookie_path == request_path:
return True
if not request_path.startswith(cookie_path):
return False
if cookie_path.endswith("/"):
return True
return len(request_path) > len(cookie_path) \
and request_path[len(cookie_path)] == "/"
def split(url):
scheme, rest = url.split("://", 1)
host, path = rest.split("/", 1)
return scheme, host, "/" + path
def select(url, top_site):
scheme, host, path = split(url)
sent, dropped = [], {}
for c in COOKIES:
if c["expires"] is not None and c["expires"] <= NOW:
reason = "expired"
elif c["secure"] and scheme != "https":
reason = "secure-flag"
elif not domain_match(host, c["domain"]):
reason = "domain-mismatch"
elif not path_match(path, c["path"]):
reason = "path-mismatch"
elif c["partitioned"] and c["key"] != top_site:
reason = "partition-mismatch"
else:
sent.append(c)
continue
dropped[c["label"]] = reason
# 5.4: longer path first, then older creation time first
sent.sort(key=lambda c: (-len(c["path"]), c["created"]))
return sent, dropped
def main():
for url, top in REQUESTS:
sent, dropped = select(url, top)
header = "; ".join(f"{c['label']}={c['value']}" for c in sent)
print(f"{url} top={top}")
print(f" send ({len(sent)}): {header or '(no Cookie header)'}")
for label in sorted(dropped):
print(f" drop {label}: {dropped[label]}")
print()
# What a naive reuse does: replay the whole store, order of storage ignored.
print("naive replay of the raw store, ignoring every rule:")
naive = "; ".join(f"{c['label']}={c['value']}" for c in COOKIES)
print(f" {naive}")
print(f" {len(COOKIES)} cookies sent to https://shop.example.com/cart "
f"instead of {len(select(REQUESTS[0][0], REQUESTS[0][1])[0])}")
main()
https://shop.example.com/cart top=https://shop.example.com
send (5): sid@/cart=c3d4; tier=platinum; consent=eu; sid=a1b2; csrf=t9
drop _ga: domain-mismatch
drop legacy: expired
drop psid: partition-mismatch
http://shop.example.com/checkout/step2 top=https://shop.example.com
send (1): csrf=t9
drop _ga: domain-mismatch
drop consent: secure-flag
drop legacy: expired
drop psid: secure-flag
drop sid: secure-flag
drop sid@/cart: secure-flag
drop tier: secure-flag
https://shop.example.com/checkout/step2 top=https://news.example
send (5): tier=platinum; consent=eu; sid=a1b2; csrf=t9; psid=p1
drop _ga: domain-mismatch
drop legacy: expired
drop sid@/cart: path-mismatch
https://img.partner.net/i.png top=https://news.example
send (1): _ga=GA1.2.9
drop consent: domain-mismatch
drop csrf: domain-mismatch
drop legacy: expired
drop psid: domain-mismatch
drop sid: domain-mismatch
drop sid@/cart: domain-mismatch
drop tier: domain-mismatch
https://tier.partner.net/t.js top=https://shop.example.com
send (1): _ga=GA1.2.9
drop consent: domain-mismatch
drop csrf: domain-mismatch
drop legacy: expired
drop psid: domain-mismatch
drop sid: domain-mismatch
drop sid@/cart: domain-mismatch
drop tier: domain-mismatch
naive replay of the raw store, ignoring every rule:
sid=a1b2; sid@/cart=c3d4; csrf=t9; consent=eu; tier=platinum; psid=p1; legacy=stale; _ga=GA1.2.9
8 cookies sent to https://shop.example.com/cart instead of 5
Storage State Beyond Cookies
Cookies are the part that is easy to serialise and the part that matters least to a sophisticated detector. What else the browser holds:
localStorageis per-origin and persists. Device-scoped counters, feature flags, an app's own "seen the tour" marker, and occasionally a session mirror. The storage and supercookie probes lesson covers what a detector reads out of here.sessionStorageis per-origin and per-tab, and dies with the tab. It is the right place for anything that must not survive, and if your snapshot captures it you have probably mis-scoped your identity.- IndexedDB is per-origin, structured, and frequently holds the application's real state: drafts, cached lists, an offline queue. A
storage_stateround trip carries it, which means a restored context can render a page your client never actually fetched, and the network tab will look suspiciously empty. - Cache API and Service Worker registrations are not in the snapshot but they are in a persistent profile. A service worker intercepting your own navigation is both a functional hazard and a strong signal.
- HSTS and the client certificate store are HTTP-state, not DOM-state, and only a real profile keeps them. Forcing TLS on a host that a fresh profile has never seen is a small but real difference, and hardening Playwright with CDP covers the protocol-layer work where that kind of value has to be set.
Making Persistence a Policy, Not a Habit
Rules that survive a bad week.
- One file per identity, named by the profile. Never a shared
cookies.json; two workers writing one file produces a jar that belongs to no one. - Store the context topology with the snapshot. Top-level versus iframe, and the top-level site for partitioned cookies. A snapshot without it is not replayable.
- Monotonic history, never reset. Storage only grows and only changes forward. A profile whose history went backwards is a signal more interesting than any single field in it.
- Re-capture on schedule, not on failure. A retry that reloads a snapshot from two hours ago because that is what was on disk is how a stale jar gets replayed at scale.
- Validate after restore, every time. One request, one assertion that the response is the page you expected. A restore that is not validated is an outage that has not happened yet.
- Treat the jar as a credential. It authenticates. Encrypt it, restrict it, log access, and never paste it into a shared channel.
The discipline that ties this together is that persistence is a claim you are making to the server: this is the same device, it has been here before, and here is the evidence. Every attribute in the RFC is a way that claim can be contradicted, and every one of them is checked in about eighty lines.