Common Crawl and Web Archives
Pages You Do Not Have to Fetch
Some of the pages you want have already been fetched by someone else. Common Crawl publishes a crawl of billions of pages roughly every month or two, free to download. The Internet Archive's Wayback Machine has stored snapshots of pages since the 1990s. Both expose an index you can query by URL, and both let you download a single stored page without touching the original site.
That changes what is possible. You can get pages from sites that no longer exist, see how a page looked in 2019, build a corpus across millions of domains without running a crawler, or seed a crawl with URLs the sitemap forgot. The search engine scraping lesson used the Common Crawl index to list a domain's URLs; this lesson goes the rest of the way to the stored pages themselves, and covers the Wayback Machine's API, change history, and the limits of both.
Archive or Live Crawl?
| Need | Live crawl | Common Crawl | Wayback Machine |
|---|---|---|---|
| Today's price or stock | Yes | No (weeks old at best) | No |
| Pages from many thousands of domains | Expensive | Built for it | Not meant for bulk |
| History of one page over years | Impossible | Samples from each crawl, indexed since 2013 | Often many snapshots |
| Sites that are gone or changed | Impossible | If it was crawled | If it was captured |
| Complete coverage of one site | Yes, if allowed | Partial sample | Partial, uneven |
| JavaScript-rendered content | With a browser | Raw HTML only | Mostly raw HTML |
| Load on the target site | Yours | None | None |
Use archives for breadth and history, and live crawling for freshness and completeness. They combine well: an archive gives you the URL inventory and a baseline, then a small live crawl refreshes only what you need (see incremental scraping).
How Common Crawl Is Organised
Each crawl has an id such as CC-MAIN-2026-39 (year and ISO week) and consists of three file types:
- WARC: the raw HTTP responses, headers and bodies. This is what you want for scraping.
- WAT: JSON metadata per page (HTTP headers, links, meta tags). Good for link graphs.
- WET: extracted plain text. Convenient, but crude; run your own article text extraction on WARC records if quality matters.
The files are hundreds of megabytes each, in the tens of thousands per crawl. You never download them whole for targeted work. Instead you look up a URL in the CDX index, which tells you which WARC file holds the capture and at which byte offset, and then download just those bytes with an HTTP range request.
Step 1: Query the CDX Index
The index server takes the same query style as other CDX servers:
| Parameter | Meaning |
|---|---|
url |
URL or pattern, e.g. example.com/* for a path prefix, *.example.com for subdomains |
matchType |
exact, prefix, host or domain instead of wildcards |
output=json |
One JSON object per line |
filter |
e.g. =status:200 (exact match) or mime:text/html; ! negates |
fl |
Return only these fields |
showNumPages=true, page=N |
Results are paged; ask how many pages, then fetch each |
import json
import time
import httpx
client = httpx.Client(timeout=120, headers={"User-Agent": "research-bot/1.0 (you@example.org)"})
def get_with_backoff(url, params, tries=6):
for attempt in range(tries):
r = client.get(url, params=params)
if r.status_code in (429, 500, 502, 503, 504):
time.sleep(min(2 ** attempt * 5, 300)) # 5, 10, 20, 40 ... seconds
continue
return r
r.raise_for_status()
def cc_captures(crawl, pattern):
api = f"https://index.commoncrawl.org/{crawl}-index"
base = {"url": pattern, "output": "json", "filter": ["=status:200", "mime:text/html"]}
pages = get_with_backoff(api, {**base, "showNumPages": "true"}).json()["pages"]
for page in range(pages):
r = get_with_backoff(api, {**base, "page": page})
if r.status_code == 404: # no captures on this page
continue
for line in r.text.splitlines():
if line.strip():
yield json.loads(line)
for cap in cc_captures("CC-MAIN-2026-39", "example.com/blog/*"):
print(cap["url"], cap["filename"], cap["offset"], cap["length"])
One index line looks like this (values shortened):
{"urlkey": "com,example)/blog/night-trains", "timestamp": "20260911083012",
"url": "https://example.com/blog/night-trains", "mime": "text/html", "status": "200",
"digest": "B2YHMQ...", "length": "18342", "offset": "412210988",
"filename": "crawl-data/CC-MAIN-2026-39/segments/.../warc/CC-MAIN-...-00123.warc.gz",
"languages": "eng", "encoding": "UTF-8"}
filename, offset and length are what you need for step 2. digest is a hash of the body, useful for spotting identical content across captures.
The backoff is not optional. The index server is a shared, heavily used service and regularly answers with 503 or 504 under load; a script without retries fails on its first busy minute. Keep one request at a time against the index, cache every answer on disk, and never re-query what you already have.
For questions across whole domains or languages ("every German-language page on these 5,000 hosts"), the per-URL API is the wrong tool. Common Crawl also publishes the index as Parquet files (the columnar index) under s3://commoncrawl/cc-index/table/cc-main/warc/, which you can query with SQL engines such as Amazon Athena, Spark or DuckDB:
SELECT url, warc_filename, warc_record_offset, warc_record_length
FROM ccindex
WHERE crawl = 'CC-MAIN-2026-39' AND subset = 'warc'
AND url_host_registered_domain = 'example.com'
AND fetch_status = 200 AND content_mime_type = 'text/html';
The result has the same filename, offset and length triples, ready for step 2.
Step 2: Range-Read One Record
Each record in a Common Crawl WARC file is compressed as its own gzip member. That is the trick that makes the whole system work: the bytes from offset to offset + length - 1 are a complete, independently decompressible file containing exactly one page.
from io import BytesIO
from warcio.archiveiterator import ArchiveIterator # pip install warcio
def fetch_range(filename, offset, length):
end = int(offset) + int(length) - 1
r = client.get(f"https://data.commoncrawl.org/{filename}",
headers={"Range": f"bytes={offset}-{end}"})
if r.status_code != 206: # 200 would mean the whole file
raise RuntimeError(f"range request failed: {r.status_code}")
return r.content
def parse_record(raw):
for rec in ArchiveIterator(BytesIO(raw)):
return {
"url": rec.rec_headers.get_header("WARC-Target-URI"),
"date": rec.rec_headers.get_header("WARC-Date"),
"status": rec.http_headers.get_statuscode(),
"type": rec.http_headers.get_header("Content-Type"),
"truncated": rec.rec_headers.get_header("WARC-Truncated"),
"body": rec.content_stream().read(),
}
Test the parser offline first. Write a small WARC with warcio's WARCWriter(f, gzip=True), note each record's offset and length (warcio's ArchiveIterator reports them via get_record_offset() and get_record_length()), then read one back exactly as the range request would:
with open("sample.warc.gz", "rb") as f:
f.seek(360) # offset of the second record
raw = f.read(418) # its length
print(raw[:2], len(raw))
rec = parse_record(raw)
print({k: v for k, v in rec.items() if k != "body"})
print(rec["body"].decode("utf-8"))
b'\x1f\x8b' 418
{'url': 'https://books.example.com/catalogue/a-light-in-the-attic_1000/', 'date': '2026-09-30T07:02:32Z',
'status': '200', 'type': 'text/html; charset=utf-8', 'truncated': None}
<html><title>A Light in the Attic</title><p class='price_color'>£51.77</p></html>
The slice starts with the gzip magic bytes 1f 8b, and warcio returns the original URL, capture time, HTTP status and body. From here the page is ordinary HTML: decode it using the Content-Type charset (see parsing engines and text encoding) and run your usual parser.
Two checks worth keeping:
WARC-Truncated. Crawlers cap how much of each response they store. If this header is set, the body is incomplete; do not treat a truncated PDF or long page as the real thing.- Status 206 versus 200. If a proxy or misconfigured client drops the
Rangeheader, you get a200and start downloading a file of hundreds of megabytes. Fail fast.
For bulk work on many records, sort them by filename and offset so consecutive reads hit the same file, and run a small number of concurrent range requests. If your compute runs in AWS us-east-1, reading from the commoncrawl S3 bucket directly is faster and avoids transfer through the public HTTPS endpoint.
The Wayback Machine CDX API
The Wayback Machine's index works the same way, with a few extra parameters that make it good for history:
def wayback_captures(url, start="2018", end="2026", collapse="digest"):
params = {
"url": url, "output": "json", "from": start, "to": end,
"fl": "timestamp,original,statuscode,digest",
"filter": "statuscode:200",
"collapse": collapse, # drop consecutive captures with the same body
}
r = get_with_backoff("https://web.archive.org/cdx/search/cdx", params)
rows = r.json() if r.text.strip() else []
if not rows:
return []
header, *data = rows # the first row is the field names
return [dict(zip(header, row)) for row in data]
def snapshot_url(cap):
return f"https://web.archive.org/web/{cap['timestamp']}id_/{cap['original']}"
The parts that are easy to get wrong:
- The JSON output is a list of lists whose first row is the header, not a list of objects. An empty result may be an empty body or
[]; handle both. collapse=digestremoves consecutive captures with identical content, so the result is a list of changes.len(wayback_captures(url))is roughly how many times the page changed in that period, which is exactly what you want for measuring change rates or reconstructing a price history.collapse=timestamp:8instead keeps one capture per day (the first 8 digits of the timestamp are the date).- The
id_suffix after the timestamp returns the original bytes as captured. Without it, the Wayback Machine rewrites every link to point back into the archive and injects its own toolbar, and your parser will extract archive URLs instead of the site's. - Timestamps are 14-digit UTC (
20240315092211). If you request a timestamp with no exact capture, you are redirected to the nearest one, which may be years away. Always read the timestamp of what you actually got. - Large result sets need
limitplusshowResumeKey=true; the response then ends with a resume key you pass back asresumeKeyto continue.
The Wayback Machine is a non-profit service with modest capacity and it enforces rate limits (429 Too Many Requests, or a temporary 503 page when it is struggling). Keep to a request every few seconds, back off hard on errors, and prefer Common Crawl for anything bulk.
What Archives Get Wrong
- Coverage is a sample. Neither archive has every page, and Common Crawl revisits only part of each site per crawl. Absence from the archive proves nothing about the site.
- Robots and exclusions. Common Crawl's crawler obeys robots.txt, and site owners can ask the Internet Archive to exclude their pages, so some sites are missing on purpose.
- Raw HTML only. Data loaded by JavaScript after the page rendered is usually missing, although the archived API responses sometimes exist as separate captures you can look up directly.
- Soft errors. A
200capture can still be a login wall, a cookie consent page or a bot challenge. Validate content before trusting it, just as you would with a live fetch. - Rights do not change. An archived copy is still the publisher's copyrighted work and may contain personal data. The same rules from legal considerations apply, plus the archive's own terms of use.
Practice
For a small site you are allowed to study, list its HTML captures in the two most recent Common Crawl crawls, compare the URL counts with its sitemap, and range-read five records to check status, truncation and encoding. Then use the Wayback CDX API with collapse=digest on its homepage for the last three years and plot how many distinct versions appeared per month.