Two Ways a Scraper Breaks

A scraper fails because you changed your code or because the site changed its pages. Regressions in your code are caught by fast, deterministic tests that run on every commit and never touch the network. Site changes are caught by a few live contract tests on a schedule. Mixing the two is the classic mistake: a unit test that fetches a real page fails whenever the site is slow, and a team that sees random red builds learns to ignore them.

Layout drift detection covers monitoring production runs and data cleaning and validation covers record-level checks; what follows is the pytest suite around them, using the scraping sandbox books.toscrape.com as the target.

A Test Layout That Scales

Layer What it proves Network When it runs
Parser tests on saved HTML Extraction logic is correct None Every commit
Crawl tests with mocked HTTP Pagination, retries and stopping rules work Mocked Every commit
Recorded-HTTP tests (cassettes) The full flow works on real responses Replayed Every commit
Live contract tests The site still matches your assumptions Real Daily, scheduled

The examples assume a package with scraper/parse.py (HTML to dicts, raising a LayoutDrift error when required elements are missing) and scraper/crawl.py, with tests under tests/ and saved pages in tests/fixtures/.

Save Fixtures as Bytes, With Their Story

A fixture is a real page frozen at a known moment. Three habits make fixtures trustworthy:

  • Save response.content, not text or prettified soup. Re-serialising through a parser changes whitespace and entities, so the test no longer sees what production sees. Raw bytes also keep encoding bugs reproducible: books.toscrape.com sends Content-Type: text/html with no charset, which httpx decodes as UTF-8 while requests falls back to ISO-8859-1 and turns £51.77 into £51.77. A fixture saved as decoded text hides exactly that bug.
  • Name fixtures by scenario: listing_last_page, product_out_of_stock, product_on_sale. Every production bug you fix adds a fixture reproducing it.
  • Record where it came from: URL, fetch time and content type in a sidecar file.
import json
from datetime import datetime, timezone
from pathlib import Path
import httpx

def save_fixture(url: str, name: str, folder: Path = Path("tests/fixtures")) -> None:
    r = httpx.get(url, headers={"User-Agent": "catalog-tests (you@example.com)"}, timeout=30)
    r.raise_for_status()
    (folder / f"{name}.html").write_bytes(r.content)
    meta = {"url": url, "fetched_at": datetime.now(timezone.utc).isoformat(timespec="seconds"),
            "content_type": r.headers.get("content-type")}
    (folder / f"{name}.meta.json").write_text(json.dumps(meta, indent=2))

Before committing a fixture from a logged-in or personalised page, remove names, emails, session tokens and CSRF values. Fixtures live in your repository forever.

Parser Tests: Golden Files Plus Targeted Facts

Two kinds of assertion complement each other. A golden file stores the parser's full output for a fixture and catches any change at all. Targeted facts state the few things you know must be true and explain themselves when they fail. The golden file is written on the first run and whenever you pass --update-golden, an option defined in conftest.py:

# tests/conftest.py
import pytest

def pytest_addoption(parser):
    parser.addoption("--update-golden", action="store_true", help="rewrite expected JSON")

@pytest.fixture
def update_golden(request):
    return request.config.getoption("--update-golden")
# tests/test_parse.py
import json
from decimal import Decimal
from pathlib import Path
import pytest
from scraper.parse import parse_listing, LayoutDrift

FIXTURES = Path(__file__).parent / "fixtures"
URL = "https://books.toscrape.com/catalogue/page-1.html"

def load(name):
    return (FIXTURES / name).read_bytes().decode("utf-8")   # same decode as production

@pytest.mark.parametrize("name", ["listing_page1", "listing_last_page"])
def test_listing_matches_golden(name, update_golden):
    actual = json.loads(json.dumps(parse_listing(load(f"{name}.html"), URL), default=str))
    golden = FIXTURES / f"{name}.expected.json"
    if update_golden or not golden.exists():
        golden.write_text(json.dumps(actual, indent=2, ensure_ascii=False), encoding="utf-8")
    expected = json.loads(golden.read_text(encoding="utf-8"))
    assert actual["next_url"] == expected["next_url"]
    assert len(actual["items"]) == len(expected["items"])
    for i, (got, want) in enumerate(zip(actual["items"], expected["items"])):
        assert got == want, f"item {i} differs"

def test_first_card_facts():
    first = parse_listing(load("listing_page1.html"), URL)["items"][0]
    assert first["title"] == "A Light in the Attic"
    assert first["price"] == Decimal("51.77")
    assert first["url"].endswith("/catalogue/a-light-in-the-attic_1000/index.html")

def test_renamed_class_raises_drift():
    html = load("listing_page1.html").replace("product_pod", "product-card")
    with pytest.raises(LayoutDrift, match="no product cards"):
        parse_listing(html, URL)

The item-by-item loop is deliberate. Comparing two large dicts in one assert produces a truncated diff you cannot read; comparing per item points at the exact record. Here is what the suite prints (abridged) after someone "simplifies" the parser to read the link text instead of its title attribute:

>           assert got == want, f"item {i} differs"
E           AssertionError: item 0 differs
E             Differing items:
E             {'title': 'A Light in the ...'} != {'title': 'A Light in the Attic'}

The listing truncates long titles in the visible text; only the attribute holds the full title. Fixture tests preserve this kind of site knowledge after everyone has forgotten it. The last test mutates a fixture to simulate a redesign and checks for a named error instead of empty records.

When you change the parser on purpose, run pytest --update-golden and read git diff tests/fixtures/ before committing; the diff shows exactly which records changed. Never regenerate golden files automatically in CI, or the test can never fail.

Schema Checks on Parser Output

Golden files only cover fixtures you have; a schema assertion states rules that must hold for any page:

from pydantic import BaseModel, ConfigDict, Field, TypeAdapter

class Item(BaseModel):
    model_config = ConfigDict(strict=True)
    title: str = Field(min_length=1)
    url: str = Field(pattern=r"^https://books\.toscrape\.com/")
    price: Decimal = Field(gt=0, lt=1000)
    in_stock: bool

def test_items_match_schema():
    items = parse_listing(load("listing_page1.html"), URL)["items"]
    TypeAdapter(list[Item]).validate_python(items)

Strict mode matters here: in lax mode the string "51.77" would pass as a Decimal, hiding a parser that forgot to convert.

Crawl Logic With Mocked HTTP

Pagination, retries and stop conditions are where crawlers hang or hammer a site, and they are easy to test once the network is replaced. For httpx, the respx library intercepts requests and returns responses you define; its pytest fixture respx_mock also fails the test if a request goes to a URL you did not mock. The crawler under test, crawl(client, start_url), follows next links while keeping a seen set, and its fetch helper retries 429 and 5xx responses up to three times, sleeping 2 ** attempt seconds in between (see error handling and pagination).

# tests/test_crawl.py
from pathlib import Path
import httpx
import pytest
from scraper import crawl as crawl_mod

FIXTURES = Path(__file__).parent / "fixtures"
BASE = "https://books.toscrape.com/catalogue/"
PAGE1 = (FIXTURES / "listing_page1.html").read_bytes()
LAST = (FIXTURES / "listing_last_page.html").read_bytes()

@pytest.fixture(autouse=True)
def slept(monkeypatch):
    calls = []
    monkeypatch.setattr(crawl_mod.time, "sleep", calls.append)   # record, never wait
    return calls

def test_follows_next_link_and_stops(respx_mock):
    respx_mock.get(BASE + "page-1.html").respond(200, content=PAGE1)
    respx_mock.get(BASE + "page-2.html").respond(200, content=LAST)
    with httpx.Client() as client:
        items = crawl_mod.crawl(client, BASE + "page-1.html")
    assert len(items) == 40 and respx_mock.calls.call_count == 2

def test_retries_with_backoff(respx_mock, slept):
    route = respx_mock.get(BASE + "page-1.html")
    route.side_effect = [httpx.Response(429), httpx.Response(503), httpx.Response(200, content=LAST)]
    with httpx.Client() as client:
        crawl_mod.crawl(client, BASE + "page-1.html")
    assert route.call_count == 3 and slept == [1, 2]

def test_self_referencing_next_link_does_not_loop(respx_mock):
    looping = PAGE1.replace(b'href="page-2.html"', b'href="page-1.html"')
    respx_mock.get(BASE + "page-1.html").respond(200, content=looping)
    with httpx.Client() as client:
        crawl_mod.crawl(client, BASE + "page-1.html")
    assert respx_mock.calls.call_count == 1

Patching time.sleep keeps the suite fast and lets you assert the backoff schedule itself. side_effect with a list returns one response per call, which is how you script "rate limited, then server error, then success". The loop test covers a real failure: sites sometimes link the last page's next back to itself, and a crawler without a seen set requests it forever.

Recording Real Traffic With Cassettes

Hand-written mocks encode your assumptions about the server. Cassettes record what the server actually sent. The pytest-recording plugin (built on vcrpy, which supports httpx, requests and aiohttp) records every interaction of a test into a YAML file and replays it afterwards:

@pytest.fixture(scope="module")
def vcr_config():
    return {"filter_headers": ["authorization", "cookie", "set-cookie"]}

@pytest.mark.vcr
def test_crawl_last_two_pages():
    with httpx.Client(headers={"User-Agent": "catalog-tests (you@example.com)"}) as client:
        items = crawl_mod.crawl(client, BASE + "page-49.html")
    assert len(items) == 40
pytest tests/test_crawl.py --record-mode=once   # first run: real requests, writes tests/cassettes/...
pytest tests/test_crawl.py                      # later runs: replayed, no network

The default record mode is none, so a test without a cassette fails with CannotOverwriteExistingCassetteException instead of silently going to the network. Filter credentials and cookies before the first recording, since cassettes are committed. Cassettes go stale as the site evolves; re-record them deliberately, and review the diff like a golden file.

Live Contract Tests

A contract test asks one question: does the live site still satisfy the assumptions the parser relies on? Assert invariants rather than values, because prices and stock legitimately change:

# tests/test_live.py
import httpx
import pytest
from scraper.parse import parse_listing

@pytest.mark.live
def test_listing_contract():
    url = "https://books.toscrape.com/catalogue/page-1.html"
    r = httpx.get(url, headers={"User-Agent": "catalog-tests (you@example.com)"}, timeout=30)
    assert r.status_code == 200
    page = parse_listing(r.text, url)
    assert 10 <= len(page["items"]) <= 50
    assert all(item["price"] > 0 for item in page["items"])
    assert page["next_url"] is not None

Exclude them from normal runs in pytest.ini, so pytest stays offline and pytest -m live (a later -m overrides the default) runs only the contract tests:

[pytest]
addopts = -m "not live"
markers =
    live: hits real websites

Then run them on a schedule, separately from pull-request checks:

name: scraper-contract
on:
  schedule:
    - cron: "17 6 * * *"
  workflow_dispatch:
jobs:
  live:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: "3.12"
      - run: pip install -r requirements.txt
      - run: pytest -m live

Keep contract tests polite: a handful of URLs, once a day, an honest User-Agent, only on pages you are allowed to scrape. When one fails, save the fresh page as a fixture, reproduce the failure offline and fix the parser, keeping the old fixture so pages still served in the old layout keep working. Route the failure to the same alerting as your runs (see scheduling and monitoring).

Mistakes That Make Scraper Tests Useless

Mistake Consequence Fix
Fixtures saved via soup.prettify() Tests pass on markup production never sees Save raw bytes
Only happy-path fixtures Missing price or last page crashes in production One fixture per scenario, one per production bug
Parser returns None on missing fields Tests pass on empty records Raise a named error and test that it is raised
Real sleep in retry tests Suite takes minutes Patch time.sleep and assert the calls

Practice

Take a scraper you have written and save five fixtures from a site you are allowed to scrape, including one edge case. Add a golden test, a schema test and a drift test that renames one class. Then write a respx test proving that your crawler stops after three consecutive 503 responses, and time the whole suite: it should run in seconds with the network disabled.