fake data Part of the simpleprog practice sandbox. Every product, person, price and number here is made up. Scraping it is allowed and encouraged: keep it to about one request a second.

practice sandbox

Web scraping practice sandbox

A small fake website that exists to be scraped. Eight exercises, one technique each, all on made-up data you have permission to scrape.

Reading about pagination or CSRF tokens only gets you so far: you learn scraping by pointing code at a real server and watching what comes back. Practising on someone else's site raises questions about terms of service, load and legality that a beginner should not have to answer first. So simpleprog runs its own target. Everything under /practice is fake data generated from a fixed seed: the 90 products, the job board, the feed, the login, the tables. Scraping it is allowed and encouraged.

Each exercise below says what it practises, where to start, a few hints, and which lesson teaches the technique. A sample solution in Python is folded away under each card; try the exercise first, then compare. Every solution runs as-is against https://simpleprog.com/practice and prints exactly what the card shows, because the data never changes.

Exercises

01

Paginated product catalogue

beginner

Following numbered pagination across 90 fake products (12 per page), narrowing the list with a ?category= query string, and visiting detail pages.

start at /practice/products

hints
  1. Follow the a[rel=next] link until it disappears instead of hard-coding the page count.
  2. Filter with ?category=audio: the page count changes, but the next link still works.
  3. Prices are text like $12.99; strip the symbol before converting to a number.

learn it Handling PaginationCSS Selectors with BeautifulSoup

sample solution (Python)

requests + BeautifulSoup

import time
from urllib.parse import urljoin

import requests
from bs4 import BeautifulSoup

BASE = "https://simpleprog.com/practice"

url, products = f"{BASE}/products", []
with requests.Session() as s:
    while url:
        soup = BeautifulSoup(s.get(url, timeout=10).text, "html.parser")
        for card in soup.select("article.product"):
            products.append({
                "name": card.select_one("h2 a").get_text(strip=True),
                "price": float(card.select_one(".price").get_text(strip=True).lstrip("$")),
                "rating": int(card.select_one(".rating")["data-rating"]),
                "in_stock": "in-stock" in card.select_one(".stock")["class"],
            })
        next_link = soup.select_one("a[rel=next]")  # absent on the last page
        url = urljoin(url, next_link["href"]) if next_link else None
        time.sleep(0.5)  # be polite

cheapest = min(products, key=lambda p: p["price"])
print(f"{len(products)} products")
print(f"cheapest: {cheapest['name']} at ${cheapest['price']:.2f}")
print(f"out of stock: {sum(not p['in_stock'] for p in products)}")

prints

90 products
cheapest: Granite Seed Tray at $5.83
out of stock: 11
02

JavaScript list with a hidden JSON API

beginner

Spotting that a page is rendered by JavaScript, finding the JSON endpoint it calls, and paging through that API directly (or rendering the page with a browser).

start at /practice/jobs

hints
  1. View the page source: the job list is empty. The jobs arrive afterwards through fetch().
  2. Open DevTools, go to Network, filter by Fetch/XHR and reload to see the request URL.
  3. The JSON tells you how many pages exist (total_pages), so you know when to stop.

learn it Finding the Hidden JSON APIsScraping JavaScript-Rendered Pages

sample solution (Python)

requests: call the API

import requests

BASE = "https://simpleprog.com/practice"

# The page itself is an empty shell; DevTools > Network > Fetch/XHR shows
# the JSON endpoint its JavaScript calls. Call that endpoint directly.
jobs, page, total_pages = [], 1, 1
with requests.Session() as s:
    while page <= total_pages:
        r = s.get(f"{BASE}/api/jobs", params={"page": page}, timeout=10)
        r.raise_for_status()
        data = r.json()
        jobs += data["results"]
        total_pages = data["total_pages"]
        page += 1

top = max(jobs, key=lambda j: j["salary_max"])
print(f"{len(jobs)} jobs, {sum(j['remote'] for j in jobs)} remote")
print(f"best paid: {top['title']} at {top['company']} (up to ${top['salary_max']:,})")

Playwright: render the page

from playwright.sync_api import sync_playwright

BASE = "https://simpleprog.com/practice"

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page()
    page.goto(f"{BASE}/jobs")
    page.wait_for_selector("#jobs .job")  # rendered by JavaScript after fetch()
    jobs = page.locator("#jobs .job")
    print(jobs.count(), "jobs on the first page")
    for job in jobs.all()[:3]:
        print(job.locator(".job-title").inner_text(), "|", job.locator(".job-company").inner_text())
    browser.close()

prints

# requests
30 jobs, 6 remote
best paid: Lead Data Engineer at Glassbird (up to $153,000)

# Playwright
10 jobs on the first page
Senior Frontend Developer | Redkite Digital
Lead Backend Developer | Redkite Digital
Senior DevOps Engineer | Pinecrest Labs
03

Infinite scroll with a cursor API

intermediate

Reproducing an infinite-scroll feed (with a load-more fallback) by walking a cursor-paginated JSON API until the cursor runs out.

start at /practice/feed

hints
  1. Scroll with the Network tab open: each new batch is one request to /practice/api/feed.
  2. Send next_cursor back as ?cursor= until it is null. Cursors are opaque, so never build them yourself.
  3. limit goes up to 25: fewer, larger requests are kinder to the server.

learn it Infinite Scroll and Lazy LoadingConsuming REST APIs Cleanly

sample solution (Python)

requests: follow the cursor

import requests

BASE = "https://simpleprog.com/practice"

posts, cursor, calls = [], None, 0
with requests.Session() as s:
    while True:
        params = {"limit": 25} | ({"cursor": cursor} if cursor else {})
        r = s.get(f"{BASE}/api/feed", params=params, timeout=10)
        r.raise_for_status()
        data = r.json()
        calls += 1
        posts += data["items"]
        cursor = data["next_cursor"]  # None after the last batch
        if not cursor:
            break

print(f"{len(posts)} posts in {calls} requests")
print("most liked:", max(posts, key=lambda p: p["likes"])["text"])

prints

60 posts in 3 requests
most liked: Still debugging the retry wrapper and now it respects robots.txt. (fake post #58)
04

Login form with a CSRF token

intermediate

Reading a hidden CSRF field, posting a login form with a persistent session, and scraping a members-only page. Log in with username demo and password demo.

start at /practice/login

login demo / demo

hints
  1. The hidden csrf_token only works together with the cookie set by the same GET, so use one requests.Session for everything.
  2. A missing or wrong token gets 403; a wrong password gets 401.
  3. Landing on the login form instead of /practice/members means your session cookie was not sent.

learn it Form Submission and Multi-Step FlowsLogins, Tokens & Long-Lived SessionsCookies, Sessions & Persistent State

sample solution (Python)

requests.Session + BeautifulSoup

import requests
from bs4 import BeautifulSoup

BASE = "https://simpleprog.com/practice"

with requests.Session() as s:  # one session: the CSRF cookie must travel with the token
    form = BeautifulSoup(s.get(f"{BASE}/login", timeout=10).text, "html.parser")
    token = form.select_one('input[name="csrf_token"]')["value"]
    r = s.post(f"{BASE}/login", timeout=10,
               data={"username": "demo", "password": "demo", "csrf_token": token})
    r.raise_for_status()  # 403 = bad token, 401 = bad password
    page = BeautifulSoup(r.text, "html.parser")  # followed the redirect to /members

print(r.url)
print(page.select_one("h1").get_text(strip=True))
for row in page.select("table#codes tbody tr")[:3]:
    print([td.get_text(strip=True) for td in row.find_all("td")])

prints

https://simpleprog.com/practice/members
Welcome back, demo
['FAKE-PEBBLE-BOO-01', 'Pebble Bookends', '10%', '2026-12-11']
['FAKE-EMBER-WATE-02', 'Ember Watering Can', '15%', '2026-12-12']
['FAKE-GRANITE-SE-03', 'Granite Seed Tray', '20%', '2026-12-13']
05

Rate-limited endpoint

intermediate

Handling HTTP 429: reading Retry-After and the X-RateLimit headers, waiting, and retrying without losing your place.

start at /practice/api/limited?page=1

hints
  1. The limit is 5 requests per 10 seconds per IP address; the 6th gets 429 with a Retry-After header.
  2. Watch X-RateLimit-Remaining to slow down before you hit the wall.
  3. Rejected requests are not counted, so sleep for Retry-After and ask for the same page again.

learn it Rate Limiting and Polite CrawlingConsuming REST APIs Cleanly

sample solution (Python)

requests + Retry-After

import time

import requests

BASE = "https://simpleprog.com/practice"

items, page = [], 1
with requests.Session() as s:
    while page:
        r = s.get(f"{BASE}/api/limited", params={"page": page}, timeout=10)
        if r.status_code == 429:
            wait = int(r.headers.get("Retry-After", "10"))
            print(f"429 on page {page}, sleeping {wait}s")
            time.sleep(wait)
            continue  # retry the same page
        r.raise_for_status()
        data = r.json()
        items += data["items"]
        page = data["next_page"]  # None after the last page

print(f"collected {len(items)} items")

prints

429 on page 6, sleeping 10s
collected 90 items
06

HTML tables with rowspan, colspan and units

beginner

Turning HTML tables into records: a two-row header with colspan, brand cells with rowspan, and values with currency symbols, units and a missing entry.

start at /practice/tables

hints
  1. Each brand cell spans several rows (rowspan), so most body rows have 5 cells, not 6.
  2. The header has two rows because Specs spans three columns (colspan).
  3. Values carry units ($1,299.00, 1.35 kg, 1 TB) and one is missing (—).

learn it Extracting Data from HTML TablesParsing Dates, Prices and Units

sample solution (Python)

requests + BeautifulSoup

import re

import requests
from bs4 import BeautifulSoup

BASE = "https://simpleprog.com/practice"


def number(text):
    """'$1,299.00' -> 1299.0, '1.35 kg' -> 1.35, and None for a missing value like '—'."""
    m = re.search(r"\d[\d,]*(?:\.\d+)?", text)
    return float(m.group().replace(",", "")) if m else None


soup = BeautifulSoup(requests.get(f"{BASE}/tables", timeout=10).text, "html.parser")
laptops, brand = [], None
for tr in soup.select("table#laptops tbody tr"):
    cells = [td.get_text(strip=True) for td in tr.find_all("td")]
    if len(cells) == 6:  # first row of a brand: it carries the rowspan brand cell
        brand, cells = cells[0], cells[1:]
    model, price, weight, battery, storage = cells
    laptops.append({
        "brand": brand, "model": model, "price": number(price), "weight_kg": number(weight),
        "battery_h": number(battery), "storage_gb": number(storage) * (1000 if "TB" in storage else 1),
    })

print(f"{len(laptops)} laptops")
for laptop in sorted(laptops, key=lambda l: l["price"])[:3]:
    print(laptop)

prints

8 laptops
{'brand': 'Quillmark', 'model': 'Q13 Air', 'price': 649.0, 'weight_kg': 1.12, 'battery_h': 11.0, 'storage_gb': 256.0}
{'brand': 'Brassbay', 'model': 'Harbor 13', 'price': 719.0, 'weight_kg': 1.25, 'battery_h': 16.0, 'storage_gb': 512.0}
{'brand': 'Tinlark', 'model': 'T-Book 14', 'price': 899.99, 'weight_kg': 1.48, 'battery_h': None, 'storage_gb': 512.0}
07

Structured data in JSON-LD

beginner

Reading the schema.org Product data a page publishes for machines, which is more stable than its HTML and holds fields the page never shows.

start at /practice/products/pebble-bookends

hints
  1. Look for <script type="application/ld+json">. There are two; only one is a Product.
  2. The GTIN barcode and brand are only in the JSON-LD, not in the visible page.
  3. offers.availability is a schema.org URL such as https://schema.org/InStock.

learn it Structured Data: JSON-LD, Microdata and Open Graph

sample solution (Python)

requests + json

import json
import time
from urllib.parse import urljoin

import requests
from bs4 import BeautifulSoup

BASE = "https://simpleprog.com/practice"

with requests.Session() as s:
    listing = BeautifulSoup(s.get(f"{BASE}/products", timeout=10).text, "html.parser")
    links = [urljoin(BASE, a["href"]) for a in listing.select("article.product h2 a")]
    for url in links[:5]:
        soup = BeautifulSoup(s.get(url, timeout=10).text, "html.parser")
        for script in soup.select('script[type="application/ld+json"]'):
            data = json.loads(script.string)
            if data.get("@type") != "Product":
                continue  # skip the BreadcrumbList
            offer = data["offers"]
            print(data["sku"], data["gtin13"], data["name"], offer["price"], offer["priceCurrency"],
                  offer["availability"].rsplit("/", 1)[-1], data["aggregateRating"]["ratingValue"])
        time.sleep(0.5)

prints

SP-OFF-1007 2351565487859 Pebble Bookends 107.40 USD InStock 2
SP-GAR-1014 2348625915452 Ember Watering Can 186.76 USD InStock 4
SP-GAR-1021 2382500918015 Granite Seed Tray 5.83 USD InStock 3
SP-KIT-1028 2150896881356 Umber Coffee Grinder 239.71 USD InStock 1
SP-AUD-1035 2030799475671 Russet Headphones 88.19 USD InStock 4
08

Flaky endpoint: retries and backoff

intermediate

Retrying transient 503 errors with Retry-After and exponential backoff, while failing fast on errors that will never succeed.

start at /practice/api/flaky/1

hints
  1. Every 4th request to this endpoint returns 503 with Retry-After: 1. Station ids run from 1 to 20.
  2. Retry only errors that can succeed later (429 and 5xx); a 404 for station 21 never will.
  3. Cap the attempts so a real outage cannot loop forever.

learn it Retries, Errors & TimeoutsRate Limiting and Polite Crawling

sample solution (Python)

requests + manual retry

import time

import requests

BASE = "https://simpleprog.com/practice"
RETRYABLE = {429, 500, 502, 503, 504}


def get_json(session, url, attempts=5):
    for attempt in range(attempts):
        r = session.get(url, timeout=10)
        if r.status_code not in RETRYABLE:
            r.raise_for_status()  # a 404 will never succeed: fail at once
            return r.json()
        wait = float(r.headers.get("Retry-After", 2 ** attempt))  # else exponential backoff
        print(f"  {r.status_code} on {url.rsplit('/', 1)[-1]}, retrying in {wait:g}s")
        time.sleep(wait)
    raise RuntimeError(f"gave up on {url} after {attempts} attempts")


with requests.Session() as s:
    readings = [get_json(s, f"{BASE}/api/flaky/{i}") for i in range(1, 21)]

warmest = max(readings, key=lambda r: r["temperature_c"])
print(f"{len(readings)} readings, warmest: {warmest['station']} at {warmest['temperature_c']} C")

prints

  503 on 4, retrying in 1s
  503 on 7, retrying in 1s
  ... (a few more; which ids fail depends on who else is using the endpoint)
20 readings, warmest: Airfield (fake) at 32.0 C

Ground rules: scrape it like a good citizen

  • Keep it to about one request a second. The sandbox is cheap to serve, but it shares a server with the rest of simpleprog. A time.sleep(0.5) between pages costs you nothing.
  • Only the rate-limit exercise is throttled on purpose (5 requests per 10 seconds per IP address). Treat everything else as if it had the same limit and you will build good habits.
  • Identify yourself. Setting a descriptive User-Agent is the polite default on any site, including this one.
  • Stay inside /practice. The lessons, news and papers elsewhere on simpleprog are real content, not exercise material.

How to use the sandbox

Open the start URL in a browser first and look at it the way a scraper does: view the source, open DevTools, watch the Network tab while you click around. Write the smallest script that fetches one page and pulls one value, then grow it until it handles every page. When you are stuck, read a hint, then the linked lesson, and only then the sample solution.

The JSON endpoints live under /practice/api/, return real status codes (400 for a bad parameter, 404 for a page that does not exist, 429 and 503 where the exercise says so) and carry a notice field reminding you the data is fake.

Common questions and pitfalls

Is it really legal to scrape this?

Yes. simpleprog owns these pages, the data is invented, and you have explicit permission to scrape anything under /practice for learning. The ethics and legal lessons explain what changes when you scrape other sites.

My scraper finds an empty list on the job board or the feed

Those pages are rendered by JavaScript, so the HTML that requests downloads contains no data. Either call the JSON endpoint the page uses (faster and more reliable) or render the page in a real browser with Playwright.

The login keeps returning 403

The CSRF token in the form is tied to a cookie set when you loaded the form. Fetch the form and post it with the same session, and send the token you just read rather than one copied from an old page.

Will the data change under me?

No. All data is generated from fixed random seeds, so the same URL always returns the same products, prices and posts. That makes the sandbox good for testing scrapers too: a changed result means your code changed, not the site.

Can I use Scrapy, Node.js or another tool?

Of course. The sample solutions use Python with requests, BeautifulSoup and Playwright because that is what the course teaches, but every exercise works with any HTTP client. See scraping with Node.js or the Scrapy lesson.

Why are the exercise pages hidden from search engines?

They hold fake products and fake jobs that nobody should find in a search. Only this page is indexed; the exercise pages are marked noindex and the JSON endpoints send X-Robots-Tag: noindex.