practice sandbox
Web scraping practice sandbox
A small fake website that exists to be scraped. Eight exercises, one technique each, all on made-up data you have permission to scrape.
Reading about pagination or CSRF tokens only gets you so far: you learn scraping by pointing code at a real server and watching what comes back. Practising on someone else's site raises questions about terms of service, load and legality that a beginner should not have to answer first. So simpleprog runs its own target. Everything under /practice is fake data generated from a fixed seed: the 90 products, the job board, the feed, the login, the tables. Scraping it is allowed and encouraged.
Each exercise below says what it practises, where to start, a few hints, and which lesson teaches the technique. A sample solution in Python is folded away under each card; try the exercise first, then compare. Every solution runs as-is against https://simpleprog.com/practice and prints exactly what the card shows, because the data never changes.
Ground rules: scrape it like a good citizen
- Keep it to about one request a second. The sandbox is cheap to serve, but it shares a server with the rest of simpleprog. A
time.sleep(0.5) between pages costs you nothing.
- Only the rate-limit exercise is throttled on purpose (5 requests per 10 seconds per IP address). Treat everything else as if it had the same limit and you will build good habits.
- Identify yourself. Setting a descriptive
User-Agent is the polite default on any site, including this one.
- Stay inside
/practice. The lessons, news and papers elsewhere on simpleprog are real content, not exercise material.
How to use the sandbox
Open the start URL in a browser first and look at it the way a scraper does: view the source, open DevTools, watch the Network tab while you click around. Write the smallest script that fetches one page and pulls one value, then grow it until it handles every page. When you are stuck, read a hint, then the linked lesson, and only then the sample solution.
The JSON endpoints live under /practice/api/, return real status codes (400 for a bad parameter, 404 for a page that does not exist, 429 and 503 where the exercise says so) and carry a notice field reminding you the data is fake.
Common questions and pitfalls
Is it really legal to scrape this?
Yes. simpleprog owns these pages, the data is invented, and you have explicit permission to scrape anything under /practice for learning. The ethics and legal lessons explain what changes when you scrape other sites.
My scraper finds an empty list on the job board or the feed
Those pages are rendered by JavaScript, so the HTML that requests downloads contains no data. Either call the JSON endpoint the page uses (faster and more reliable) or render the page in a real browser with Playwright.
The login keeps returning 403
The CSRF token in the form is tied to a cookie set when you loaded the form. Fetch the form and post it with the same session, and send the token you just read rather than one copied from an old page.
Will the data change under me?
No. All data is generated from fixed random seeds, so the same URL always returns the same products, prices and posts. That makes the sandbox good for testing scrapers too: a changed result means your code changed, not the site.
Can I use Scrapy, Node.js or another tool?
Of course. The sample solutions use Python with requests, BeautifulSoup and Playwright because that is what the course teaches, but every exercise works with any HTTP client. See scraping with Node.js or the Scrapy lesson.
Why are the exercise pages hidden from search engines?
They hold fake products and fake jobs that nobody should find in a search. Only this page is indexed; the exercise pages are marked noindex and the JSON endpoints send X-Robots-Tag: noindex.