Project: Build a Price Monitor
What You Will Build
A price monitor that checks a catalogue on a schedule, keeps the full price history in SQLite, and sends one alert message when prices drop or items come back in stock. The target is books.toscrape.com, a sandbox built for scraping practice: 1,000 books across 50 listing pages. It is static, so prices never change on their own; you will simulate changes to test the alert path.
The capstone project builds a general catalogue monitor with sitemaps, conditional requests and a crash-safe run loop. This project is narrower and goes deeper on what makes a price monitor useful: a change-log data model, alert rules that are worth reading, and guards that stop your own bugs from reaching users. The finished script is about 130 lines of Python using httpx, BeautifulSoup and the standard library.
Design Decisions First
| Decision | Choice | Why |
|---|---|---|
| Which pages to fetch | The 50 listing pages | They show price and stock for 20 books each: 50 requests per run instead of 1,050 |
| Price type | Integer pence | Exact comparison, no float noise |
| Storage | Current state plus an append-only change log | "What is the price now?" and "how did it move?" are both one query |
| Alerts | One digest per run, drops of 5% or more, stock changes | People mute monitors that send twenty messages for pennies |
| Safety | Reject runs where too many prices change at once | A parser bug looks like a sitewide price change |
Fetching the cheapest page that carries the fields you monitor is the biggest optimisation available; product pages add detail a price monitor does not need on every run.
Step 1: Crawl the Listing Pages
import os
import sqlite3
import sys
import time
from datetime import datetime, timezone
from decimal import Decimal
from urllib.parse import urljoin
import httpx
from bs4 import BeautifulSoup
START = "https://books.toscrape.com/catalogue/page-1.html"
HEADERS = {"User-Agent": "price-monitor-tutorial (you@example.com)"}
DELAY = 1.0 # seconds between requests
DB_PATH = "prices.db"
MIN_DROP_PCT = 5 # alert on drops of at least 5%
MAX_CHANGED_SHARE = 0.2 # more than 20% changed in one run looks like a bug
def parse_listing(html, page_url):
soup = BeautifulSoup(html, "lxml")
cards = soup.select("article.product_pod")
if not cards:
raise RuntimeError(f"no product cards on {page_url}: layout changed?")
items = []
for card in cards:
link = card.select_one("h3 a")
price = Decimal(card.select_one("p.price_color").get_text(strip=True).removeprefix("£"))
items.append({
"url": urljoin(page_url, link["href"]),
"title": link["title"], # the link text is truncated
"price_minor": int(price * 100),
"in_stock": int("In stock" in card.select_one("p.availability").get_text()),
})
nxt = soup.select_one("li.next a")
return items, (urljoin(page_url, nxt["href"]) if nxt else None)
def crawl(client):
url, pages, items = START, 0, []
while url:
r = client.get(url)
r.raise_for_status()
page_items, url = parse_listing(r.text, str(r.url))
items.extend(page_items)
pages += 1
time.sleep(DELAY)
return pages, items
The product URL is the key: it is stable and unique on this site. Titles are not unique in real catalogues, and the visible link text is cut off with "..." for long titles, which is why the title attribute is used. Any exception aborts the run before anything is written, so a half-crawled catalogue never overwrites good data. For sites with other number formats, swap the Decimal line for the locale-aware parsing in parsing dates, prices and units.
Step 2: Current State Plus a Change Log
SCHEMA = '''
CREATE TABLE IF NOT EXISTS products (
url TEXT PRIMARY KEY, title TEXT NOT NULL,
price_minor INTEGER NOT NULL, in_stock INTEGER NOT NULL,
first_seen TEXT NOT NULL, last_seen TEXT NOT NULL
);
CREATE TABLE IF NOT EXISTS observations (
url TEXT NOT NULL, price_minor INTEGER NOT NULL,
in_stock INTEGER NOT NULL, seen_at TEXT NOT NULL
);
CREATE INDEX IF NOT EXISTS obs_url ON observations(url, seen_at);
CREATE TABLE IF NOT EXISTS runs (
started TEXT, finished TEXT, pages INTEGER, items INTEGER, changes INTEGER, status TEXT
);
'''
def record(db, item, now):
row = db.execute("SELECT price_minor, in_stock FROM products WHERE url = ?",
(item["url"],)).fetchone()
if row is None:
db.execute("INSERT INTO products VALUES (?, ?, ?, ?, ?, ?)",
(item["url"], item["title"], item["price_minor"], item["in_stock"], now, now))
db.execute("INSERT INTO observations VALUES (?, ?, ?, ?)",
(item["url"], item["price_minor"], item["in_stock"], now))
return None # a first sighting is not a change
db.execute("UPDATE products SET last_seen = ?, title = ? WHERE url = ?",
(now, item["title"], item["url"]))
if row == (item["price_minor"], item["in_stock"]):
return None
db.execute("UPDATE products SET price_minor = ?, in_stock = ? WHERE url = ?",
(item["price_minor"], item["in_stock"], item["url"]))
db.execute("INSERT INTO observations VALUES (?, ?, ?, ?)",
(item["url"], item["price_minor"], item["in_stock"], now))
return {"title": item["title"], "old": row[0], "new": item["price_minor"],
"was_in_stock": row[1], "in_stock": item["in_stock"]}
products answers "what is true now", observations gets a row only when something changed, and last_seen moves on every run. That split keeps the history small (1,000 rows after the first run, then only real changes) while still telling you which products vanished: anything whose last_seen is older than the latest run. Storing every observation is also valid, at 1,000 rows per run; incremental scraping discusses the same trade-off.
Step 3: Alerts Worth Reading
def worth_alerting(c):
if c["was_in_stock"] != c["in_stock"]:
return True
return (c["old"] - c["new"]) * 100 >= MIN_DROP_PCT * c["old"] # integer maths, no floats
def describe(c):
if c["was_in_stock"] != c["in_stock"]:
return f"{c['title']}: {'back in stock' if c['in_stock'] else 'sold out'}"
pct = (c["old"] - c["new"]) * 100 / c["old"]
return f"{c['title']}: £{c['old'] / 100:.2f} -> £{c['new'] / 100:.2f} (-{pct:.0f}%)"
def send_alert(lines):
text = "Price monitor\n" + "\n".join(f"- {line}" for line in lines)
webhook = os.environ.get("ALERT_WEBHOOK_URL")
if webhook:
httpx.post(webhook, json={"text": text}, timeout=10).raise_for_status()
else:
print(text)
All changes are recorded, but only meaningful ones are announced: drops of at least 5% and stock transitions. Price rises are stored and not sent. The {"text": ...} payload is the format Slack-style incoming webhooks accept; for email, replace the POST with smtplib. Keep the webhook URL in an environment variable, never in the code.
Step 4: Guard Against Your Own Bugs
def main():
db = sqlite3.connect(DB_PATH)
db.executescript(SCHEMA)
started = datetime.now(timezone.utc).isoformat(timespec="seconds")
with httpx.Client(headers=HEADERS, timeout=30, follow_redirects=True) as client:
pages, items = crawl(client)
known = db.execute("SELECT COUNT(*) FROM products").fetchone()[0]
changes = [c for item in items if (c := record(db, item, started))]
if known and len(changes) > MAX_CHANGED_SHARE * known:
db.rollback() # keep the last trusted prices
status = "suspicious"
send_alert([f"{len(changes)} of {known} products changed in one run; check the parser"])
else:
status = "ok"
alerts = [describe(c) for c in changes if worth_alerting(c)]
if alerts:
send_alert(alerts) # if this raises, nothing is committed
db.execute("INSERT INTO runs VALUES (?, ?, ?, ?, ?, ?)",
(started, datetime.now(timezone.utc).isoformat(timespec="seconds"),
pages, len(items), len(changes), status))
db.commit()
print(f"{pages} pages, {len(items)} items, {len(changes)} changes, status={status}")
if __name__ == "__main__":
sys.exit(main())
Python's sqlite3 opens a transaction implicitly at the first write, so every record call belongs to one uncommitted transaction until commit(). That gives two guarantees. If more than 20% of known products change at once, which is what a pence-versus-pounds bug or a switch to another currency looks like, rollback() discards the run and one warning goes out instead of 1,000 bogus alerts. And the alert is sent before the commit: if the webhook is down, the exception leaves nothing committed, the next run sees the same changes, and the alert is retried. A genuine sitewide sale also trips the guard, which is the right default; confirm it by hand and rerun with the limit raised.
Step 5: Run It and Simulate a Change
python monitor.py
50 pages, 1000 items, 0 changes, status=ok
The first run takes about 90 seconds, most of it the one-second delay. Since the site never changes, edit the database to pretend the previous prices were different, then run again:
UPDATE products SET price_minor = 5999 WHERE title = 'A Light in the Attic';
UPDATE observations SET price_minor = 5999 WHERE url LIKE '%/a-light-in-the-attic_1000/%';
UPDATE products SET price_minor = 5474 WHERE title = 'Tipping the Velvet';
UPDATE products SET in_stock = 0 WHERE title = 'Soumission';
Price monitor
- A Light in the Attic: £59.99 -> £51.77 (-14%)
- Soumission: back in stock
50 pages, 1000 items, 3 changes, status=ok
Three changes were recorded, two were announced. "Tipping the Velvet" went from £54.74 to £53.74, a 1.8% drop below the alert threshold, so it is in the history but not in the message. To test the guard, multiply every stored price by 100 and rerun: the run ends with status=suspicious, a single warning, and unchanged prices.
Step 6: Query the History
SQLite's window functions turn the change log into price movements:
SELECT seen_at, printf('%.2f', price_minor / 100.0) AS price,
price_minor - LAG(price_minor) OVER (ORDER BY seen_at) AS change_minor
FROM observations
WHERE url = 'https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html'
ORDER BY seen_at;
2026-09-30T07:19:57+00:00|59.99|
2026-09-30T07:23:22+00:00|51.77|-822
Two other queries earn a place in any monitor. Products that disappeared from the catalogue: SELECT title FROM products WHERE last_seen < (SELECT MAX(started) FROM runs). And run health, from the runs table: a run whose items count falls well below the previous runs, or no run in the last day, deserves its own alert (see scheduling and monitoring).
Step 7: Schedule It
Four runs a day is plenty for a catalogue that reprices daily, and it keeps the load at 200 requests per day. On Linux, cron with flock prevents overlapping runs:
17 */6 * * * cd /opt/price-monitor && flock -n /tmp/price-monitor.lock python monitor.py >> monitor.log 2>&1
On Windows, Task Scheduler does the same job, and by default it does not start a new instance while the previous one is still running:
schtasks /create /tn PriceMonitor /sc hourly /mo 6 /tr "C:\price-monitor\run.cmd"
An odd minute such as :17 avoids everyone else's top-of-the-hour jobs. Keep the delay even on a sandbox; on a real site, use the rate its terms and robots.txt allow (see rate limiting and politeness).
Practice
Extend the monitor in three steps. Add a watchlist table with a target price per URL and alert when a watched book falls below its target, regardless of the 5% rule. Save two listing pages as fixtures and write tests for parse_listing, worth_alerting and the rollback guard, following testing your scrapers. Finally, add a product-page fetch only for watched books, sending If-None-Match with the stored ETag so unchanged pages cost a 304 instead of a full download.