How to use these cheat sheets

Each sheet is one page of tables you can scan in seconds: the syntax, what it does, and a working example. Open one in a tab while you write a scraper, type in the filter box (or press /) to narrow every table to the rows you need, and click any snippet in a table to copy it.

To get a paper or PDF copy, use the Print / save as PDF button at the top of a sheet. The print layout drops the site navigation, switches to black on white and fits the tables to the page width, so "Save as PDF" in the print dialog gives you a clean offline copy.

What makes them different

  • Tested, not copied. Every regex, XPath expression and CSS selector was run in Python (the re module, lxml and BeautifulSoup's soupsieve) before it went on a sheet. Where a pattern has a known false positive, the caveat sits right next to it.
  • Written for scraping. The sheets cover the questions that come up in real scrapers: why a selector copied from DevTools finds nothing, which status codes to retry, and when regex on HTML will bite you.
  • Linked to the full lessons. Each sheet points to the lessons that explain the topic properly, from CSS selectors with BeautifulSoup to common ports and protocols.

FAQ

Are the cheat sheets free?

Yes. They are free to read, print and share, with no sign-up, like every course on simpleprog.

Do the selectors work in BeautifulSoup, lxml and the browser?

The CSS sheet has a column-by-column comparison of what BeautifulSoup (soupsieve) supports and what browsers support. XPath in lxml and in browsers is the same XPath 1.0, so the XPath sheet applies to both, and it lists the XPath 2.0 functions you cannot use.

Where should I start if I am new to scraping?

Start with the HTTP basics lesson and the HTTP status codes sheet, then move to CSS selectors. The full web scraping course walks through everything in order.