Where you use CSS selectors

Tool Call Notes
BeautifulSoup soup.select("css"), soup.select_one("css") Powered by soupsieve. select_one returns None when nothing matches; select(css, limit=5) stops early
lxml tree.cssselect("css") Needs the separate cssselect package; it translates CSS to XPath
parsel / Scrapy response.css("h1::text").get() Adds the non-standard ::text and ::attr(name) pseudo-elements
Playwright page.locator("css") Adds :has-text(), :text(), :text-is(), :visible; pierces open shadow DOM
Browser DevTools document.querySelectorAll("css"), $$("css") Test a selector against the live page in the console

Basic selectors

Selector Matches Example
* any element div.card > *
tag elements by tag name h2
.cls class list contains cls .price
.a.b has both classes (no space) li.item.featured
#id element with that id #main
A, B matches A or B, returned in document order h1, h2, h3
tag.cls tag with the class span.price

Attribute selectors

Selector Attribute value... Matches Does not match
[href] exists (even empty) <a href=""> <a>
[data-state="open"] equals exactly open opened, Open
[class~="ad"] whitespace-separated list contains the word class="item ad" class="adv"
[lang|="en"] equals en or starts with en- en, en-US english
[href^="https"] starts with https://x.com /https
[href$=".pdf"] ends with /a.pdf /a.pdf?x=1
[href*="/product/"] contains the substring /shop/product/42 /products
[href$=".pdf" i] i flag: ASCII case-insensitive .PDF, .Pdf
[data-x="v" s] s flag: force case-sensitive soupsieve supports it; browser support is patchy
  • Quote values that contain spaces or punctuation: [data-x="has space"]. Unquoted, soupsieve raises SelectorSyntaxError.
  • [class="item"] compares the whole attribute string: it matches class="item" but not class="item ad". Use .item (or [class~="item"]) instead.
  • [class^="Card_"] only works if that class comes first in the attribute. For hashed CSS-module classes such as ProductCard_title__x9Az, use [class*="ProductCard_title"].

Combinators

Selector Meaning Example
A B B anywhere inside A div.card a
A > B B is a direct child of A ul.results > li
A + B B immediately follows A (same parent) h2 + p
A ~ B B is any later sibling of A h2 ~ p

CSS has no parent or previous-sibling combinator. Use :has() instead: div:has(> h2) selects the div that contains the h2, h2:has(+ ul) selects the h2 sitting right before a list. Or switch to XPath (.., preceding-sibling::).

:nth-child and friends

Counting is 1-based and counts element siblings only (text nodes are ignored). An+B is evaluated for n = 0, 1, 2, ...

Selector Selects
:first-child / :last-child first / last child of its parent
:only-child the parent's only element child
:nth-child(3) the 3rd child
:nth-child(odd) = :nth-child(2n+1) 1st, 3rd, 5th, ...
:nth-child(even) = :nth-child(2n) 2nd, 4th, 6th, ...
:nth-child(n+3) 3rd child onwards (skip the first two)
:nth-child(-n+3) first three children
:nth-child(3n) every third child
:nth-last-child(2) second from the end
:nth-child(2 of .item) 2nd sibling among those matching .item
:nth-of-type(2) 2nd sibling of the same tag
:first-of-type, :last-of-type, :only-of-type same, by tag
:empty no children (see whitespace note below)
:root the html element

Classic trap: li.item:nth-child(2) means "the 2nd child, if it is an li.item", not "the 2nd li.item". If an ad or a header sits first, you get the wrong row or nothing. Use :nth-child(2 of li.item) or index in Python: soup.select("li.item")[1].

Logical pseudo-classes

Selector Selects Example
:not(A, B) matches none of the list li:not(.ad, .sponsored)
:is(A, B) matches any of the list :is(h1, h2, h3) > a
:where(A, B) like :is(), zero specificity (irrelevant when scraping) :where(.a, .b)
:has(> img) has a direct child img li:has(> img)
:has(img) has a descendant img div.card:has(.price)
:has(+ .x) immediately followed by .x h2:has(+ ul)
:not(:has(a)) contains no link li:not(:has(a))

BeautifulSoup (soupsieve) vs browsers

Checked against soupsieve 2.6. Browser columns describe current Chrome, Firefox and Safari.

Feature soupsieve (BeautifulSoup) Browsers
All attribute operators + i flag yes yes
s flag yes not in every browser
:is(), :where(), :not(A, B) yes yes
:has() yes yes (Chrome 105+, Safari 15.4+, Firefox 121+)
:nth-child(An+B of S) yes yes in current versions
Text matching :-soup-contains() none in CSS (Playwright adds :has-text())
Pseudo-elements (::before, ::text) NotImplementedError styling only, never selectable
:hover, :focus, :visited never match live state
:empty on whitespace-only content matches (whitespace ignored) does not match
#123 (id starts with a digit) SelectorSyntaxError SyntaxError too
Escaping [id="123"], #\31 23, or soupsieve.escape("123") CSS.escape("123")
<tbody> present only if in the source (or with the html5lib parser) always inserted

The tbody trap: browsers insert <tbody> into every table, so a selector copied from DevTools (table > tbody > tr) finds nothing in the raw HTML parsed by html.parser or lxml. Use table tr (descendant) or table tr:has(td), which work either way.

Scraping recipes

Goal Selector Then
Product cards, skip ads div.product:not(.ad, .sponsored) loop
Card title link div.product h2 > a a.get("href")
Absolute links only a[href^="http"] urljoin the rest
PDF links (any case) a[href$=".pdf" i]
Product page links a[href*="/product/"]
Next page a[rel="next"], link[rel="next"] .get("href")
Canonical URL link[rel="canonical"] .get("href")
Open Graph title / image meta[property="og:title"], meta[property="og:image"] .get("content")
Meta description meta[name="description"] .get("content")
JSON-LD blocks script[type="application/ld+json"] json.loads(tag.string)
Lazy-loaded images img[data-src], img[srcset] prefer data-src over the placeholder src
Table data rows (skip header) table.specs tr:has(td) works with or without tbody
Second column of a table table tr > td:nth-child(2)
Value next to a label (dl) dt:-soup-contains("Price") + dd soupsieve only
Value next to a label (th/td) tr:has(> th:-soup-contains("Weight")) > td soupsieve only
Hashed / CSS-module class [class*="ProductCard_title"] survives rebuilds better
Stable hooks [data-testid="price"], [itemprop="price"] itemprop often has content
Cards that have a price div.card:has(.price)
Last breadcrumb nav.breadcrumb li:last-child

Python pattern

from bs4 import BeautifulSoup

soup = BeautifulSoup(html, "lxml")              # or "html.parser"
rows = []
for card in soup.select("div.product:not(.ad)"):
    link = card.select_one("h2 a")
    price = card.select_one("[itemprop=price], .price")
    if link is None:                            # select_one -> None on no match
        continue
    rows.append({
        "name": link.get_text(" ", strip=True),
        "url": link.get("href"),                # .get() -> None instead of KeyError
        "price": price.get_text(strip=True) if price else None,
    })

Pitfalls

  • get_text(strip=True) strips each text node and glues them: <p>Hello <b>big</b> world</p> becomes Hellobigworld. Use get_text(" ", strip=True).
  • DevTools "Copy selector" gives brittle paths like #root > div:nth-child(3) > div > span. Rewrite them with classes, attributes or :has().
  • Content missing? If a selector works in DevTools but not on requests HTML, the data is rendered by JavaScript. Check view-source, then look for the JSON API or use a browser.
  • Searching inside a card: card.select("a") is scoped to the card. Calling soup.select(...) inside the loop searches the whole page again.
  • Class names with a colon or other special characters need escaping: .a\:b (or soupsieve.escape("a:b")).