Every selector you reach for when pulling data out of HTML, with the BeautifulSoup (soupsieve) behaviour spelled out where it differs from a browser. All examples were run against BeautifulSoup 4.13 with soupsieve 2.6.
Powered by soupsieve. select_one returns None when nothing matches; select(css, limit=5) stops early
lxml
tree.cssselect("css")
Needs the separate cssselect package; it translates CSS to XPath
parsel / Scrapy
response.css("h1::text").get()
Adds the non-standard ::text and ::attr(name) pseudo-elements
Playwright
page.locator("css")
Adds :has-text(), :text(), :text-is(), :visible; pierces open shadow DOM
Browser DevTools
document.querySelectorAll("css"), $$("css")
Test a selector against the live page in the console
Basic selectors
Selector
Matches
Example
*
any element
div.card > *
tag
elements by tag name
h2
.cls
class list contains cls
.price
.a.b
has both classes (no space)
li.item.featured
#id
element with that id
#main
A, B
matches A or B, returned in document order
h1, h2, h3
tag.cls
tag with the class
span.price
Attribute selectors
Selector
Attribute value...
Matches
Does not match
[href]
exists (even empty)
<a href="">
<a>
[data-state="open"]
equals exactly
open
opened, Open
[class~="ad"]
whitespace-separated list contains the word
class="item ad"
class="adv"
[lang|="en"]
equals en or starts with en-
en, en-US
english
[href^="https"]
starts with
https://x.com
/https
[href$=".pdf"]
ends with
/a.pdf
/a.pdf?x=1
[href*="/product/"]
contains the substring
/shop/product/42
/products
[href$=".pdf" i]
i flag: ASCII case-insensitive
.PDF, .Pdf
[data-x="v" s]
s flag: force case-sensitive
soupsieve supports it; browser support is patchy
Quote values that contain spaces or punctuation: [data-x="has space"]. Unquoted, soupsieve raises SelectorSyntaxError.
[class="item"] compares the whole attribute string: it matches class="item" but not class="item ad". Use .item (or [class~="item"]) instead.
[class^="Card_"] only works if that class comes first in the attribute. For hashed CSS-module classes such as ProductCard_title__x9Az, use [class*="ProductCard_title"].
Combinators
Selector
Meaning
Example
A B
B anywhere inside A
div.card a
A > B
B is a direct child of A
ul.results > li
A + B
B immediately follows A (same parent)
h2 + p
A ~ B
B is any later sibling of A
h2 ~ p
CSS has no parent or previous-sibling combinator. Use :has() instead: div:has(> h2) selects the div that contains the h2, h2:has(+ ul) selects the h2 sitting right before a list. Or switch to XPath (.., preceding-sibling::).
:nth-child and friends
Counting is 1-based and counts element siblings only (text nodes are ignored). An+B is evaluated for n = 0, 1, 2, ...
Selector
Selects
:first-child / :last-child
first / last child of its parent
:only-child
the parent's only element child
:nth-child(3)
the 3rd child
:nth-child(odd) = :nth-child(2n+1)
1st, 3rd, 5th, ...
:nth-child(even) = :nth-child(2n)
2nd, 4th, 6th, ...
:nth-child(n+3)
3rd child onwards (skip the first two)
:nth-child(-n+3)
first three children
:nth-child(3n)
every third child
:nth-last-child(2)
second from the end
:nth-child(2 of .item)
2nd sibling among those matching .item
:nth-of-type(2)
2nd sibling of the same tag
:first-of-type, :last-of-type, :only-of-type
same, by tag
:empty
no children (see whitespace note below)
:root
the html element
Classic trap:li.item:nth-child(2) means "the 2nd child, if it is an li.item", not "the 2nd li.item". If an ad or a header sits first, you get the wrong row or nothing. Use :nth-child(2 of li.item) or index in Python: soup.select("li.item")[1].
Logical pseudo-classes
Selector
Selects
Example
:not(A, B)
matches none of the list
li:not(.ad, .sponsored)
:is(A, B)
matches any of the list
:is(h1, h2, h3) > a
:where(A, B)
like :is(), zero specificity (irrelevant when scraping)
:where(.a, .b)
:has(> img)
has a direct child img
li:has(> img)
:has(img)
has a descendant img
div.card:has(.price)
:has(+ .x)
immediately followed by .x
h2:has(+ ul)
:not(:has(a))
contains no link
li:not(:has(a))
Text, form and link pseudo-classes
Selector
Where
Selects
:-soup-contains("Price")
soupsieve only
text anywhere inside contains the substring (case-sensitive); several strings mean "any of"
:-soup-contains-own("Price")
soupsieve only
only the element's own text nodes, not its children's
language from the nearest lang attribute (matches en-US too)
:hover, :focus, :visited
browsers
parse in soupsieve but never match: static HTML has no live state
BeautifulSoup (soupsieve) vs browsers
Checked against soupsieve 2.6. Browser columns describe current Chrome, Firefox and Safari.
Feature
soupsieve (BeautifulSoup)
Browsers
All attribute operators + i flag
yes
yes
s flag
yes
not in every browser
:is(), :where(), :not(A, B)
yes
yes
:has()
yes
yes (Chrome 105+, Safari 15.4+, Firefox 121+)
:nth-child(An+B of S)
yes
yes in current versions
Text matching
:-soup-contains()
none in CSS (Playwright adds :has-text())
Pseudo-elements (::before, ::text)
NotImplementedError
styling only, never selectable
:hover, :focus, :visited
never match
live state
:empty on whitespace-only content
matches (whitespace ignored)
does not match
#123 (id starts with a digit)
SelectorSyntaxError
SyntaxError too
Escaping
[id="123"], #\31 23, or soupsieve.escape("123")
CSS.escape("123")
<tbody>
present only if in the source (or with the html5lib parser)
always inserted
The tbody trap: browsers insert <tbody> into every table, so a selector copied from DevTools (table > tbody > tr) finds nothing in the raw HTML parsed by html.parser or lxml. Use table tr (descendant) or table tr:has(td), which work either way.
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, "lxml") # or "html.parser"
rows = []
for card in soup.select("div.product:not(.ad)"):
link = card.select_one("h2 a")
price = card.select_one("[itemprop=price], .price")
if link is None: # select_one -> None on no match
continue
rows.append({
"name": link.get_text(" ", strip=True),
"url": link.get("href"), # .get() -> None instead of KeyError
"price": price.get_text(strip=True) if price else None,
})
Pitfalls
get_text(strip=True) strips each text node and glues them: <p>Hello <b>big</b> world</p> becomes Hellobigworld. Use get_text(" ", strip=True).
DevTools "Copy selector" gives brittle paths like #root > div:nth-child(3) > div > span. Rewrite them with classes, attributes or :has().
Content missing? If a selector works in DevTools but not on requests HTML, the data is rendered by JavaScript. Check view-source, then look for the JSON API or use a browser.
Searching inside a card:card.select("a") is scoped to the card. Calling soup.select(...) inside the loop searches the whole page again.
Class names with a colon or other special characters need escaping: .a\:b (or soupsieve.escape("a:b")).
Nothing on this sheet matches.
Source: simpleprog.com/cheatsheets/css-selectors · free courses at simpleprog.com