cheat sheet
XPath Cheat Sheet
The XPath you actually use when scraping, with the traps that cost hours (text() vs string value, //li[1] vs (//li)[1], absolute paths inside loops). Every expression here was run with lxml 6, which implements XPath 1.0 just like browsers do.
by Muhammad Usman · updated Oct 08, 2026
Where you use XPath
Tool
Call
Returns
lxml
tree.xpath("//h1/text()")
a list (or float / str / bool for scalar expressions)
parsel / Scrapy
response.xpath("//h1/text()").get() / .getall()
first string or None / list of strings
Playwright
page.locator("xpath=//h1")
locator (strings starting with // are treated as XPath)
Selenium
driver.find_elements(By.XPATH, "//h1")
elements
Browser console
$x("//h1")
array of nodes
BeautifulSoup
not supported
parse with lxml.html instead
Path syntax
Expression
Means
/html/body
absolute path from the document root
//div
any div at any depth
.//a
any a below the current node
..
parent
.
current node
*
any element
@href / @*
an attribute / all attributes
text()
the node's direct child text nodes
node()
any node (element, text, comment)
//a/@href
the href values themselves (strings)
//p//text()
every text node inside each p
Loop trap: inside a loop, card.xpath("//a") searches the whole document again, because // starts at the root. Use card.xpath(".//a").
Axes
Write axis::nodetest. On reverse axes (ancestor, preceding*), [1] is the nearest node.
Axis
Selects
Example
child:: (default)
children
//ul/li = //ul/child::li
descendant::
children, grandchildren, ...
//div[@id="main"]/descendant::b
descendant-or-self::
the node and all descendants
// is short for /descendant-or-self::node()/
parent:: (..)
the parent
//b/parent::p
ancestor::
parent, grandparent, ...
//b/ancestor::*[@id][1] (nearest ancestor with an id)
ancestor-or-self::
the node and its ancestors
ancestor-or-self::*[@lang][1]/@lang (effective language)
following-sibling::
later siblings
//dt[.="Price"]/following-sibling::dd[1]
preceding-sibling::
earlier siblings
preceding-sibling::*[1] (the one right before)
following::
everything after the node (not its descendants)
//h2/following::table[1]
preceding::
everything before the node (not its ancestors)
//li[2]/preceding::h1
self::
the node itself, filtered by name
//*[self::h2 or self::h3]
attribute:: (@)
attributes
//img/attribute::src = //img/@src
Predicates
Predicate
Keeps
[1]
the first (XPath is 1-based)
[last()], [last()-1]
the last, the second to last
[position() <= 3]
the first three
[position() > 1]
all but the first (skip a header row)
[@class] / [not(@class)]
has / lacks the attribute
[@id="main"]
attribute equals
[a] / [.//img]
has a child a / has a descendant img
[@data-price > 5]
numeric comparison (the string is converted to a number)
[@a and @b], [@a or @b]
boolean logic
[@class="x"][2]
the 2nd among those with class="x"
[2][@class="x"]
the 2nd, only if it has class="x"
Position trap: //li[1] is the first li in each list . (//li)[1] is the first li in the document. Same for [last()].
Functions (XPath 1.0)
Function
Does
Example
contains(s, sub)
substring test, case-sensitive
//a[contains(@href, "/product/")]
starts-with(s, prefix)
prefix test
//a[starts-with(@href, "http")]
normalize-space(s)
trim + collapse whitespace; no argument = current node
//a[normalize-space()="Next"]
string(node)
full text of the node and its descendants
string(//p[@class="price"])
text()
direct text nodes only
//h1/text()
concat(a, b, ...)
join strings
concat(//h1, " - ", //h2)
substring(s, start, len?)
1-based slice
substring(@datetime, 1, 10)
substring-before(s, sep) / substring-after(s, sep)
split once
substring-after(@href, "page=")
string-length(s)
length; no argument = current node
//p[string-length() > 200]
translate(s, from, to)
character replace (used for lowercasing)
see below
count(nodes)
number of nodes (a float in lxml: 5.0)
//ul[count(li) > 2]
position(), last()
index, size of the context
//tr[position() > 1]
not(x)
negation
//p[not(ancestor::aside)]
sum(), number(), round(), floor(), ceiling()
maths
sum(//span/@data-price)
name(), local-name()
element name (without prefix)
//*[local-name()="svg"]
text() vs string value
For <p>Price: <b>$10</b> today</p>:
Expression
Result
//p/text()
['Price: ', ' today'] (two text nodes, the b is skipped)
//p//text()
['Price: ', '$10', ' today']
string(//p)
'Price: $10 today'
//p[contains(text(), "today")]
nothing : in XPath 1.0 only the first text node is tested
//p[contains(., "today")]
the p (. is the full string value)
//a[text()="Next"] on <a> Next </a>
nothing : whitespace counts
//a[normalize-space()="Next"]
the link
Rule of thumb: test with . or normalize-space(), extract with text() or string().
Classes, case and "ends with"
Need
XPath
class token price (exact, not price-old)
//*[contains(concat(" ", normalize-space(@class), " "), " price ")]
class contains text (also matches price-old)
//*[contains(@class, "price")]
case-insensitive compare
//a[contains(translate(., "ABCDEFGHIJKLMNOPQRSTUVWXYZ", "abcdefghijklmnopqrstuvwxyz"), "next")]
ends with .pdf
//a[substring(@href, string-length(@href) - 3) = ".pdf"] (3 = length of ".pdf" minus 1)
ends with, any case
//a[translate(substring(@href, string-length(@href) - 3), "PDF", "pdf") = ".pdf"]
Unions
Expression
Returns
//h1 | //h2
all h1 and h2, in document order
(//h1 | //h2)[1]
the first heading of either kind
//*[self::h2 or self::h3]
same idea, single path
//img/@src | //img/@data-src
both attributes
Scraping recipes
Goal
XPath
Value after a label (dl)
//dt[normalize-space()="Price"]/following-sibling::dd[1]/text()
Value after a label (th/td)
//th[normalize-space()="Weight"]/following-sibling::td[1]
Other cell in the same row
//td[.="SKU-42"]/../td[3]/text()
Column by header name
//tr[td]/td[count(../../tr/th[.="Price"]/preceding-sibling::th) + 1]/text()
Data rows, skip the header
//table//tr[position() > 1]
Next-page link
//a[@rel="next"]/@href | //a[normalize-space()="Next"]/@href
Query param from a link
substring-after(//a[@rel="next"]/@href, "p=")
Open Graph title
//meta[@property="og:title"]/@content
JSON-LD
//script[@type="application/ld+json"]/text()
Main text, minus sidebars
//article//p[not(ancestor::aside)]
Nearest container with an id
ancestor::*[@id][1]
Lists with more than 2 items
//ul[count(li) > 2]
A list containing a given item
//ul[li[.="e"]]
Python pattern
from lxml import html
tree = html.fromstring(page_html)
tree.make_links_absolute("https://example.com/") # resolve relative hrefs
for row in tree.xpath('//table[@id="specs"]//tr[td]'):
label = row.xpath("string(./th)").strip() # note the leading "./"
value = row.xpath("normalize-space(./td[1])")
print(label, value)
titles = tree.xpath("//h2/a/text()") # list of strings
count = tree.xpath("count(//li)") # 5.0 (float)
has_h1 = tree.xpath("boolean(//h1)") # True / False
Text results are _ElementUnicodeResult, a str subclass with .getparent(), so you can step back to the element a string came from.
XPath 1.0 limits (lxml and browsers)
lxml (libxml2) and every browser's document.evaluate implement XPath 1.0 only . These XPath 2.0+ functions raise XPathEvalError: Unregistered function in lxml:
Missing
Workaround
ends-with()
substring(s, string-length(s) - n + 1) = "suffix"
lower-case(), upper-case()
translate() with both alphabets
matches(), replace(), tokenize()
lxml EXSLT regex (below), or filter in Python
string-join()
" ".join(el.xpath(".//text()")) in Python
if ... then ... else, sequences, distinct-values()
do it in Python
lxml's EXSLT regular expressions:
NS = {"re": "http://exslt.org/regular-expressions"}
pdfs = tree.xpath(r'//a[re:test(@href, "\.pdf$", "i")]/@href', namespaces=NS)
In the browser: document.evaluate('//h1', document, null, XPathResult.ORDERED_NODE_SNAPSHOT_TYPE, null), then snapshotItem(i).
Namespaces
HTML parsed with lxml.html has no namespaces: //div just works.
XML with a default namespace (sitemaps, RSS extensions, XHTML) does not match //loc. Map a prefix, or ignore namespaces with local-name():
from lxml import etree
root = etree.fromstring(sitemap_bytes)
root.xpath("//loc/text()") # [] - namespaced
root.xpath("//sm:loc/text()",
namespaces={"sm": "http://www.sitemaps.org/schemas/sitemap/0.9"})
root.xpath('//*[local-name()="loc"]/text()') # works too
In browsers, inline <svg> elements live in the SVG namespace, so //svg misses them; use //*[local-name()="svg"].
CSS to XPath
CSS
XPath
div
//div
div > p
//div/p
div p
//div//p
#main
//*[@id="main"]
.price
//*[contains(concat(" ", normalize-space(@class), " "), " price ")]
a[href]
//a[@href]
a[href="/x"]
//a[@href="/x"]
a[href^="http"]
//a[starts-with(@href, "http")]
a[href$=".pdf"]
//a[substring(@href, string-length(@href) - 3) = ".pdf"]
a[href*="shop"]
//a[contains(@href, "shop")]
h2 + p
//h2/following-sibling::*[1][self::p]
h2 ~ p
//h2/following-sibling::p
h2, h3
//h2 | //h3
li:first-child
//li[not(preceding-sibling::*)]
li:last-child
//li[not(following-sibling::*)]
li:nth-child(3)
//*[3][self::li]
li:nth-child(odd)
//li[count(preceding-sibling::*) mod 2 = 0]
li:first-of-type
//li[1]
li:nth-of-type(3)
//li[3]
li:not(.ad)
//li[not(contains(concat(" ", normalize-space(@class), " "), " ad "))]
ul:has(> li)
//ul[li]
div:has(img)
//div[.//img]
:-soup-contains("x")
//*[contains(., "x")]
(no CSS) parent
//span/..
(no CSS) previous sibling
//h2/preceding-sibling::*[1]
(no CSS) exact text
//a[normalize-space()="Next"]
The rows that have a CSS equivalent were checked to return the same elements in lxml as the selector does in soupsieve.
Nothing on this sheet matches. Clear the filter