Where you use XPath

Tool Call Returns
lxml tree.xpath("//h1/text()") a list (or float / str / bool for scalar expressions)
parsel / Scrapy response.xpath("//h1/text()").get() / .getall() first string or None / list of strings
Playwright page.locator("xpath=//h1") locator (strings starting with // are treated as XPath)
Selenium driver.find_elements(By.XPATH, "//h1") elements
Browser console $x("//h1") array of nodes
BeautifulSoup not supported parse with lxml.html instead

Path syntax

Expression Means
/html/body absolute path from the document root
//div any div at any depth
.//a any a below the current node
.. parent
. current node
* any element
@href / @* an attribute / all attributes
text() the node's direct child text nodes
node() any node (element, text, comment)
//a/@href the href values themselves (strings)
//p//text() every text node inside each p

Loop trap: inside a loop, card.xpath("//a") searches the whole document again, because // starts at the root. Use card.xpath(".//a").

Axes

Write axis::nodetest. On reverse axes (ancestor, preceding*), [1] is the nearest node.

Axis Selects Example
child:: (default) children //ul/li = //ul/child::li
descendant:: children, grandchildren, ... //div[@id="main"]/descendant::b
descendant-or-self:: the node and all descendants // is short for /descendant-or-self::node()/
parent:: (..) the parent //b/parent::p
ancestor:: parent, grandparent, ... //b/ancestor::*[@id][1] (nearest ancestor with an id)
ancestor-or-self:: the node and its ancestors ancestor-or-self::*[@lang][1]/@lang (effective language)
following-sibling:: later siblings //dt[.="Price"]/following-sibling::dd[1]
preceding-sibling:: earlier siblings preceding-sibling::*[1] (the one right before)
following:: everything after the node (not its descendants) //h2/following::table[1]
preceding:: everything before the node (not its ancestors) //li[2]/preceding::h1
self:: the node itself, filtered by name //*[self::h2 or self::h3]
attribute:: (@) attributes //img/attribute::src = //img/@src

Predicates

Predicate Keeps
[1] the first (XPath is 1-based)
[last()], [last()-1] the last, the second to last
[position() <= 3] the first three
[position() > 1] all but the first (skip a header row)
[@class] / [not(@class)] has / lacks the attribute
[@id="main"] attribute equals
[a] / [.//img] has a child a / has a descendant img
[@data-price > 5] numeric comparison (the string is converted to a number)
[@a and @b], [@a or @b] boolean logic
[@class="x"][2] the 2nd among those with class="x"
[2][@class="x"] the 2nd, only if it has class="x"

Position trap: //li[1] is the first li in each list. (//li)[1] is the first li in the document. Same for [last()].

Functions (XPath 1.0)

Function Does Example
contains(s, sub) substring test, case-sensitive //a[contains(@href, "/product/")]
starts-with(s, prefix) prefix test //a[starts-with(@href, "http")]
normalize-space(s) trim + collapse whitespace; no argument = current node //a[normalize-space()="Next"]
string(node) full text of the node and its descendants string(//p[@class="price"])
text() direct text nodes only //h1/text()
concat(a, b, ...) join strings concat(//h1, " - ", //h2)
substring(s, start, len?) 1-based slice substring(@datetime, 1, 10)
substring-before(s, sep) / substring-after(s, sep) split once substring-after(@href, "page=")
string-length(s) length; no argument = current node //p[string-length() > 200]
translate(s, from, to) character replace (used for lowercasing) see below
count(nodes) number of nodes (a float in lxml: 5.0) //ul[count(li) > 2]
position(), last() index, size of the context //tr[position() > 1]
not(x) negation //p[not(ancestor::aside)]
sum(), number(), round(), floor(), ceiling() maths sum(//span/@data-price)
name(), local-name() element name (without prefix) //*[local-name()="svg"]

text() vs string value

For <p>Price: <b>$10</b> today</p>:

Expression Result
//p/text() ['Price: ', ' today'] (two text nodes, the b is skipped)
//p//text() ['Price: ', '$10', ' today']
string(//p) 'Price: $10 today'
//p[contains(text(), "today")] nothing: in XPath 1.0 only the first text node is tested
//p[contains(., "today")] the p (. is the full string value)
//a[text()="Next"] on <a> Next </a> nothing: whitespace counts
//a[normalize-space()="Next"] the link

Rule of thumb: test with . or normalize-space(), extract with text() or string().

Classes, case and "ends with"

Need XPath
class token price (exact, not price-old) //*[contains(concat(" ", normalize-space(@class), " "), " price ")]
class contains text (also matches price-old) //*[contains(@class, "price")]
case-insensitive compare //a[contains(translate(., "ABCDEFGHIJKLMNOPQRSTUVWXYZ", "abcdefghijklmnopqrstuvwxyz"), "next")]
ends with .pdf //a[substring(@href, string-length(@href) - 3) = ".pdf"] (3 = length of ".pdf" minus 1)
ends with, any case //a[translate(substring(@href, string-length(@href) - 3), "PDF", "pdf") = ".pdf"]

Unions

Expression Returns
//h1 | //h2 all h1 and h2, in document order
(//h1 | //h2)[1] the first heading of either kind
//*[self::h2 or self::h3] same idea, single path
//img/@src | //img/@data-src both attributes

Scraping recipes

Goal XPath
Value after a label (dl) //dt[normalize-space()="Price"]/following-sibling::dd[1]/text()
Value after a label (th/td) //th[normalize-space()="Weight"]/following-sibling::td[1]
Other cell in the same row //td[.="SKU-42"]/../td[3]/text()
Column by header name //tr[td]/td[count(../../tr/th[.="Price"]/preceding-sibling::th) + 1]/text()
Data rows, skip the header //table//tr[position() > 1]
Next-page link //a[@rel="next"]/@href | //a[normalize-space()="Next"]/@href
Query param from a link substring-after(//a[@rel="next"]/@href, "p=")
Open Graph title //meta[@property="og:title"]/@content
JSON-LD //script[@type="application/ld+json"]/text()
Main text, minus sidebars //article//p[not(ancestor::aside)]
Nearest container with an id ancestor::*[@id][1]
Lists with more than 2 items //ul[count(li) > 2]
A list containing a given item //ul[li[.="e"]]

Python pattern

from lxml import html

tree = html.fromstring(page_html)
tree.make_links_absolute("https://example.com/")      # resolve relative hrefs

for row in tree.xpath('//table[@id="specs"]//tr[td]'):
    label = row.xpath("string(./th)").strip()          # note the leading "./"
    value = row.xpath("normalize-space(./td[1])")
    print(label, value)

titles = tree.xpath("//h2/a/text()")                   # list of strings
count = tree.xpath("count(//li)")                      # 5.0 (float)
has_h1 = tree.xpath("boolean(//h1)")                   # True / False

Text results are _ElementUnicodeResult, a str subclass with .getparent(), so you can step back to the element a string came from.

XPath 1.0 limits (lxml and browsers)

lxml (libxml2) and every browser's document.evaluate implement XPath 1.0 only. These XPath 2.0+ functions raise XPathEvalError: Unregistered function in lxml:

Missing Workaround
ends-with() substring(s, string-length(s) - n + 1) = "suffix"
lower-case(), upper-case() translate() with both alphabets
matches(), replace(), tokenize() lxml EXSLT regex (below), or filter in Python
string-join() " ".join(el.xpath(".//text()")) in Python
if ... then ... else, sequences, distinct-values() do it in Python

lxml's EXSLT regular expressions:

NS = {"re": "http://exslt.org/regular-expressions"}
pdfs = tree.xpath(r'//a[re:test(@href, "\.pdf$", "i")]/@href', namespaces=NS)

In the browser: document.evaluate('//h1', document, null, XPathResult.ORDERED_NODE_SNAPSHOT_TYPE, null), then snapshotItem(i).

Namespaces

  • HTML parsed with lxml.html has no namespaces: //div just works.
  • XML with a default namespace (sitemaps, RSS extensions, XHTML) does not match //loc. Map a prefix, or ignore namespaces with local-name():
from lxml import etree
root = etree.fromstring(sitemap_bytes)
root.xpath("//loc/text()")                                     # [] - namespaced
root.xpath("//sm:loc/text()",
           namespaces={"sm": "http://www.sitemaps.org/schemas/sitemap/0.9"})
root.xpath('//*[local-name()="loc"]/text()')                   # works too
  • In browsers, inline <svg> elements live in the SVG namespace, so //svg misses them; use //*[local-name()="svg"].

CSS to XPath

CSS XPath
div //div
div > p //div/p
div p //div//p
#main //*[@id="main"]
.price //*[contains(concat(" ", normalize-space(@class), " "), " price ")]
a[href] //a[@href]
a[href="/x"] //a[@href="/x"]
a[href^="http"] //a[starts-with(@href, "http")]
a[href$=".pdf"] //a[substring(@href, string-length(@href) - 3) = ".pdf"]
a[href*="shop"] //a[contains(@href, "shop")]
h2 + p //h2/following-sibling::*[1][self::p]
h2 ~ p //h2/following-sibling::p
h2, h3 //h2 | //h3
li:first-child //li[not(preceding-sibling::*)]
li:last-child //li[not(following-sibling::*)]
li:nth-child(3) //*[3][self::li]
li:nth-child(odd) //li[count(preceding-sibling::*) mod 2 = 0]
li:first-of-type //li[1]
li:nth-of-type(3) //li[3]
li:not(.ad) //li[not(contains(concat(" ", normalize-space(@class), " "), " ad "))]
ul:has(> li) //ul[li]
div:has(img) //div[.//img]
:-soup-contains("x") //*[contains(., "x")]
(no CSS) parent //span/..
(no CSS) previous sibling //h2/preceding-sibling::*[1]
(no CSS) exact text //a[normalize-space()="Next"]

The rows that have a CSS equivalent were checked to return the same elements in lxml as the selector does in soupsieve.