How to use it
- Paste the HTML you want to scrape. For a real page, use the raw HTML your scraper downloads (
view-source:orprint(response.text)), not the DevTools Elements panel, which shows the page after JavaScript ran. - Choose CSS selector or XPath and type an expression. Results update as you type.
- Each match shows its text and its HTML. Below the count you get the same query for BeautifulSoup, lxml or Scrapy, ready to paste.
CSS or XPath?
| Task | CSS | XPath |
|---|---|---|
| By class | .price | //*[contains(concat(' ', @class, ' '), ' price ')] |
| By attribute | [data-id="102"] | //*[@data-id="102"] |
| Descendant / child | article a / h2 > a | //article//a / //h2/a |
| Nth item | article:nth-of-type(2) | (//article)[2] |
| By text | not in standard CSS | //a[contains(text(), "hub")] |
| Parent of a match | article:has(.out) | //span[@class="stock out"]/.. |
| Attribute value | select, then read ["href"] | //a/@href |
CSS is shorter and easier to read; XPath can move up the tree, match on text and return attributes or counts directly. Most scrapers use CSS for most fields and XPath for the awkward ones.
Common mistakes
- Testing against the rendered DOM. If the data appears only after JavaScript runs, your HTTP scraper will not see it. Look for the JSON API the page calls instead.
tbodysurprises. Browsers (and this tester) add<tbody>to tables; lxml does not. Prefertable trovertable > tbody > tr.- Matching the class attribute exactly in XPath.
@class="stock"fails onclass="stock in". Usecontains()or a CSS class selector. - Brittle, generated class names. Classes like
css-1x2y3zchange on every deploy. Anchor on stable attributes (data-*,itemprop, ids) or on text labels. - Forgetting XPath indexes start at 1.
(//article)[1]is the first match, and//article[1]means "every article that is the first article child of its parent".
FAQ
Why does my selector work here but not in my scraper?
The usual cause is that the page you scrape is not the page you see: content added by JavaScript is missing from the raw HTML that requests downloads. Paste the HTML from view-source: (or response.text), not from the DevTools Elements panel.
Does the browser's parser match BeautifulSoup and lxml?
Mostly. This tester uses the browser's HTML5 parser, like BeautifulSoup with html5lib. lxml and html.parser repair broken HTML differently, and browsers insert tbody into tables, so a selector with tbody may fail in lxml.
Can I select attributes or text with XPath?
Yes. //a/@href returns attribute values and //h2/text() returns text nodes; both are shown as plain values. Functions such as count(//a) return a single number.
Does CSS support :contains() or :has()?
:has() works in modern browsers and in soupsieve (BeautifulSoup). :contains() is not standard CSS; soupsieve has :-soup-contains(), and in XPath you would use contains(text(), '...').
Is my HTML sent anywhere?
No. Parsing and matching happen in JavaScript in this page, with DOMParser, querySelectorAll and document.evaluate. Scripts in the pasted HTML are not executed.