Scraping with Node.js: Cheerio and Playwright
The Same Fundamentals in JavaScript
None of this course's fundamentals are Python-specific. Node.js is a strong alternative when your team already writes JavaScript, when the scraper lives inside a Node backend, or when you need to run a site's own JavaScript. Here is the core toolkit in Node: fetch plus Cheerio for static pages, p-limit for bounded concurrency, Playwright for pages that need a browser, and the Node-specific traps that catch Python developers.
The examples use books.toscrape.com and quotes.toscrape.com, sandbox sites built for scraping practice. They need Node 18 or newer (for built-in fetch) and ES modules:
mkdir scraper && cd scraper
npm init -y
npm pkg set type=module # allow import syntax and top-level await
npm install cheerio p-limit playwright
npx playwright install chromium
p-limit is published only as an ES module, which is one reason to set "type": "module" from the start rather than mixing require and import.
Static Pages: fetch and Cheerio
Cheerio parses HTML into a tree and gives you a jQuery-style API over it. It does not run JavaScript or compute layout, which is why it is fast.
import * as cheerio from "cheerio";
const BASE = "https://books.toscrape.com/";
const UA = "simpleprog-tutorial/1.0 (contact@example.org)";
async function getHtml(url) {
const res = await fetch(url, {
headers: { "User-Agent": UA, Accept: "text/html" },
signal: AbortSignal.timeout(15_000),
});
if (!res.ok) throw new Error(`${res.status} ${res.statusText} for ${url}`);
return res.text();
}
const $ = cheerio.load(await getHtml(BASE));
const books = $("article.product_pod").map((_, el) => {
const card = $(el);
const link = card.find("h3 a");
return {
title: link.attr("title"),
url: new URL(link.attr("href"), BASE).href,
price: card.find(".price_color").text().trim(),
rating: card.find("p.star-rating").attr("class").replace("star-rating", "").trim(),
};
}).get();
console.log(books.length, books.slice(0, 2));
const next = $("li.next a").attr("href");
console.log("next page:", next && new URL(next, BASE).href);
20 [
{
title: 'A Light in the Attic',
url: 'https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html',
price: '£51.77',
rating: 'Three'
},
{
title: 'Tipping the Velvet',
url: 'https://books.toscrape.com/catalogue/tipping-the-velvet_999/index.html',
price: '£53.74',
rating: 'One'
}
]
next page: https://books.toscrape.com/catalogue/page-2.html
The lines that matter:
AbortSignal.timeout(15_000).fetchhas no default timeout. Without a signal, one stalled server hangs that request forever.if (!res.ok). Likerequests,fetchdoes not throw on 404 or 500; it only rejects on network errors. Check the status yourself.new URL(href, base)is theurljoinof JavaScript. It resolves../and relative paths correctly..map(...).get(). Cheerio's.mapreturns a Cheerio object;.get()turns it into a plain array.
Two Cheerio behaviours surprise people coming from BeautifulSoup. .text() on a selection of several elements returns all their text concatenated, not a list, so $(".price_color").text() on a listing gives £51.77£53.74.... And .attr() returns the attribute of the first element only. Iterate with .map or .each whenever you expect more than one match.
Cheerio supports CSS selectors only; for XPath you need a different library, one area where Python's lxml and parsel are ahead.
The Encoding Trap in fetch
res.text() always decodes the body as UTF-8, whatever the Content-Type header or <meta charset> says. On a site still served in windows-1252 you get replacement characters, silently:
import * as cheerio from "cheerio";
// "Café – 5 €" encoded as windows-1252, with a matching <meta charset>
const bytes = Buffer.from([
...Buffer.from('<html><head><meta charset="windows-1252"></head><body><p>Caf'),
0xe9, 0x20, 0x96, 0x20, 0x35, 0x20, 0x80,
...Buffer.from("</p></body></html>"),
]);
const res = new Response(bytes, { headers: { "Content-Type": "text/html" } });
console.log("res.text(): ", cheerio.load(await res.clone().text())("p").text());
console.log("loadBuffer():", cheerio.loadBuffer(Buffer.from(await res.arrayBuffer()))("p").text());
res.text(): Caf� � 5 �
loadBuffer(): Café – 5 €
cheerio.loadBuffer takes raw bytes and sniffs the encoding from a byte order mark or <meta charset>. If the response header names a charset, pass it too: cheerio.loadBuffer(buf, { encoding: { transportLayerEncodingLabel: "windows-1252" } }). The rules behind this are in parsing engines and text encoding. For crawlers that touch many sites, use arrayBuffer() plus loadBuffer everywhere and never res.text().
Bounded Concurrency With p-limit
Node is asynchronous by default, which makes it easy to overload a site by accident: await Promise.all(urls.map(u => fetch(u))) starts every request at once. p-limit caps how many run at the same time. Combined with a delay and retries it gives you a small, polite crawler:
import * as cheerio from "cheerio";
import pLimit from "p-limit";
import { setTimeout as sleep } from "node:timers/promises";
const UA = "simpleprog-tutorial/1.0 (contact@example.org)";
const limit = pLimit(2); // at most 2 requests in flight
const DELAY_MS = 500; // pause each slot after every request
async function fetchPage(url, attempts = 4) {
for (let i = 0; i < attempts; i++) {
try {
const res = await fetch(url, { headers: { "User-Agent": UA }, signal: AbortSignal.timeout(15_000) });
if (res.status === 429 || res.status >= 500) throw new Error(`retryable ${res.status}`);
if (!res.ok) throw Object.assign(new Error(`HTTP ${res.status}`), { fatal: true });
return cheerio.loadBuffer(Buffer.from(await res.arrayBuffer()));
} catch (err) {
if (err.fatal || i === attempts - 1) throw err;
await sleep(1000 * 2 ** i); // 1s, 2s, 4s
}
}
}
async function scrapeListing(url) {
const $ = await fetchPage(url);
await sleep(DELAY_MS);
return $("article.product_pod").map((_, el) => ({
title: $(el).find("h3 a").attr("title"),
price: $(el).find(".price_color").text(),
})).get();
}
const urls = [1, 2, 3, 4, 5, 999].map(n => `https://books.toscrape.com/catalogue/page-${n}.html`);
const results = await Promise.allSettled(urls.map(u => limit(() => scrapeListing(u))));
results.forEach((r, i) => {
if (r.status === "fulfilled") console.log(urls[i], r.value.length, "books");
else console.log(urls[i], "FAILED:", r.reason.message);
});
https://books.toscrape.com/catalogue/page-1.html 20 books
https://books.toscrape.com/catalogue/page-2.html 20 books
https://books.toscrape.com/catalogue/page-3.html 20 books
https://books.toscrape.com/catalogue/page-4.html 20 books
https://books.toscrape.com/catalogue/page-5.html 20 books
https://books.toscrape.com/catalogue/page-999.html FAILED: HTTP 404
Why it is built this way:
limit(() => scrapeListing(u))wraps a function, not a promise.limit(scrapeListing(u))would start the request immediately.- The sleep sits inside the limited function, so each of the two slots pauses after its request. That caps the rate at roughly two requests per (latency + 0.5 s). Concurrency alone is not a rate limit; see rate limiting and politeness.
Promise.allSettled, notPromise.all.Promise.allrejects on the first failure and you lose the other results; withallSettledthe 404 is one failed row, not a failed crawl.- 404 is fatal, 429 and 5xx are retried with exponential backoff, the same policy as the Python lessons.
For crawls with a growing frontier (URLs discovered while crawling), mapping over a fixed list does not work. Use a queue with a fixed number of workers, exactly the pattern from parallel scraping with asyncio, or a framework such as Crawlee, which is to Node roughly what Scrapy is to Python.
Pages That Need a Browser: Playwright for Node
Playwright is written in TypeScript and its Node API is the reference version; the Python API from browser automation mirrors it almost name for name, with await in place of sync calls and camelCase names. Two techniques cover most cases: read the rendered DOM, or capture the JSON the page loads.
import { chromium } from "playwright";
const browser = await chromium.launch({ headless: true });
try {
const context = await browser.newContext({ userAgent: "simpleprog-tutorial/1.0 (contact@example.org)" });
await context.route("**/*.{png,jpg,jpeg,gif,woff,woff2}", route => route.abort());
const page = await context.newPage();
// 1. Rendered DOM: the quotes only exist after the page's script runs
await page.goto("https://quotes.toscrape.com/js/", { waitUntil: "domcontentloaded" });
await page.locator("div.quote").first().waitFor();
const quotes = await page.locator("div.quote").evaluateAll(nodes =>
nodes.map(n => ({
text: n.querySelector(".text")?.textContent.trim(),
author: n.querySelector(".author")?.textContent.trim(),
})));
console.log(quotes.length, quotes[0].author);
// 2. Network capture: the infinite-scroll page loads JSON from an API
const batches = [];
page.on("response", async res => {
if (res.url().includes("/api/quotes") && res.ok()) batches.push(await res.json());
});
await page.goto("https://quotes.toscrape.com/scroll", { waitUntil: "networkidle" });
for (let i = 0; i < 2; i++) {
const next = page.waitForResponse(r => r.url().includes("/api/quotes"));
await page.mouse.wheel(0, 20_000);
await next;
}
await page.waitForTimeout(500); // let the async response handlers finish
console.log(batches.map(b => `page ${b.page}: ${b.quotes.length} quotes`));
} finally {
await browser.close();
}
10 Albert Einstein
[ 'page 1: 10 quotes', 'page 2: 10 quotes', 'page 3: 10 quotes', 'page 4: 10 quotes' ]
The number of batches varies between runs (one large scroll can trigger more than one load), so stop on the API's own has_next field in real code rather than on a scroll count.
Notes:
evaluateAllruns in the page. The callback is serialised into the browser, so it cannot see your Node variables (pass them as a second argument). One call returning plain objects beats a round trip per element.waitForResponseis created before the scroll. Creating it after could miss a response that arrives quickly.try/finallyaround the browser. An exception that skipsbrowser.close()leaves a Chromium process running.- Once you have seen
/api/quotes?page=N, you may not need the browser at all. Call that endpoint withfetchand skip rendering entirely (see XHR and API reverse engineering). The browser was the discovery tool.
Node Pitfalls Worth Knowing
- A missing
awaitgives you a pending Promise instead of data, and an error inside it becomes an unhandled rejection that can crash the process. Lint with a rule that flags floating promises. - No cookie jar. Node's
fetchdoes not keep cookies between requests. For sessions, add a jar (for example thefetch-cookiepackage withtough-cookie) or copySet-Cookievalues yourself. - Proxies. Node's
fetchis built on undici; route it through a proxy withfetch(url, { dispatcher: new ProxyAgent("http://host:port") }), importingProxyAgentfrom theundicipackage. - CPU-heavy parsing blocks the event loop, just as in asyncio. Parsing a 5 MB page with Cheerio stalls every other request in the process; move heavy parsing to worker threads if it matters.
Node or Python?
| Consideration | Node.js | Python |
|---|---|---|
| Parsing | Cheerio (CSS only, fast) | lxml, parsel, selectolax, BeautifulSoup; XPath available |
| Browser automation | Playwright and Puppeteer, first-class | Playwright, same API |
| Concurrency | Async everywhere by default | asyncio, opt-in |
| Frameworks | Crawlee | Scrapy |
| Data work after scraping | Limited | pandas, pydantic, notebooks |
| Running a site's own JavaScript | Native (for signing functions and the like) | Needs an embedded JS engine |
| Fits best when | The team and the backend are already JavaScript | Data processing and analysis follow the scrape |
The last-but-one row matters for reversing JavaScript signatures: Node can load an extracted signing function and call it directly. Do that in an isolated process or container; node:vm is not a security sandbox, and code copied from a website is untrusted.
Practice
Port one of your Python scrapers to Node: crawl every catalogue page of books.toscrape.com by following the "next" link, with p-limit(2), a 500 ms delay, loadBuffer for decoding and allSettled for results. Write the books to a JSON Lines file, then compare the item count and run time with your Python version.