The Same Fundamentals in JavaScript

None of this course's fundamentals are Python-specific. Node.js is a strong alternative when your team already writes JavaScript, when the scraper lives inside a Node backend, or when you need to run a site's own JavaScript. Here is the core toolkit in Node: fetch plus Cheerio for static pages, p-limit for bounded concurrency, Playwright for pages that need a browser, and the Node-specific traps that catch Python developers.

The examples use books.toscrape.com and quotes.toscrape.com, sandbox sites built for scraping practice. They need Node 18 or newer (for built-in fetch) and ES modules:

mkdir scraper && cd scraper
npm init -y
npm pkg set type=module          # allow import syntax and top-level await
npm install cheerio p-limit playwright
npx playwright install chromium

p-limit is published only as an ES module, which is one reason to set "type": "module" from the start rather than mixing require and import.

Static Pages: fetch and Cheerio

Cheerio parses HTML into a tree and gives you a jQuery-style API over it. It does not run JavaScript or compute layout, which is why it is fast.

import * as cheerio from "cheerio";

const BASE = "https://books.toscrape.com/";
const UA = "simpleprog-tutorial/1.0 (contact@example.org)";

async function getHtml(url) {
  const res = await fetch(url, {
    headers: { "User-Agent": UA, Accept: "text/html" },
    signal: AbortSignal.timeout(15_000),
  });
  if (!res.ok) throw new Error(`${res.status} ${res.statusText} for ${url}`);
  return res.text();
}

const $ = cheerio.load(await getHtml(BASE));

const books = $("article.product_pod").map((_, el) => {
  const card = $(el);
  const link = card.find("h3 a");
  return {
    title: link.attr("title"),
    url: new URL(link.attr("href"), BASE).href,
    price: card.find(".price_color").text().trim(),
    rating: card.find("p.star-rating").attr("class").replace("star-rating", "").trim(),
  };
}).get();

console.log(books.length, books.slice(0, 2));
const next = $("li.next a").attr("href");
console.log("next page:", next && new URL(next, BASE).href);
20 [
  {
    title: 'A Light in the Attic',
    url: 'https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html',
    price: '£51.77',
    rating: 'Three'
  },
  {
    title: 'Tipping the Velvet',
    url: 'https://books.toscrape.com/catalogue/tipping-the-velvet_999/index.html',
    price: '£53.74',
    rating: 'One'
  }
]
next page: https://books.toscrape.com/catalogue/page-2.html

The lines that matter:

  • AbortSignal.timeout(15_000). fetch has no default timeout. Without a signal, one stalled server hangs that request forever.
  • if (!res.ok). Like requests, fetch does not throw on 404 or 500; it only rejects on network errors. Check the status yourself.
  • new URL(href, base) is the urljoin of JavaScript. It resolves ../ and relative paths correctly.
  • .map(...).get(). Cheerio's .map returns a Cheerio object; .get() turns it into a plain array.

Two Cheerio behaviours surprise people coming from BeautifulSoup. .text() on a selection of several elements returns all their text concatenated, not a list, so $(".price_color").text() on a listing gives £51.77£53.74.... And .attr() returns the attribute of the first element only. Iterate with .map or .each whenever you expect more than one match.

Cheerio supports CSS selectors only; for XPath you need a different library, one area where Python's lxml and parsel are ahead.

The Encoding Trap in fetch

res.text() always decodes the body as UTF-8, whatever the Content-Type header or <meta charset> says. On a site still served in windows-1252 you get replacement characters, silently:

import * as cheerio from "cheerio";
// "Café – 5 €" encoded as windows-1252, with a matching <meta charset>
const bytes = Buffer.from([
  ...Buffer.from('<html><head><meta charset="windows-1252"></head><body><p>Caf'),
  0xe9, 0x20, 0x96, 0x20, 0x35, 0x20, 0x80,
  ...Buffer.from("</p></body></html>"),
]);
const res = new Response(bytes, { headers: { "Content-Type": "text/html" } });

console.log("res.text():  ", cheerio.load(await res.clone().text())("p").text());
console.log("loadBuffer():", cheerio.loadBuffer(Buffer.from(await res.arrayBuffer()))("p").text());
res.text():   Caf� � 5 �
loadBuffer(): Café – 5 €

cheerio.loadBuffer takes raw bytes and sniffs the encoding from a byte order mark or <meta charset>. If the response header names a charset, pass it too: cheerio.loadBuffer(buf, { encoding: { transportLayerEncodingLabel: "windows-1252" } }). The rules behind this are in parsing engines and text encoding. For crawlers that touch many sites, use arrayBuffer() plus loadBuffer everywhere and never res.text().

Bounded Concurrency With p-limit

Node is asynchronous by default, which makes it easy to overload a site by accident: await Promise.all(urls.map(u => fetch(u))) starts every request at once. p-limit caps how many run at the same time. Combined with a delay and retries it gives you a small, polite crawler:

import * as cheerio from "cheerio";
import pLimit from "p-limit";
import { setTimeout as sleep } from "node:timers/promises";

const UA = "simpleprog-tutorial/1.0 (contact@example.org)";
const limit = pLimit(2);               // at most 2 requests in flight
const DELAY_MS = 500;                  // pause each slot after every request

async function fetchPage(url, attempts = 4) {
  for (let i = 0; i < attempts; i++) {
    try {
      const res = await fetch(url, { headers: { "User-Agent": UA }, signal: AbortSignal.timeout(15_000) });
      if (res.status === 429 || res.status >= 500) throw new Error(`retryable ${res.status}`);
      if (!res.ok) throw Object.assign(new Error(`HTTP ${res.status}`), { fatal: true });
      return cheerio.loadBuffer(Buffer.from(await res.arrayBuffer()));
    } catch (err) {
      if (err.fatal || i === attempts - 1) throw err;
      await sleep(1000 * 2 ** i);        // 1s, 2s, 4s
    }
  }
}

async function scrapeListing(url) {
  const $ = await fetchPage(url);
  await sleep(DELAY_MS);
  return $("article.product_pod").map((_, el) => ({
    title: $(el).find("h3 a").attr("title"),
    price: $(el).find(".price_color").text(),
  })).get();
}

const urls = [1, 2, 3, 4, 5, 999].map(n => `https://books.toscrape.com/catalogue/page-${n}.html`);
const results = await Promise.allSettled(urls.map(u => limit(() => scrapeListing(u))));

results.forEach((r, i) => {
  if (r.status === "fulfilled") console.log(urls[i], r.value.length, "books");
  else console.log(urls[i], "FAILED:", r.reason.message);
});
https://books.toscrape.com/catalogue/page-1.html 20 books
https://books.toscrape.com/catalogue/page-2.html 20 books
https://books.toscrape.com/catalogue/page-3.html 20 books
https://books.toscrape.com/catalogue/page-4.html 20 books
https://books.toscrape.com/catalogue/page-5.html 20 books
https://books.toscrape.com/catalogue/page-999.html FAILED: HTTP 404

Why it is built this way:

  • limit(() => scrapeListing(u)) wraps a function, not a promise. limit(scrapeListing(u)) would start the request immediately.
  • The sleep sits inside the limited function, so each of the two slots pauses after its request. That caps the rate at roughly two requests per (latency + 0.5 s). Concurrency alone is not a rate limit; see rate limiting and politeness.
  • Promise.allSettled, not Promise.all. Promise.all rejects on the first failure and you lose the other results; with allSettled the 404 is one failed row, not a failed crawl.
  • 404 is fatal, 429 and 5xx are retried with exponential backoff, the same policy as the Python lessons.

For crawls with a growing frontier (URLs discovered while crawling), mapping over a fixed list does not work. Use a queue with a fixed number of workers, exactly the pattern from parallel scraping with asyncio, or a framework such as Crawlee, which is to Node roughly what Scrapy is to Python.

Pages That Need a Browser: Playwright for Node

Playwright is written in TypeScript and its Node API is the reference version; the Python API from browser automation mirrors it almost name for name, with await in place of sync calls and camelCase names. Two techniques cover most cases: read the rendered DOM, or capture the JSON the page loads.

import { chromium } from "playwright";

const browser = await chromium.launch({ headless: true });
try {
  const context = await browser.newContext({ userAgent: "simpleprog-tutorial/1.0 (contact@example.org)" });
  await context.route("**/*.{png,jpg,jpeg,gif,woff,woff2}", route => route.abort());
  const page = await context.newPage();

  // 1. Rendered DOM: the quotes only exist after the page's script runs
  await page.goto("https://quotes.toscrape.com/js/", { waitUntil: "domcontentloaded" });
  await page.locator("div.quote").first().waitFor();
  const quotes = await page.locator("div.quote").evaluateAll(nodes =>
    nodes.map(n => ({
      text: n.querySelector(".text")?.textContent.trim(),
      author: n.querySelector(".author")?.textContent.trim(),
    })));
  console.log(quotes.length, quotes[0].author);

  // 2. Network capture: the infinite-scroll page loads JSON from an API
  const batches = [];
  page.on("response", async res => {
    if (res.url().includes("/api/quotes") && res.ok()) batches.push(await res.json());
  });
  await page.goto("https://quotes.toscrape.com/scroll", { waitUntil: "networkidle" });
  for (let i = 0; i < 2; i++) {
    const next = page.waitForResponse(r => r.url().includes("/api/quotes"));
    await page.mouse.wheel(0, 20_000);
    await next;
  }
  await page.waitForTimeout(500);       // let the async response handlers finish
  console.log(batches.map(b => `page ${b.page}: ${b.quotes.length} quotes`));
} finally {
  await browser.close();
}
10 Albert Einstein
[ 'page 1: 10 quotes', 'page 2: 10 quotes', 'page 3: 10 quotes', 'page 4: 10 quotes' ]

The number of batches varies between runs (one large scroll can trigger more than one load), so stop on the API's own has_next field in real code rather than on a scroll count.

Notes:

  • evaluateAll runs in the page. The callback is serialised into the browser, so it cannot see your Node variables (pass them as a second argument). One call returning plain objects beats a round trip per element.
  • waitForResponse is created before the scroll. Creating it after could miss a response that arrives quickly.
  • try/finally around the browser. An exception that skips browser.close() leaves a Chromium process running.
  • Once you have seen /api/quotes?page=N, you may not need the browser at all. Call that endpoint with fetch and skip rendering entirely (see XHR and API reverse engineering). The browser was the discovery tool.

Node Pitfalls Worth Knowing

  • A missing await gives you a pending Promise instead of data, and an error inside it becomes an unhandled rejection that can crash the process. Lint with a rule that flags floating promises.
  • No cookie jar. Node's fetch does not keep cookies between requests. For sessions, add a jar (for example the fetch-cookie package with tough-cookie) or copy Set-Cookie values yourself.
  • Proxies. Node's fetch is built on undici; route it through a proxy with fetch(url, { dispatcher: new ProxyAgent("http://host:port") }), importing ProxyAgent from the undici package.
  • CPU-heavy parsing blocks the event loop, just as in asyncio. Parsing a 5 MB page with Cheerio stalls every other request in the process; move heavy parsing to worker threads if it matters.

Node or Python?

Consideration Node.js Python
Parsing Cheerio (CSS only, fast) lxml, parsel, selectolax, BeautifulSoup; XPath available
Browser automation Playwright and Puppeteer, first-class Playwright, same API
Concurrency Async everywhere by default asyncio, opt-in
Frameworks Crawlee Scrapy
Data work after scraping Limited pandas, pydantic, notebooks
Running a site's own JavaScript Native (for signing functions and the like) Needs an embedded JS engine
Fits best when The team and the backend are already JavaScript Data processing and analysis follow the scrape

The last-but-one row matters for reversing JavaScript signatures: Node can load an extracted signing function and call it directly. Do that in an isolated process or container; node:vm is not a security sandbox, and code copied from a website is untrusted.

Practice

Port one of your Python scrapers to Node: crawl every catalogue page of books.toscrape.com by following the "next" link, with p-limit(2), a 500 ms delay, loadBuffer for decoding and allSettled for results. Write the books to a JSON Lines file, then compare the item count and run time with your Python version.