The Article Is a Small Part of the Page

A news or blog page is mostly not the article. Around the body text sit a cookie banner, navigation, ad slots, "related stories" boxes, a most-read list, a newsletter form, comments and a footer. For a search index, a text corpus or an LLM input, all of that is noise, and the related headlines are worse than noise: real text about other stories, now attributed to this one.

Removing it is called boilerplate removal or main-content extraction. Unlike product scraping, one selector rarely does the job, because an article corpus spans many sites and templates. What follows covers the general-purpose tools, how their heuristics work, how to measure them, and when to write per-site rules instead.

What the Extractor Has to Decide

For every block on the page, an extractor answers one question: is this part of the article? The signals it can use are:

Signal Main text tends to have Boilerplate tends to have
Text length per block Long paragraphs Short fragments
Link density Few links relative to text Mostly link text (menus, related lists)
Tag and position <p> inside <article> or <main> <nav>, <aside>, <footer>, <form>
Class and id names content, story, post-body sidebar, related, comment, ad-
Repetition across pages Unique to this URL Identical on every page of the site

No single signal is reliable: <article> also wraps teaser cards, class names are often generated hashes, and some sites use no <p> tags at all. Good extractors combine signals and fall back to another method when the first yields too little.

Start With trafilatura

For general-purpose work, trafilatura is the usual first choice in Python. It combines its own tree-based rules with fallbacks (a port of the readability algorithm and jusText), and it also extracts metadata. Install it with pip install trafilatura.

The examples run against a saved test page, article.html: an <article> with a headline, byline, five paragraphs, an ad slot and a related-links box, surrounded by a cookie banner, nav, comments, a most-read sidebar, a newsletter form and a footer.

import json
import trafilatura

html = open("article.html", encoding="utf-8").read()
result = trafilatura.extract(
    html,
    url="https://news.example.com/travel/night-trains",
    include_comments=False,
    with_metadata=True,
    output_format="json",
)
doc = json.loads(result)
print({k: doc[k] for k in ("title", "author", "date")})
print(doc["text"])
{'title': 'Why Night Trains Are Coming Back', 'author': 'Dana Reyes', 'date': '2026-09-14'}
Why Night Trains Are Coming Back
Ten years ago most European operators were cutting their sleeper services. Carriages were old, ...
That trend has reversed. New operators have leased refurbished coaches, state railways have ...
What changed
Three things moved at once. Travellers began counting the climate cost of short flights, ...
Operators also learned to sell the product differently. A couchette is no longer marketed ...
Not every route works. Night trains need a distance long enough to fill a night but short ...

(Lines trimmed with ....) Everything except the article is gone, including the byline. The author came from <meta name="author"> and the date from article:published_time.

Options you will actually use:

  • include_comments defaults to True. Turn it off unless you want reader comments in your corpus.
  • favor_precision / favor_recall shift the thresholds: precision for training data, where noise hurts; recall for search, where missing text hurts.
  • output_format accepts "txt", "markdown", "json", "xml" and others. Markdown keeps headings and lists, which helps chunking.
  • prune_xpath drops known junk first, for example prune_xpath="//div[@class='ad-slot']".
  • fast=True skips the fallback algorithms: quicker, but loses some difficult pages.

extract accepts bytes as well as strings; see parsing engines and text encoding if the output shows mojibake.

How the Heuristics Work

You will rarely write your own extractor, but knowing the classic readability-style algorithm explains the failures of every tool. Score each paragraph, give the points to its parent and grandparent, adjust by class names, and punish containers made of links.

import re
from lxml import html as lh

POSITIVE = re.compile(r"article|body|content|entry|main|post|story|text", re.I)
NEGATIVE = re.compile(r"ad-|banner|comment|cookie|footer|menu|nav|newsletter|related|share|sidebar|social", re.I)

def text_len(el):
    return len(" ".join(el.text_content().split()))

def link_density(el):
    total = text_len(el)
    links = sum(text_len(a) for a in el.iter("a"))
    return links / total if total else 1.0

def score_candidates(root):
    scores = {}
    for p in root.iter("p"):
        n = text_len(p)
        if n < 25:                       # skip bylines, captions, "Advertisement"
            continue
        points = 1 + n / 100 + p.text_content().count(",")
        parent = p.getparent()
        grand = parent.getparent() if parent is not None else None
        for node, share in ((parent, 1.0), (grand, 0.5)):
            if node is not None:
                scores[node] = scores.get(node, 0) + points * share
    for node in scores:
        hint = f"{node.get('class', '')} {node.get('id', '')}"
        if POSITIVE.search(hint):
            scores[node] += 10
        if NEGATIVE.search(hint):
            scores[node] -= 25
        scores[node] *= 1 - link_density(node)
    return scores

def extract_main(html_text):
    root = lh.fromstring(html_text)
    for bad in root.xpath("//script|//style|//noscript|//nav|//footer|//form"):
        bad.drop_tree()
    scores = score_candidates(root)
    best = max(scores, key=scores.get)
    for child in list(best):             # prune link lists and flagged boxes inside the winner
        hint = f"{child.get('class', '')} {child.get('id', '')}"
        if link_density(child) > 0.5 or NEGATIVE.search(hint):
            child.drop_tree()
    return best, scores

best, scores = extract_main(open("article.html", encoding="utf-8").read())
for node, s in sorted(scores.items(), key=lambda kv: -kv[1])[:3]:
    print(f"{s:6.1f}  <{node.tag}>")
  24.9  <article>
  12.5  <main>

The non-obvious parts:

  • Points go to parent and grandparent, so the true container wins even when paragraphs sit inside extra wrapper <div>s.
  • Commas count. Prose has commas; menus and headlines rarely do.
  • Multiplying by 1 - link_density kills navigation and related lists: a box that is 90 percent link text keeps 10 percent of its score.
  • Pruning inside the winner removes the related box embedded mid-article, which scoring alone would keep.

Mozilla's Readability (Firefox's reader view) and its Python port readability-lxml work along these lines with many more rules. jusText instead classifies each paragraph by length, link density and the share of stopwords ("the", "and", "of"): prose is full of them, menus are not. That makes it language-dependent; it needs the stoplist for the page's language.

Measure Extraction Quality

"It looks right on the page I checked" is not an evaluation. Save 30 to 50 pages across your sites, paste the correct article text for each into a .txt file (the gold text), and score every extractor against them. Token overlap is good enough to compare methods:

import re
from collections import Counter

def tokens(text):
    return Counter(re.findall(r"\w+", text.lower()))

def overlap(extracted, gold):
    e, g = tokens(extracted or ""), tokens(gold)
    hit = sum((e & g).values())
    return round(hit / max(sum(e.values()), 1), 3), round(hit / max(sum(g.values()), 1), 3)

html = open("article.html", encoding="utf-8").read()
gold = open("article.gold.txt", encoding="utf-8").read()
whole_page = lh.fromstring(html).text_content()
best, _ = extract_main(html)
for name, text in [("whole page", whole_page),
                   ("density heuristic", best.text_content()),
                   ("trafilatura", trafilatura.extract(html, include_comments=False))]:
    print(f"{name:18} precision/recall = {overlap(text, gold)}")
whole page         precision/recall = (0.636, 1.0)
density heuristic  precision/recall = (0.969, 1.0)
trafilatura        precision/recall = (1.0, 1.0)

Precision is the share of extracted words that belong to the article (low means boilerplate leaked in). Recall is the share of article words captured (low means paragraphs went missing). The heuristic lost a little precision by keeping the byline. One synthetic page proves little; across a varied sample of real pages, these numbers are how you choose tools, tune favor_precision, and notice when a site change breaks extraction.

How Extraction Fails

Two production failure modes, reproduced with trafilatura 2.x:

shell = '<html><body><nav><a href=/>Home</a><a href=/x>News</a></nav><div id=app></div></body></html>'
print(repr(trafilatura.extract(shell)))
listing = "<html><body><main>" + "".join(
    f"<div class=card><h2><a href=/p{i}>Story {i} headline</a></h2>"
    f"<p>A short teaser sentence for story {i}, written to entice a click.</p></div>"
    for i in range(12)) + "</main></body></html>"
print(trafilatura.extract(listing)[:80])
'Home\nNews'
Story 0 headline
A short teaser sentence for story 0, written to entice a click.
  • JavaScript-rendered pages return junk, not None. The article was never in the HTML, so the extractor returned the menu. Fetch the rendered page or its data API instead (see JavaScript pages).
  • Listing pages look like articles. Twelve teasers are twelve paragraphs of prose. Extract only from pages that declare themselves articles: og:type of article, or a JSON-LD @type such as NewsArticle (see structured data).
  • Short posts fall below trafilatura's MIN_EXTRACTED_SIZE (250 characters by default) and trigger fallbacks that can pull in navigation.
  • Paywalls, consent walls and multi-page articles leave you with a fragment that looks like success.
  • Odd markup: <br>-separated text instead of <p>, duplicated pull quotes, captions.

Guard every result before storing it:

def accept(text, html_len):
    if not text or len(text) < 500:
        return False, "too short"
    long_paras = [line for line in text.splitlines() if len(line) >= 150]
    if len(long_paras) < 2:
        return False, "no long paragraphs (menu, listing or teaser?)"
    if len(text) / max(html_len, 1) < 0.01:
        return False, "tiny text-to-HTML ratio"
    return True, "ok"

for page in (html, listing, shell):
    print(accept(trafilatura.extract(page, include_comments=False), len(page)))
(True, 'ok')
(False, 'no long paragraphs (menu, listing or teaser?)')
(False, 'too short')

Tune the thresholds on your gold sample so real short articles still pass, and log rejections per site. A site whose acceptance rate drops from 95 to 20 percent overnight changed its template or started rendering client-side, the same signal the layout drift detection lesson builds on.

When to Write Site-Specific Rules

For a handful of important sources, explicit rules beat any general extractor on accuracy and debuggability. Keep both: rules where you have them, the general extractor everywhere else.

from urllib.parse import urlsplit
from bs4 import BeautifulSoup

SITE_RULES = {
    "news.example.com": {"body": "main article", "drop": [".ad-slot", "aside.related", "p.byline"]},
}

def extract(url, html):
    rules = SITE_RULES.get(urlsplit(url).hostname)
    if rules:
        soup = BeautifulSoup(html, "lxml")
        body = soup.select_one(rules["body"])
        if body is not None:
            for sel in rules["drop"]:
                for el in body.select(sel):
                    el.decompose()
            text = "\n".join(t.get_text(" ", strip=True)
                             for t in body.select("h1, h2, h3, p, li, blockquote"))
            if len(text) > 200:
                return text, "rules"
    text = trafilatura.extract(html, url=url, include_comments=False)
    return text, "trafilatura" if text else "failed"

Counting how often rules fall through to the generic path tells you when a rule has gone stale.

Situation Approach
Thousands of sites, text for search or a corpus General extractor, precision/recall checked on a sample
A few high-value sources you read daily Site rules, generic fallback
Need exact fields (author, section, dates, tags) Structured data first, then rules
Layouts vary wildly and fields are fuzzy Extract text, then an LLM with a schema (LLM-assisted extraction)
Same story syndicated across many sites Extract, then deduplicate (near-duplicate detection)

Store the raw HTML next to the extracted text: re-extracting after a rule fix is cheap, re-fetching old articles is not.

Practice

Save 30 article pages from five sites you are allowed to crawl, plus three listing pages and one JavaScript-rendered page. Write gold text for the articles, score the density heuristic and trafilatura (default and favor_precision=True) per site, and confirm accept rejects all four non-article pages.