Structured Data: JSON-LD, Microdata and Open Graph
The Copy of the Data Written for Machines
Sites that want rich search results (ratings, prices, event dates) describe their pages in schema.org markup. It is often the cleanest data on the page: typed fields with stable names, identical across templates, and rarely touched by redesigns because the SEO team depends on it. The HTML structure lesson showed the basic json.loads on a JSON-LD block. This lesson covers what real pages throw at you: several syntaxes, @graph containers, @id references, broken JSON, and values that disagree with the visible page.
The Formats
| Format | Where it lives | Typical content | Parse with |
|---|---|---|---|
| JSON-LD | <script type="application/ld+json"> |
Full schema.org objects | json plus your own normaliser, or extruct |
| Microdata | itemscope, itemtype, itemprop attributes on visible elements |
Same vocabulary, woven into the HTML | extruct |
| RDFa | vocab, typeof, property attributes |
Same idea, older and rarer | extruct |
| Open Graph | <meta property="og:..."> |
Title, image, URL, type for link previews | A few lines of XPath, or extruct |
| Twitter/X cards | <meta name="twitter:..."> |
Similar preview fields | A few lines of XPath |
JSON-LD is the most common and easiest. Open Graph is nearly universal but shallow. Schema.org types worth knowing by name: Product, ProductGroup (variants via hasVariant), Offer, AggregateOffer, AggregateRating, NewsArticle/BlogPosting, Recipe, Event, JobPosting, LocalBusiness, Organization and BreadcrumbList (the category path, which is often hard to get any other way).
Survey a Page With extruct
extruct (pip install extruct) pulls every syntax out of a page in one call. The test page product.html has an Open Graph block, one JSON-LD block using @graph, a second JSON-LD block wrapped in an HTML comment (an old habit some templates keep), and a microdata review.
import extruct
html = open("product.html", encoding="utf-8").read()
data = extruct.extract(
html,
base_url="https://shop.example.com/p/trail-runner-2",
syntaxes=["json-ld", "microdata", "opengraph"],
uniform=True, # microdata and Open Graph come back in JSON-LD-like shape
errors="log", # log a failing syntax instead of raising
)
print(list(data))
print(data["microdata"])
Failed to extract json-ld, raises Expecting value: line 2 column 1 (char 1)
['microdata', 'opengraph']
[{'@type': 'Review', '@context': 'https://schema.org', 'author': {'@type': 'Person', 'name': 'Sam K.'},
'reviewRating': {'@type': 'Rating', 'ratingValue': '5'}, 'reviewBody': 'Great grip on wet rock.'}]
JSON-LD is missing entirely. The commented-out block failed to parse (extruct copes with trailing commas, but not a multi-line <!-- ... --> wrapper), and because extruct processes a syntax as a unit, the healthy @graph block with the product was lost too. With the default errors="strict" the whole call raises.
Keep extruct for microdata and RDFa. For JSON-LD, parse each block yourself so one failure costs one block.
A Tolerant JSON-LD Reader
import json
import re
from lxml import html as lh
WRAPPERS = re.compile(r"^\s*(?:<!--|<!\[CDATA\[)|(?:-->|\]\]>)\s*$")
TRAILING_COMMA = re.compile(r",\s*([}\]])")
def load_jsonld_block(raw):
text = WRAPPERS.sub("", raw).strip()
for attempt in (text, TRAILING_COMMA.sub(r"\1", text)):
try:
return json.loads(attempt, strict=False) # strict=False tolerates raw newlines in strings
except json.JSONDecodeError:
continue
return None
def jsonld_nodes(html_text):
root = lh.fromstring(html_text)
nodes, failed = [], 0
for script in root.xpath('//script[@type="application/ld+json"]'):
data = load_jsonld_block(script.text or "")
if data is None:
failed += 1
continue
stack = data if isinstance(data, list) else [data]
while stack:
item = stack.pop(0)
if isinstance(item, dict) and "@graph" in item:
stack.extend(item["@graph"])
elif isinstance(item, dict):
nodes.append(item)
return nodes, failed
def types(node):
t = node.get("@type", [])
return {x.rsplit("/", 1)[-1] for x in ([t] if isinstance(t, str) else t)}
nodes, failed = jsonld_nodes(html)
print("blocks that failed:", failed)
for n in nodes:
print(sorted(types(n)), n.get("@id", ""))
blocks that failed: 0
['Organization'] https://shop.example.com/#org
['BreadcrumbList']
['IndividualProduct', 'Product'] https://shop.example.com/p/trail-runner-2#product
['WebPage']
What each piece handles:
- Wrappers (
<!-- -->,<![CDATA[ ]]>) are stripped before parsing. - Trailing commas are removed only on the second attempt, since the regex could alter a string containing
, }. Valid JSON is never touched. - Lists and
@graph. A block can be an object, a list, or an object whose@graphholds the real nodes; the loop flattens all three. @typeis a string or a list, sometimes a full URL.types()returns short names, so"Product" in types(n)always works.
Count failed per site: a block that breaks after a site deploy is an early warning, like the rejection rate in JSON payloads.
Resolve @id References
In a @graph, nodes refer to each other by @id. The product's brand above is just {"@id": "https://shop.example.com/#org"}; the name lives on the Organization node. Build an index and substitute references, guarding against cycles (a WebPage that points to a WebSite that points back is common):
def resolve(value, index, seen=frozenset()):
if isinstance(value, dict):
ref = value.get("@id")
if len(value) == 1 and ref in index and ref not in seen:
return resolve(index[ref], index, seen | {ref})
return {k: resolve(v, index, seen) for k, v in value.items()}
if isinstance(value, list):
return [resolve(v, index, seen) for v in value]
return value
Only dicts whose sole key is @id are treated as references; a full node that happens to carry an @id is kept as it is.
Normalise Into a Record
Schema.org is permissive: offers as a dict or a list, AggregateOffer with lowPrice, prices as numbers or strings, availability as "InStock" or a http/https schema.org URL, brand as a string or an object. Collapse them in one function:
from decimal import Decimal, InvalidOperation
def as_list(value):
return [] if value is None else value if isinstance(value, list) else [value]
def to_decimal(value):
try:
return Decimal(str(value)) if value not in (None, "") else None
except InvalidOperation:
return None
def product_from_jsonld(nodes):
index = {n["@id"]: n for n in nodes if "@id" in n}
product = next((n for n in nodes if "Product" in types(n)), None)
if product is None:
return None
p = resolve(product, index)
brand = p.get("brand")
offers = []
for o in as_list(p.get("offers")):
if "AggregateOffer" in types(o):
offers.append({"price": to_decimal(o.get("lowPrice")),
"currency": o.get("priceCurrency"), "availability": None})
continue
offers.append({"price": to_decimal(o.get("price")),
"currency": o.get("priceCurrency"),
"availability": str(o.get("availability", "")).rsplit("/", 1)[-1] or None})
rating = p.get("aggregateRating") or {}
return {
"name": p.get("name"), "sku": p.get("sku"),
"gtin": p.get("gtin13") or p.get("gtin") or p.get("gtin12") or p.get("gtin14"),
"brand": brand.get("name") if isinstance(brand, dict) else brand,
"offers": offers,
"rating": to_decimal(rating.get("ratingValue")),
"reviews": int(rating["reviewCount"]) if rating.get("reviewCount") else None,
}
print(product_from_jsonld(nodes))
{'name': 'Trail Runner 2', 'sku': 'TR2-42', 'gtin': '4006381333931', 'brand': 'Example Outdoor',
'offers': [{'price': Decimal('129.00'), 'currency': 'EUR', 'availability': 'InStock'},
{'price': Decimal('119'), 'currency': 'EUR', 'availability': 'OutOfStock'}],
'rating': Decimal('4.6'), 'reviews': 212}
Prices go through Decimal(str(...)) so 129.00 and "129.00" give the same exact value. Keep the GTIN: it is the identifier that lets you match the same product across retailers without fuzzy matching (see entity resolution). Pages with several products (a ProductGroup, or a listing with an ItemList) need you to return all matches instead of the first.
Microdata and RDFa
Microdata puts the same vocabulary on visible elements. itemscope starts an object, itemtype names its type, and itemprop names a property whose value comes from the element: the content attribute on <meta>, href on <a> and <link>, src on <img>, datetime on <time>, and the text for everything else. A nested itemscope with an itemprop becomes a nested object, which is why the review above has an author object. These value rules are fiddly to reimplement, so let extruct do it. RDFa follows the same pattern with vocab, typeof and property; extruct returns it as expanded JSON-LD with full property URLs, which needs extra mapping, so reach for it only when a site uses nothing else.
Open Graph and Twitter Cards
Open Graph uses property=, Twitter cards use name=, and extruct's opengraph syntax only reads the first. A few lines collect both:
def meta_tags(root, prefixes=("og", "twitter", "article", "product")):
out = {}
for m in root.xpath("//meta[@content]"):
key = m.get("property") or m.get("name")
if key and key.split(":")[0] in prefixes:
out.setdefault(key, m.get("content")) # first occurrence wins
return out
print(meta_tags(lh.fromstring(html)))
{'og:type': 'product', 'og:title': 'Trail Runner 2 Running Shoe', 'og:image': 'https://shop.example.com/img/tr2-1200.jpg',
'og:url': 'https://shop.example.com/p/trail-runner-2', 'product:price:amount': '129.00',
'product:price:currency': 'EUR', 'twitter:card': 'summary_large_image', 'twitter:title': 'Trail Runner 2'}
Preview titles are written for sharing and differ from the JSON-LD name. Use og:type to classify pages and og:image for the main image; take names from JSON-LD.
Verify Against the Visible Page
Structured data is a claim made for search engines, and it can be wrong:
- Stale values from a cache or nightly feed can lag the visible price by hours.
- Several offers. Our page lists 129.00 and 119 (an out-of-stock variant); "the price" is a decision, not a field.
- Injected by JavaScript. Tag managers often add JSON-LD after load, so check the rendered DOM before concluding a site has none (see JavaScript pages).
- Placeholders: empty strings,
"price": "0", the site name as the product name.
Cross-check one visible field per record:
visible = root.xpath('string(//p[@class="price"])').strip() # '€129,00'
shown = Decimal(re.sub(r"[^\d,]", "", visible).replace(",", "."))
prices = {o["price"] for o in product_from_jsonld(nodes)["offers"]}
print(shown, shown in prices)
129.00 True
The comma-to-dot swap is only correct for this locale; the dates, prices and units lesson handles separators properly. Log mismatches per site rather than failing records: a mismatch rate near zero means you can trust the structured data there, and a sudden rise means it went stale.
Choosing a Source per Field
| Field | Best source | Fallback |
|---|---|---|
| Name, SKU, GTIN, brand | JSON-LD | Microdata, then the visible page |
| Current price and stock | Visible page or the site's JSON API, cross-checked with JSON-LD | JSON-LD alone |
| Canonical URL | <link rel="canonical"> |
og:url |
| Main image | og:image or JSON-LD image |
Largest <img> in the content |
| Page type | JSON-LD @type, og:type |
URL pattern |
| Category path | JSON-LD BreadcrumbList |
Breadcrumb links |
| Article body | Neither: use article text extraction |
Practice
Pick three sites you are allowed to scrape: a shop, a news site and a recipe or events site. For 20 pages each, run both extruct and jsonld_nodes, count pages where extruct lost JSON-LD that your reader recovered, list every @type seen, and measure how often the JSON-LD price or headline matches the visible one.