The Same Number Means Different Things

1,299 is one thousand two hundred and ninety-nine on a US shop and one point two nine nine on a German one. 03/04/2026 is 4 March in New York and 3 April in London. Today, 09:15 is a different instant depending on whose clock the site uses. None of these raise an exception when you get them wrong; they produce plausible values that are off by a factor of a thousand, a month, or a few hours.

The data cleaning and validation lesson introduced heuristic price and date parsers, and the regular expressions lesson showed how to pull raw strings out of text. Every heuristic eventually meets a string it misreads, so the step beyond them is configuration: record per source (or per source and region, see geo-targeted scraping) its locale such as de_DE, its currency, its IANA time zone and its date order, then use parsers that refuse anything that does not fit.

Locale-Aware Numbers With Babel

The Babel library ships the Unicode CLDR data that browsers and operating systems use, so it knows every locale's decimal and grouping symbols, including the ones people forget: French groups thousands with a narrow no-break space (U+202F), Swiss German with a right single quote (1’299.50), and Indian English groups by two digits after the first three.

from babel.numbers import parse_decimal, NumberFormatError

cases = [("1.299,50", "de_DE"), ("1\u202f299,50", "fr_FR"), ("1,23,456.75", "en_IN"),
         ("1,299", "en_US"), ("1,299", "de_DE"), ("12,34", "en_US")]
for text, loc in cases:
    print(f"{text!r:18} {loc}  ->  {parse_decimal(text, locale=loc)}")

try:
    parse_decimal("12,34", locale="en_US", strict=True)
except NumberFormatError as e:
    print(e)
'1.299,50'         de_DE  ->  1299.50
'1\u202f299,50'    fr_FR  ->  1299.50
'1,23,456.75'      en_IN  ->  123456.75
'1,299'            en_US  ->  1299
'1,299'            de_DE  ->  1.299
'12,34'            en_US  ->  1234
'12,34' is not a properly formatted decimal number. Did you mean '1,234'? Or maybe '12.34'?

The last two lines are the reason to pass strict=True. Lax mode simply drops grouping symbols, so 12,34 becomes 1234; strict mode checks that they sit where the locale puts them. A NumberFormatError on a source that normally parses cleanly usually means the page is not in the locale you configured, such as a region switch, which is exactly what you want to hear about.

Decimal, Rounding and Minor Units

parse_decimal returns a Decimal, and it should stay one. Build Decimals from strings, never from floats: Decimal(19.99) is 19.98999999999999843.... Choose the rounding mode explicitly, because Decimal.quantize defaults to banker's rounding (0.125 becomes 0.12) while most shops round half up. Watch for floats sneaking in through APIs: json.loads(text, parse_float=Decimal) keeps JSON payloads exact.

Not every currency has two decimal places. The Japanese yen has none and the Kuwaiti dinar has three. Babel knows each currency's minor units, so you can store integers in the smallest unit without hard-coding * 100:

from decimal import Decimal, ROUND_HALF_UP
from babel.numbers import get_currency_precision

def to_minor(amount: Decimal, currency: str) -> int:
    places = get_currency_precision(currency)        # USD 2, JPY 0, KWD 3
    step = Decimal(1).scaleb(-places)                # 0.01, 1, 0.001
    return int(amount.quantize(step, rounding=ROUND_HALF_UP) * 10 ** places)

print(to_minor(Decimal("1299.5"), "EUR"), to_minor(Decimal("1500"), "JPY"),
      to_minor(Decimal("12.345"), "KWD"))
# 129950 1500 12345

An integer in minor units plus an ISO 4217 code (EUR) survives any database, JSON library or language without precision loss.

A Price String Is Not Always One Price

Price elements hold sale prices, ranges and bundle offers. Generic extractors such as the price-parser package handle simple strings well but return the first number they find, so Was £30, now £24.99 yields 30, £10 - £15 yields 10 and 3 for £5 yields 3. Classify the string before picking a number:

import re

AMOUNT = re.compile(r"\d[\d.,'\u00a0\u202f\u2019 ]*\d|\d")

def amounts(text, locale):
    try:
        return [parse_decimal(m.group().strip(), locale=locale, strict=True)
                for m in AMOUNT.finditer(text)]
    except NumberFormatError:
        return None                      # one bad number makes the whole string suspect

def parse_offer(text, locale, currency):
    low = text.lower()
    if re.search(r"\b\d+\s*(for|x)\s", low):
        return {"kind": "multibuy", "raw": text}
    nums = amounts(text, locale)
    if not nums:
        return None
    if re.search(r"\d\s*(-|–|to|bis)\s*\D{0,3}\d", low):
        return {"kind": "range", "min": to_minor(nums[0], currency),
                "max": to_minor(nums[-1], currency)}
    if len(nums) == 2 and re.search(r"\b(was|now|rrp|statt|jetzt)\b", low):
        return {"kind": "sale", "price": to_minor(min(nums), currency),
                "was": to_minor(max(nums), currency)}
    if len(nums) == 1:
        return {"kind": "single", "price": to_minor(nums[0], currency)}
    return {"kind": "ambiguous", "raw": text}

for text, loc, cur in [("Was £30, now £24.99", "en_GB", "GBP"), ("£10 - £15", "en_GB", "GBP"),
                       ("3 for £5", "en_GB", "GBP"), ("statt 49,99 € jetzt 39,99 €", "de_DE", "EUR")]:
    print(f"{text!r:32} -> {parse_offer(text, loc, cur)}")
'Was £30, now £24.99'            -> {'kind': 'sale', 'price': 2499, 'was': 3000}
'£10 - £15'                      -> {'kind': 'range', 'min': 1000, 'max': 1500}
'3 for £5'                       -> {'kind': 'multibuy', 'raw': '3 for £5'}
'statt 49,99 € jetzt 39,99 €'    -> {'kind': 'sale', 'price': 3999, 'was': 4999}

The keyword lists are per language; the important design choice is the ambiguous and None outcomes, which you store with the raw text and count per run instead of forcing a number. Better still, avoid the problem: sale and regular prices usually sit in separate elements (<del>), and product pages often carry price and priceCurrency in structured data.

Dates: Relative, Localised and Partial

The dateparser package reads dates in over two hundred language locales, including relative forms like vor 3 Stunden and hace 2 días. Its defaults are tuned for convenience, and several quietly invent information:

Input Default behaviour Fix
03/04/2026 Month first for English Set DATE_ORDER per source
March 2026 Fills in today's day of the month PREFER_DAY_OF_MONTH: "first", and record month precision
Posted: 5 hours ago Returns None: the label confuses it Strip known labels first
5 hours ago Resolved against your machine's clock Pass the fetch time as RELATIVE_BASE

The function below applies these fixes. It also interprets clock times in the site's time zone, and returns a plain date when the text has no time, because dateparser otherwise attaches the current time of day to yesterday and stores a day-precision value as a fake timestamp:

from datetime import datetime, timezone
from zoneinfo import ZoneInfo
import dateparser

HAS_TIME = re.compile(r"\d{1,2}:\d{2}|\b(sec|min|hour|hr|stunde|minute|heure|hora)", re.I)
LABEL = re.compile(r"^\s*(posted|updated|published|aktualisiert|publié)\s*(on|am|le)?\s*:?\s*", re.I)

def parse_site_date(raw, fetched_at, site_tz, languages, date_order=None):
    tz = ZoneInfo(site_tz)
    text = LABEL.sub("", raw)
    settings = {
        "RELATIVE_BASE": fetched_at.astimezone(tz).replace(tzinfo=None),  # "now" on the site's clock
        "PREFER_DAY_OF_MONTH": "first",
        "PREFER_DATES_FROM": "past",
    }
    if date_order:
        settings["DATE_ORDER"] = date_order
    naive = dateparser.parse(text, languages=languages, settings=settings)
    if naive is None:
        return None
    if not HAS_TIME.search(text):
        return naive.date()
    return naive.replace(tzinfo=tz).astimezone(timezone.utc)

fetched = datetime(2026, 9, 30, 14, 5, tzinfo=timezone.utc)
for raw, langs, order in [("Today, 09:15", ["en"], None), ("Posted: 5 hours ago", ["en"], None),
                          ("Aktualisiert am 12. März 2026, 18:30", ["de"], None),
                          ("03/04/2026", ["en"], "DMY"), ("gestern", ["de"], None)]:
    print(f"{raw!r:40} -> {parse_site_date(raw, fetched, 'Europe/Berlin', langs, order)}")
'Today, 09:15'                           -> 2026-09-30 07:15:00+00:00
'Posted: 5 hours ago'                    -> 2026-09-30 09:05:00+00:00
'Aktualisiert am 12. März 2026, 18:30'   -> 2026-03-12 17:30:00+00:00
'03/04/2026'                             -> 2026-04-03
'gestern'                                -> 2026-09-29

Notice the offsets: 09:15 in Berlin in September is 07:15 UTC (summer time), while 18:30 in March is 17:30 UTC (winter time). A fixed +01:00 would be wrong half the year. Passing languages also matters: without it dateparser guesses, which is slower and sometimes wrong for short strings. And 3 weeks ago looks exact but is accurate to a week at best, so prefer a <time datetime="..."> attribute or datePublished in JSON-LD when the page has one.

Time Zone Traps

  • Store UTC, keep the zone name. Compare and store in UTC; keep the source's IANA zone to display local time later.
  • Never attach a zone with pytz's constructor. datetime(..., tzinfo=pytz.timezone("Europe/Berlin")) gives the historical Local Mean Time offset of +00:53. The standard library's zoneinfo does not have this trap; on Windows it needs the tzdata package.
  • The autumn DST hour happens twice. 02:30 in Berlin on 25 October 2026 is 00:30 UTC, or 01:30 UTC with fold=1. For event or financial data, find a value with an explicit offset.
  • Epochs come in seconds or milliseconds. A 13-digit value is almost certainly milliseconds; divide by 1000 before datetime.fromtimestamp(value, tz=timezone.utc), or the call fails because the value lands tens of thousands of years in the future.

Units and Pack Sizes

Price per kilogram or litre is often what people actually compare. The pint library knows the conversion factors, including those that differ by country (US pint about 473 ml, imperial pint about 568 ml), and refuses conversions between incompatible dimensions:

import pint

ureg = pint.UnitRegistry(non_int_type=Decimal)       # keep Decimal arithmetic
UNITS = {"g": "gram", "kg": "kilogram", "ml": "milliliter", "cl": "centiliter",
         "l": "liter", "lb": "pound", "fl oz": "fluid_ounce"}
SIZE = re.compile(r"(?:(\d+)\s*[x×]\s*)?(\d+(?:[.,]\d+)?)\s*(fl\.?\s?oz|kg|g|ml|cl|l|lb)\b", re.I)

def parse_size(text):
    m = SIZE.search(text)
    if not m:
        return None
    count, value, unit = m.groups()
    total = Decimal(value.replace(",", ".")) * int(count or 1)
    q = ureg.Quantity(total, UNITS[re.sub(r"[.\s]+", " ", unit.lower())])
    base = "gram" if q.check("[mass]") else "milliliter"
    return q.to(base).magnitude.quantize(Decimal("0.1")), base

for t in ["Cola 6 x 330ml", "Oat drink 1,5 l", "Shampoo 12 fl oz", "Size M"]:
    print(f"{t!r:20} -> {parse_size(t)}")
'Cola 6 x 330ml'     -> (Decimal('1980.0'), 'milliliter')
'Oat drink 1,5 l'    -> (Decimal('1500.0'), 'milliliter')
'Shampoo 12 fl oz'   -> (Decimal('354.9'), 'milliliter')
'Size M'             -> None

The price per kilogram is then price_minor * 1000 / grams, still in Decimal. Weak spots worth a test case each: fluid_ounce in pint is the US fluid ounce, so a UK product needs imperial_fluid_ounce; the comma replacement would read 1,500 g on a US site as 1.5 g, so use parse_decimal with the source locale instead; and Pack of 2, 500g each states the count in words the regex does not see. Prefer a dedicated size field from the spec table when one exists.

Finally, store price_raw, date_raw and size_raw next to the parsed fields, plus a parser version, so a locale mistake found weeks later can be re-parsed without re-crawling. Track the share of None and ambiguous results per source per run; a jump means the site changed its format or served you a different region.

Practice

Collect thirty raw price strings and fifteen raw date strings from each of two permitted sites with different number formats. Write their configs and a table of expected outputs covering a sale price, a range, a relative date and a date-only value. Then swap the two locales and confirm that strict parsing fails loudly instead of returning wrong numbers.