Syntax

Token Matches
. any character except newline (any at all with re.S)
\d / \D digit / non-digit (Unicode digits in Python, see gotchas)
\w / \W word character (letters, digits, _, Unicode letters) / not
\s / \S whitespace (includes newline, tab, non-breaking space) / not
[abc], [a-z], [^abc] one of / range / none of
^, $ start / end of string (of each line with re.M)
\A, \Z start / very end of string, ignoring re.M
\b, \B word boundary / not a boundary
*, +, ? 0+, 1+, 0 or 1 (greedy)
{3}, {2,}, {2,5} exact count / at least / range
*?, +?, ??, {2,5}? non-greedy (lazy) versions
*+, ++, (?>...) possessive / atomic, no backtracking (Python 3.11+)
a|b alternation
(...) capturing group
(?:...) non-capturing group
(?P<name>...) named group
\1, (?P=name) backreference to a group
(?=...), (?!...) lookahead / negative lookahead
(?<=...), (?<!...) lookbehind / negative lookbehind (fixed width in re)
\., \$, \( escape special characters (re.escape() does it for you)

Python re functions

Call Returns
re.search(p, s) first match anywhere, or None
re.match(p, s) match only at the start of s, or None
re.fullmatch(p, s) match only if the whole string matches
re.findall(p, s) list of strings; list of group tuples if the pattern has groups
re.finditer(p, s) iterator of match objects (use for named groups)
re.sub(p, repl, s) replaced string; repl can be a function
re.split(p, s) list of pieces
re.compile(p, flags) reusable compiled pattern
m.group(), m[1], m["name"] whole match / group 1 / named group
m.groups(), m.groupdict() all groups as a tuple / dict
m.start(), m.end(), m.span() positions in the string
import re

re.findall(r"\d+-\d+", "1-2 3-4")        # ['1-2', '3-4']
re.findall(r"(\d+)-(\d+)", "1-2 3-4")    # [('1', '2'), ('3', '4')]  groups change the result
re.match(r"\d+", "abc 123")              # None: match() anchors at the start
re.search(r"\d+", "abc 123").group()     # '123'

m = re.search(r"(?P<year>\d{4})-(?P<month>\d{2})", "2024-05")
m.groupdict()                            # {'year': '2024', 'month': '05'}

re.sub(r"(?P<y>\d{4})-(?P<m>\d{2})-(?P<d>\d{2})", r"\g<d>/\g<m>/\g<y>", "2024-05-06")
# '06/05/2024'
re.sub(r"\d+", lambda m: str(int(m[0]) * 2), "a1 b20")   # 'a2 b40'

Always write patterns as raw strings (r"\d+") so Python does not eat the backslashes.

Flags

Flag Inline Effect
re.I / re.IGNORECASE (?i) case-insensitive
re.M / re.MULTILINE (?m) ^ and $ match at every line
re.S / re.DOTALL (?s) . also matches newline (needed across lines of HTML)
re.X / re.VERBOSE (?x) ignore whitespace, allow # comments in the pattern
re.A / re.ASCII (?a) \d, \w, \s, \b become ASCII-only
  • Combine with |: re.compile(p, re.I | re.S).
  • Scoped inline flags: (?i:h)ello matches Hello and hello, but not HELLO.
  • Since Python 3.11, a global inline flag such as (?i) must be at the very start of the pattern, otherwise re.error.

Greedy vs non-greedy

s = "<b>one</b> and <b>two</b>"
re.findall(r"<b>(.*)</b>", s)     # ['one</b> and <b>two']   greedy: runs to the last </b>
re.findall(r"<b>(.*?)</b>", s)    # ['one', 'two']           lazy: stops at the first

A negated class is often clearer and faster than lazy dots: "([^"]*)" grabs a quoted string.

Lookarounds

Pattern Text Result
(?<=\$)\d+ $12 USD 5 12 (digits preceded by $)
\d+(?= items) 12 items, 5 pages 12
(?:(?<=\$)|(?<=USD ))\d+ $12 USD 5 12, 5
(?<=\$|USD )\d+ re.error: lookbehind branches must have equal width
(?<=\$\s*)\d+ re.error: no variable-width lookbehind in re

The third-party regex module allows variable-width lookbehind and \p{L} style Unicode classes; it is a drop-in replacement for most of re.

Ready-made patterns

Each was tested; read the caveat before trusting it on real data.

Data Pattern Caveats
Email [A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,} Accepts foo@bar..com and logo@2x.png (image file names). Misses name [at] site.com. Prefer a[href^="mailto:"] when there is a link
Price with symbol (?P<cur>[$€£¥])\s?(?P<amount>\d+(?:[.,\s]\d{3})*(?:[.,]\d{1,2})?) Handles $1,299.99, € 1.234,56, £5. \s as a thousands separator can merge $5 100 items into 5 100. Symbol-after formats (10,50 €) need a second pattern
ISO date \b(\d{4})-(0[1-9]|1[0-2])-(0[1-9]|[12]\d|3[01])\b Accepts 2025-02-30: validate with datetime.strptime. Misses 2024-12-01T10:00 because \b fails before T; drop the final \b for timestamps
Month-name date \b(?:Jan(?:uary)?|Feb(?:ruary)?|Mar(?:ch)?|Apr(?:il)?|May|June?|July?|Aug(?:ust)?|Sep(?:t(?:ember)?)?|Oct(?:ober)?|Nov(?:ember)?|Dec(?:ember)?)\.?\s+\d{1,2}(?:st|nd|rd|th)?,?\s+\d{4}\b with re.I Matches March 3rd, 2025, Sept. 14 2024. Accepts June 31: validate. English only
Numeric date \b(\d{1,2})[/.-](\d{1,2})[/.-](\d{2}|\d{4})\b 03/04/2025 is March 4 (US) or 3 April (EU): you must know the site's locale
URL in text https?://[^\s"'<>]+ Picks up trailing . , ): clean it (below). For links in HTML, read href with a parser and urljoin
US phone (?:\+?1[\s.-]?)?(?:\(\d{3}\)|\b\d{3})[\s.-]?\d{3}[\s.-]?\d{4}\b Matches (555) 123-4567, +1 555.123.4567, but also any 10-digit run such as order numbers. For real work use the phonenumbers library
E.164 phone \+[1-9]\d{6,14}\b Only normalised +<country><number> strings
IPv4 \b(?:(?:25[0-5]|2[0-4]\d|1?\d?\d)\.){3}(?:25[0-5]|2[0-4]\d|1?\d?\d)\b Rejects 256.1.1.1, but takes 1.2.3.4 out of 1.2.3.4.5 and accepts 01.2.3.4: confirm with ipaddress.ip_address()
Number with thousands -?\d{1,3}(?:,\d{3})+(?:\.\d+)?|-?\d+(?:\.\d+)? US format; 12,34 splits into 12 and 34
Percentage (\d+(?:\.\d+)?)\s?% 15%, 7.5 %
Collapse whitespace re.sub(r"\s+", " ", s).strip() Also replaces non-breaking spaces (\xa0), which is usually what you want

Helpers for the messy cases

import re
from decimal import Decimal

def to_decimal(text):
    # '1,299.99' / '1.299,99' / '1 299' -> Decimal.
    # The last separator followed by 1-2 digits is the decimal mark.
    s = re.sub(r"[^\d.,]", "", text)
    m = re.search(r"[.,](\d{1,2})$", s)
    whole, frac = (s[:m.start()], m[1]) if m else (s, "0")
    return Decimal(re.sub(r"[.,]", "", whole) + "." + frac)

to_decimal("1,299.99")   # Decimal('1299.99')
to_decimal("1.299,99")   # Decimal('1299.99')
to_decimal("10,5")       # Decimal('10.5')
# Caveat: "1,234" -> 1234 (3 digits are read as thousands, never decimals)

def clean_url(u):
    # Strip trailing punctuation, but keep a ')' that closes a '(' in the URL.
    u = u.rstrip(".,;:!?")
    if u.endswith(")") and u.count("(") < u.count(")"):
        u = u[:-1]
    return u

[clean_url(u) for u in re.findall(r"https?://[^\s\"'<>]+", "(see https://ex.com/a).")]
# ['https://ex.com/a']

Money: parse into Decimal, never float.

JSON inside script tags

Use regex only to find where the JSON starts, then let the JSON decoder find where it ends. A lazy ({.*?}); stops at the first };, even one inside a string.

import json, re

m = re.search(r"window\.__INITIAL_STATE__\s*=\s*", html)
data, end = json.JSONDecoder().raw_decode(html, m.end())

Python gotchas

  • \d matches any Unicode digit: re.findall(r"\d+", "٣٤") finds Arabic-Indic digits. Use [0-9] or re.ASCII when you mean 0-9 (int() accepts those digits, other code may not).
  • \w matches accented letters (café); with re.A it stops at caf.
  • $ also matches before a trailing newline: re.findall(r"\d+$", "123\n") finds 123. Use \Z or fullmatch() for strict checks.
  • re.findall with groups returns only the groups. Use (?:...) for grouping you do not want back, or finditer.
  • re.match is not "match anywhere": that is re.search.
  • Compile patterns used in a loop (PRICE = re.compile(...)); re caches some, but explicit is clearer.

Catastrophic backtracking

Nested quantifiers such as (a+)+$ on a string that almost matches take exponential time: on "a" * 24 + "b" it took about 0.8 s, and each extra character roughly doubles it. Fixes:

  • remove the nesting: a+$
  • make the inner part atomic or possessive (Python 3.11+): (?>a+)+$, a++$
  • anchor and bound repeats: \d{1,10} instead of \d+ when you know the limits
  • never run user-supplied patterns over large pages without a timeout

When not to use regex on HTML

Regex sees characters, not elements. This "simple" link extractor misses two of three links and returns an escaped URL:

html = '''<a class="x" href='/one'>1</a><a href="/two?a=1&amp;b=2">2</a><a
href="/three">3</a>'''

re.findall(r'<a href="(.*?)"', html)   # ['/two?a=1&amp;b=2']

from lxml import html as lh
lh.fromstring(f"<div>{html}</div>").xpath("//a/@href")
# ['/one', '/two?a=1&b=2', '/three']

What breaks regex on HTML: attribute order, single vs double quotes, newlines inside tags, entities (&amp;), comments and scripts that contain markup, and nesting.

Use a parser (CSS / XPath) for Use regex for
finding elements and attributes pulling a number, date or ID out of a text node
tables, lists, nested structure locating JSON or variables inside script tags
anything where markup may change URLs and tokens inside JavaScript strings
cleaning text: whitespace, currency symbols, units

Best of both: select the element with a parser, then run the regex on element.get_text().