Syntax
| Token | Matches |
|---|---|
. |
any character except newline (any at all with re.S) |
\d / \D |
digit / non-digit (Unicode digits in Python, see gotchas) |
\w / \W |
word character (letters, digits, _, Unicode letters) / not |
\s / \S |
whitespace (includes newline, tab, non-breaking space) / not |
[abc], [a-z], [^abc] |
one of / range / none of |
^, $ |
start / end of string (of each line with re.M) |
\A, \Z |
start / very end of string, ignoring re.M |
\b, \B |
word boundary / not a boundary |
*, +, ? |
0+, 1+, 0 or 1 (greedy) |
{3}, {2,}, {2,5} |
exact count / at least / range |
*?, +?, ??, {2,5}? |
non-greedy (lazy) versions |
*+, ++, (?>...) |
possessive / atomic, no backtracking (Python 3.11+) |
a|b |
alternation |
(...) |
capturing group |
(?:...) |
non-capturing group |
(?P<name>...) |
named group |
\1, (?P=name) |
backreference to a group |
(?=...), (?!...) |
lookahead / negative lookahead |
(?<=...), (?<!...) |
lookbehind / negative lookbehind (fixed width in re) |
\., \$, \( |
escape special characters (re.escape() does it for you) |
Python re functions
| Call | Returns |
|---|---|
re.search(p, s) |
first match anywhere, or None |
re.match(p, s) |
match only at the start of s, or None |
re.fullmatch(p, s) |
match only if the whole string matches |
re.findall(p, s) |
list of strings; list of group tuples if the pattern has groups |
re.finditer(p, s) |
iterator of match objects (use for named groups) |
re.sub(p, repl, s) |
replaced string; repl can be a function |
re.split(p, s) |
list of pieces |
re.compile(p, flags) |
reusable compiled pattern |
m.group(), m[1], m["name"] |
whole match / group 1 / named group |
m.groups(), m.groupdict() |
all groups as a tuple / dict |
m.start(), m.end(), m.span() |
positions in the string |
import re
re.findall(r"\d+-\d+", "1-2 3-4") # ['1-2', '3-4']
re.findall(r"(\d+)-(\d+)", "1-2 3-4") # [('1', '2'), ('3', '4')] groups change the result
re.match(r"\d+", "abc 123") # None: match() anchors at the start
re.search(r"\d+", "abc 123").group() # '123'
m = re.search(r"(?P<year>\d{4})-(?P<month>\d{2})", "2024-05")
m.groupdict() # {'year': '2024', 'month': '05'}
re.sub(r"(?P<y>\d{4})-(?P<m>\d{2})-(?P<d>\d{2})", r"\g<d>/\g<m>/\g<y>", "2024-05-06")
# '06/05/2024'
re.sub(r"\d+", lambda m: str(int(m[0]) * 2), "a1 b20") # 'a2 b40'
Always write patterns as raw strings (r"\d+") so Python does not eat the backslashes.
Flags
| Flag | Inline | Effect |
|---|---|---|
re.I / re.IGNORECASE |
(?i) |
case-insensitive |
re.M / re.MULTILINE |
(?m) |
^ and $ match at every line |
re.S / re.DOTALL |
(?s) |
. also matches newline (needed across lines of HTML) |
re.X / re.VERBOSE |
(?x) |
ignore whitespace, allow # comments in the pattern |
re.A / re.ASCII |
(?a) |
\d, \w, \s, \b become ASCII-only |
- Combine with
|:re.compile(p, re.I | re.S). - Scoped inline flags:
(?i:h)ellomatchesHelloandhello, but notHELLO. - Since Python 3.11, a global inline flag such as
(?i)must be at the very start of the pattern, otherwisere.error.
Greedy vs non-greedy
s = "<b>one</b> and <b>two</b>"
re.findall(r"<b>(.*)</b>", s) # ['one</b> and <b>two'] greedy: runs to the last </b>
re.findall(r"<b>(.*?)</b>", s) # ['one', 'two'] lazy: stops at the first
A negated class is often clearer and faster than lazy dots: "([^"]*)" grabs a quoted string.
Lookarounds
| Pattern | Text | Result |
|---|---|---|
(?<=\$)\d+ |
$12 USD 5 |
12 (digits preceded by $) |
\d+(?= items) |
12 items, 5 pages |
12 |
(?:(?<=\$)|(?<=USD ))\d+ |
$12 USD 5 |
12, 5 |
(?<=\$|USD )\d+ |
re.error: lookbehind branches must have equal width |
|
(?<=\$\s*)\d+ |
re.error: no variable-width lookbehind in re |
The third-party regex module allows variable-width lookbehind and \p{L} style Unicode classes; it is a drop-in replacement for most of re.
Ready-made patterns
Each was tested; read the caveat before trusting it on real data.
| Data | Pattern | Caveats |
|---|---|---|
[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,} |
Accepts foo@bar..com and logo@2x.png (image file names). Misses name [at] site.com. Prefer a[href^="mailto:"] when there is a link |
|
| Price with symbol | (?P<cur>[$€£¥])\s?(?P<amount>\d+(?:[.,\s]\d{3})*(?:[.,]\d{1,2})?) |
Handles $1,299.99, € 1.234,56, £5. \s as a thousands separator can merge $5 100 items into 5 100. Symbol-after formats (10,50 €) need a second pattern |
| ISO date | \b(\d{4})-(0[1-9]|1[0-2])-(0[1-9]|[12]\d|3[01])\b |
Accepts 2025-02-30: validate with datetime.strptime. Misses 2024-12-01T10:00 because \b fails before T; drop the final \b for timestamps |
| Month-name date | \b(?:Jan(?:uary)?|Feb(?:ruary)?|Mar(?:ch)?|Apr(?:il)?|May|June?|July?|Aug(?:ust)?|Sep(?:t(?:ember)?)?|Oct(?:ober)?|Nov(?:ember)?|Dec(?:ember)?)\.?\s+\d{1,2}(?:st|nd|rd|th)?,?\s+\d{4}\b with re.I |
Matches March 3rd, 2025, Sept. 14 2024. Accepts June 31: validate. English only |
| Numeric date | \b(\d{1,2})[/.-](\d{1,2})[/.-](\d{2}|\d{4})\b |
03/04/2025 is March 4 (US) or 3 April (EU): you must know the site's locale |
| URL in text | https?://[^\s"'<>]+ |
Picks up trailing . , ): clean it (below). For links in HTML, read href with a parser and urljoin |
| US phone | (?:\+?1[\s.-]?)?(?:\(\d{3}\)|\b\d{3})[\s.-]?\d{3}[\s.-]?\d{4}\b |
Matches (555) 123-4567, +1 555.123.4567, but also any 10-digit run such as order numbers. For real work use the phonenumbers library |
| E.164 phone | \+[1-9]\d{6,14}\b |
Only normalised +<country><number> strings |
| IPv4 | \b(?:(?:25[0-5]|2[0-4]\d|1?\d?\d)\.){3}(?:25[0-5]|2[0-4]\d|1?\d?\d)\b |
Rejects 256.1.1.1, but takes 1.2.3.4 out of 1.2.3.4.5 and accepts 01.2.3.4: confirm with ipaddress.ip_address() |
| Number with thousands | -?\d{1,3}(?:,\d{3})+(?:\.\d+)?|-?\d+(?:\.\d+)? |
US format; 12,34 splits into 12 and 34 |
| Percentage | (\d+(?:\.\d+)?)\s?% |
15%, 7.5 % |
| Collapse whitespace | re.sub(r"\s+", " ", s).strip() |
Also replaces non-breaking spaces (\xa0), which is usually what you want |
Helpers for the messy cases
import re
from decimal import Decimal
def to_decimal(text):
# '1,299.99' / '1.299,99' / '1 299' -> Decimal.
# The last separator followed by 1-2 digits is the decimal mark.
s = re.sub(r"[^\d.,]", "", text)
m = re.search(r"[.,](\d{1,2})$", s)
whole, frac = (s[:m.start()], m[1]) if m else (s, "0")
return Decimal(re.sub(r"[.,]", "", whole) + "." + frac)
to_decimal("1,299.99") # Decimal('1299.99')
to_decimal("1.299,99") # Decimal('1299.99')
to_decimal("10,5") # Decimal('10.5')
# Caveat: "1,234" -> 1234 (3 digits are read as thousands, never decimals)
def clean_url(u):
# Strip trailing punctuation, but keep a ')' that closes a '(' in the URL.
u = u.rstrip(".,;:!?")
if u.endswith(")") and u.count("(") < u.count(")"):
u = u[:-1]
return u
[clean_url(u) for u in re.findall(r"https?://[^\s\"'<>]+", "(see https://ex.com/a).")]
# ['https://ex.com/a']
Money: parse into Decimal, never float.
JSON inside script tags
Use regex only to find where the JSON starts, then let the JSON decoder find where it ends. A lazy ({.*?}); stops at the first };, even one inside a string.
import json, re
m = re.search(r"window\.__INITIAL_STATE__\s*=\s*", html)
data, end = json.JSONDecoder().raw_decode(html, m.end())
Python gotchas
\dmatches any Unicode digit:re.findall(r"\d+", "٣٤")finds Arabic-Indic digits. Use[0-9]orre.ASCIIwhen you mean 0-9 (int()accepts those digits, other code may not).\wmatches accented letters (café); withre.Ait stops atcaf.$also matches before a trailing newline:re.findall(r"\d+$", "123\n")finds123. Use\Zorfullmatch()for strict checks.re.findallwith groups returns only the groups. Use(?:...)for grouping you do not want back, orfinditer.re.matchis not "match anywhere": that isre.search.- Compile patterns used in a loop (
PRICE = re.compile(...));recaches some, but explicit is clearer.
Catastrophic backtracking
Nested quantifiers such as (a+)+$ on a string that almost matches take exponential time: on "a" * 24 + "b" it took about 0.8 s, and each extra character roughly doubles it. Fixes:
- remove the nesting:
a+$ - make the inner part atomic or possessive (Python 3.11+):
(?>a+)+$,a++$ - anchor and bound repeats:
\d{1,10}instead of\d+when you know the limits - never run user-supplied patterns over large pages without a timeout
When not to use regex on HTML
Regex sees characters, not elements. This "simple" link extractor misses two of three links and returns an escaped URL:
html = '''<a class="x" href='/one'>1</a><a href="/two?a=1&b=2">2</a><a
href="/three">3</a>'''
re.findall(r'<a href="(.*?)"', html) # ['/two?a=1&b=2']
from lxml import html as lh
lh.fromstring(f"<div>{html}</div>").xpath("//a/@href")
# ['/one', '/two?a=1&b=2', '/three']
What breaks regex on HTML: attribute order, single vs double quotes, newlines inside tags, entities (&), comments and scripts that contain markup, and nesting.
| Use a parser (CSS / XPath) for | Use regex for |
|---|---|
| finding elements and attributes | pulling a number, date or ID out of a text node |
| tables, lists, nested structure | locating JSON or variables inside script tags |
| anything where markup may change | URLs and tokens inside JavaScript strings |
| cleaning text: whitespace, currency symbols, units |
Best of both: select the element with a parser, then run the regex on element.get_text().