What each status code means and, more usefully, what your scraper or API client should do about it. "Retry" always means retry with backoff and a cap, and only for requests that are safe to repeat.
Success, but still check the body: some sites serve a block page or "no results" with 200
301, 308
Follow, and store the new URL (permanent)
302, 303, 307
Follow, keep the original URL (temporary)
304
Use your cached copy (only appears if you sent If-None-Match / If-Modified-Since)
400, 404, 405, 410, 422
Do not retry: the request or URL is wrong
401, 407
Fix credentials (log in again, refresh the token, check proxy auth), then retry once
403
Usually a block when scraping. Slow down, fix headers / fingerprint; do not hammer
408, 425
Retry with backoff
429
Back off: honour Retry-After, lower the rate for the whole host
500, 502, 503, 504
Retry with exponential backoff + jitter; honour Retry-After on 503
501, 505
Do not retry: the server does not support what you asked for
1xx Informational
You rarely see these: HTTP libraries handle them before handing you the final response.
Code
Meaning
Client action
100 Continue
Server agrees to receive the request body (after Expect: 100-continue)
Handled by the library
101 Switching Protocols
Upgrade accepted, e.g. to WebSocket
Expected when opening a WebSocket
103 Early Hints
Preload hints (Link headers) before the real response
Ignore; wait for the final status
2xx Success
Code
Meaning
Client action
200 OK
Request succeeded
Parse it, but validate: a CAPTCHA or "access denied" page can come back as 200
201 Created
Resource created (POST/PUT)
Read the Location header for the new resource
202 Accepted
Queued for async processing, not done yet
Poll the status URL given in the body or Location
204 No Content
Success with an empty body
Do not call .json(): there is nothing to decode
206 Partial Content
Only the requested byte range was sent
Expected with Range:; use it to resume downloads
3xx Redirection
Code
Meaning
Method on redirect
Client action
301 Moved Permanently
URL changed for good
Clients may turn POST into GET
Update stored URLs / dedupe keys
302 Found
Temporary redirect
Clients turn POST into GET in practice
Follow; keep the original URL
303 See Other
"Fetch the result with GET over there"
Always GET
Follow (typical after a form POST)
304 Not Modified
Your cached copy is still valid
n/a
Reuse the cached body; saves bandwidth on re-crawls
307 Temporary Redirect
Temporary, method and body kept
Same method
Follow, re-send the body
308 Permanent Redirect
Permanent, method and body kept
Same method
Follow and update stored URLs
requests follows redirects by default (except for requests.head()); a 302 or 303 turns POST into GET, and so does a 301 for POST.
httpx does not follow redirects unless you pass follow_redirects=True.
Redirect loops raise TooManyRedirects. A redirect to a login or consent page usually means your session or cookies are missing.
response.history (requests) lists every hop; check it when a "200" page looks wrong.
4xx Client errors
Code
Meaning
Retry?
Scraper / API client action
400 Bad Request
Malformed request
No
Check the body, JSON, query params and encoding
401 Unauthorized
Not authenticated (missing / expired credentials)
After fixing
Refresh the token or log in again; read WWW-Authenticate
403 Forbidden
Refused. When scraping, usually a WAF / bot block
Not blindly
Slow down; send realistic headers; check TLS fingerprint, IP reputation, cookies
404 Not Found
No resource at this URL
No
Mark the URL dead. Watch for soft 404s: 200 pages that say "not found"
405 Method Not Allowed
Wrong HTTP method
No
Read the Allow header
406 Not Acceptable
Cannot satisfy your Accept headers
No
Loosen or fix Accept / Accept-Language
407 Proxy Authentication Required
The proxy wants credentials
After fixing
Check proxy username and password
408 Request Timeout
Server gave up waiting for your request
Yes
Retry with backoff
409 Conflict
Clashes with current state (edit conflict, duplicate)
Rarely
Re-fetch the state, then decide
410 Gone
Removed permanently
No
Drop the URL from your frontier for good
412 Precondition Failed
If-Match / If-Unmodified-Since did not hold
No
Re-fetch, then retry with the new ETag
413 Content Too Large
Request body too big
No
Shrink or chunk the upload
414 URI Too Long
URL too long
No
Move parameters into a POST body
415 Unsupported Media Type
Wrong Content-Type
No
Send JSON or form data, whichever the API expects
418 I'm a teapot
Joke code (RFC 2324); some sites return it to bots
No
Treat as a block
421 Misdirected Request
Wrong server for this host (HTTP/2 connection reuse)
Yes
Retry on a new connection
422 Unprocessable Content
Well-formed but fails validation
No
Read the error body (FastAPI, Rails and others list the bad fields)
425 Too Early
Server will not process a TLS early-data request
Yes
Retry after the handshake completes
428 Precondition Required
Server requires a conditional request
After fixing
Send If-Match with the current ETag
429 Too Many Requests
Rate limited
Yes, later
Honour Retry-After; slow the whole host down, not just this URL
431 Request Header Fields Too Large
Headers too big, often a bloated cookie jar
No
Trim cookies and headers
451 Unavailable For Legal Reasons
Blocked for legal reasons, often by region
No
Respect it; geo-blocking may be involved
5xx Server errors
Code
Meaning
Retry?
Scraper / API client action
500 Internal Server Error
Server crashed on your request
Yes, a few times
Back off. If it is always the same URL, your input may trigger the bug
501 Not Implemented
Server does not support the method
No
Use a different method or endpoint
502 Bad Gateway
Proxy / CDN got a bad answer from the origin
Yes
Back off and retry; also check your own proxy
503 Service Unavailable
Overloaded or in maintenance; also used by bot challenges
Yes
Honour Retry-After; if the body is a challenge page, retrying will not help
504 Gateway Timeout
Proxy / CDN timed out waiting for the origin
Yes
Back off; ask for smaller pages if the API allows it
505 HTTP Version Not Supported
Protocol version refused
No
Use HTTP/1.1
511 Network Authentication Required
Captive portal (hotel / airport Wi-Fi)
No
Log in to the network
Non-standard codes you will meet
Code
Who
Meaning
499
nginx
Client closed the connection before the response; seen in server logs, never received by the client
520
Cloudflare
Origin returned an unknown or empty response
521
Cloudflare
Origin web server is down (refused the connection)
522
Cloudflare
Connection to the origin timed out
523
Cloudflare
Origin unreachable
524
Cloudflare
Origin accepted the connection but did not answer in time
525 / 526
Cloudflare
TLS handshake with the origin failed / invalid origin certificate
999
LinkedIn
Request denied to suspected bots
Cloudflare 52x codes are origin problems: retry later with backoff. A Cloudflare challenge usually comes as 403 or 503 with a cf-mitigated: challenge header or a challenge HTML page.
Retry rules
Retry only idempotent requests (GET, HEAD, PUT, DELETE, OPTIONS) automatically. Retrying a POST can create duplicates unless the API supports an Idempotency-Key header.
Use exponential backoff with jitter and a cap on attempts and on total wait.
Retry-After is either seconds (120) or an HTTP date. Honour it on 429 and 503.
Back off per host: when one URL gets a 429, every worker hitting that host should slow down.
Always set a timeout. requests waits forever by default; httpx defaults to 5 seconds.
requests + urllib3 Retry
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
retry = Retry(
total=5,
backoff_factor=1, # urllib3 2.x sleeps 0, 2, 4, 8, 16 s (max 120 s)
status_forcelist=[429, 500, 502, 503, 504],
allowed_methods=["GET", "HEAD", "OPTIONS"], # POST is never retried here
respect_retry_after_header=True, # default; applies to 413, 429, 503
)
session = requests.Session()
session.mount("https://", HTTPAdapter(max_retries=retry))
session.mount("http://", HTTPAdapter(max_retries=retry))
resp = session.get(url, timeout=(5, 30)) # (connect, read) seconds
resp.raise_for_status() # turn 4xx/5xx into exceptions
When the retries run out on a listed status, requests raises requests.exceptions.RetryError.
Parsing Retry-After yourself
import random
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime
def retry_after_seconds(value, default=None):
# Retry-After is delta-seconds ("120") or an HTTP-date.
if not value:
return default
if value.strip().isdigit():
return int(value)
try:
when = parsedate_to_datetime(value)
except (TypeError, ValueError):
return default
return max(0.0, (when - datetime.now(timezone.utc)).total_seconds())
def backoff(attempt, base=1.0, cap=60.0):
# Exponential backoff with full jitter.
return random.uniform(0, min(cap, base * 2 ** attempt))
Nothing on this sheet matches.
Source: simpleprog.com/cheatsheets/http-status-codes · free courses at simpleprog.com