Quick decision table

Status Default action
2xx Success, but still check the body: some sites serve a block page or "no results" with 200
301, 308 Follow, and store the new URL (permanent)
302, 303, 307 Follow, keep the original URL (temporary)
304 Use your cached copy (only appears if you sent If-None-Match / If-Modified-Since)
400, 404, 405, 410, 422 Do not retry: the request or URL is wrong
401, 407 Fix credentials (log in again, refresh the token, check proxy auth), then retry once
403 Usually a block when scraping. Slow down, fix headers / fingerprint; do not hammer
408, 425 Retry with backoff
429 Back off: honour Retry-After, lower the rate for the whole host
500, 502, 503, 504 Retry with exponential backoff + jitter; honour Retry-After on 503
501, 505 Do not retry: the server does not support what you asked for

1xx Informational

You rarely see these: HTTP libraries handle them before handing you the final response.

Code Meaning Client action
100 Continue Server agrees to receive the request body (after Expect: 100-continue) Handled by the library
101 Switching Protocols Upgrade accepted, e.g. to WebSocket Expected when opening a WebSocket
103 Early Hints Preload hints (Link headers) before the real response Ignore; wait for the final status

2xx Success

Code Meaning Client action
200 OK Request succeeded Parse it, but validate: a CAPTCHA or "access denied" page can come back as 200
201 Created Resource created (POST/PUT) Read the Location header for the new resource
202 Accepted Queued for async processing, not done yet Poll the status URL given in the body or Location
204 No Content Success with an empty body Do not call .json(): there is nothing to decode
206 Partial Content Only the requested byte range was sent Expected with Range:; use it to resume downloads

3xx Redirection

Code Meaning Method on redirect Client action
301 Moved Permanently URL changed for good Clients may turn POST into GET Update stored URLs / dedupe keys
302 Found Temporary redirect Clients turn POST into GET in practice Follow; keep the original URL
303 See Other "Fetch the result with GET over there" Always GET Follow (typical after a form POST)
304 Not Modified Your cached copy is still valid n/a Reuse the cached body; saves bandwidth on re-crawls
307 Temporary Redirect Temporary, method and body kept Same method Follow, re-send the body
308 Permanent Redirect Permanent, method and body kept Same method Follow and update stored URLs
  • requests follows redirects by default (except for requests.head()); a 302 or 303 turns POST into GET, and so does a 301 for POST.
  • httpx does not follow redirects unless you pass follow_redirects=True.
  • Redirect loops raise TooManyRedirects. A redirect to a login or consent page usually means your session or cookies are missing.
  • response.history (requests) lists every hop; check it when a "200" page looks wrong.

4xx Client errors

Code Meaning Retry? Scraper / API client action
400 Bad Request Malformed request No Check the body, JSON, query params and encoding
401 Unauthorized Not authenticated (missing / expired credentials) After fixing Refresh the token or log in again; read WWW-Authenticate
403 Forbidden Refused. When scraping, usually a WAF / bot block Not blindly Slow down; send realistic headers; check TLS fingerprint, IP reputation, cookies
404 Not Found No resource at this URL No Mark the URL dead. Watch for soft 404s: 200 pages that say "not found"
405 Method Not Allowed Wrong HTTP method No Read the Allow header
406 Not Acceptable Cannot satisfy your Accept headers No Loosen or fix Accept / Accept-Language
407 Proxy Authentication Required The proxy wants credentials After fixing Check proxy username and password
408 Request Timeout Server gave up waiting for your request Yes Retry with backoff
409 Conflict Clashes with current state (edit conflict, duplicate) Rarely Re-fetch the state, then decide
410 Gone Removed permanently No Drop the URL from your frontier for good
412 Precondition Failed If-Match / If-Unmodified-Since did not hold No Re-fetch, then retry with the new ETag
413 Content Too Large Request body too big No Shrink or chunk the upload
414 URI Too Long URL too long No Move parameters into a POST body
415 Unsupported Media Type Wrong Content-Type No Send JSON or form data, whichever the API expects
418 I'm a teapot Joke code (RFC 2324); some sites return it to bots No Treat as a block
421 Misdirected Request Wrong server for this host (HTTP/2 connection reuse) Yes Retry on a new connection
422 Unprocessable Content Well-formed but fails validation No Read the error body (FastAPI, Rails and others list the bad fields)
425 Too Early Server will not process a TLS early-data request Yes Retry after the handshake completes
428 Precondition Required Server requires a conditional request After fixing Send If-Match with the current ETag
429 Too Many Requests Rate limited Yes, later Honour Retry-After; slow the whole host down, not just this URL
431 Request Header Fields Too Large Headers too big, often a bloated cookie jar No Trim cookies and headers
451 Unavailable For Legal Reasons Blocked for legal reasons, often by region No Respect it; geo-blocking may be involved

5xx Server errors

Code Meaning Retry? Scraper / API client action
500 Internal Server Error Server crashed on your request Yes, a few times Back off. If it is always the same URL, your input may trigger the bug
501 Not Implemented Server does not support the method No Use a different method or endpoint
502 Bad Gateway Proxy / CDN got a bad answer from the origin Yes Back off and retry; also check your own proxy
503 Service Unavailable Overloaded or in maintenance; also used by bot challenges Yes Honour Retry-After; if the body is a challenge page, retrying will not help
504 Gateway Timeout Proxy / CDN timed out waiting for the origin Yes Back off; ask for smaller pages if the API allows it
505 HTTP Version Not Supported Protocol version refused No Use HTTP/1.1
511 Network Authentication Required Captive portal (hotel / airport Wi-Fi) No Log in to the network

Non-standard codes you will meet

Code Who Meaning
499 nginx Client closed the connection before the response; seen in server logs, never received by the client
520 Cloudflare Origin returned an unknown or empty response
521 Cloudflare Origin web server is down (refused the connection)
522 Cloudflare Connection to the origin timed out
523 Cloudflare Origin unreachable
524 Cloudflare Origin accepted the connection but did not answer in time
525 / 526 Cloudflare TLS handshake with the origin failed / invalid origin certificate
999 LinkedIn Request denied to suspected bots

Cloudflare 52x codes are origin problems: retry later with backoff. A Cloudflare challenge usually comes as 403 or 503 with a cf-mitigated: challenge header or a challenge HTML page.

Retry rules

  • Retry only idempotent requests (GET, HEAD, PUT, DELETE, OPTIONS) automatically. Retrying a POST can create duplicates unless the API supports an Idempotency-Key header.
  • Use exponential backoff with jitter and a cap on attempts and on total wait.
  • Retry-After is either seconds (120) or an HTTP date. Honour it on 429 and 503.
  • Back off per host: when one URL gets a 429, every worker hitting that host should slow down.
  • Always set a timeout. requests waits forever by default; httpx defaults to 5 seconds.

requests + urllib3 Retry

import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

retry = Retry(
    total=5,
    backoff_factor=1,                     # urllib3 2.x sleeps 0, 2, 4, 8, 16 s (max 120 s)
    status_forcelist=[429, 500, 502, 503, 504],
    allowed_methods=["GET", "HEAD", "OPTIONS"],   # POST is never retried here
    respect_retry_after_header=True,      # default; applies to 413, 429, 503
)
session = requests.Session()
session.mount("https://", HTTPAdapter(max_retries=retry))
session.mount("http://", HTTPAdapter(max_retries=retry))

resp = session.get(url, timeout=(5, 30))  # (connect, read) seconds
resp.raise_for_status()                   # turn 4xx/5xx into exceptions

When the retries run out on a listed status, requests raises requests.exceptions.RetryError.

Parsing Retry-After yourself

import random
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime

def retry_after_seconds(value, default=None):
    # Retry-After is delta-seconds ("120") or an HTTP-date.
    if not value:
        return default
    if value.strip().isdigit():
        return int(value)
    try:
        when = parsedate_to_datetime(value)
    except (TypeError, ValueError):
        return default
    return max(0.0, (when - datetime.now(timezone.utc)).total_seconds())

def backoff(attempt, base=1.0, cap=60.0):
    # Exponential backoff with full jitter.
    return random.uniform(0, min(cap, base * 2 ** attempt))