Rate Limiting and Brute-Force Protection
Making Guessing Too Slow to Pay Off
Every endpoint that checks a secret can be attacked by trying secrets until one works: the login form, the password reset token, the six-digit MFA code, the gift card balance checker. Rate limiting makes guessing slow enough that the attacker gives up or gets caught. Doing it well means choosing the algorithm, the key to count by, and the response when the limit trips, without locking out real users. Password security introduces login throttling and API security lists unrestricted resource consumption as a top risk; this lesson builds the limiter itself.
The Arithmetic of Brute Force
Limits turn attacks from minutes into centuries. A six-digit one-time code has 1,000,000 possibilities, so an attacker needs about 500,000 guesses on average. At 100 requests per second, unlimited, that is roughly 83 minutes. Allow 5 attempts per code and then invalidate it, and the chance of success per code drops to 5 in 1,000,000, or 1 in 200,000. The same logic applies to reset tokens (make them long and random so the space is enormous), and to passwords, where the defence against online guessing is a limit per account, not password complexity.
Four Algorithms and Where They Break
| Algorithm | How it counts | Weakness |
|---|---|---|
| Fixed window | Counter per key per calendar minute | Allows double the limit across a window boundary |
| Sliding window log | Stores every request timestamp | Exact, but memory grows with the limit |
| Sliding window counter | Current count plus the previous window's count, weighted by overlap | Approximate, but cheap and close enough |
| Token bucket | Tokens refill at a steady rate up to a capacity; each request spends one | Needs two values per key (tokens, last update) |
Here are three of them with an injectable clock, so you can see the boundary problem rather than take it on trust:
class FixedWindow:
def __init__(self, limit, window):
self.limit, self.window, self.counts = limit, window, {}
def allow(self, key, now):
slot = (key, int(now // self.window))
self.counts[slot] = self.counts.get(slot, 0) + 1
return self.counts[slot] <= self.limit
class SlidingWindowCounter:
'''Weights the previous window's count by how much of it still overlaps.'''
def __init__(self, limit, window):
self.limit, self.window, self.counts = limit, window, {}
def allow(self, key, now):
cur = int(now // self.window)
elapsed = (now % self.window) / self.window
prev = self.counts.get((key, cur - 1), 0)
this = self.counts.get((key, cur), 0)
if prev * (1 - elapsed) + this + 1 > self.limit:
return False
self.counts[(key, cur)] = this + 1
return True
class TokenBucket:
'''Holds up to `capacity` tokens, refilled at `rate` per second.'''
def __init__(self, capacity, rate):
self.capacity, self.rate, self.state = capacity, rate, {}
def allow(self, key, now):
tokens, last = self.state.get(key, (self.capacity, now))
tokens = min(self.capacity, tokens + (now - last) * self.rate)
if tokens < 1:
self.state[key] = (tokens, now)
return False
self.state[key] = (tokens - 1, now)
return True
# Policy: 10 login attempts per minute. The attacker fires 10 at t=59s and 10 at t=61s.
times = [59.0 + i * 0.05 for i in range(10)] + [61.0 + i * 0.05 for i in range(10)]
for limiter in (FixedWindow(10, 60), SlidingWindowCounter(10, 60), TokenBucket(10, 10 / 60)):
allowed = sum(limiter.allow("alice", t) for t in times)
print(f"{type(limiter).__name__:21} allowed {allowed} of 20 in ~2.5 seconds")
FixedWindow allowed 20 of 20 in ~2.5 seconds
SlidingWindowCounter allowed 10 of 20 in ~2.5 seconds
TokenBucket allowed 10 of 20 in ~2.5 seconds
The fixed window resets at t=60, so twice the policy gets through in two and a half seconds. For login and code verification, use a token bucket or sliding window counter. The token bucket's parameters map neatly onto policy: capacity is the quick burst a user may make (mistyping twice is normal), rate the long-run allowance.
Making It Work Across Servers
The dictionaries above live in one process. With four app servers each running eight workers, an in-memory limit of 10 becomes 320. Production limiters keep state in a shared store, and the read-modify-write must be atomic or two concurrent requests both see "one token left". A Lua script runs atomically inside Redis:
import time
import redis
TOKEN_BUCKET = '''
local key = KEYS[1]
local capacity = tonumber(ARGV[1])
local rate = tonumber(ARGV[2]) -- tokens per second
local now = tonumber(ARGV[3])
local state = redis.call("HMGET", key, "tokens", "ts")
local tokens = tonumber(state[1]) or capacity
local ts = tonumber(state[2]) or now
tokens = math.min(capacity, tokens + (now - ts) * rate)
local allowed = 0
if tokens >= 1 then
tokens = tokens - 1
allowed = 1
end
redis.call("HSET", key, "tokens", tokens, "ts", now)
redis.call("EXPIRE", key, math.ceil(capacity / rate) + 1)
local retry_after = 0
if allowed == 0 then retry_after = math.ceil((1 - tokens) / rate) end
return {allowed, retry_after}
'''
r = redis.Redis(host="localhost", port=6379)
take_token = r.register_script(TOKEN_BUCKET)
def allow(key: str, capacity: int, per_second: float) -> tuple[bool, int]:
allowed, retry_after = take_token(keys=[f"rl:{key}"],
args=[capacity, per_second, time.time()])
return bool(allowed), int(retry_after)
for i in range(7): # 5 attempts per 5 minutes for one account
print(i + 1, allow("login:acct:alice@example.com", capacity=5, per_second=5 / 300))
1 (True, 0)
2 (True, 0)
3 (True, 0)
4 (True, 0)
5 (True, 0)
6 (False, 60)
7 (False, 60)
retry_after is ready for a Retry-After header, and EXPIRE lets idle keys disappear so memory tracks active clients. The script trusts the app server's clock, so keep servers on NTP. Decide in advance what happens when Redis is down: failing open keeps the site up without protection, failing closed locks everyone out. Many teams fail open for general traffic and closed for OTP verification.
Choosing the Key
What you count by decides which attacks you stop:
| Key | Stops | Misses or harms |
|---|---|---|
| Per IP | One machine hammering you | Botnets and proxy pools; punishes users behind shared NAT, offices, mobile carriers |
| Per account | Guessing one user's password | Credential stuffing (one attempt per account); lets attackers lock victims out |
| Per IP + account | Repeated guessing of one account from one place | Distributed attacks |
| Per API key or token | A single client over-using an API | Nothing about who holds the key |
| Global, per endpoint | Sudden aggregate spikes in login failures | Cannot tell attacker from user; use as an alarm or circuit breaker |
Real deployments layer several. For IPv6, a single subscriber typically receives a whole /64, so count per /64, and unwrap IPv4-mapped addresses from dual-stack sockets. Sample inputs and keys:
import ipaddress
def ip_key(ip: str) -> str:
addr = ipaddress.ip_address(ip)
if addr.version == 6 and addr.ipv4_mapped: # ::ffff:203.0.113.9 from dual-stack sockets
addr = addr.ipv4_mapped
if addr.version == 6:
return str(ipaddress.ip_network(f"{addr}/64", strict=False))
return str(addr)
203.0.113.9 -> 203.0.113.9
2001:db8:1:2::1 -> 2001:db8:1:2::/64
2001:db8:1:2:aaaa:bbbb:cccc:dddd -> 2001:db8:1:2::/64
::ffff:203.0.113.9 -> 203.0.113.9
Without the mapped-address check, every IPv4 client would share the key ::/64 and one attacker would rate-limit your whole user base. And behind a proxy or CDN, take the client address from the forwarded header only when the request came from your proxy's own addresses; otherwise attackers set X-Forwarded-For to a fresh value on each request and never hit a limit.
Lockout Versus Throttling
A hard lockout ("account locked after 5 failures") looks strong but hands attackers a denial-of-service switch: fail five times against every username you can find and your customers are locked out. Better patterns:
- Progressive delay per account: after a few failures, require a growing wait between attempts, capped at minutes, not permanent.
- Step-up instead of stop: after N failures, require a CAPTCHA or an emailed link, so a human can still get in.
- Trust known devices: a signed cookie set after a previous successful login lets the real user bypass the per-account throttle while strangers stay slowed.
- Tell the user: an email saying "several failed sign-in attempts" turns them into a detector.
- Keep responses uniform: the same message and similar timing for wrong password, unknown user and throttled, so limits do not leak which accounts exist.
Credential Stuffing Needs Different Signals
Stuffing replays username and password pairs leaked from other sites, one attempt per account, spread across thousands of IPs. Per-IP and per-account limits each see almost nothing. What changes is the aggregate:
- The login success rate drops sharply, because most leaked pairs are wrong for your site.
- Failures for usernames that do not exist spike; real users rarely mistype their username repeatedly, but leaked lists contain many accounts you never had.
- Many distinct usernames per IP or per device fingerprint, and many IPs per username.
Alert on these global signals and respond on the login flow: challenges for new devices, blocking known-bad IP ranges, and rejecting passwords that appear in breach corpuses at signup and at login. Multi-factor authentication makes a correct stolen password insufficient, which is the strongest mitigation of all. Recognising the automation itself is covered in bot management for site owners.
nginx limit_req, Precisely
The DDoS lesson shows a basic limit_req setup. For login endpoints, the details decide whether real users get hurt:
limit_req_zone $binary_remote_addr zone=login_ip:10m rate=10r/m;
limit_req_zone $server_name zone=login_all:1m rate=50r/s;
location = /login {
limit_req zone=login_ip burst=5 delay=2;
limit_req zone=login_all burst=100;
limit_req_status 429;
limit_req_log_level warn;
# limit_req_dry_run on; # log what would be rejected, reject nothing
proxy_pass http://app;
}
nginx implements a leaky bucket with millisecond precision: rate=10r/m means one request every six seconds, so without burst a second click within six seconds is rejected. burst=5 lets up to five excess requests queue; delay=2 serves the first two excess immediately and paces the rest; requests beyond the burst are rejected with the status you set (the default is 503, which monitoring will misread as an outage). Multiple limit_req lines all apply, so the second zone, keyed on a constant, acts as a global ceiling. Enable limit_req_dry_run first and read the error log for a week before enforcing.
Edge limits are coarse because nginx does not know which account a request targets or whether the password was correct. Keep per-account and failure-based limits in the application, and treat the web server limits as the outer fence.
Practice
On your own staging app, write a script that sends 20 login attempts in three seconds across a minute boundary, and check whether your limiter lets more than the policy through. Then send attempts for 50 non-existent usernames from one IP and confirm that something, an alert or a log line, notices.