Rate Limiter

A rate limiter caps how many requests a client may make in a period, answering the excess with 429 Too Many Requests — to protect a service from overload and abuse.

Token bucket

Bucket
5.0 tokens · +1/s
0s5s10s15s20s
  • 200 OK
  • 429 Too Many Requests

A bucket holds up to 5 tokens and refills 1 per second. Each request takes one; no token → 429 Too Many Requests.

Step 1 / 22
Allowed
0
429 rejected
0

Step by step

The example above, written out — the same steps the animation plays.

  1. 1Fixed window. Count requests in fixed 10-second windows; at most 5 per window.
  2. 2t = 0s → 200 OK. 0 of 5 used in this window: allowed.
  3. 3t = 0s → 429. 5 of 5 used in this window: rejected.
  4. 4t = 3s → 429. 5 of 5 used in this window: rejected.
  5. 5t = 4s → 429. 5 of 5 used in this window: rejected.
  6. 6t = 4.5s → 429. 5 of 5 used in this window: rejected.
  7. 7t = 19.5s → 200 OK. 0 of 5 used in this window: allowed.
  8. 8t = 19.6s → 200 OK. 1 of 5 used in this window: allowed.
  9. 9t = 19.7s → 200 OK. 2 of 5 used in this window: allowed.
  10. 10t = 19.8s → 200 OK. 3 of 5 used in this window: allowed.
  11. 11t = 19.9s → 200 OK. 4 of 5 used in this window: allowed.
  12. 12t = 20s → 200 OK. 0 of 5 used in this window: allowed.
  13. 13t = 20.1s → 200 OK. 1 of 5 used in this window: allowed.
  14. 14t = 20.2s → 200 OK. 2 of 5 used in this window: allowed.

…and 3 more steps — press Play above to watch them all.

What's happening?

  1. Token bucket: a bucket holds up to N tokens and refills at a steady rate; each request spends one, no token means 429.
  2. Fixed window: count requests per window (say 10 seconds) and reset the count when the window ends.
  3. A fixed window lets a client send a full window's worth at the end of one window and again at the start of the next — double the limit in a moment.

Complexity

Time
O(1) per request
Space
O(1) per client

Where you'll meet it

Public APIs (GitHub, Stripe), login and OTP endpoints against brute force, and expensive AI calls per user.

Common mistake

Rate-limiting in each server's memory behind a load balancer: every server allows the full limit. Keep the counters in a shared store such as Redis.

FAQ

Token bucket or fixed window?

Token bucket allows short bursts but smooths sustained traffic; fixed window is simplest but has the boundary burst. Sliding windows fix that at a small cost.

What should a 429 include?

A Retry-After header telling the client when to try again.

Per user, per IP or per key?

Whatever you are protecting: per API key or user for fair use, per IP for anonymous endpoints like login.