Rate Limits, 429 and Retries
Try it: Rate Limits, 429 and Retries
How an API enforces a per-key rate limit (token bucket or fixed window), answers 429 Too Many Requests with Retry-After, and how a client's retry policy — exponential backoff with a cap, honouring Retry-After, and a timeout for slow responses — decides when each request finally succeeds or gives up.
How it works
- The server reads the API key and looks it up; an unknown key gets 401 and uses no quota.
- Each known key has its own limiter: a token bucket (one token per request, refilled at a fixed rate up to the capacity) or a fixed window (at most L requests per window).
- With no quota left the server answers 429 with Retry-After: the whole seconds until quota returns.
- The client waits min(cap, base × multiplier^(attempt-1)) — or longer, if Retry-After says so — and retries 429s and timeouts up to the maximum number of attempts; a 401 is never retried.
- A response slower than the client's timeout counts as a failed attempt even though the server did the work.
Default run (15 steps): 8 requests are scheduled. The server gives each API key its own token bucket (capacity 3, +1 token every 1.00 s); the client retries 429s and timeouts up to 4 attempts in total. … All requests finished by 6.70 s. 7/8 succeeded · 0 gave up · 1 unauthorized · 13 attempts (3× 429, 2× timeout).
Simplified: Deterministic model on a simulated millisecond clock: fixed server latency (longer inside one slow period), no jitter in the backoff, no network delay for the request itself, and two API keys plus one invalid key. Real APIs vary in how they compute Retry-After and which errors they retry.
Loading the simulation…