An API without limits is one buggy client loop away from an outage. Rate limiting caps how many requests a caller may make in a period; quotas cap total usage over a billing cycle. In this lesson you will learn the common limiting algorithms, how to communicate limits with status codes and headers, and how to add them to an Express API.
Limits are keyed by API key or user id for authenticated traffic and by IP address for anonymous traffic.
| Algorithm | How it works | Trade-off | |---|---|---| | Fixed window | Count requests per calendar minute; reset at :00 | Simple, but allows 2x bursts across the boundary | | Sliding window | Weight the previous window's count by its overlap | Smooths the boundary problem at small extra cost | | Token bucket | Tokens refill at a steady rate; each request spends one | Allows bursts up to the bucket size, then a steady rate | | Leaky bucket | Requests queue and drain at a fixed rate | Smooth output; extra latency under load |
Token bucket is the usual choice for public APIs because it tolerates natural burstiness while enforcing an average rate.
When a client exceeds its limit, respond with 429 Too Many Requests and a Retry-After header. On every response, advertise the current state so clients can pace themselves. The IETF draft standard uses RateLimit and RateLimit-Policy; older APIs send X-RateLimit-* headers:
HTTP/1.1 429 Too Many Requests
Content-Type: application/json
Retry-After: 22
RateLimit-Policy: 100;w=60
RateLimit: limit=100, remaining=0, reset=22
{ "error": "rate_limited", "detail": "Too many requests, slow down." }RateLimit-Policy: 100;w=60 reads as "100 requests per 60-second window", and reset=22 means the window resets in 22 seconds.
The express-rate-limit package implements a fixed-window counter with all of the headers above:
const { rateLimit, ipKeyGenerator } = require("express-rate-limit");
const apiLimiter = rateLimit({
windowMs: 60 * 1000,
limit: 100,
standardHeaders: "draft-7",
legacyHeaders: false,
keyGenerator: (req) => req.auth?.sub ?? ipKeyGenerator(req.ip),
message: { error: "rate_limited", detail: "Too many requests, slow down." },
});
const loginLimiter = rateLimit({ windowMs: 15 * 60 * 1000, limit: 5 });
app.use("/api", apiLimiter);
app.post("/api/auth/login", loginLimiter, loginHandler);Two details matter in production. The default in-memory store counts per process, so behind a load balancer use a shared store such as rate-limit-redis to apply limits globally. And behind a proxy, set app.set("trust proxy", 1) so req.ip is the real client address.
A quota is a long-window limit tied to a plan: 10,000 requests per day, 1 GB of uploads per month. Count usage in a database or Redis keyed by account and period, and answer overruns with the same 429 but a distinct error type such as quota_exceeded. Expose usage so customers can monitor themselves before they hit the wall:
GET /account/usage
{ "plan": "team", "period": "2026-09", "requests": { "used": 8120, "limit": 10000 }, "resetsAt": "2026-10-01T00:00:00Z" }In larger systems both limits and quotas often move to an API gateway in front of the services, keeping the policy in one place.
Retry-After. Without it clients guess, usually too short, and hammer the endpoint so the window never clears.Which status code tells a client it has exceeded its rate limit?
429, Retry-After and RateLimit headers so clients can pace themselves.express-rate-limit needs a shared store and trust proxy behind a load balancer.Next lesson: File Uploads and Downloads — accept multipart uploads safely and stream files back efficiently.