Rate Limiting and Quotas

Intermediate
12 min

Rate Limiting and Quotas

An API without limits is one buggy client loop away from an outage. Rate limiting caps how many requests a caller may make in a period; quotas cap total usage over a billing cycle. In this lesson you will learn the common limiting algorithms, how to communicate limits with status codes and headers, and how to add them to an Express API.

Why Limit at All

  • Stability: one misbehaving client cannot exhaust database connections for everyone.
  • Security: brute-force logins and credential stuffing become impractical.
  • Fairness and cost: paying tiers get the capacity they paid for; free tiers stay affordable.

Limits are keyed by API key or user id for authenticated traffic and by IP address for anonymous traffic.

Limiting Algorithms

| Algorithm | How it works | Trade-off | |---|---|---| | Fixed window | Count requests per calendar minute; reset at :00 | Simple, but allows 2x bursts across the boundary | | Sliding window | Weight the previous window's count by its overlap | Smooths the boundary problem at small extra cost | | Token bucket | Tokens refill at a steady rate; each request spends one | Allows bursts up to the bucket size, then a steady rate | | Leaky bucket | Requests queue and drain at a fixed rate | Smooth output; extra latency under load |

Token bucket is the usual choice for public APIs because it tolerates natural burstiness while enforcing an average rate.

Communicating Limits to Clients

When a client exceeds its limit, respond with 429 Too Many Requests and a Retry-After header. On every response, advertise the current state so clients can pace themselves. The IETF draft standard uses RateLimit and RateLimit-Policy; older APIs send X-RateLimit-* headers:

http
HTTP/1.1 429 Too Many Requests Content-Type: application/json Retry-After: 22 RateLimit-Policy: 100;w=60 RateLimit: limit=100, remaining=0, reset=22 { "error": "rate_limited", "detail": "Too many requests, slow down." }

RateLimit-Policy: 100;w=60 reads as "100 requests per 60-second window", and reset=22 means the window resets in 22 seconds.

Rate Limiting in Express

The express-rate-limit package implements a fixed-window counter with all of the headers above:

javascript
const { rateLimit, ipKeyGenerator } = require("express-rate-limit"); const apiLimiter = rateLimit({ windowMs: 60 * 1000, limit: 100, standardHeaders: "draft-7", legacyHeaders: false, keyGenerator: (req) => req.auth?.sub ?? ipKeyGenerator(req.ip), message: { error: "rate_limited", detail: "Too many requests, slow down." }, }); const loginLimiter = rateLimit({ windowMs: 15 * 60 * 1000, limit: 5 }); app.use("/api", apiLimiter); app.post("/api/auth/login", loginLimiter, loginHandler);

Two details matter in production. The default in-memory store counts per process, so behind a load balancer use a shared store such as rate-limit-redis to apply limits globally. And behind a proxy, set app.set("trust proxy", 1) so req.ip is the real client address.

Quotas

A quota is a long-window limit tied to a plan: 10,000 requests per day, 1 GB of uploads per month. Count usage in a database or Redis keyed by account and period, and answer overruns with the same 429 but a distinct error type such as quota_exceeded. Expose usage so customers can monitor themselves before they hit the wall:

json
GET /account/usage { "plan": "team", "period": "2026-09", "requests": { "used": 8120, "limit": 10000 }, "resetsAt": "2026-10-01T00:00:00Z" }

In larger systems both limits and quotas often move to an API gateway in front of the services, keeping the policy in one place.

Common Mistakes

  • Limiting only by IP. Corporate networks and mobile carriers put thousands of users behind one address.
  • Applying one global limit. Login and password-reset endpoints need far stricter limits than reads.
  • Forgetting Retry-After. Without it clients guess, usually too short, and hammer the endpoint so the window never clears.
Quick Quiz
Question 1 of 2

Which status code tells a client it has exceeded its rate limit?

Key Takeaways

  • Rate limits protect stability, deter brute force and keep capacity fair; quotas cap usage per billing period.
  • Token bucket allows bursts with a bounded average; fixed windows are simplest but leak at boundaries.
  • Answer with 429, Retry-After and RateLimit headers so clients can pace themselves.
  • express-rate-limit needs a shared store and trust proxy behind a load balancer.
  • Key limits by user or API key, and give sensitive endpoints stricter limits.

Next lesson: File Uploads and Downloads — accept multipart uploads safely and stream files back efficiently.

Rate Limiting and Quotas - REST APIs | CodeYourCraft | CodeYourCraft