Availability is one third of the CIA triad, and the easiest third to attack: no vulnerability is required, only more requests than you can serve, or a few unusually expensive ones. Rate limiting caps how much any one client can ask for; resilience engineering makes sure that when the cap is exceeded, or when traffic is simply larger than expected, the service degrades gracefully instead of falling over.
In this lesson you will learn which limiting algorithm to choose, how to apply tiered limits in Express with a shared store, how to bound every expensive resource, and which architectural layers absorb attacks your application never sees.
The first three are application-layer problems; rate limits plus resource bounds address them directly.
| Algorithm | How it works | Trade-off |
|---|---|---|
| Fixed window | Count requests per calendar minute | Simple; allows a double burst at the window boundary |
| Sliding window | Weighted count across the current and previous window | Smooth and accurate; the default in express-rate-limit |
| Token bucket | Tokens refill at a steady rate; each request spends one | Allows short bursts while enforcing an average rate |
For most APIs a sliding window per client is enough. Use a token bucket where legitimate clients are bursty, such as a mobile app syncing after being offline.
The sample code applies two tiers backed by Redis, which matters as soon as you run more than one instance: an in-memory store would give every instance its own counter and multiply the effective limit. The response for an exceeded limit is 429 Too Many Requests with RateLimit-* headers and a Retry-After value so well-behaved clients back off.
Key the limit by the most specific identity available: user ID, then API key, then IP as the fallback. trust proxy must be set correctly or every request appears to come from your load balancer.
Rate limiting counts requests; it does not know that one request is a thousand times more expensive than another. Set explicit ceilings:
app.use(express.json({ limit: "100kb" })); // request bodies
const pageSize = Math.min(Number(req.query.limit) || 20, 100); // pagination caps
const upload = multer({ limits: { fileSize: 5 * 1024 * 1024 } }); // uploadsserver.requestTimeout = 30_000; server.headersTimeout = 10_000; // slow-client protectionAdd database statement timeouts, GraphQL depth and complexity limits, worker pools for CPU-heavy work, and length limits before any regular expression. Every unbounded input is a denial-of-service vector waiting to be noticed.
# Layers that absorb load before it reaches Node.js
cdn_or_waf: caches static and cacheable API responses; blocks known bad traffic; absorbs volumetric floods
load_balancer: connection limits, health checks, removes unhealthy instances
autoscaling: adds instances under load, with a budget cap so cost cannot spiral
queue: expensive jobs (exports, emails) run asynchronously with backpressureInside the application, use circuit breakers so a slow dependency fails fast, prefer degraded responses (cached data, a reduced feature) over errors, and keep a priority lane for authenticated users when capacity is scarce. Load-test before an attacker does; knowing your real capacity tells you where the limits belong.
Why must the rate limiter use a shared store such as Redis when the API runs on several instances?
Retry-After.Next lesson: Security Testing: SAST, DAST and OWASP ZAP Basics — finding the flaws in this course automatically, before release.