A single Node process runs JavaScript on one thread. That is fast for I/O-bound work, but one slow synchronous call stalls every request, and one process uses one CPU core while the other seven idle. After this lesson you will be able to measure an endpoint's throughput, shrink responses with compression, keep the event loop free, run one worker per core with the cluster module or PM2, and shut down without dropping requests.
autocannon fires concurrent requests and reports latency and requests per second:
npx autocannon -c 50 -d 10 http://localhost:3000/api/todosRun it before and after every change; a "performance improvement" that does not move these numbers is not one. For CPU hot spots, start Node with --cpu-prof and open the generated profile in Chrome DevTools.
Text responses (JSON, HTML, CSS) shrink by 70 to 90 percent with gzip or Brotli. The compression middleware handles content negotiation with the Accept-Encoding header:
npm install compressionimport compression from "compression";
app.use(compression({ threshold: 1024 })); // skip bodies under 1 KBPlace it before routes and static files. Already-compressed formats (images, video, zip) are skipped automatically by content type. When Nginx or a CDN sits in front of the app, let it compress instead; it is faster and frees Node for application work.
The event loop processes one callback at a time. Anything synchronous and slow blocks every connection:
| Avoid in request handlers | Use instead |
| --- | --- |
| fs.readFileSync, child_process.execSync | node:fs/promises, async exec |
| Large JSON.parse / JSON.stringify of multi-megabyte bodies | Body size limits, streaming, pagination |
| Image resizing, PDF generation, hashing loops | worker_threads or a job queue such as BullMQ |
| bcrypt with a very high cost factor | Cost 10 to 12; it runs in the libuv thread pool but still takes time |
Database habits matter more than micro-optimisations: .lean() queries, indexes on filtered fields, Promise.all for independent calls and pagination from earlier lessons typically give the biggest wins.
The cluster module forks the process once per core; the primary accepts connections and distributes them, and each worker runs your full Express app:
// src/cluster.js
import cluster from "node:cluster";
import { availableParallelism } from "node:os";
if (cluster.isPrimary) {
for (let i = 0; i < availableParallelism(); i++) cluster.fork();
cluster.on("exit", (worker, code) => {
console.log(`worker ${worker.process.pid} died (${code}); forking a new one`);
cluster.fork();
});
} else {
await import("./server.js");
}Workers share nothing in memory. Anything that must be consistent across them (sessions, rate-limit counters, Socket.IO rooms, in-process caches) has to live in Redis, which is why earlier lessons kept pointing there.
PM2 wraps the same cluster mechanism in a process manager with logs, restarts and zero-downtime reloads, and it is the standard way to run Express on a virtual machine.
npm install -g pm2
pm2 start src/server.js -i max --name api # one worker per core
pm2 reload api # rolling restart, no dropped requests
pm2 logs api
pm2 monit
pm2 startup && pm2 save # relaunch on server rebootAn ecosystem.config.cjs file keeps these settings in the repository (the .cjs extension is required in an ES-module project):
module.exports = {
apps: [{
name: "api",
script: "src/server.js",
instances: "max",
exec_mode: "cluster",
max_memory_restart: "400M",
env_production: { NODE_ENV: "production", PORT: 3000 },
}],
};Start it with pm2 start ecosystem.config.cjs --env production. In containers, skip PM2 and run one process per container, letting the orchestrator scale replicas.
pm2 reload, Docker and Kubernetes all stop a process by sending SIGTERM. Finish in-flight requests before exiting:
const server = app.listen(config.port);
process.on("SIGTERM", () => {
server.close(async () => { // stop accepting, wait for open requests
await mongoose.disconnect();
process.exit(0);
});
setTimeout(() => process.exit(1), 10_000).unref(); // hard deadline
});Behind a load balancer, also raise server.keepAliveTimeout above the balancer's idle timeout (for example 65_000 ms for a 60-second balancer) to avoid sporadic 502 errors from connections closed mid-reuse.
Why does a single Node process not use all CPU cores?
autocannon before and after every optimisation.compression() shrinks text responses; prefer compressing at the proxy or CDN when one exists.node:cluster and PM2 cluster mode run one worker per core; move shared state to Redis first.SIGTERM with server.close() and tune keepAliveTimeout behind load balancers.Next lesson: Deploying an Express App and Best Practices — take the application to a real server with environment configuration, process management and a production checklist.