How I Stopped Dropping Requests on Deploy With Graceful Shutdown in Node.js
Handling SIGTERM, draining keep-alive connections and closing Prisma in the right order so Docker and Kubernetes deploys stop causing 502s.
Jason
6 min read·
For months, every deploy of one of my Node.js APIs produced a small burst of 502s. Nothing dramatic, a handful of failed requests per rollout, but enough that the error-rate graph had a little spike every time I shipped. The code was fine. The problem was how the process died.
When a container platform replaces your app, it doesn't wait politely for requests to finish. It sends a signal, waits a bit, and then kills the process. If your server doesn't react to that signal, anything in flight gets cut off mid-response. In this post I'll walk through how I fixed it: handling SIGTERM properly, draining HTTP connections, closing the Prisma connection pool, and making sure Docker actually delivers the signal to Node in the first place.
What actually happens when a container stops
Whether it's docker stop, a Kubernetes rolling update, or most PaaS deploys, the sequence is roughly the same:
- The platform sends SIGTERM to the main process in the container (PID 1).
- It waits for a grace period. For
docker stopthe default is 10 seconds; for a Kubernetes pod,terminationGracePeriodSecondsdefaults to 30. - If the process is still alive, it sends SIGKILL, which cannot be caught. Open sockets are dropped instantly.
So the contract is simple: when you get SIGTERM, stop taking new work, finish what you have, clean up, and exit before the deadline. A default Node.js HTTP server does none of that by itself.
Bug #1: my process never received SIGTERM
The first thing I discovered was embarrassing. My shutdown handler wasn't even running. The Dockerfile ended with:
CMD npm start
That's the shell form of CMD. Docker runs it as /bin/sh -c "npm start", so PID 1 is a shell, which starts npm, which starts Node. The signal goes to PID 1, and depending on the shell and npm version it may never reach your Node process. The container just sits there until the grace period expires and gets SIGKILLed. That's why my deploys always seemed to take exactly ten seconds per container.
There's a second subtlety: the Linux kernel treats PID 1 specially. Signals sent to PID 1 are ignored unless the process has explicitly installed a handler for them. Node running as PID 1 with no SIGTERM listener won't exit on SIGTERM the way it would in your terminal.
The fix is to run Node directly in exec form, and install your own handler:
FROM node:22-slim
WORKDIR /app
COPY package*.json ./
RUN npm ci --omit=dev
COPY dist ./dist
# Exec form: node is PID 1 and receives SIGTERM directly.
CMD ["node", "dist/server.js"]
# Avoid these, they put a shell or npm between Docker and your process:
# CMD node dist/server.js
# CMD ["npm", "start"]
If your app spawns child processes or you can't control the entrypoint, run the container with docker run --init (or add tini) so a tiny init process sits at PID 1, forwards signals, and reaps zombies.
Bug #2: server.close() isn't enough on its own
Once the signal arrived, my first handler looked like what most tutorials show:
process.on("SIGTERM", () => {
server.close(() => process.exit(0));
});
server.close() stops the server from accepting new connections and calls the callback once every existing connection has ended. The catch is HTTP keep-alive. Browsers, load balancers and HTTP clients reuse connections, and an idle keep-alive socket counts as an open connection. Depending on your Node version and keep-alive timeout, close() can sit waiting on sockets that aren't doing anything, right up until SIGKILL arrives.
Node's http.Server has two methods (added in the Node 18 line) that solve this:
server.closeIdleConnections()closes connections that aren't currently handling a request. In-flight requests are left alone.server.closeAllConnections()forcibly closes everything, including active requests. This is your last resort when the drain takes too long.
Whether close() also closes idle connections for you has changed across Node releases, so I don't rely on it. I call closeIdleConnections() explicitly; on versions where it's redundant, it's harmless.
The full shutdown sequence
Here's the handler I ended up with. It's TypeScript with a plain node:http server wrapping an Express-style app, but the same idea works with Fastify (fastify.close()) or any framework that exposes the underlying server.
// server.ts
import http from "node:http";
import { PrismaClient } from "@prisma/client";
import { app } from "./app";
const prisma = new PrismaClient();
const server = http.createServer(app);
let shuttingDown = false;
server.listen(Number(process.env.PORT ?? 3000), () => {
console.log("listening");
});
async function shutdown(signal: string) {
if (shuttingDown) return; // a second SIGTERM shouldn't restart the sequence
shuttingDown = true;
console.log(`${signal} received, draining connections...`);
// Hard deadline: stay safely under the orchestrator's grace period.
const forceExit = setTimeout(() => {
console.error("Drain timed out, forcing connections closed");
server.closeAllConnections();
process.exit(1);
}, 8_000);
forceExit.unref();
// 1. Stop accepting new connections; callback fires when all are gone.
const closed = new Promise<void>((resolve, reject) =>
server.close((err) => (err ? reject(err) : resolve()))
);
// 2. Kick idle keep-alive sockets so they don't hold close() open.
server.closeIdleConnections();
try {
await closed; // in-flight requests finish here
await prisma.$disconnect(); // 3. release the DB pool last
console.log("Clean shutdown");
process.exit(0);
} catch (err) {
console.error("Error during shutdown", err);
process.exit(1);
}
}
process.on("SIGTERM", () => void shutdown("SIGTERM"));
process.on("SIGINT", () => void shutdown("SIGINT"));
The order matters:
- Stop accepting connections with
server.close(). - Close idle sockets so keep-alive connections don't stall the drain.
- Wait for in-flight requests to finish, because they may still need the database.
- Disconnect Prisma only after the server is fully closed. Disconnecting first would make every still-running request fail with a connection error, which is exactly what I was trying to avoid.
- Exit explicitly. Open handles like timers, queue consumers or Redis clients can keep the event loop alive forever, so I don't count on Node exiting by itself.
The setTimeout is the safety net. I set it a couple of seconds below the platform's grace period (8 seconds against Docker's 10) so that if something hangs, my code decides how to fail and logs it, instead of a silent SIGKILL. Calling .unref() means the timer itself won't keep the process alive if everything else finishes early.
Bug #3: the load balancer kept sending traffic
With the handler in place, the 502s dropped but didn't disappear. The remaining ones came from a race: in Kubernetes, removing a pod from the Service endpoints and sending SIGTERM happen in parallel. For a brief window, kube-proxy and ingress controllers may still route new connections to a pod that has already stopped listening.
Two changes closed that gap. First, a readiness endpoint that starts failing as soon as shutdown begins, plus a Connection: close header so keep-alive clients reconnect somewhere healthy:
// app.ts (Express-style handlers)
import { isShuttingDown } from "./lifecycle";
app.get("/healthz/ready", (_req, res) => {
if (isShuttingDown()) {
res.status(503).json({ ready: false });
return;
}
res.json({ ready: true });
});
app.use((req, res, next) => {
if (isShuttingDown()) {
// Tell keep-alive clients to open a new connection elsewhere.
res.setHeader("Connection", "close");
}
next();
});
Here lifecycle.ts is just a tiny module that exports the shuttingDown flag the handler above sets. Second, a short preStop hook. Kubernetes runs it before sending SIGTERM, which gives the endpoint removal time to propagate while the app keeps serving normally:
spec:
terminationGracePeriodSeconds: 30
containers:
- name: api
lifecycle:
preStop:
exec:
command: ["sleep", "5"]
readinessProbe:
httpGet:
path: /healthz/ready
port: 3000
periodSeconds: 2
Keep the arithmetic in mind: the preStop sleep counts against terminationGracePeriodSeconds. A 5-second sleep plus my 8-second drain deadline fits comfortably inside 30. Note that the sleep binary must exist in your image; very minimal or distroless images may not have it, in which case recent Kubernetes versions offer a native sleep action for preStop.
Testing it locally
You don't need a cluster to check this. I start the container, fire a slow request at it, and stop the container while that request is running:
docker run -d --name api -p 3000:3000 my-api
curl -s localhost:3000/slow-endpoint & # takes ~3s
sleep 1 && docker stop api
docker logs api
If shutdown works, the curl completes with a normal response, the logs show "SIGTERM received" followed by "Clean shutdown", and docker stop returns in about two seconds instead of ten. If docker stop always takes the full ten seconds, your signal isn't reaching Node; go back and check the CMD.
Things that still need care
- Long-lived connections. WebSockets and Server-Sent Events won't end on their own. Track them and close them during shutdown, ideally with a message telling the client to reconnect.
- Background work. Queue consumers should stop pulling new jobs on
SIGTERMand finish (or release) the current one. Idempotent job handlers make this much less scary. - Serverless and managed Next.js hosting. If the platform manages the process lifecycle for you, much of this is handled by the platform; it matters most when you run your own Node server in a container, including a self-hosted Next.js standalone build.
Key takeaways
- Platforms send SIGTERM, wait a grace period (10s for
docker stop, 30s default in Kubernetes), then SIGKILL. Your app must finish inside that window. - Use the exec form of
CMD(["node", "server.js"]), notnpm startor shell form, or use--init/tini, so the signal actually reaches Node. - Call
server.close()andserver.closeIdleConnections(), with a timedcloseAllConnections()fallback below the grace period. - Disconnect Prisma and other clients after in-flight requests finish, then exit explicitly.
- In Kubernetes, fail readiness on shutdown and add a short
preStopsleep to cover the endpoint-removal race.
Written by Jason
Published October 7, 2026 · Updated Oct 9, 2026