Graceful Shutdown & Zero-Dropped Requests: Mastering SIGTERM in Docker, Nginx & Gunicorn

Rolling deployments often silently drop in-flight HTTP requests and abort active transactions during container restarts. Master POSIX signal forwarding, Nginx upstream retries, and Gunicorn graceful timeouts for true zero-downtime releases.

The Anatomy of a Dropped Request During Deployments

Modern Continuous Delivery (CD) pipelines deploy code updates dozens of times a week. Teams celebrate their automated GitHub Actions workflows, yet users occasionally experience mysterious 502 Bad Gateway errors, dropped WebSocket streams, or aborted Stripe webhook callbacks right in the middle of a release.

The root cause is almost always an improperly handled SIGTERM signal lifecycle. When Docker, Kubernetes, or systemd stops an existing application container to spin up a new revision, it sends a SIGTERM signal. If your entrypoint script, process manager, or web server is not configured to trap, propagate, and gracefully drain in-flight connections, the container abruptly terminates active threads, leaving Nginx holding a severed TCP socket.

1. The Signal Propagation Chain

A typical production web request travels through multiple layers before executing Python application code. When a shutdown occurs, every link in the chain must coordinate in reverse order:

Layer Signal Handling Responsibility Common Failure Mode
Reverse Proxy (Nginx) Stop routing new requests to stopping container; retry backup upstreams Continues dispatching new connections to terminating sockets (502 error)
Container PID 1 (Docker) Trap SIGTERM and forward immediately to child process (Gunicorn) Shell wrapper eats signal; Docker waits 10s then sends brutal SIGKILL
WSGI Master (Gunicorn) Stop listening on socket; allow active workers --graceful-timeout to finish Worker pool killed instantly mid-transaction
Application Thread Complete active database write and commit transaction cleanly Database transaction aborted; state corrupted or duplicate retry triggered

2. The Docker Entrypoint Trap Anti-Pattern

A prevalent mistake in production Dockerfiles is using shell form instead of exec form for the ENTRYPOINT or CMD directive:

# ❌ WRONG: Shell form spawns /bin/sh -c as PID 1, which ignores SIGTERM!
CMD python manage.py runserver

# ❌ WRONG: Shell wrapper script without exec
ENTRYPOINT ["/entrypoint.sh"]
# Inside entrypoint.sh:
# gunicorn devmanue_project.wsgi:application (spawns child; shell ignores SIGTERM)

Because standard Unix shells do not forward signals to child processes by default, Docker's SIGTERM hits /bin/sh, which does nothing. Docker waits for its default 10-second timeout, gives up, and issues a SIGKILL (Signal 9), instantly obliterating all in-flight requests. Always use exec in entrypoint scripts to replace the shell process with the actual server executable:

#!/bin/sh
# entrypoint.sh - Production-grade signal forwarding entrypoint
set -e

# Run database migrations or asset checks if required
python manage.py collectstatic --noinput

# 'exec' replaces the PID 1 shell process with Gunicorn
exec gunicorn devmanue_project.wsgi:application     --bind 0.0.0.0:8000     --workers 4     --timeout 60     --graceful-timeout 30     --keep-alive 5

3. Tuning Gunicorn Graceful Timeout

In Gunicorn, the --graceful-timeout parameter defines how many seconds active worker processes are given to finish serving current requests after receiving a SIGTERM. Once the signal arrives, the master process stops accepting new connections on the listening socket, sends SIGTERM to its workers, and waits:

# gunicorn.conf.py
import multiprocessing

bind = "unix:/run/devmanue/devmanue.sock"
workers = multiprocessing.cpu_count() * 2 + 1
timeout = 60            # Max duration for an individual request
graceful_timeout = 30   # Time allotted to drain in-flight requests on shutdown
keepalive = 5

def on_starting(server):
    server.log.info("Gunicorn master booting with graceful signal hooks.")

def worker_abort(worker):
    worker.log.error("Worker received SIGABRT or exceeded timeout limit.")

4. Nginx Upstream Retry Configuration

Even with graceful draining, there is a sub-millisecond race condition between when Gunicorn closes its listening socket and when Nginx detects the backend is unavailable. To achieve 100.0% zero-dropped releases, configure Nginx to automatically retry failed requests on a backup upstream:

# /etc/nginx/sites-available/devmanue.conf
upstream devmanue_app {
    server unix:/run/devmanue/devmanue.sock max_fails=3 fail_timeout=10s;
    server unix:/run/devmanue/devmanue_standby.sock backup;
    keepalive 32;
}

server {
    listen 443 ssl http2;
    server_name devmanue.com;

    location / {
        proxy_pass http://devmanue_app;
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;

        # Transparent Upstream Failover & Zero-Drop Retry Rules
        proxy_next_upstream error timeout invalid_header http_502 http_503;
        proxy_next_upstream_tries 3;
        proxy_next_upstream_timeout 10s;

        proxy_connect_timeout 5s;
        proxy_read_timeout 60s;
        proxy_send_timeout 60s;
    }
}
Architectural Continuity & Deep Dives

For related production architectures and system implementations, explore these companion guides:

Production Engineering Takeaways

  • Always replace shell with exec: Ensure Gunicorn runs as PID 1 inside the container to receive signals without shell interception.
  • Align timeout budgets: Set Docker stop timeout (docker stop -t 35) higher than Gunicorn's graceful-timeout (30s) so Docker never kills a worker that is legitimately finishing a long request.
  • Deploy upstream failover in Nginx: proxy_next_upstream http_502 seamlessly routes transient disconnects to standby sockets with zero user perception. See our full blueprint on Zero-Downtime Django Deployments.
All Insights
Chat on WhatsApp