The Anatomy of a Dropped Request During Deployments
Modern Continuous Delivery (CD) pipelines deploy code updates dozens of times a week. Teams celebrate their automated GitHub Actions workflows, yet users occasionally experience mysterious 502 Bad Gateway errors, dropped WebSocket streams, or aborted Stripe webhook callbacks right in the middle of a release.
The root cause is almost always an improperly handled SIGTERM signal lifecycle. When Docker, Kubernetes, or systemd stops an existing application container to spin up a new revision, it sends a SIGTERM signal. If your entrypoint script, process manager, or web server is not configured to trap, propagate, and gracefully drain in-flight connections, the container abruptly terminates active threads, leaving Nginx holding a severed TCP socket.
1. The Signal Propagation Chain
A typical production web request travels through multiple layers before executing Python application code. When a shutdown occurs, every link in the chain must coordinate in reverse order:
| Layer | Signal Handling Responsibility | Common Failure Mode |
|---|---|---|
| Reverse Proxy (Nginx) | Stop routing new requests to stopping container; retry backup upstreams | Continues dispatching new connections to terminating sockets (502 error) |
| Container PID 1 (Docker) | Trap SIGTERM and forward immediately to child process (Gunicorn) |
Shell wrapper eats signal; Docker waits 10s then sends brutal SIGKILL |
| WSGI Master (Gunicorn) | Stop listening on socket; allow active workers --graceful-timeout to finish |
Worker pool killed instantly mid-transaction |
| Application Thread | Complete active database write and commit transaction cleanly | Database transaction aborted; state corrupted or duplicate retry triggered |
2. The Docker Entrypoint Trap Anti-Pattern
A prevalent mistake in production Dockerfiles is using shell form instead of exec form for the ENTRYPOINT or CMD directive:
# ❌ WRONG: Shell form spawns /bin/sh -c as PID 1, which ignores SIGTERM!
CMD python manage.py runserver
# ❌ WRONG: Shell wrapper script without exec
ENTRYPOINT ["/entrypoint.sh"]
# Inside entrypoint.sh:
# gunicorn devmanue_project.wsgi:application (spawns child; shell ignores SIGTERM)
Because standard Unix shells do not forward signals to child processes by default, Docker's SIGTERM hits /bin/sh, which does nothing. Docker waits for its default 10-second timeout, gives up, and issues a SIGKILL (Signal 9), instantly obliterating all in-flight requests. Always use exec in entrypoint scripts to replace the shell process with the actual server executable:
#!/bin/sh
# entrypoint.sh - Production-grade signal forwarding entrypoint
set -e
# Run database migrations or asset checks if required
python manage.py collectstatic --noinput
# 'exec' replaces the PID 1 shell process with Gunicorn
exec gunicorn devmanue_project.wsgi:application --bind 0.0.0.0:8000 --workers 4 --timeout 60 --graceful-timeout 30 --keep-alive 5
3. Tuning Gunicorn Graceful Timeout
In Gunicorn, the --graceful-timeout parameter defines how many seconds active worker processes are given to finish serving current requests after receiving a SIGTERM. Once the signal arrives, the master process stops accepting new connections on the listening socket, sends SIGTERM to its workers, and waits:
# gunicorn.conf.py
import multiprocessing
bind = "unix:/run/devmanue/devmanue.sock"
workers = multiprocessing.cpu_count() * 2 + 1
timeout = 60 # Max duration for an individual request
graceful_timeout = 30 # Time allotted to drain in-flight requests on shutdown
keepalive = 5
def on_starting(server):
server.log.info("Gunicorn master booting with graceful signal hooks.")
def worker_abort(worker):
worker.log.error("Worker received SIGABRT or exceeded timeout limit.")
4. Nginx Upstream Retry Configuration
Even with graceful draining, there is a sub-millisecond race condition between when Gunicorn closes its listening socket and when Nginx detects the backend is unavailable. To achieve 100.0% zero-dropped releases, configure Nginx to automatically retry failed requests on a backup upstream:
# /etc/nginx/sites-available/devmanue.conf
upstream devmanue_app {
server unix:/run/devmanue/devmanue.sock max_fails=3 fail_timeout=10s;
server unix:/run/devmanue/devmanue_standby.sock backup;
keepalive 32;
}
server {
listen 443 ssl http2;
server_name devmanue.com;
location / {
proxy_pass http://devmanue_app;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
# Transparent Upstream Failover & Zero-Drop Retry Rules
proxy_next_upstream error timeout invalid_header http_502 http_503;
proxy_next_upstream_tries 3;
proxy_next_upstream_timeout 10s;
proxy_connect_timeout 5s;
proxy_read_timeout 60s;
proxy_send_timeout 60s;
}
}
For related production architectures and system implementations, explore these companion guides:
- Zero-Downtime Django Deployments & Atomic Releases — Coordinate rolling reloads with Gunicorn master processes for seamless deployment transitions.
- Taming Redis & Celery Worker Pools in Production — Enforce clean worker task completion and warm shutdowns on Celery task queues during releases.
- Linux cgroups v2 & Memory Pressure Stalling — Prevent abrupt OOM supervisor termination by sizing container memory headroom defensively.
Production Engineering Takeaways
- Always replace shell with
exec: Ensure Gunicorn runs as PID 1 inside the container to receive signals without shell interception. - Align timeout budgets: Set Docker stop timeout (
docker stop -t 35) higher than Gunicorn'sgraceful-timeout(30s) so Docker never kills a worker that is legitimately finishing a long request. - Deploy upstream failover in Nginx:
proxy_next_upstream http_502seamlessly routes transient disconnects to standby sockets with zero user perception. See our full blueprint on Zero-Downtime Django Deployments.