The Silent Killer: Page Cache Reclamation vs. OOM Termination
In high-density production environments, container memory sizing is almost universally approached as a binary problem: either a container stays alive within its allocated threshold, or the Linux Out-Of-Memory (OOM) killer abruptly terminates the process with exit code 137. However, the most destructive container performance degradations occur long before the OOM killer ever fires.
Under Linux control groups (cgroups), when a container's memory consumption approaches its configured limit, the kernel's memory management subsystem begins aggressively reclaiming clean executable pages, file-backed page caches, and inactive buffers to stave off an OOM crisis. As the available page cache collapses, standard disk reads, dynamic library references, and shared memory lookups immediately suffer synchronous disk cache misses. The CPU spends upward of 80% of its execution cycles stuck in iowait, thrashing the kernel memory subsystem. To external monitoring dashboards tracking basic CPU and RAM utilization percentages, the container appears to be operating normally within limits—yet API response times spike from 25 milliseconds to over 8 seconds.
Deciphering Pressure Stall Information (PSI): /proc/pressure/memory
To expose this invisible thrashing, modern Linux kernels (5.x+) equipped with cgroups v2 introduce Pressure Stall Information (PSI). Rather than reporting static byte allocations, PSI measures the exact percentage of time tasks are blocked waiting for memory resources.
Inspecting /proc/pressure/memory (or a container's local memory.pressure cgroup file) yields two fundamental metrics:
some avg10=X.XX avg60=X.XX avg300=X.XX: Represents the percentage of wall-clock time during which at least one active process thread was completely stalled waiting for memory allocation, page reclaim, or swap-in. In healthy systems,some avg10should remain below2.00%. Values above15.00%signal severe throughput degradation.full avg10=X.XX avg60=X.XX avg300=X.XX: Represents the percentage of time during which all non-idle threads in the container were stalled simultaneously. Any non-zero value on thefullmetric indicates complete system livelock where application processing has essentially ground to a halt.
Configuring memory.high vs. memory.max in cgroups v2
Legacy cgroups v1 offered only a single hard ceiling (memory.limit_in_bytes). Crossing this ceiling triggered instantaneous process termination. In contrast, cgroups v2 provides a sophisticated two-tier memory control mechanism that enables graceful degradation:
# Systemd Service Unit: /etc/systemd/system/gunicorn-api.service.d/override.conf
[Service]
# Hard limit: Kernel OOM killer terminates tasks if breached
MemoryMax=4G
# Soft throttling ceiling: Kernel aggressively reclaims and throttles CPU allocations
MemoryHigh=3200M
# Minimum protected memory: Protected from global host page cache reclaim
MemoryMin=512M
# Proactively enforce cgroups v2 memory limits
CPUWeight=100
IOWeight=100
When a container crosses MemoryHigh, the kernel exerts proportional backpressure. It throttles the allocating processes during system calls, gently slowing down incoming work rather than crashing the process. This dynamic backpressure buys time for background garbage collectors, Python object deallocations, and connection pool cleanups to stabilize memory footprints.
Production Python PSI Monitor with epoll
Linux allows user-space daemons to register kernel event triggers on PSI files using epoll, providing sub-millisecond alerts when memory pressure spikes. Below is a lightweight Python watchdog daemon:
import select
import sys
import os
import logging
logging.basicConfig(level=logging.INFO, format="%(asctime)s [%(levelname)s] %(message)s")
logger = logging.getLogger("psi.monitor")
PRESSURE_PATH = "/proc/pressure/memory"
# Trigger alert if 'some' tasks stall for more than 150ms (150000us) within a 1-second window (1000000us)
TRIGGER_CONFIG = "some 150000 1000000"
def monitor_memory_stalls():
if not os.path.exists(PRESSURE_PATH):
logger.error("Kernel Pressure Stall Information (PSI) is not enabled on this host.")
sys.exit(1)
with open(PRESSURE_PATH, "w+") as fd:
fd.write(TRIGGER_CONFIG)
fd.flush()
epoll = select.epoll()
# Register for high-priority polling events
epoll.register(fd.fileno(), select.EPOLLPRI)
logger.info(f"Active kernel PSI watchdog listening on {PRESSURE_PATH}...")
while True:
events = epoll.poll(timeout=5.0)
for fileno, event in events:
if event & select.EPOLLPRI:
logger.warning(
"CRITICAL MEMORY PRESSURE: Kernel stall duration exceeded 150ms in 1s window! "
"Triggering automated memory shedder, clearing cache pools..."
)
execute_memory_shedding()
def execute_memory_shedding():
# Application-specific shedding: close idle DB connections, evict local caches
pass
if __name__ == "__main__":
monitor_memory_stalls()
For related production architectures and system implementations, explore these companion guides:
- Battle-Tested Production Dockerfile for Python & Django — Containerize application runtimes with precise memory boundaries and non-root UID execution.
- Memory Profiling in Production Python with Memray — Pinpoint native C-extension leaks and Python heap allocations before triggering cgroup OOM thresholds.
- Chaos Engineering on a Shoestring Budget — Simulate artificial kernel memory pressure and OOM conditions to verify system resilience.
Operational Takeaways & Sizing Rules
When sizing production Docker containers and systemd services, adhere to these three rules:
- Always configure
MemoryHighat 75–80% ofMemoryMax: This cushion allows the kernel to exert gradual CPU throttling and shed cache before invoking catastrophic OOM termination. - Graph PSI in Grafana: Ingest
memory.pressuremetrics alongside traditional RAM utilization to identify thrashing hours before it causes an incident. - Verify cgroups v2 mounting: Run
mount | grep cgroupand verify thatcgroup2 on /sys/fs/cgroupis active. Upgrading from cgroups v1 to v2 is the single highest-ROI infrastructure improvement you can deploy to protect production latency SLAs.