Linux cgroups v2 & Memory Pressure Stalling: Diagnosing Kernel Thrashing and Sizing Container Limits in Production

Containers frequently suffer debilitating tail-latency spikes long before triggering OOM kills because the Linux kernel thrashes page cache allocations under pressure. Learn to interpret /proc/pressure/memory and configure memory.high in cgroups v2.

The Silent Killer: Page Cache Reclamation vs. OOM Termination

In high-density production environments, container memory sizing is almost universally approached as a binary problem: either a container stays alive within its allocated threshold, or the Linux Out-Of-Memory (OOM) killer abruptly terminates the process with exit code 137. However, the most destructive container performance degradations occur long before the OOM killer ever fires.

Under Linux control groups (cgroups), when a container's memory consumption approaches its configured limit, the kernel's memory management subsystem begins aggressively reclaiming clean executable pages, file-backed page caches, and inactive buffers to stave off an OOM crisis. As the available page cache collapses, standard disk reads, dynamic library references, and shared memory lookups immediately suffer synchronous disk cache misses. The CPU spends upward of 80% of its execution cycles stuck in iowait, thrashing the kernel memory subsystem. To external monitoring dashboards tracking basic CPU and RAM utilization percentages, the container appears to be operating normally within limits—yet API response times spike from 25 milliseconds to over 8 seconds.

Deciphering Pressure Stall Information (PSI): /proc/pressure/memory

To expose this invisible thrashing, modern Linux kernels (5.x+) equipped with cgroups v2 introduce Pressure Stall Information (PSI). Rather than reporting static byte allocations, PSI measures the exact percentage of time tasks are blocked waiting for memory resources.

Inspecting /proc/pressure/memory (or a container's local memory.pressure cgroup file) yields two fundamental metrics:

  • some avg10=X.XX avg60=X.XX avg300=X.XX: Represents the percentage of wall-clock time during which at least one active process thread was completely stalled waiting for memory allocation, page reclaim, or swap-in. In healthy systems, some avg10 should remain below 2.00%. Values above 15.00% signal severe throughput degradation.
  • full avg10=X.XX avg60=X.XX avg300=X.XX: Represents the percentage of time during which all non-idle threads in the container were stalled simultaneously. Any non-zero value on the full metric indicates complete system livelock where application processing has essentially ground to a halt.

Configuring memory.high vs. memory.max in cgroups v2

Legacy cgroups v1 offered only a single hard ceiling (memory.limit_in_bytes). Crossing this ceiling triggered instantaneous process termination. In contrast, cgroups v2 provides a sophisticated two-tier memory control mechanism that enables graceful degradation:

# Systemd Service Unit: /etc/systemd/system/gunicorn-api.service.d/override.conf
[Service]
# Hard limit: Kernel OOM killer terminates tasks if breached
MemoryMax=4G

# Soft throttling ceiling: Kernel aggressively reclaims and throttles CPU allocations
MemoryHigh=3200M

# Minimum protected memory: Protected from global host page cache reclaim
MemoryMin=512M

# Proactively enforce cgroups v2 memory limits
CPUWeight=100
IOWeight=100

When a container crosses MemoryHigh, the kernel exerts proportional backpressure. It throttles the allocating processes during system calls, gently slowing down incoming work rather than crashing the process. This dynamic backpressure buys time for background garbage collectors, Python object deallocations, and connection pool cleanups to stabilize memory footprints.

Production Python PSI Monitor with epoll

Linux allows user-space daemons to register kernel event triggers on PSI files using epoll, providing sub-millisecond alerts when memory pressure spikes. Below is a lightweight Python watchdog daemon:

import select
import sys
import os
import logging

logging.basicConfig(level=logging.INFO, format="%(asctime)s [%(levelname)s] %(message)s")
logger = logging.getLogger("psi.monitor")

PRESSURE_PATH = "/proc/pressure/memory"
# Trigger alert if 'some' tasks stall for more than 150ms (150000us) within a 1-second window (1000000us)
TRIGGER_CONFIG = "some 150000 1000000"

def monitor_memory_stalls():
    if not os.path.exists(PRESSURE_PATH):
        logger.error("Kernel Pressure Stall Information (PSI) is not enabled on this host.")
        sys.exit(1)

    with open(PRESSURE_PATH, "w+") as fd:
        fd.write(TRIGGER_CONFIG)
        fd.flush()

        epoll = select.epoll()
        # Register for high-priority polling events
        epoll.register(fd.fileno(), select.EPOLLPRI)
        logger.info(f"Active kernel PSI watchdog listening on {PRESSURE_PATH}...")

        while True:
            events = epoll.poll(timeout=5.0)
            for fileno, event in events:
                if event & select.EPOLLPRI:
                    logger.warning(
                        "CRITICAL MEMORY PRESSURE: Kernel stall duration exceeded 150ms in 1s window! "
                        "Triggering automated memory shedder, clearing cache pools..."
                    )
                    execute_memory_shedding()

def execute_memory_shedding():
    # Application-specific shedding: close idle DB connections, evict local caches
    pass

if __name__ == "__main__":
    monitor_memory_stalls()
Architectural Continuity & Deep Dives

For related production architectures and system implementations, explore these companion guides:

Operational Takeaways & Sizing Rules

When sizing production Docker containers and systemd services, adhere to these three rules:

  1. Always configure MemoryHigh at 75–80% of MemoryMax: This cushion allows the kernel to exert gradual CPU throttling and shed cache before invoking catastrophic OOM termination.
  2. Graph PSI in Grafana: Ingest memory.pressure metrics alongside traditional RAM utilization to identify thrashing hours before it causes an incident.
  3. Verify cgroups v2 mounting: Run mount | grep cgroup and verify that cgroup2 on /sys/fs/cgroup is active. Upgrading from cgroups v1 to v2 is the single highest-ROI infrastructure improvement you can deploy to protect production latency SLAs.
All Insights
Chat on WhatsApp