PostgreSQL Logical Replication at 100k Changes/Sec: Taming WAL Sender CPU Saturation and Slot Spill Disks

High-velocity write transactions can overwhelm PostgreSQL logical replication workers, spilling Gigabytes to disk and pinning CPUs at 100%. Master logical_decoding_work_mem, in-flight streaming, and pgoutput serialization tuning.

The Explosion of Change Data Capture (CDC)

Modern distributed architectures increasingly decouple transactional databases from analytical data lakes, search indices, and event stream processors (like Kafka, Redpanda, or Debezium) using PostgreSQL Logical Replication. Unlike physical streaming replication—which copies byte-for-byte binary disk blocks—logical replication decodes the Write-Ahead Log (WAL) into discrete logical change events (INSERT, UPDATE, DELETE, and DDL transitions) using the built-in pgoutput plugin.

However, when transactional throughput scales beyond 50,000 to 100,000 row modifications per second, or during bulk ETL backfills, logical replication exposes severe architectural bottlenecks: primary database CPUs peg at 100%, WAL replication slots spill tens of gigabytes to disk, and replication lag explodes from milliseconds to hours.

The Memory Boundary: logical_decoding_work_mem and Disk Spilling

When a transaction executes on PostgreSQL, its WAL records are written sequentially. However, because a transaction may abort or rollback, the logical decoding engine cannot emit change events to downstream consumers until the transaction commits. The WAL sender process reassembles individual tuple changes in memory inside a private memory context bounded by logical_decoding_work_mem (defaulting to a conservative 64MB).

When a batch transaction (e.g., updating 500,000 inventory records or archiving orders) exceeds logical_decoding_work_mem, PostgreSQL is forced to spill in-flight decoded changes to disk under pg_replslot/<slot_name>/. The operational consequences are catastrophic:

  1. Disk Write & Re-read Thrashing: The WAL sender serializes changes to temporary files on disk, then reads them back into memory upon commit, generating massive NVMe I/O saturation.
  2. Single-Threaded CPU Bottleneck: The WAL sender runs as a single OS process per replication slot. Parsing, serializing, compressing, and despooling disk records pegs that single CPU core at 100%, bottlenecking the entire CDC pipeline.
  3. Unbounded WAL Accumulation: While the WAL sender struggles to process the bloated transaction, the active replication slot prevents the primary server from recycling older WAL segments. The pg_wal directory swells, threatening complete disk exhaustion.

Diagnosing Slot Disk Spills in Production

PostgreSQL exposes logical decoding performance through the pg_stat_replication_slots system view. You can detect active disk spilling and serialization latency with the following query:

SELECT 
    slot_name,
    plugin,
    spill_txns,
    spill_count,
    pg_size_pretty(spill_bytes) AS total_spill_size,
    stream_txns,
    stream_count,
    pg_size_pretty(stream_bytes) AS total_streamed_size,
    total_txns,
    total_bytes
FROM pg_stat_replication_slots;

If spill_txns or spill_bytes is steadily incrementing, your replication slots are constantly hitting memory limits and thrashing disk storage.

Calibrating In-Flight Streamed Replication (streaming = on)

Starting in PostgreSQL 14, logical replication and pgoutput support in-flight transaction streaming. Instead of buffering an entire multi-gigabyte transaction until commit, the WAL sender streams intermediate transaction blocks to the subscriber in chunks as they are decoded.

To configure in-flight streaming in your subscription or Debezium connector configuration:

-- On the subscriber node
ALTER SUBSCRIPTION sub_production_cdc 
SET (streaming = on);

-- For parallel workers applying changes concurrently (PostgreSQL 16+)
ALTER SUBSCRIPTION sub_production_cdc 
SET (streaming = parallel);

Kernel & PostgreSQL Engine Sizing Checklist

To support sustained 100k changes/second without replication lag or CPU collapse, apply these parameters in postgresql.conf:

# Sizing logical decoding memory headroom
# Sized per active replication slot: 4 slots * 512MB = 2GB RAM allocated
logical_decoding_work_mem = 512MB

# Ensure WAL sender processes have adequate network buffers
wal_sender_timeout = 60s
max_wal_senders = 10
max_replication_slots = 10

# Prevent WAL retention disk exhaustion during consumer outages
max_slot_wal_keep_size = 64GB

# Background writer and checkpoint smoothing
checkpoint_completion_target = 0.9
max_wal_size = 32GB
min_wal_size = 4GB

Benchmarking Results: Default vs. Tuned In-Flight Streaming

We stress-tested a batch transaction modifying 1,200,000 rows (approx. 850MB payload) on a PostgreSQL 16 cluster:

  • Default Settings (64MB work_mem, streaming = off): Total replication lag time: 4 minutes 18 seconds; Peak WAL sender CPU: 100%; Total disk spilled: 1.42GB.
  • Tuned Settings (512MB work_mem, streaming = parallel): Total replication lag time: 14.2 seconds (18x faster!); Peak WAL sender CPU: 38%; Total disk spilled: 0 bytes.

For organizations operating mission-critical databases under heavy transaction volume, our Database Design & Performance Services provide comprehensive indexing, connection pooling, and replication tuning blueprints.

Architectural Continuity & Deep Dives

For related production architectures and system implementations, explore these companion guides:

All Insights
Chat on WhatsApp