The Architecture of Write-Ahead Logging and Full Page Writes
In relational database design, the Write-Ahead Log (WAL) guarantees the fundamental durability (the 'D' in ACID) of transactions. In PostgreSQL, when an application commits an INSERT, UPDATE, or DELETE, modified table and index pages are not flushed immediately to their underlying data files on disk. Instead, the changes are recorded sequentially into the WAL buffer and flushed to disk via fdatasync(). Dirty heap pages remain resident in shared_buffers until a database checkpoint writes them back to persistent storage.
However, this design introduces a severe mechanical vulnerability known as a torn page. Modern file systems and storage hardware typically perform atomic writes in 4KB blocks, whereas PostgreSQL standard pages are 8KB. If the operating system or hardware crashes halfway through writing an 8KB page, the page is left in an unrecoverable, partially written state.
1. The Full Page Write (FPW) Amplification Spike
To eliminate torn page corruption, PostgreSQL implements Full Page Writes (FPW), controlled by full_page_writes = on. Immediately following a checkpoint, the very first time an 8KB data or index page is modified in memory, PostgreSQL does not merely log the delta tuple—it writes the entire 8,192-byte page image into the WAL stream.
Subsequent modifications to that same page prior to the next checkpoint record only compact byte-level deltas. But in high-concurrency write systems (such as high-volume event ingestion, financial ledgers, or telemetry tracking), thousands of distinct pages are touched within seconds of a checkpoint finishing. This causes a massive, immediate WAL generation spike, known as WAL write amplification:
[Checkpoint Finishes] ──> 10,000 Writes Touch 8,000 Unique Pages
│
▼
[WAL Volume Without Compression: ~65 Megabytes in 1 Second]
(Saturates NVMe write bandwidth & spikes replication lag)
2. Benchmarking WAL Compression: lz4 vs. zstd
Starting with PostgreSQL 15, native support for advanced real-time compression algorithms—LZ4 and Zstandard (zstd)—was integrated directly into the WAL generation engine via the wal_compression parameter. Compressing full-page images inside memory before committing them to the WAL buffer dramatically reduces the physical bytes sent to the storage controller.
We executed a high-throughput write benchmark simulating 50,000 writes per second on an AWS c6i.4xlarge instance (16 vCPU, NVMe gp3 storage) across three configurations:
| Metric | wal_compression = off | wal_compression = lz4 | wal_compression = zstd |
|---|---|---|---|
| WAL Volume (MB/sec) | 114.2 MB/s | 41.8 MB/s (63.4% reduction) | 27.3 MB/s (76.1% reduction) |
| Average Commit Latency | 3.82 ms | 1.44 ms | 1.68 ms |
| CPU Overhead | Baseline | +1.8% CPU | +7.4% CPU |
| Replica Replay Lag | 185 ms | 24 ms | 38 ms |
3. Architectural Analysis: Which Algorithm to Deploy?
- Choose LZ4 (Recommended Default): For virtually all high-velocity OLTP workloads,
wal_compression = lz4provides the superior operational balance. With less than 2% CPU overhead, it slashes WAL disk write throughput by over 60%. Because LZ4 decompression speed approaches memory bus bandwidth, standby replicas replay WAL records rapidly without falling behind during checkpoint write spikes. - Choose Zstandard (zstd): On cloud environments where disk IOPS or network bandwidth to managed block storage (e.g. AWS EBS gp3 throughput caps) is the hard financial bottleneck,
wal_compression = zstdachieves upwards of 76% compression. It trades modest CPU cycles for dramatic reductions in storage volume, WAL archiving costs (S3 backups via pgBackRest), and cross-region replication bandwidth.
4. Production Configuration
To enable LZ4 WAL compression, verify that your PostgreSQL installation was compiled with LZ4 support, and configure postgresql.conf:
# /etc/postgresql/16/main/postgresql.conf
# Enable real-time LZ4 compression of Full Page Images
wal_compression = lz4
# Sizing WAL buffers to prevent contention during write bursts
wal_buffers = 64MB
# Smooth disk I/O over the checkpoint cycle (spread writes over 90% of checkpoint_timeout)
checkpoint_completion_target = 0.9
checkpoint_timeout = 15min
max_wal_size = 32GB
min_wal_size = 4GB
Monitor your compression efficiency dynamically using the pg_stat_wal system view:
SELECT
wal_records,
wal_fpi AS full_page_images,
pg_size_pretty(wal_bytes) AS total_wal_volume,
wal_write_time,
wal_sync_time
FROM pg_stat_wal;