The Peril of Naive Automated Failover
In enterprise systems engineering, establishing high availability for relational databases is fraught with architectural danger. When a primary database node crashes, the system must promote a standby replica to resume writes as quickly as possible. However, the most catastrophic failure mode in database operations is not downtime—it is split-brain data divergence.
A split-brain scenario occurs when a network partition separates database nodes, leading a standby to believe the primary is dead and promote itself, while the original primary continues accepting writes from an isolated segment of application servers. Once both nodes accept conflicting writes independently, automated reconciliation is mathematically impossible, resulting in permanent data corruption.
1. Distributed Consensus with etcd and Patroni
To achieve reliable, automated failover without split-brain risk, modern database architectures decouple state consensus from the database engine itself. Patroni is an open-source orchestration daemon that manages PostgreSQL instances through an external Distributed Configuration Store (DCS)—most commonly etcd running the Raft consensus algorithm.
In a 3-node Patroni cluster, the primary node must continuously acquire and renew a timed leader lock key in etcd (e.g. /service/batman/leader) within a strict lease window (e.g. 10 seconds). If the primary fails to renew its lease before expiration—whether due to a kernel panic, hardware crash, or network partition—the DCS consensus election opens.
┌───────────────────────────────────────────────┐
│ etcd Raft Consensus Cluster │
│ (Holds DCS Leader Key) │
└───────────────┬───────────────────────────────┘
│
┌─────────┴─────────┐
▼ ▼
┌───────────────┐ ┌───────────────┐
│ Node 1 (Pri) │ │ Node 2 (Sby) │
│ Patroni + PG │ │ Patroni + PG │
└───────┬───────┘ └───────┬───────┘
│ │
└─────────►─────────┘
Synchronous Physical Replication
(Zero Data Loss)
2. Guaranteeing Zero Data Loss with Synchronous Replication
Failover automation without strict durability guarantees can result in silent data loss (RPO > 0). If an unpromoted standby was lagging behind the primary by 50ms of un-replayed WAL at the moment of failover, promoting that standby irrevocably discards those transactions. Patroni eliminates this by enforcing synchronous replication management:
# patroni.yml configuration
scope: postgres-prod
namespace: /service
postgresql:
parameters:
synchronous_commit: "on"
wal_level: replica
max_wal_senders: 10
max_replication_slots: 10
synchronous_mode: true
synchronous_mode_strict: false
# Lease Timing Windows
loop_wait: 10
ttl: 30
retry_timeout: 10
With synchronous_mode: true, Patroni dynamically manages PostgreSQL's synchronous_standby_names. The primary guarantees that every committed transaction is physically flushed to disk on at least one standby before acknowledging success to the application client. If the primary crashes, the standby is mathematically guaranteed to possess 100% of committed transactions.
3. Hardware and Kernel Watchdog Fencing (STONITH)
What prevents an isolated primary from continuing to serve writes if it loses contact with etcd? Patroni implements shoot-the-other-node-in-the-head (STONITH) fencing via the Linux kernel's hardware or software watchdog (/dev/watchdog or softdog).
While holding the leader lock, Patroni periodically sends a heartbeat (pings) to the kernel watchdog device. If Patroni fails to refresh the watchdog within the timeout window (because network isolation prevented lease renewal), the Linux kernel immediately resets the machine at the hardware level. The zombie primary is terminated before a replacement standby can ever be promoted.
4. Self-Healing with pg_rewind
When a failed former primary node boots back up, it cannot simply resume as a standby: its local timeline may have advanced past the new primary's timeline during the crash. Historically, recovering an old primary required copying hundreds of gigabytes of basebackups across the network.
Patroni automates pg_rewind integration. When the old node rejoins, Patroni analyzes the WAL divergence point between the two servers, rewinds only the conflicting data blocks back to the shared fork point, and re-attaches the node as a streaming standby replica in seconds.