Replicas and failover
A PostgreSQL database on CloudNativePG can run 1 to 3 instances, each on a different server, under Instances and replication on its page. More instances need more servers.
This is not high availability
The cluster's control plane runs only on the main server. If the main server is lost, nothing is failed over. See What Orbit does not promise.
The preview
It picks a server for each instance, starting with the primary's, and refuses:
- when there are not enough servers with room for the database's memory;
- synchronous replication with two instances, because losing either server would stop every write.
Replication
- Asynchronous (the default): commits do not wait for replicas, so a failover can lose the last writes.
- Synchronous (three instances): each commit waits until a replica has it, and only a replica that holds every committed write is promoted, so a failover loses none. Writes are slower, and they stop if two of the three servers are lost.
Failover
When the primary's server is lost, a replica is promoted and the database's address follows the new primary, so applications keep their variables. When that server comes back, its instance rejoins as a replica.
A primary that loses its network but keeps running reaches neither the cluster nor its replicas, so it restarts itself and ends its sessions about half a minute after the cut. That is before a replica is promoted, so only one primary accepts writes at a time.
Make primary on a replica promotes it on purpose (a switchover). It took 9 to 10 seconds on a test cluster.
What Orbit shows
The page lists each instance with its server, its role and how far a replica is behind. Every failover and switchover is recorded in the activity and on the database's timeline with what Orbit measured:
- How long the database had no primary: from the moment the old one was last known to work (its server's last heartbeat or its last replicated commit) to the promotion.
- What was kept: with synchronous replication, every committed write. With asynchronous replication, the time of the newest write the new primary has, so anything the old primary committed after it was lost.
Backups and point-in-time recovery keep working after a failover, since the archive follows the primary.
Tested under failure
Orbit's acceptance test breaks a three-node cluster on purpose: the main server and two more, running as containers on one machine. Two PostgreSQL 17 databases run there with three instances each, one per server. One replicates asynchronously, the other synchronously and archives its WAL to S3-compatible storage. An application writes to each through the database's address five times a second and records which writes were acknowledged.
| Scenario | Measured |
|---|---|
| The primary's server stops | 63 to 64 s without a primary; writes back after 67 to 70 s. No acknowledged write lost, asynchronous or synchronous. |
| The primary's server loses its network | 96 s without a primary; writes back after 101 to 103 s. The isolated asynchronous primary stopped accepting writes 34 s after the cut, 61 s before the promotion, and the writes it took in those 34 s were lost. The synchronous one accepted none. Never two writable primaries. |
A destructive change (delete without where) | Restored to the second before it in 63 s. |
| Every instance and volume lost | Rebuilt from the archive alone: writes back 51 s after the loss. The 41 s of writes that had not reached the archive were lost, and Orbit reported exactly up to which commit it recovered. |
| A configuration that cannot survive losing a server | The preview refuses it. |
What remains is listed in What Orbit does not promise.