CNPG Day-2 Operations¶
Day-2 health and lifecycle runbook for the CloudNativePG (CNPG)
Postgres fleet in databases — 21 Cluster resources under
kubernetes/apps/databases/cloudnative-pg/config/<app>/, all sharing
one operator HelmRelease at
kubernetes/apps/databases/cloudnative-pg/app/.
This is not a recovery runbook. For an actual restore, stop here and go to:
cnpg_restore.md— per-cluster Barman-cloud restore (PITR, total cluster loss, Garage substrate loss).immich_cnpg.md— the alternatepg_dump/SQL import path, used when you have an Immich-format dump rather than a Barman backup.cluster_rebuild.md§ "CNPG recovery from Garage" — the full-cluster-teardown case.
Storage-class rationale (why PGData is always ceph-block, why
backups land in Garage) lives in
.agents/instructions/storage-class.instructions.md
— read that before questioning the pattern below, don't relitigate it
here. Underlying volume health (Ceph OSD state, PVC binding) is
storage_operations.md (Ceph/Longhorn day-2 doc; if it doesn't exist
yet in your checkout, rook_ceph_dr.md covers the Ceph side today).
The pattern, in one paragraph¶
Every cluster is named postgres-<app> and lives in the databases
namespace. spec.storage and spec.walStorage are ceph-block.
Every cluster runs the barman-cloud.cloudnative-pg.io plugin as its
WAL archiver, backed by a paired barmancloud.cnpg.io/ObjectStore
(garage-<app>) pointing at
http://garage.storage.svc.cluster.local:3900, bucket
s3://postgres-<app>-backup/. A ScheduledBackup named
postgres-<app>-backup runs @weekly with immediate: true and
30-day Barman retention (bzip2 compression on both base backups and
WAL). All of this is generated by the shared
kubernetes/components/cnpg-app-database kustomize Component — a new
cluster gets it for free; don't hand-roll a bespoke
ObjectStore/ScheduledBackup pair unless you have a real reason to
deviate from the fleet.
Almost every cluster runs 3 instances (1 primary + 2 async standbys)
on PostgreSQL 17.7. Two exceptions to know about before you assume
uniformity: postgres-nametag is still on PG16.11, and
postgres-pump is on 17.5 — always check the live imageName before
planning a coordinated bump. Three clusters (postgres-immich,
postgres-khoj, postgres-langgraph-memory) run the
cloudnative-vectorchord image instead of stock
cloudnative-pg/postgresql — they need pgvector/vectorchord
extensions; treat their minor-version bumps as a separate track from
the stock-image fleet.
Apps connect via the CNPG-managed services —
postgres-<app>-rw.databases.svc.cluster.local for read-write,
-ro for read-only replica traffic, -r for any-instance read
traffic. There is no Pooler (pgbouncer) resource anywhere in this
fleet today; every app holds direct connections to the service.
postgres-immich is the one exception with an app-level need — it
additionally exposes a LoadBalancer Service (postgres in
kubernetes/apps/databases/cloudnative-pg/config/immich/service.yaml)
selecting cnpg.io/cluster: postgres-immich, role: primary, for a
consumer that needs to reach Postgres from outside the cluster
network path used by everything else.
Cluster health checks¶
The cnpg kubectl plugin is the primary tool. It's already on the
laptop and any host with kubectl configured against this cluster.
# Fleet-wide one-liner: instances, readiness, current primary
kubectl -n databases get clusters.postgresql.cnpg.io
# Full per-cluster detail: replication lag, LSNs, cert expiry, instance placement
kubectl cnpg -n databases status postgres-<app>
# Add -v for postgresql.conf / pg_hba / full replication-slot detail
kubectl cnpg -n databases status postgres-<app> -v
A healthy cluster's status block reads Cluster in healthy state,
Ready instances equals Instances, and every row in "Streaming
Replication status" shows State: streaming with Write Lag /
Flush Lag / Replay Lag at or near 00:00:00.
Known plugin-output gotcha: kubectl cnpg status prints
Continuous Backup status: Not configured for every cluster in this
fleet, even though WAL archiving and scheduled backups are both
running. That line only reflects CNPG's native (non-plugin) backup
config; ours is 100% plugin-based (barman-cloud.cloudnative-pg.io),
which the plugin's summary view doesn't surface. Don't read that line
as a problem — check the Cluster conditions instead (below).
For the ground truth on backup/archiving health, read the Cluster
resource's own conditions rather than the plugin summary:
kubectl -n databases get cluster postgres-<app> \
-o jsonpath='{range .status.conditions[*]}{.type}{"="}{.status}{" ("}{.reason}{") "}{end}'
Four conditions matter day-to-day:
| Condition | Healthy value | Meaning |
|---|---|---|
Ready |
True |
Cluster accepting traffic |
ContinuousArchiving |
True |
WAL archiver has succeeded at least once — not a freshness signal, see below |
LastBackupSucceeded |
True |
Most recent ScheduledBackup/Backup completed |
ConsistentSystemID |
True |
All instances agree on one PostgreSQL system identifier (a mismatch means a replica is on the wrong timeline — see Gotchas) |
Per-instance role and lag are also visible directly on the metrics pipeline:
# Replication lag in seconds, per standby
cnpg_pg_replication_lag{namespace="databases"}
# Which pod is primary right now
cnpg_collector_pg_replication_role{namespace="databases"}
cnpg-app-database-generated alerts already watch this fleet — see
PGReplication (lag > 300s for 10m) and the archiving/backup alerts
below in the cloudnative-pg-rules PrometheusRule
(kubernetes/apps/databases/cloudnative-pg/app/prometheusrule.yaml).
If Alertmanager hasn't paged, you likely don't have a live incident —
but run the checks above before believing that in a postmortem.
Failover and switchover¶
CNPG distinguishes a planned switchover (you choose the new primary, old primary shuts down cleanly first) from an unplanned failover (primary died, operator promotes automatically). Prefer switchover whenever you have the choice.
Planned switchover — before draining a node¶
If you're about to drain/reboot a node that's hosting a cluster's current primary, move the primary off it first so the app sees zero write downtime instead of a failover gap:
# Confirm current primary and its node
kubectl cnpg -n databases status postgres-<app> | grep -A1 'Primary instance'
kubectl -n databases get pod postgres-<app>-<N> -o wide
# Promote a specific standby (must already be a healthy, caught-up replica)
kubectl cnpg promote -n databases postgres-<app> postgres-<app>-<replica-N>
kubectl cnpg promote triggers a planned switchover — CNPG
demotes the current primary to standby only after confirming the
target replica is caught up, so there's no window where the cluster
has zero primaries. Watch it land:
kubectl cnpg -n databases status postgres-<app>
# Primary instance: should now show the promoted pod
Then drain the now-standby-holding node as normal.
Unplanned failover¶
If the primary pod or its node dies outright, the operator promotes the most-caught-up standby automatically — no action needed on your part beyond watching it happen and understanding what you're seeing:
kubectl -n databases get events --field-selector involvedObject.name=postgres-<app> \
--sort-by=.lastTimestamp | tail -20
kubectl cnpg -n databases status postgres-<app>
spec.failoverDelay: 30 (set on every cluster in this fleet) means
the operator waits 30s after losing contact with the primary before
declaring failover — this absorbs a brief network blip or kubelet
restart without triggering an unnecessary promotion. If the primary
comes back inside that window, nothing happens. primaryUpdateStrategy:
unsupervised means the operator won't wait for your manual go-ahead
during a rolling image update either — see Lifecycle below.
After any failover (planned or not), verify the old primary rejoined as a healthy standby rather than getting stuck:
kubectl cnpg -n databases status postgres-<app>
# the old primary's row under "Instances status" should read
# "Standby (async)" / "OK" — not missing, not "Standby (sync)"-stuck-catching-up
If it doesn't rejoin cleanly, see "old-timeline replica" under Gotchas.
Backup verification¶
Where backups land¶
Barman writes to Garage (garage.storage.svc.cluster.local:3900),
bucket postgres-<app>-backup, with this layout inside the bucket:
base/<backup-id>/— full base backupswals/— continuous WAL archive
30-day retention (retentionPolicy: 30d on the ObjectStore) — older
base backups and their now-unneeded WAL age out automatically. This
is the on-cluster-loss recovery path; it does not protect against
Garage substrate loss (see garage_restore.md for that layer, and
offsite_recovery.md for the small subset of apps — Immich, Paperless
— with an additional offsite copy).
Checking backup recency¶
# Fleet-wide: LAST BACKUP column is time-since-last-successful-run
kubectl -n databases get scheduledbackups.postgresql.cnpg.io
# Per-cluster backup history
kubectl -n databases get backups.postgresql.cnpg.io -l cnpg.io/cluster=postgres-<app>
Every ScheduledBackup in this fleet runs @weekly — expect LAST
BACKUP to read up to ~7 days, never much more. If a cluster's most
recent Backup object shows PHASE: failed, check its .status.error
field and the operator pod's logs
(kubectl -n cnpg-system logs deploy/cloudnative-pg).
The LastBackupSucceeded condition on the Cluster (see Health
checks above) is the authoritative single-value signal — it flips
False immediately on a failed run and back to True on the next
success, independent of how the plugin's status view renders.
Do not trust cnpg_collector_last_available_backup_timestamp for
this fleet. It's frozen at each cluster's plugin-migration timestamp
rather than advancing per successful backup — a known plugin-vs-native
metric mismatch (tracked in the cloudnative-pg-rules
PrometheusRule comments, upstream home-ops#12075). The
BackupStale alert in that same rule file works around it by
watching cnpg_cluster_condition_last_transition_time{type=
"LastBackupSucceeded"} instead, with an 8-day threshold (one full
week + a day of grace, since every cluster is on a weekly cadence)
— replicate that query if you're spot-checking staleness by hand.
Spotting a stalled WAL archive¶
An idle cluster with no writes will legitimately stop advancing
WAL forever — Postgres won't force a segment switch just because
archive_timeout elapsed if there's genuinely nothing to flush. So
"time since last archive" is not a usable staleness signal on this
fleet (confirmed on postgres-netbox, which sat WAL-quiet for weeks
while completely healthy). What actually indicates a stuck archiver is
a growing backlog of produced-but-unarchived segments:
# Per-cluster count of WAL segments sitting in pg_wal/archive_status
# marked .ready (produced, not yet shipped). Healthy clusters sit at 0.
max by (job, namespace) (cnpg_collector_pg_wal_archive_status{value="ready"})
The WALArchivingStalled alert fires at >= 10 sustained for 15
minutes — that threshold comes from a 7-day replay across the whole
fleet where the highest transient blip observed (under burst write
load) was 7, on postgres-home-assistant. Anything that clears 10 and
stays there for 15+ minutes is a genuinely stuck archiver, not a
catch-up blip. If you see this, check the Barman plugin sidecar's logs
on the primary pod and confirm Garage is reachable from the
databases namespace:
First-recoverability-point check¶
Before treating a cluster as DR-ready after any manual intervention (a promote, a restore test, a retention-policy change), confirm the earliest point you could recover to actually moved forward as expected:
kubectl -n databases get backups.postgresql.cnpg.io \
-l cnpg.io/cluster=postgres-<app> \
-o custom-columns=NAME:.metadata.name,PHASE:.status.phase,STARTED:.status.startedAt \
--sort-by=.status.startedAt
The oldest completed entry still inside the 30-day retention window
is your current recovery floor; anything older has already been
pruned from Garage by the ObjectStore's retention policy.
Lifecycle operations¶
Minor-version image bumps¶
Every cluster's spec.imageName is pinned explicitly in its
cluster.yaml — Renovate tracks these like any other image reference.
primaryUpdateStrategy: unsupervised (set fleet-wide) means once you
merge a bump, the operator rolls it out on its own schedule without
waiting for a manual unlock: standbys first, then a switchover to
apply the new image to the last (former-primary) instance. Expect one
short primary-switchover blip per bump, not a hard outage.
# After merging a Renovate/manual imageName bump, watch the rollout
kubectl -n databases get pods -l cnpg.io/cluster=postgres-<app> -w
# Confirm the new version landed on every instance
kubectl cnpg -n databases status postgres-<app> | grep 'PostgreSQL Image'
Remember the two version outliers noted above (nametag on PG16,
pump on 17.5) before assuming a fleet-wide bump PR covers every
cluster — check each cluster.yaml's current imageName, don't copy
a version string across all 21 blind.
Major-version bumps (e.g. 16→17) are not a simple imageName edit
— CNPG doesn't in-place major-upgrade PGData. That's an online
upgrade/pg_upgrade project of its own; don't attempt one by editing
imageName and hoping.
Scaling instances¶
Scaling up adds a new standby that clones from the current primary and
joins streaming replication — no primary disruption. Scaling down
removes the newest-ordinal instance(s) first; CNPG picks which pod to
retire, you don't choose. Every cluster in this fleet already runs 3
instances (1 primary + 2 standbys); going below 2 total loses the
ContinuousArchiving/promotion safety margin CNPG relies on for
failoverDelay-gated automatic promotion, so don't scale a
production-serving cluster to 1 instance without a specific reason and
a plan for the gap.
Connecting to a cluster¶
No pgbouncer/Pooler layer exists in this fleet — connect directly to
the CNPG-managed Service:
postgres-<app>-rw.databases.svc.cluster.local:5432 # read-write, always the primary
postgres-<app>-ro.databases.svc.cluster.local:5432 # read-only, load-balanced across standbys
postgres-<app>-r.databases.svc.cluster.local:5432 # any instance, primary or standby
Interactive access from the plugin (uses the superuser secret automatically):
kubectl cnpg -n databases psql postgres-<app>
kubectl cnpg -n databases psql postgres-<app> -- -c '\l'
Common stuck-cluster situations¶
Cluster stuck mid-rollout / a pod won't come Ready. Check the instance manager logs on the stuck pod first, then the operator:
kubectl -n databases logs postgres-<app>-<N> -c postgres --tail=200
kubectl -n cnpg-system logs deploy/cloudnative-pg --tail=200
A replica won't join / stays "Standby (sync)" forever without
catching up. Usually means the new/recovering replica can't pull
WAL fast enough or the replication slot backing it was dropped. Check
max_slot_wal_keep_size (set per-cluster in postgresql.parameters
— e.g. 10GB on postgres-home-assistant, 2GB on postgres-atuin)
against how far behind the replica actually is; a replica that falls
behind that limit has its slot invalidated and needs a fresh base
backup, not just more time.
Old-timeline replica after a failover/restore. If Cluster's
ConsistentSystemID condition goes False, or a specific instance
sits at Status: OK but never reaches streaming in the replication
table, that instance is very likely still on the pre-failover timeline
and can't reconcile. The fix is to let CNPG re-clone it, not to debug
Postgres internals by hand:
# Confirm which instance is the odd one out first — do not guess
kubectl cnpg -n databases status postgres-<app> -v
# Delete the stuck instance's pod AND its PVCs to force a fresh clone
# from the current primary. This is destructive to that replica's
# local data (not the cluster's data — the primary is untouched) —
# propose before running against anything with only 2 instances left.
kubectl -n databases delete pod postgres-<app>-<N>
kubectl -n databases delete pvc postgres-<app>-<N>,postgres-<app>-<N>-wal
CNPG's StatefulSet-like controller notices the missing instance and
re-provisions it from a fresh pg_basebackup against the primary.
Watch it rejoin via kubectl cnpg status as in the failover section
above.
Gotchas summary¶
kubectl cnpg status's "Continuous Backup status: Not configured" line is a plugin-vs-native blind spot, not a real problem — checkClusterconditions instead.cnpg_collector_last_available_backup_timestampis stale/frozen on every cluster in this fleet (plugin-method backups aren't reflected) — use theLastBackupSucceededcondition transition time.- WAL-archive staleness by clock time is a false signal on an idle
database — watch the
.ready-segment backlog count, not time-since- last-archive. nametag(PG16) andpump(17.5) are the two clusters off the fleet-standard 17.7 image — verify before a coordinated bump.- No
Poolerexists anywhere in this fleet; apps connect straight to the-rw/-ro/-rservices. bootstrap.recoveryonly fires on first cluster creation — it is not a lever for day-2 operations. That'scnpg_restore.mdterritory, not this doc's.
See also¶
cnpg_restore.md— Barman-cloud restore (PITR, total loss, Garage substrate loss).immich_cnpg.md— SQL-dump import path.cluster_rebuild.md— full-cluster teardown and CNPG recovery from Garage.garage_restore.md— restoring the Barman backup target itself.rook_ceph_dr.md— underlying Ceph/PGData volume recovery.debugging.md§ CNPG — quick-reference command cheatsheet this doc expands on..agents/instructions/storage-class.instructions.md— why PGData isceph-blockand backups are Garage, not the other way around.