Skip to content

CNPG Day-2 Operations

Day-2 health and lifecycle runbook for the CloudNativePG (CNPG) Postgres fleet in databases — 21 Cluster resources under kubernetes/apps/databases/cloudnative-pg/config/<app>/, all sharing one operator HelmRelease at kubernetes/apps/databases/cloudnative-pg/app/.

This is not a recovery runbook. For an actual restore, stop here and go to:

  • cnpg_restore.md — per-cluster Barman-cloud restore (PITR, total cluster loss, Garage substrate loss).
  • immich_cnpg.md — the alternate pg_dump/SQL import path, used when you have an Immich-format dump rather than a Barman backup.
  • cluster_rebuild.md § "CNPG recovery from Garage" — the full-cluster-teardown case.

Storage-class rationale (why PGData is always ceph-block, why backups land in Garage) lives in .agents/instructions/storage-class.instructions.md — read that before questioning the pattern below, don't relitigate it here. Underlying volume health (Ceph OSD state, PVC binding) is storage_operations.md (Ceph/Longhorn day-2 doc; if it doesn't exist yet in your checkout, rook_ceph_dr.md covers the Ceph side today).

The pattern, in one paragraph

Every cluster is named postgres-<app> and lives in the databases namespace. spec.storage and spec.walStorage are ceph-block. Every cluster runs the barman-cloud.cloudnative-pg.io plugin as its WAL archiver, backed by a paired barmancloud.cnpg.io/ObjectStore (garage-<app>) pointing at http://garage.storage.svc.cluster.local:3900, bucket s3://postgres-<app>-backup/. A ScheduledBackup named postgres-<app>-backup runs @weekly with immediate: true and 30-day Barman retention (bzip2 compression on both base backups and WAL). All of this is generated by the shared kubernetes/components/cnpg-app-database kustomize Component — a new cluster gets it for free; don't hand-roll a bespoke ObjectStore/ScheduledBackup pair unless you have a real reason to deviate from the fleet.

Almost every cluster runs 3 instances (1 primary + 2 async standbys) on PostgreSQL 17.7. Two exceptions to know about before you assume uniformity: postgres-nametag is still on PG16.11, and postgres-pump is on 17.5 — always check the live imageName before planning a coordinated bump. Three clusters (postgres-immich, postgres-khoj, postgres-langgraph-memory) run the cloudnative-vectorchord image instead of stock cloudnative-pg/postgresql — they need pgvector/vectorchord extensions; treat their minor-version bumps as a separate track from the stock-image fleet.

Apps connect via the CNPG-managed services — postgres-<app>-rw.databases.svc.cluster.local for read-write, -ro for read-only replica traffic, -r for any-instance read traffic. There is no Pooler (pgbouncer) resource anywhere in this fleet today; every app holds direct connections to the service. postgres-immich is the one exception with an app-level need — it additionally exposes a LoadBalancer Service (postgres in kubernetes/apps/databases/cloudnative-pg/config/immich/service.yaml) selecting cnpg.io/cluster: postgres-immich, role: primary, for a consumer that needs to reach Postgres from outside the cluster network path used by everything else.

Cluster health checks

The cnpg kubectl plugin is the primary tool. It's already on the laptop and any host with kubectl configured against this cluster.

# Fleet-wide one-liner: instances, readiness, current primary
kubectl -n databases get clusters.postgresql.cnpg.io

# Full per-cluster detail: replication lag, LSNs, cert expiry, instance placement
kubectl cnpg -n databases status postgres-<app>

# Add -v for postgresql.conf / pg_hba / full replication-slot detail
kubectl cnpg -n databases status postgres-<app> -v

A healthy cluster's status block reads Cluster in healthy state, Ready instances equals Instances, and every row in "Streaming Replication status" shows State: streaming with Write Lag / Flush Lag / Replay Lag at or near 00:00:00.

Known plugin-output gotcha: kubectl cnpg status prints Continuous Backup status: Not configured for every cluster in this fleet, even though WAL archiving and scheduled backups are both running. That line only reflects CNPG's native (non-plugin) backup config; ours is 100% plugin-based (barman-cloud.cloudnative-pg.io), which the plugin's summary view doesn't surface. Don't read that line as a problem — check the Cluster conditions instead (below).

For the ground truth on backup/archiving health, read the Cluster resource's own conditions rather than the plugin summary:

kubectl -n databases get cluster postgres-<app> \
  -o jsonpath='{range .status.conditions[*]}{.type}{"="}{.status}{" ("}{.reason}{") "}{end}'

Four conditions matter day-to-day:

Condition Healthy value Meaning
Ready True Cluster accepting traffic
ContinuousArchiving True WAL archiver has succeeded at least once — not a freshness signal, see below
LastBackupSucceeded True Most recent ScheduledBackup/Backup completed
ConsistentSystemID True All instances agree on one PostgreSQL system identifier (a mismatch means a replica is on the wrong timeline — see Gotchas)

Per-instance role and lag are also visible directly on the metrics pipeline:

# Replication lag in seconds, per standby
cnpg_pg_replication_lag{namespace="databases"}

# Which pod is primary right now
cnpg_collector_pg_replication_role{namespace="databases"}

cnpg-app-database-generated alerts already watch this fleet — see PGReplication (lag > 300s for 10m) and the archiving/backup alerts below in the cloudnative-pg-rules PrometheusRule (kubernetes/apps/databases/cloudnative-pg/app/prometheusrule.yaml). If Alertmanager hasn't paged, you likely don't have a live incident — but run the checks above before believing that in a postmortem.

Failover and switchover

CNPG distinguishes a planned switchover (you choose the new primary, old primary shuts down cleanly first) from an unplanned failover (primary died, operator promotes automatically). Prefer switchover whenever you have the choice.

Planned switchover — before draining a node

If you're about to drain/reboot a node that's hosting a cluster's current primary, move the primary off it first so the app sees zero write downtime instead of a failover gap:

# Confirm current primary and its node
kubectl cnpg -n databases status postgres-<app> | grep -A1 'Primary instance'
kubectl -n databases get pod postgres-<app>-<N> -o wide

# Promote a specific standby (must already be a healthy, caught-up replica)
kubectl cnpg promote -n databases postgres-<app> postgres-<app>-<replica-N>

kubectl cnpg promote triggers a planned switchover — CNPG demotes the current primary to standby only after confirming the target replica is caught up, so there's no window where the cluster has zero primaries. Watch it land:

kubectl cnpg -n databases status postgres-<app>
# Primary instance: should now show the promoted pod

Then drain the now-standby-holding node as normal.

Unplanned failover

If the primary pod or its node dies outright, the operator promotes the most-caught-up standby automatically — no action needed on your part beyond watching it happen and understanding what you're seeing:

kubectl -n databases get events --field-selector involvedObject.name=postgres-<app> \
  --sort-by=.lastTimestamp | tail -20
kubectl cnpg -n databases status postgres-<app>

spec.failoverDelay: 30 (set on every cluster in this fleet) means the operator waits 30s after losing contact with the primary before declaring failover — this absorbs a brief network blip or kubelet restart without triggering an unnecessary promotion. If the primary comes back inside that window, nothing happens. primaryUpdateStrategy: unsupervised means the operator won't wait for your manual go-ahead during a rolling image update either — see Lifecycle below.

After any failover (planned or not), verify the old primary rejoined as a healthy standby rather than getting stuck:

kubectl cnpg -n databases status postgres-<app>
# the old primary's row under "Instances status" should read
# "Standby (async)" / "OK" — not missing, not "Standby (sync)"-stuck-catching-up

If it doesn't rejoin cleanly, see "old-timeline replica" under Gotchas.

Backup verification

Where backups land

Barman writes to Garage (garage.storage.svc.cluster.local:3900), bucket postgres-<app>-backup, with this layout inside the bucket:

  • base/<backup-id>/ — full base backups
  • wals/ — continuous WAL archive

30-day retention (retentionPolicy: 30d on the ObjectStore) — older base backups and their now-unneeded WAL age out automatically. This is the on-cluster-loss recovery path; it does not protect against Garage substrate loss (see garage_restore.md for that layer, and offsite_recovery.md for the small subset of apps — Immich, Paperless — with an additional offsite copy).

Checking backup recency

# Fleet-wide: LAST BACKUP column is time-since-last-successful-run
kubectl -n databases get scheduledbackups.postgresql.cnpg.io

# Per-cluster backup history
kubectl -n databases get backups.postgresql.cnpg.io -l cnpg.io/cluster=postgres-<app>

Every ScheduledBackup in this fleet runs @weekly — expect LAST BACKUP to read up to ~7 days, never much more. If a cluster's most recent Backup object shows PHASE: failed, check its .status.error field and the operator pod's logs (kubectl -n cnpg-system logs deploy/cloudnative-pg).

The LastBackupSucceeded condition on the Cluster (see Health checks above) is the authoritative single-value signal — it flips False immediately on a failed run and back to True on the next success, independent of how the plugin's status view renders.

Do not trust cnpg_collector_last_available_backup_timestamp for this fleet. It's frozen at each cluster's plugin-migration timestamp rather than advancing per successful backup — a known plugin-vs-native metric mismatch (tracked in the cloudnative-pg-rules PrometheusRule comments, upstream home-ops#12075). The BackupStale alert in that same rule file works around it by watching cnpg_cluster_condition_last_transition_time{type= "LastBackupSucceeded"} instead, with an 8-day threshold (one full week + a day of grace, since every cluster is on a weekly cadence) — replicate that query if you're spot-checking staleness by hand.

Spotting a stalled WAL archive

An idle cluster with no writes will legitimately stop advancing WAL forever — Postgres won't force a segment switch just because archive_timeout elapsed if there's genuinely nothing to flush. So "time since last archive" is not a usable staleness signal on this fleet (confirmed on postgres-netbox, which sat WAL-quiet for weeks while completely healthy). What actually indicates a stuck archiver is a growing backlog of produced-but-unarchived segments:

# Per-cluster count of WAL segments sitting in pg_wal/archive_status
# marked .ready (produced, not yet shipped). Healthy clusters sit at 0.
max by (job, namespace) (cnpg_collector_pg_wal_archive_status{value="ready"})

The WALArchivingStalled alert fires at >= 10 sustained for 15 minutes — that threshold comes from a 7-day replay across the whole fleet where the highest transient blip observed (under burst write load) was 7, on postgres-home-assistant. Anything that clears 10 and stays there for 15+ minutes is a genuinely stuck archiver, not a catch-up blip. If you see this, check the Barman plugin sidecar's logs on the primary pod and confirm Garage is reachable from the databases namespace:

kubectl -n databases logs postgres-<app>-<primary-N> -c plugin-barman-cloud --tail=100

First-recoverability-point check

Before treating a cluster as DR-ready after any manual intervention (a promote, a restore test, a retention-policy change), confirm the earliest point you could recover to actually moved forward as expected:

kubectl -n databases get backups.postgresql.cnpg.io \
  -l cnpg.io/cluster=postgres-<app> \
  -o custom-columns=NAME:.metadata.name,PHASE:.status.phase,STARTED:.status.startedAt \
  --sort-by=.status.startedAt

The oldest completed entry still inside the 30-day retention window is your current recovery floor; anything older has already been pruned from Garage by the ObjectStore's retention policy.

Lifecycle operations

Minor-version image bumps

Every cluster's spec.imageName is pinned explicitly in its cluster.yaml — Renovate tracks these like any other image reference. primaryUpdateStrategy: unsupervised (set fleet-wide) means once you merge a bump, the operator rolls it out on its own schedule without waiting for a manual unlock: standbys first, then a switchover to apply the new image to the last (former-primary) instance. Expect one short primary-switchover blip per bump, not a hard outage.

# After merging a Renovate/manual imageName bump, watch the rollout
kubectl -n databases get pods -l cnpg.io/cluster=postgres-<app> -w

# Confirm the new version landed on every instance
kubectl cnpg -n databases status postgres-<app> | grep 'PostgreSQL Image'

Remember the two version outliers noted above (nametag on PG16, pump on 17.5) before assuming a fleet-wide bump PR covers every cluster — check each cluster.yaml's current imageName, don't copy a version string across all 21 blind.

Major-version bumps (e.g. 16→17) are not a simple imageName edit — CNPG doesn't in-place major-upgrade PGData. That's an online upgrade/pg_upgrade project of its own; don't attempt one by editing imageName and hoping.

Scaling instances

spec:
  instances: 3   # bump or drop this in the cluster.yaml

Scaling up adds a new standby that clones from the current primary and joins streaming replication — no primary disruption. Scaling down removes the newest-ordinal instance(s) first; CNPG picks which pod to retire, you don't choose. Every cluster in this fleet already runs 3 instances (1 primary + 2 standbys); going below 2 total loses the ContinuousArchiving/promotion safety margin CNPG relies on for failoverDelay-gated automatic promotion, so don't scale a production-serving cluster to 1 instance without a specific reason and a plan for the gap.

Connecting to a cluster

No pgbouncer/Pooler layer exists in this fleet — connect directly to the CNPG-managed Service:

postgres-<app>-rw.databases.svc.cluster.local:5432   # read-write, always the primary
postgres-<app>-ro.databases.svc.cluster.local:5432   # read-only, load-balanced across standbys
postgres-<app>-r.databases.svc.cluster.local:5432    # any instance, primary or standby

Interactive access from the plugin (uses the superuser secret automatically):

kubectl cnpg -n databases psql postgres-<app>
kubectl cnpg -n databases psql postgres-<app> -- -c '\l'

Common stuck-cluster situations

Cluster stuck mid-rollout / a pod won't come Ready. Check the instance manager logs on the stuck pod first, then the operator:

kubectl -n databases logs postgres-<app>-<N> -c postgres --tail=200
kubectl -n cnpg-system logs deploy/cloudnative-pg --tail=200

A replica won't join / stays "Standby (sync)" forever without catching up. Usually means the new/recovering replica can't pull WAL fast enough or the replication slot backing it was dropped. Check max_slot_wal_keep_size (set per-cluster in postgresql.parameters — e.g. 10GB on postgres-home-assistant, 2GB on postgres-atuin) against how far behind the replica actually is; a replica that falls behind that limit has its slot invalidated and needs a fresh base backup, not just more time.

Old-timeline replica after a failover/restore. If Cluster's ConsistentSystemID condition goes False, or a specific instance sits at Status: OK but never reaches streaming in the replication table, that instance is very likely still on the pre-failover timeline and can't reconcile. The fix is to let CNPG re-clone it, not to debug Postgres internals by hand:

# Confirm which instance is the odd one out first — do not guess
kubectl cnpg -n databases status postgres-<app> -v

# Delete the stuck instance's pod AND its PVCs to force a fresh clone
# from the current primary. This is destructive to that replica's
# local data (not the cluster's data — the primary is untouched) —
# propose before running against anything with only 2 instances left.
kubectl -n databases delete pod postgres-<app>-<N>
kubectl -n databases delete pvc postgres-<app>-<N>,postgres-<app>-<N>-wal

CNPG's StatefulSet-like controller notices the missing instance and re-provisions it from a fresh pg_basebackup against the primary. Watch it rejoin via kubectl cnpg status as in the failover section above.

Gotchas summary

  • kubectl cnpg status's "Continuous Backup status: Not configured" line is a plugin-vs-native blind spot, not a real problem — check Cluster conditions instead.
  • cnpg_collector_last_available_backup_timestamp is stale/frozen on every cluster in this fleet (plugin-method backups aren't reflected) — use the LastBackupSucceeded condition transition time.
  • WAL-archive staleness by clock time is a false signal on an idle database — watch the .ready-segment backlog count, not time-since- last-archive.
  • nametag (PG16) and pump (17.5) are the two clusters off the fleet-standard 17.7 image — verify before a coordinated bump.
  • No Pooler exists anywhere in this fleet; apps connect straight to the -rw/-ro/-r services.
  • bootstrap.recovery only fires on first cluster creation — it is not a lever for day-2 operations. That's cnpg_restore.md territory, not this doc's.

See also

  • cnpg_restore.md — Barman-cloud restore (PITR, total loss, Garage substrate loss).
  • immich_cnpg.md — SQL-dump import path.
  • cluster_rebuild.md — full-cluster teardown and CNPG recovery from Garage.
  • garage_restore.md — restoring the Barman backup target itself.
  • rook_ceph_dr.md — underlying Ceph/PGData volume recovery.
  • debugging.md § CNPG — quick-reference command cheatsheet this doc expands on.
  • .agents/instructions/storage-class.instructions.md — why PGData is ceph-block and backups are Garage, not the other way around.