Databases · 2026-10-06 · backups / disaster-recovery / postgres / silent-failure / monitoring · medium

The nightly backup job ran for years, but the bucket was empty

01Symptom

Your product runs on one primary PostgreSQL server and one hot standby. Around 23:00 UTC, standby replication breaks, and the primary has already discarded the required log segments. You start rebuilding the standby. At about 23:30, someone wipes a data directory on the primary. Roughly 300 GB is gone, the standby was already emptied for its rebuild, and production is down. When you reach for backup, the S3 bucket is empty.

02Constraints

  • The standby exists for failover; it is not a disaster-recovery copy
  • On paper there are four recovery mechanisms: a daily logical dump in S3, a daily disk snapshot in staging, disk snapshots, and replication
  • Cloud disk snapshots are enabled for the file servers, but not for the database hosts
  • PostgreSQL 9.6 is in use; the packaging supports 9.2 and 9.6 and picks binaries by reading the local data directory version
  • The backup cron job runs on an application server, not on the database host
  • Cron failure notices are sent by email
  • Staging uses cheap, slow disks in a different region

03Evidence

  • The logical-dump S3 bucket is empty, and there is no recent dump anywhere on disk
  • The job is enabled and still scheduled, but it has never produced a usable file
  • The application server has no PostgreSQL data directory, so the packaging falls back to its default PostgreSQL 9.2 binaries, while the server is 9.6
  • The dump tool failed every run on the major-version mismatch
  • Cron sent a failure notice, but the receiving mail server rejected it because the sender was not DMARC-signed
  • The standby was emptied for the rebuild, leaving no failover target
  • Cloud disk snapshots were never enabled for the database hosts
  • The newest usable copy is a hand-made snapshot about six hours old; the scheduled one is nearly 24 hours old
  • Restoring means copying roughly the whole data directory from staging over throttled disks topping out around 60 Mbps
  • Nobody was ever assigned to test restores
→

Recent

Previously diagnosed.