Automating PostgreSQL/MongoDB Backups in Kubernetes
A Kubernetes backup system must protect three things: database data, database-level metadata, and the restore process itself. A scheduled backup that has never been restored is only an assumption, so design the workflow around a measurable recovery point objective (RPO) and recovery time objective (RTO).
A robust pattern is:
-
A Kubernetes
CronJobstarts a short-lived, non-root backup Pod. -
The Pod reads credentials from a Secret mounted as environment variables or files.
-
It connects to PostgreSQL or MongoDB through an internal Service using least-privilege backup credentials.
-
It creates a compressed, encrypted backup artifact.
-
It uploads the artifact to object storage or a remote backup repository.
-
It writes logs, emits metrics, and removes temporary local files.
-
A separate scheduled restore test validates the latest backup in an isolated namespace.
Do not depend only on a PersistentVolume mounted beside the database. If the cluster, storage class, credentials, or namespace are compromised together, both the production data and “backup” may be lost.
PostgreSQL Backup Strategy
For logical backups, pg_dump backs up one PostgreSQL database, while pg_dumpall covers an entire cluster and preserves global objects such as roles and tablespaces. For individual application databases, use the custom archive format (-Fc): it is compressed, works with pg_restore, and supports selective restore operations.
A basic PostgreSQL backup command is:
pg_dump \
--host="$PGHOST" \
--port="$PGPORT" \
--username="$PGUSER" \
--format=custom \
--no-owner \
--file="/work/${PGDATABASE}-${STAMP}.dump" \
"$PGDATABASE"The --no-owner option helps when restoring into a different environment where source role names do not exist. If your recovery requires roles and tablespaces too, run a separate global metadata export:
pg_dumpall \
--host="$PGHOST" \
--username="$PGUSER" \
--globals-only \
--file="/work/${PGDATABASE}-${STAMP}-globals.sql"For large, high-change PostgreSQL systems, logical dumps may not meet your RPO/RTO. Use physical backups plus WAL archiving through a PostgreSQL-aware solution such as pgBackRest, WAL-G, or an operator-supported backup mechanism. The important distinction is that pg_dump captures a logical snapshot, while point-in-time recovery requires archived WAL files in addition to a suitable base backup.
For self-managed MongoDB, mongodump creates a binary export and mongorestore restores that export to a running MongoDB deployment. Use a single compressed archive for easier handling and avoid a large tree of loose BSON files:
mongodump \
--uri="$MONGODB_URI" \
--archive="/work/mongodb-${STAMP}.archive.gz" \
--gzipmongodump can back up a full deployment, a database, a collection, or query-filtered content, depending on options and operational needs. For a replica set, ensure you understand consistency requirements and test restoration against the exact MongoDB topology you operate. For managed MongoDB Atlas, platform-native continuous backups and point-in-time recovery are often preferable to running mongodump as the sole protection mechanism.
This example uses PostgreSQL. The same container pattern works for MongoDB by replacing the image and backup command. It writes to /work; in production, add an upload command to S3-compatible storage, Azure Blob, GCS, or a hardened backup server.
apiVersion: batch/v1
kind: CronJob
metadata:
name: postgresql-backup
namespace: data
spec:
schedule: "15 2 * * *"
timeZone: "Europe/Bucharest"
concurrencyPolicy: Forbid
startingDeadlineSeconds: 600
successfulJobsHistoryLimit: 3
failedJobsHistoryLimit: 5
jobTemplate:
spec:
backoffLimit: 2
ttlSecondsAfterFinished: 86400
template:
spec:
restartPolicy: Never
serviceAccountName: database-backup
securityContext:
runAsNonRoot: true
runAsUser: 10001
fsGroup: 10001
containers:
- name: backup
image: postgres:16
imagePullPolicy: IfNotPresent
envFrom:
- secretRef:
name: postgres-backup-credentials
command:
- /bin/sh
- -ec
- |
umask 077
STAMP="$(date -u +%Y%m%dT%H%M%SZ)"
OUT="/work/${PGDATABASE}-${STAMP}.dump"
pg_dump \
–host=”$PGHOST” \
–port=”${PGPORT:-5432}” \
–username=”$PGUSER” \
–format=custom \
–no-owner \
–file=”$OUT” \
“$PGDATABASE”
test -s “$OUT”
sha256sum “$OUT” > “${OUT}.sha256”
# Upload OUT and OUT.sha256 to remote immutable storage here.
# Delete local files only after upload verification succeeds.
volumeMounts:
– name: backup-work
mountPath: /work
volumes:
– name: backup-work
emptyDir:
sizeLimit: 20Gi
concurrencyPolicy: Forbid prevents a slow backup from overlapping with the next scheduled run. Set startingDeadlineSeconds so an overdue backup does not start at an operationally unsafe time. CronJobs are appropriate for recurring work, but Kubernetes documents that scheduling can be approximate, so make the job idempotent and record success externally rather than trusting schedule timing alone.
Use a dedicated database account whose permissions are sufficient to read backup data but not to modify production records. Store connection strings, passwords, TLS certificates, and object-storage credentials in a secret-management system; restrict access with namespace RBAC, separate service accounts, and cloud workload identity where available.
Apply these controls:
-
Use TLS for database connections and object-storage uploads.
-
Encrypt backup files before leaving the Pod, or use a destination that provides strong server-side encryption with independently controlled keys.
-
Keep backup storage in a separate account, project, or tenant from the production Kubernetes cluster.
-
Enable versioning and retention locks/object immutability to resist accidental deletion and ransomware.
-
Include a checksum or manifest containing backup timestamp, database/version, artifact size, checksum, and tool version.
-
Never place database URIs, passwords, or cloud keys directly in CronJob YAML, Git repositories, or CI logs.
-
Restrict egress from the backup Pod to only the database endpoint, DNS/time dependencies, and the backup destination.
Retention, Monitoring, and Restores
Define retention before automation. A typical scheme might keep daily backups for 14–30 days, weekly backups for several months, and monthly backups for a year, but the right policy follows business, legal, and storage requirements. Monitor at least: last-success timestamp, job duration, output size, upload success, checksum verification, and age of the newest restorable artifact.
Restore testing should be automated in a disposable namespace: provision an empty Postgres or MongoDB instance, download a recent backup, restore it, run integrity checks and representative application queries, then destroy the environment. PostgreSQL custom-format dumps restore with pg_restore; MongoDB archive dumps restore with mongorestore. A successful restore test is the most meaningful backup metric.mongodb+1
As a practical first step, define whether your most important database needs daily logical recovery or point-in-time recovery, then explain what maximum acceptable data loss in minutes or hours would be for that workload.
[mai mult...]