Backups & Recovery¶
Every mechanism below is described by what it can get back, and by what it cannot.
Current state¶
| Data | Backed up | How |
|---|---|---|
| Kubernetes objects | Nightly, 02:00 | Velero to the Ceph object store, 14 day TTL |
| Ceph RBD volumes (PVCs) | Nightly, 02:00 | Velero CSI snapshot, moved into the object store |
| Grafana dashboards | Nightly | Its PVC is covered by the above |
| etcd (raw) | Nightly, 01:00 | CronJob to the object store, last 14 kept |
| OpenBao secrets | Manual | Raft snapshot |
| OpenBao unseal keys | Manual, off-cluster | Printed once at bao operator init |
| Prometheus metrics | No | 10 day retention, then gone |
Everything in payload/ |
Yes | It is in Git |
What is not covered¶
Every automated backup lands in the same cluster's Ceph. That protects
against operator error (a deleted PVC, a bad prune, a corrupted database) and
not against losing the cluster, because the backups go with it. Off-site
replication is the missing piece: RGW supports bucket replication and Velero a
second BackupStorageLocation, and neither is configured. Until then, disaster
recovery is Git, the OpenBao unseal keys, and a copy of the latest OpenBao
and etcd snapshots kept off-cluster; the nightly backups protect against
mistakes.
Velero's PrometheusRule (VeleroBackupFailures critical,
VeleroBackupPartialFailures warning) is what notices a backup that silently
stopped, and it only reaches whoever Alertmanager is configured to tell — see
Monitoring.
Velero¶
Velero backs up Kubernetes objects and volume data nightly. Chart values and why each is set are in Velero.
| Property | Value |
|---|---|
| Schedule | 0 2 * * *, TTL 336h |
| Scope | All namespaces except kube-system (rebuilt from Git, and large), minus events (they expire anyway) |
| Destination | S3 bucket velero in the Ceph object store |
| Volume data | CSI snapshot, streamed into the bucket by the data mover (Kopia), snapshot deleted; the durable copy is the one in the bucket |
The velero CLI defaults to the velero namespace; this install is in
backup, so pass -n backup every time.
kubectl -n backup get backups.velero.io
kubectl -n backup get backupstoragelocation # should be Available
velero -n backup backup create manual-$(date +%s) --include-namespaces my-app
Backups in PartiallyFailed with no volume data mean Velero found no
VolumeSnapshotClass labelled velero.io/csi-volumesnapshot-class: "true"
and skipped the volumes silently — check
payload/platform/backup/volumesnapshotclass.yaml is applied.
etcd¶
Velero restores through the API server, which is no use with no API server
left (lost quorum, a corrupted data directory). For that a CronJob in
payload/platform/backup/etcd-backup.yaml snapshots etcd nightly, an hour
before Velero runs.
| Property | Value |
|---|---|
| Schedule | 0 1 * * * |
| Where it runs | Any control-plane node, on the host network: etcd listens on 127.0.0.1 and its client certs are on the node |
| Verification | etcdutl snapshot status before upload, so a truncated snapshot fails the job instead of replacing a good backup (etcdctl snapshot status was removed in etcd 3.6) |
| Destination | S3 bucket etcd-backup, separate from Velero's because it is restored by different means; newest 14 kept |
Two settings there are easy to undo by accident: the DNS policy for a
host-network pod (the node's resolv.conf cannot resolve the RGW Service), and
the AWS CLI config file selecting path-style addressing (no environment
variable sets it).
Taking one by hand¶
kubectl -n kube-system exec -it etcd-<node> -- etcdctl \
--cacert /etc/kubernetes/pki/etcd/ca.crt \
--cert /etc/kubernetes/pki/etcd/server.crt \
--key /etc/kubernetes/pki/etcd/server.key \
snapshot save /var/lib/etcd/snapshot.db
kubectl cp kube-system/etcd-<node>:/var/lib/etcd/snapshot.db ./etcd-snapshot.db
Copy it off the cluster immediately.
OpenBao¶
kubectl -n openbao exec -it openbao-0 -- bao operator raft snapshot save /tmp/snapshot.bao
kubectl -n openbao cp openbao-0:/tmp/snapshot.bao ./openbao-snapshot.bao
The snapshot holds all KV data plus policies, roles and mounts. It does not hold the unseal keys and is useless without them — see OpenBao.
Ceph volumes¶
PVCs are covered by Velero. Ceph's own replication is not a backup:
it spreads each block across OSDs, which survives a disk or node failing and
nothing else. Almost every Application runs with prune: true (kube-vip and
security are the deliberate exceptions), so removing a
PersistentVolumeClaim from Git deletes the volume.
Restoring¶
The commands are the upstream-documented ones for the versions pinned in
payload/platform/; the etcd sequence has not been exercised on this cluster.
Velero restore¶
-
Find the backup.
-
Restore from it, scoped to what was lost; without
--include-namespaceseverything in the backup is restored, and existing objects are skipped, not overwritten. -
Watch it finish and read the warnings.
OpenBao snapshot restore¶
Needs an initialised, unsealed cluster and the root token. -force is what
lets a snapshot from a different cluster (a fresh bao operator init) load;
afterwards the pods seal and want the snapshot's keys, not the new cluster's.
-
Copy the snapshot in.
-
Restore it.
-
Unseal every replica with the snapshot's key shares and confirm.
etcd restore¶
Every etcd member restores from the same snapshot, then all three static pods
come back together. Flatcar ships no etcdutl, so run it from the etcd image
the CronJob uses. Upstream:
Restoring an etcd cluster.
- Copy the snapshot to every control-plane node as
/var/lib/etcd-restore/snapshot.db. -
On every control-plane node, stop the API server and etcd by moving their static pod manifests out, and wait until neither container is listed.
-
On every control-plane node, restore into a new data directory with the
--name,--initial-clusterand--initial-advertise-peer-urlsvalues from that node's/etc/kubernetes/manifests/etcd.yaml(the image tag is inpayload/platform/backup/etcd-backup.yaml).sudo ctr -n k8s.io run --rm \ --mount type=bind,src=/var/lib/etcd-restore,dst=/restore,options=rbind:rw \ registry.k8s.io/etcd:<tag> etcd-restore \ etcdutl snapshot restore /restore/snapshot.db --data-dir /restore/etcd \ --name <node> --initial-cluster <name=https://ip:2380,...> \ --initial-advertise-peer-urls https://<node-ip>:2380 -
Swap the data directory in, keeping the old one and its ownership.
-
Put the manifests back on every node, then confirm.
Rebuild from scratch¶
- Have the Git repository, the OpenBao unseal keys and root token, and an OpenBao snapshot; an etcd snapshot is optional and skips re-issuing certificates.
- Provision the cluster per the Quickstart from step 1.
- When the rollout pauses at
05-secrets, eithermake bao-initand restore the OpenBao snapshot as above, ormake bao-initand re-enter every secret withmake bao-secrets. - Once the platform is Synced and Healthy, restore workloads with a Velero restore from the copied-off backup.