8.2 KiB
Atlas recovery runbook
This runbook is for a replacement Rocky Linux 9 installation, not a normal playbook run. A scaled whole-OS rebuild with a disposable pool passed in an isolated VM on 2026-09-30, but no production-size whole-host recovery has been tested. The existing production pool must be imported, never created or rewritten. The provisional targets are RPO 24 hours and RTO 72 hours, for Archive and Atlas services alike. They are planning objectives, not demonstrated recovery times. The manual USB cadence may leave an older copy; a recent Borg archive is needed to meet the RPO after total pool loss.
Before an incident
- Keep an offline copy of the encrypted Ansible Vault, its unlock material, the exported Borg repository key, and the Borg passphrase. Do not store unlock material in this repository or in a recovery command line. On 2026-09-30 the operator confirmed these are available independently of Atlas and the Ansible controller; their usability has not been tested here.
- Keep the Atlas installation media and a reproducible checkout of this repository available independently of Atlas. Record the exact Git revision used for a successful deployment.
- Record the pool's current disk identities with
zpool status -P zpoolandlsblk -o NAME,SIZE,MODEL,SERIAL,FSTYPE,UUID. Compare these withatlas_zpool_disksbefore touching a replacement host. Thehost_varsvalues are historical identifiers, not evidence that a newly attached disk is the same device. - Verify that the latest hourly/daily snapshots, Borg archive, and offline USB version exist and note their timestamps. A timer being enabled is not proof that a backup completed.
Incident gate
- Identify whether the fault is the OS disk, one or more pool disks, accidental deletion, or an unavailable host. Preserve failed media when possible.
- Stop writes to affected services and capture the last known good backup
timestamps. Do not run
zpool create,zpool destroy,zfs rollback,zpool import -F,zpool import -X,zpool import -f, or disk formatting as a diagnostic shortcut. - Choose one recovery source below. Do not merge several sources into the production namespace without comparing their timestamps and content.
Rebuild the OS and import the existing pool
- Install Rocky Linux 9 on a separate system disk. Configure basic network, SSH, a temporary sudo administrator, SELinux enforcing, and the current OpenZFS kmod repository. Keep the pool drives untouched.
- Run read-only identification:
lsblk -f,zpool import, andzpool import -d /dev/disk/by-id. Check the pool GUID, vdev layout, and stable drive identities against the incident record. If any differ, stop. - Import only after matching the expected pool and host ownership. A pool
cleanly exported from the old host can be imported with
zpool import -d /dev/disk/by-id zpool. If it reports that the pool is active elsewhere or needs a rewind/force, stop and investigate rather than adding flags. Verify withzpool status -v zpool,zfs list -r zpool,zfs get -r mountpoint,canmount zpool, andfindmnt -R /zpool. - Leave
atlas_create_pool: false. Ensurehost_vars/atlas.ymlreflects the replacement host's actual SSH address and disk identities before running Ansible. Applyansible/site.yml --limit atlaswith the bootstrap admin connection override as documented in the Atlas setup section of README. This may start shares/services, so keep clients disconnected or services gated until data and permissions are verified. - Check
getenforce,zpool status -v zpool,systemctl --failed, SSH, firewalld, Cockpit, NFS, SMB, and the backup/monitoring timers. Do not report recovery complete on the basis of Ansible success alone.
Choose the data source
- Local snapshot, pool intact: inspect
zfs list -t snapshot -r zpool. Mount/access the chosen snapshot read-only and copy selected files to an empty staging directory; compare content, owner, mode, mtime, and POSIX ACL. Move into the live namespace only after an operator-approved scope review. Do not use an automatic rollback: it can discard newer changes in the dataset and descendants. - Offline USB: verify the configured LUKS and ext4 UUIDs from
host_vars/atlas.ymlbefore unlocking. Mount ext4 read-only withro,noload, use only a publishedatlas/latestversion, and restore to an empty staging directory. Compare checksums and metadata. The USB copy intentionally omits generic xattrs and SELinux labels; relabel only the restored destination. Never run the backup service to perform a restore. - Hetzner Borg: use the dedicated pinned host key, repository path,
offline exported recovery key, and Vault-backed passphrase. List archives
and extract a selected archive into an empty staging directory, never the
live
/zpooltree. A repository check and sample restore were previously performed; that does not prove this incident's archive is complete. Compare content and metadata before publication. Avoidborg break-lockwhile any backup/check job may still be active.
After publishing restored files, run the explicit Ansible restorecon tag only
for the paths actually restored, for example:
ansible-playbook ansible/site.yml --limit atlas --tags restorecon \
-e '{"atlas_restorecon_paths":["/zpool/archive"]}'
Then check ownership/ACLs, application-specific integrity, SMB/NFS client
access, backup service health, and zpool status -v zpool. Reconnect clients
only after these checks pass. Record the last recoverable timestamp (actual
RPO) and elapsed service outage (actual RTO) in the incident log.
Scaled isolated rehearsal (2026-09-30)
The lab setup, repeatable checks, and preserved VM state are recorded in
atlas-dr-lab.md.
On Ikaros, a local libvirt atlas-dr-lab VM used a 30 GiB Rocky 9.8 system
disk and four separate, disposable 4 GiB virtio data disks with stable
/dev/disk/by-id identities. The official Rocky cloud image matched its
published SHA-256. The lab inventory was separate from production, used a
fresh lab-only password hash and the operator's public SSH key, and disabled
sharing, Borg, USB backup, monitoring, media services, and the Prometheus pull.
No production disk, Vault secret, or production data was attached or copied.
- The existing
packages_rockyandprofile_atlasroles installed OpenZFS, created a RAIDZ2zpoolthrough the explicit one-time pool gate, and built all 12 declared datasets with a lab-sized 1 GiB backup reservation. The pool creation gate was set false immediately afterward. - A 4 MiB canary file was written under the lab
archivedataset and a ZFS snapshot created. The pool was cleanly exported and the VM shut down. - Only the system-disk volume was replaced by a fresh Rocky cloud image;
the four virtio data volumes were retained. Ansible reinstalled OpenZFS.
Read-only
zpool import -d /dev/disk/by-idshowed the expected RAIDZ2 topology and pool GUID8880368391795119587before an ordinary import without-f, rewind, or rollback. - The imported pool was healthy. The canary SHA-256 matched its pre-rebuild
value.
profile_atlascompleted against the imported pool and a second run reportedchanged=0. A file restored from the preserved snapshot into/var/tmpmatched SHA-256, owner, group, mode, size, and mtime; the temporary copy was removed. Final checks found SELinux Enforcing, 12 datasets, the snapshot, no failed units, and a healthy pool. The VM was shut down while retaining its disposable volumes for a future rehearsal.
This proves the sequence for a cleanly exported, small pool and the tested Ansible subset, not recovery duration or capacity at 2 TB. The earlier 2026-09-25 independent production ZFS/USB file restores and the earlier Borg temporary-directory restore remain separate evidence. The VM did not restore production USB/Borg archives, exercise services with production data, test an unclean import, or prove the provisional RPO/RTO. Before relying on 24h/72h, measure a representative full restore and service cutover in a suitably sized future change window. Never use the production Atlas pool for a rehearsal.