Document and rehearse Atlas disaster recovery

This commit is contained in:
Fabio Scotto di Santolo
2026-09-30 21:19:28 +02:00
parent 702283b430
commit 9798fe3a12
3 changed files with 223 additions and 3 deletions

View File

@@ -226,7 +226,7 @@ successfully. The first monthly scrub remains a runtime check.
read-only ZFS snapshot test restored one file to `/var/tmp`, confirmed matching contents, ownership,
mode, mtime and ACL, then removed its temporary copy and on-demand mount. This is a file-level smoke
test, not full dataset recovery. An independent USB file restore passed on 2026-09-25 with matching
content and metadata; full disaster recovery remains a separate Priority 2 task.
content and metadata; the later scaled OS-rebuild rehearsal is documented under Priority 2.
- [x] Add monitoring and alerting for pool health, scrub/resilver, SMART data, temperatures, free space,
snapshot/local-backup growth, Hetzner Storage Box quota, and failed maintenance or backup timers.
The half-hourly Atlas health monitor and systemd final-failure hooks are deployed. A live probe
@@ -237,8 +237,13 @@ successfully. The first monthly scrub remains a runtime check.
on Atlas, but a new real failure notification has not been deliberately triggered.
### Priority 2 - NAS operability and recovery
- [ ] Document and test disaster recovery: rebuild Atlas with Ansible, import the existing pool, restore
from snapshot/USB/Hetzner, preserve Vault and Borg recovery material offline, and define RPO/RTO.
- [x] Document and test disaster recovery in `docs/atlas-recovery.md`: the operator confirmed Vault
and Borg recovery material is available offline; provisional targets are RPO 24h/RTO 72h. On
2026-09-30 an isolated small Rocky VM was rebuilt with the Atlas Ansible roles, imported its
preserved RAIDZ2 pool without force/rewind, and restored a file from the preserved snapshot;
the second Ansible run was idempotent. Earlier independent production ZFS, USB, and Borg file
restore tests remain separate evidence. A production-size full restore, unclean import, and
measured 24h/72h compliance are not claimed.
- [ ] Define a controlled Rocky kernel/OpenZFS update and reboot procedure.
- [ ] Add the Atlas-initiated least-privilege Prometheus backup pull: Prometheus exposes only prepared
read-only dumps through a dedicated account and Atlas retains the private SSH key, pinned host key,

74
docs/atlas-dr-lab.md Normal file
View File

@@ -0,0 +1,74 @@
# Isolated Atlas DR lab
This is a **scaled rehearsal**, not a substitute for a full-data restore. The
`atlas-dr-lab` libvirt VM on Ikaros was left **shut off** on 2026-09-30. Its
persistent volumes are in the default libvirt pool: the current 30 GiB OS
volume `atlas-dr-lab-os-rebuild2.qcow2`, the pre-rebuild OS volume
`atlas-dr-lab-os.qcow2`, and four independent 4 GiB
`atlas-dr-lab-data{1,2,3,4}.qcow2` volumes. The VM uses libvirt's `default`
NAT network (last DHCP address `192.168.122.168`), 2 vCPU, and 4 GiB RAM.
The data disks have `virtio-atlasdrdata{1,2,3,4}` serials. No physical disk or
production Atlas storage is attached. The VM has no autostart.
## Rebuild inputs and isolation
- Use Rocky's **9.8 GenericCloud Base x86_64** image
`Rocky-9-GenericCloud-Base-9.8-20260525.0.x86_64.qcow2` from
`https://download.rockylinux.org/pub/rocky/9.8/images/x86_64/`.
Verify its `.CHECKSUM` file; the observed SHA-256 was
`92c206cc6f790c61583247eefe87890f8828420662c17cacf247cec78ab4eec8`.
- Use a dedicated lab-only inventory merged **after** the repository
inventory, and always `--limit atlas_dr_lab`. The temporary 2026-09-30
inventory/playbook and logs are in `/tmp/atlas-dr-lab-image/`; copy a
sanitized inventory to durable private storage before `/tmp` is cleared if
the lab will be repeated. Never reuse `host_vars/atlas.yml`, production
Vault secrets, or production disk by-id paths for the lab.
- The lab host belongs to `platform_rocky` and `atlas`. It uses `dradmin`
(UID/GID 1000) with the operator's **public** SSH key and a random,
unknown password hash, the libvirt DHCP address, pool `zpool`, mount root
`/zpool`, the four `virtio-atlasdrdata*` by-id paths, a 1 GiB backup
reservation, and `rocky_manage_openzfs_repo: true` with only `zfs` in
`host_packages`. The following gates remain false: sharing, firewall,
media stack, ZFS timers, Borg, USB, monitoring, and Prometheus pull.
`atlas_manage_storage` is true. Set `atlas_create_pool: true` **only for the
first disposable pool creation**, then set it false before any later run.
- A minimal lab playbook selects `atlas_dr_lab`, `become: true`, and the
existing `packages_rocky` and `profile_atlas` roles. Use a separate
`ANSIBLE_CONFIG` without the production Vault password script, and keep
host-key checking on with a lab-specific known-hosts file. The 2026-09-30
runs used `-i ansible/inventory/hosts.yml -i <lab-inventory.yml>` and
`--limit atlas_dr_lab` throughout.
## Rehearsal and narrow checks
1. Before any pool operation, compare `virsh -c qemu:///system domblklist
atlas-dr-lab` with the four intended qcow2 paths, and in the guest compare
`/dev/disk/by-id/virtio-atlasdrdata*` with `lsblk`. Do not proceed if a
physical disk or production identity appears.
2. For a first-time disposable build only, run the lab playbook with
`--tags pool` and `atlas_create_pool: true`, then immediately set the gate
false. Run the full lab playbook and check `zpool status -P zpool`,
`zfs list -r zpool`, SELinux, and failed systemd units.
3. Write a non-sensitive canary under the lab `/zpool/archive` and snapshot
it. Record the pool GUID and canary SHA-256. Export the lab pool cleanly,
shut down the VM, and replace **only the OS volume** with a fresh verified
Rocky image. Preserve all four data volumes. Reconfigure cloud-init for a
new instance; the seed CD-ROM must use **SATA**. The SCSI seed attachment
tried during this rehearsal was not detected by cloud-init and was
replaced with a SATA attachment before proceeding.
4. On the new OS, apply `packages_rocky` to reinstall OpenZFS. First run
`zpool import -d /dev/disk/by-id` **without importing**, compare GUID and
vdev membership, then use ordinary `zpool import -d /dev/disk/by-id zpool`.
Do not use `-f`, `-F`, `-X`, rollback, or pool creation.
5. Reapply `profile_atlas` with the lab gates and `atlas_create_pool: false`.
Verify the canary, restored snapshot file in an empty temporary directory,
dataset hierarchy, SELinux, and pool health. A second full playbook run
should report `changed=0`. Remove temporary restored files and shut down
the VM after testing.
The observed 2026-09-30 pool GUID was `8880368391795119587`; the canary
SHA-256 was `949701c7a95fadae1fddc21abe846c4312212dbfeb7477948f3188fc3ec34a78`.
The post-rebuild Ansible run succeeded, a repeat run reported `changed=0`,
12 datasets and the original snapshot were present, and the pool was healthy.
The snapshot-restored file matched content and basic metadata. See
[`atlas-recovery.md`](atlas-recovery.md) for the production runbook and limits.

141
docs/atlas-recovery.md Normal file
View File

@@ -0,0 +1,141 @@
# Atlas recovery runbook
This runbook is for a **replacement Rocky Linux 9 installation**, not a normal
playbook run. A scaled whole-OS rebuild with a disposable pool passed in an
isolated VM on 2026-09-30, but no production-size whole-host recovery has been
tested. The existing production pool must be imported, never created or
rewritten. The provisional targets are **RPO 24 hours**
and **RTO 72 hours**, for Archive and Atlas services alike. They are planning
objectives, not demonstrated recovery times. The manual USB cadence may leave
an older copy; a recent Borg archive is needed to meet the RPO after total
pool loss.
## Before an incident
- Keep an offline copy of the encrypted Ansible Vault, its unlock material,
the exported Borg repository key, and the Borg passphrase. Do not store
unlock material in this repository or in a recovery command line.
On 2026-09-30 the operator confirmed these are available independently of
Atlas and the Ansible controller; their usability has not been tested here.
- Keep the Atlas installation media and a reproducible checkout of this
repository available independently of Atlas. Record the exact Git revision
used for a successful deployment.
- Record the pool's current disk identities with `zpool status -P zpool` and
`lsblk -o NAME,SIZE,MODEL,SERIAL,FSTYPE,UUID`. Compare these with
`atlas_zpool_disks` before touching a replacement host. The `host_vars`
values are historical identifiers, not evidence that a newly attached disk
is the same device.
- Verify that the latest hourly/daily snapshots, Borg archive, and offline USB
version exist and note their timestamps. A timer being enabled is not proof
that a backup completed.
## Incident gate
1. Identify whether the fault is the OS disk, one or more pool disks, accidental
deletion, or an unavailable host. Preserve failed media when possible.
2. Stop writes to affected services and capture the last known good backup
timestamps. Do not run `zpool create`, `zpool destroy`, `zfs rollback`,
`zpool import -F`, `zpool import -X`, `zpool import -f`, or disk formatting
as a diagnostic shortcut.
3. Choose one recovery source below. Do not merge several sources into the
production namespace without comparing their timestamps and content.
## Rebuild the OS and import the existing pool
1. Install Rocky Linux 9 on a **separate system disk**. Configure basic network,
SSH, a temporary sudo administrator, SELinux enforcing, and the current
OpenZFS kmod repository. Keep the pool drives untouched.
2. Run read-only identification: `lsblk -f`, `zpool import`, and
`zpool import -d /dev/disk/by-id`. Check the pool GUID, vdev layout, and
stable drive identities against the incident record. If any differ, stop.
3. Import only after matching the expected pool and host ownership. A pool
cleanly exported from the old host can be imported with
`zpool import -d /dev/disk/by-id zpool`. If it reports that the pool is
active elsewhere or needs a rewind/force, stop and investigate rather than
adding flags. Verify with `zpool status -v zpool`, `zfs list -r zpool`,
`zfs get -r mountpoint,canmount zpool`, and `findmnt -R /zpool`.
4. Leave `atlas_create_pool: false`. Ensure `host_vars/atlas.yml` reflects the
replacement host's actual SSH address and disk identities before running
Ansible. Apply `ansible/site.yml --limit atlas` with the bootstrap admin
connection override as documented in the Atlas setup section of README.
This may start shares/services, so keep clients disconnected or services
gated until data and permissions are verified.
5. Check `getenforce`, `zpool status -v zpool`, `systemctl --failed`, SSH,
firewalld, Cockpit, NFS, SMB, and the backup/monitoring timers. Do not
report recovery complete on the basis of Ansible success alone.
## Choose the data source
- **Local snapshot, pool intact:** inspect `zfs list -t snapshot -r zpool`.
Mount/access the chosen snapshot read-only and copy selected files to an
empty staging directory; compare content, owner, mode, mtime, and POSIX ACL.
Move into the live namespace only after an operator-approved scope review.
Do not use an automatic rollback: it can discard newer changes in the
dataset and descendants.
- **Offline USB:** verify the configured LUKS and ext4 UUIDs from
`host_vars/atlas.yml` before unlocking. Mount ext4 read-only with `ro,noload`,
use only a published `atlas/latest` version, and restore to an empty staging
directory. Compare checksums and metadata. The USB copy intentionally omits
generic xattrs and SELinux labels; relabel only the restored destination.
Never run the backup service to perform a restore.
- **Hetzner Borg:** use the dedicated pinned host key, repository path,
offline exported recovery key, and Vault-backed passphrase. List archives
and extract a selected archive into an empty staging directory, never the
live `/zpool` tree. A repository check and sample restore were previously
performed; that does not prove this incident's archive is complete. Compare
content and metadata before publication. Avoid `borg break-lock` while any
backup/check job may still be active.
After publishing restored files, run the explicit Ansible `restorecon` tag only
for the paths actually restored, for example:
```bash
ansible-playbook ansible/site.yml --limit atlas --tags restorecon \
-e '{"atlas_restorecon_paths":["/zpool/archive"]}'
```
Then check ownership/ACLs, application-specific integrity, SMB/NFS client
access, backup service health, and `zpool status -v zpool`. Reconnect clients
only after these checks pass. Record the last recoverable timestamp (actual
RPO) and elapsed service outage (actual RTO) in the incident log.
## Scaled isolated rehearsal (2026-09-30)
The lab setup, repeatable checks, and preserved VM state are recorded in
[`atlas-dr-lab.md`](atlas-dr-lab.md).
On Ikaros, a local libvirt `atlas-dr-lab` VM used a 30 GiB Rocky 9.8 system
disk and four separate, disposable 4 GiB virtio data disks with stable
`/dev/disk/by-id` identities. The official Rocky cloud image matched its
published SHA-256. The lab inventory was separate from production, used a
fresh lab-only password hash and the operator's public SSH key, and disabled
sharing, Borg, USB backup, monitoring, media services, and the Prometheus pull.
No production disk, Vault secret, or production data was attached or copied.
1. The existing `packages_rocky` and `profile_atlas` roles installed OpenZFS,
created a RAIDZ2 `zpool` through the explicit one-time pool gate, and built
all 12 declared datasets with a lab-sized 1 GiB backup reservation. The
pool creation gate was set false immediately afterward.
2. A 4 MiB canary file was written under the lab `archive` dataset and a ZFS
snapshot created. The pool was cleanly exported and the VM shut down.
3. Only the system-disk volume was replaced by a fresh Rocky cloud image;
the four virtio data volumes were retained. Ansible reinstalled OpenZFS.
Read-only `zpool import -d /dev/disk/by-id` showed the expected RAIDZ2
topology and pool GUID `8880368391795119587` before an ordinary import
without `-f`, rewind, or rollback.
4. The imported pool was healthy. The canary SHA-256 matched its pre-rebuild
value. `profile_atlas` completed against the imported pool and a second
run reported `changed=0`. A file restored from the preserved snapshot into
`/var/tmp` matched SHA-256, owner, group, mode, size, and mtime; the temporary
copy was removed. Final checks found SELinux Enforcing, 12 datasets, the
snapshot, no failed units, and a healthy pool. The VM was shut down while
retaining its disposable volumes for a future rehearsal.
This proves the **sequence** for a cleanly exported, small pool and the tested
Ansible subset, not recovery duration or capacity at 2 TB. The earlier
2026-09-25 independent production ZFS/USB file restores and the earlier Borg
temporary-directory restore remain separate evidence. The VM did not restore
production USB/Borg archives, exercise services with production data, test an
unclean import, or prove the provisional RPO/RTO. Before relying on 24h/72h,
measure a representative full restore and service cutover in a suitably sized
future change window. Never use the production Atlas pool for a rehearsal.