diff --git a/AGENTS.md b/AGENTS.md index 4148393..c14349c 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -226,7 +226,7 @@ successfully. The first monthly scrub remains a runtime check. read-only ZFS snapshot test restored one file to `/var/tmp`, confirmed matching contents, ownership, mode, mtime and ACL, then removed its temporary copy and on-demand mount. This is a file-level smoke test, not full dataset recovery. An independent USB file restore passed on 2026-09-25 with matching - content and metadata; full disaster recovery remains a separate Priority 2 task. + content and metadata; the later scaled OS-rebuild rehearsal is documented under Priority 2. - [x] Add monitoring and alerting for pool health, scrub/resilver, SMART data, temperatures, free space, snapshot/local-backup growth, Hetzner Storage Box quota, and failed maintenance or backup timers. The half-hourly Atlas health monitor and systemd final-failure hooks are deployed. A live probe @@ -237,8 +237,13 @@ successfully. The first monthly scrub remains a runtime check. on Atlas, but a new real failure notification has not been deliberately triggered. ### Priority 2 - NAS operability and recovery -- [ ] Document and test disaster recovery: rebuild Atlas with Ansible, import the existing pool, restore - from snapshot/USB/Hetzner, preserve Vault and Borg recovery material offline, and define RPO/RTO. +- [x] Document and test disaster recovery in `docs/atlas-recovery.md`: the operator confirmed Vault + and Borg recovery material is available offline; provisional targets are RPO 24h/RTO 72h. On + 2026-09-30 an isolated small Rocky VM was rebuilt with the Atlas Ansible roles, imported its + preserved RAIDZ2 pool without force/rewind, and restored a file from the preserved snapshot; + the second Ansible run was idempotent. Earlier independent production ZFS, USB, and Borg file + restore tests remain separate evidence. A production-size full restore, unclean import, and + measured 24h/72h compliance are not claimed. - [ ] Define a controlled Rocky kernel/OpenZFS update and reboot procedure. - [ ] Add the Atlas-initiated least-privilege Prometheus backup pull: Prometheus exposes only prepared read-only dumps through a dedicated account and Atlas retains the private SSH key, pinned host key, diff --git a/docs/atlas-dr-lab.md b/docs/atlas-dr-lab.md new file mode 100644 index 0000000..93ed96b --- /dev/null +++ b/docs/atlas-dr-lab.md @@ -0,0 +1,74 @@ +# Isolated Atlas DR lab + +This is a **scaled rehearsal**, not a substitute for a full-data restore. The +`atlas-dr-lab` libvirt VM on Ikaros was left **shut off** on 2026-09-30. Its +persistent volumes are in the default libvirt pool: the current 30 GiB OS +volume `atlas-dr-lab-os-rebuild2.qcow2`, the pre-rebuild OS volume +`atlas-dr-lab-os.qcow2`, and four independent 4 GiB +`atlas-dr-lab-data{1,2,3,4}.qcow2` volumes. The VM uses libvirt's `default` +NAT network (last DHCP address `192.168.122.168`), 2 vCPU, and 4 GiB RAM. +The data disks have `virtio-atlasdrdata{1,2,3,4}` serials. No physical disk or +production Atlas storage is attached. The VM has no autostart. + +## Rebuild inputs and isolation + +- Use Rocky's **9.8 GenericCloud Base x86_64** image + `Rocky-9-GenericCloud-Base-9.8-20260525.0.x86_64.qcow2` from + `https://download.rockylinux.org/pub/rocky/9.8/images/x86_64/`. + Verify its `.CHECKSUM` file; the observed SHA-256 was + `92c206cc6f790c61583247eefe87890f8828420662c17cacf247cec78ab4eec8`. +- Use a dedicated lab-only inventory merged **after** the repository + inventory, and always `--limit atlas_dr_lab`. The temporary 2026-09-30 + inventory/playbook and logs are in `/tmp/atlas-dr-lab-image/`; copy a + sanitized inventory to durable private storage before `/tmp` is cleared if + the lab will be repeated. Never reuse `host_vars/atlas.yml`, production + Vault secrets, or production disk by-id paths for the lab. +- The lab host belongs to `platform_rocky` and `atlas`. It uses `dradmin` + (UID/GID 1000) with the operator's **public** SSH key and a random, + unknown password hash, the libvirt DHCP address, pool `zpool`, mount root + `/zpool`, the four `virtio-atlasdrdata*` by-id paths, a 1 GiB backup + reservation, and `rocky_manage_openzfs_repo: true` with only `zfs` in + `host_packages`. The following gates remain false: sharing, firewall, + media stack, ZFS timers, Borg, USB, monitoring, and Prometheus pull. + `atlas_manage_storage` is true. Set `atlas_create_pool: true` **only for the + first disposable pool creation**, then set it false before any later run. +- A minimal lab playbook selects `atlas_dr_lab`, `become: true`, and the + existing `packages_rocky` and `profile_atlas` roles. Use a separate + `ANSIBLE_CONFIG` without the production Vault password script, and keep + host-key checking on with a lab-specific known-hosts file. The 2026-09-30 + runs used `-i ansible/inventory/hosts.yml -i ` and + `--limit atlas_dr_lab` throughout. + +## Rehearsal and narrow checks + +1. Before any pool operation, compare `virsh -c qemu:///system domblklist + atlas-dr-lab` with the four intended qcow2 paths, and in the guest compare + `/dev/disk/by-id/virtio-atlasdrdata*` with `lsblk`. Do not proceed if a + physical disk or production identity appears. +2. For a first-time disposable build only, run the lab playbook with + `--tags pool` and `atlas_create_pool: true`, then immediately set the gate + false. Run the full lab playbook and check `zpool status -P zpool`, + `zfs list -r zpool`, SELinux, and failed systemd units. +3. Write a non-sensitive canary under the lab `/zpool/archive` and snapshot + it. Record the pool GUID and canary SHA-256. Export the lab pool cleanly, + shut down the VM, and replace **only the OS volume** with a fresh verified + Rocky image. Preserve all four data volumes. Reconfigure cloud-init for a + new instance; the seed CD-ROM must use **SATA**. The SCSI seed attachment + tried during this rehearsal was not detected by cloud-init and was + replaced with a SATA attachment before proceeding. +4. On the new OS, apply `packages_rocky` to reinstall OpenZFS. First run + `zpool import -d /dev/disk/by-id` **without importing**, compare GUID and + vdev membership, then use ordinary `zpool import -d /dev/disk/by-id zpool`. + Do not use `-f`, `-F`, `-X`, rollback, or pool creation. +5. Reapply `profile_atlas` with the lab gates and `atlas_create_pool: false`. + Verify the canary, restored snapshot file in an empty temporary directory, + dataset hierarchy, SELinux, and pool health. A second full playbook run + should report `changed=0`. Remove temporary restored files and shut down + the VM after testing. + +The observed 2026-09-30 pool GUID was `8880368391795119587`; the canary +SHA-256 was `949701c7a95fadae1fddc21abe846c4312212dbfeb7477948f3188fc3ec34a78`. +The post-rebuild Ansible run succeeded, a repeat run reported `changed=0`, +12 datasets and the original snapshot were present, and the pool was healthy. +The snapshot-restored file matched content and basic metadata. See +[`atlas-recovery.md`](atlas-recovery.md) for the production runbook and limits. diff --git a/docs/atlas-recovery.md b/docs/atlas-recovery.md new file mode 100644 index 0000000..4b70ff5 --- /dev/null +++ b/docs/atlas-recovery.md @@ -0,0 +1,141 @@ +# Atlas recovery runbook + +This runbook is for a **replacement Rocky Linux 9 installation**, not a normal +playbook run. A scaled whole-OS rebuild with a disposable pool passed in an +isolated VM on 2026-09-30, but no production-size whole-host recovery has been +tested. The existing production pool must be imported, never created or +rewritten. The provisional targets are **RPO 24 hours** +and **RTO 72 hours**, for Archive and Atlas services alike. They are planning +objectives, not demonstrated recovery times. The manual USB cadence may leave +an older copy; a recent Borg archive is needed to meet the RPO after total +pool loss. + +## Before an incident + +- Keep an offline copy of the encrypted Ansible Vault, its unlock material, + the exported Borg repository key, and the Borg passphrase. Do not store + unlock material in this repository or in a recovery command line. + On 2026-09-30 the operator confirmed these are available independently of + Atlas and the Ansible controller; their usability has not been tested here. +- Keep the Atlas installation media and a reproducible checkout of this + repository available independently of Atlas. Record the exact Git revision + used for a successful deployment. +- Record the pool's current disk identities with `zpool status -P zpool` and + `lsblk -o NAME,SIZE,MODEL,SERIAL,FSTYPE,UUID`. Compare these with + `atlas_zpool_disks` before touching a replacement host. The `host_vars` + values are historical identifiers, not evidence that a newly attached disk + is the same device. +- Verify that the latest hourly/daily snapshots, Borg archive, and offline USB + version exist and note their timestamps. A timer being enabled is not proof + that a backup completed. + +## Incident gate + +1. Identify whether the fault is the OS disk, one or more pool disks, accidental + deletion, or an unavailable host. Preserve failed media when possible. +2. Stop writes to affected services and capture the last known good backup + timestamps. Do not run `zpool create`, `zpool destroy`, `zfs rollback`, + `zpool import -F`, `zpool import -X`, `zpool import -f`, or disk formatting + as a diagnostic shortcut. +3. Choose one recovery source below. Do not merge several sources into the + production namespace without comparing their timestamps and content. + +## Rebuild the OS and import the existing pool + +1. Install Rocky Linux 9 on a **separate system disk**. Configure basic network, + SSH, a temporary sudo administrator, SELinux enforcing, and the current + OpenZFS kmod repository. Keep the pool drives untouched. +2. Run read-only identification: `lsblk -f`, `zpool import`, and + `zpool import -d /dev/disk/by-id`. Check the pool GUID, vdev layout, and + stable drive identities against the incident record. If any differ, stop. +3. Import only after matching the expected pool and host ownership. A pool + cleanly exported from the old host can be imported with + `zpool import -d /dev/disk/by-id zpool`. If it reports that the pool is + active elsewhere or needs a rewind/force, stop and investigate rather than + adding flags. Verify with `zpool status -v zpool`, `zfs list -r zpool`, + `zfs get -r mountpoint,canmount zpool`, and `findmnt -R /zpool`. +4. Leave `atlas_create_pool: false`. Ensure `host_vars/atlas.yml` reflects the + replacement host's actual SSH address and disk identities before running + Ansible. Apply `ansible/site.yml --limit atlas` with the bootstrap admin + connection override as documented in the Atlas setup section of README. + This may start shares/services, so keep clients disconnected or services + gated until data and permissions are verified. +5. Check `getenforce`, `zpool status -v zpool`, `systemctl --failed`, SSH, + firewalld, Cockpit, NFS, SMB, and the backup/monitoring timers. Do not + report recovery complete on the basis of Ansible success alone. + +## Choose the data source + +- **Local snapshot, pool intact:** inspect `zfs list -t snapshot -r zpool`. + Mount/access the chosen snapshot read-only and copy selected files to an + empty staging directory; compare content, owner, mode, mtime, and POSIX ACL. + Move into the live namespace only after an operator-approved scope review. + Do not use an automatic rollback: it can discard newer changes in the + dataset and descendants. +- **Offline USB:** verify the configured LUKS and ext4 UUIDs from + `host_vars/atlas.yml` before unlocking. Mount ext4 read-only with `ro,noload`, + use only a published `atlas/latest` version, and restore to an empty staging + directory. Compare checksums and metadata. The USB copy intentionally omits + generic xattrs and SELinux labels; relabel only the restored destination. + Never run the backup service to perform a restore. +- **Hetzner Borg:** use the dedicated pinned host key, repository path, + offline exported recovery key, and Vault-backed passphrase. List archives + and extract a selected archive into an empty staging directory, never the + live `/zpool` tree. A repository check and sample restore were previously + performed; that does not prove this incident's archive is complete. Compare + content and metadata before publication. Avoid `borg break-lock` while any + backup/check job may still be active. + +After publishing restored files, run the explicit Ansible `restorecon` tag only +for the paths actually restored, for example: + +```bash +ansible-playbook ansible/site.yml --limit atlas --tags restorecon \ + -e '{"atlas_restorecon_paths":["/zpool/archive"]}' +``` + +Then check ownership/ACLs, application-specific integrity, SMB/NFS client +access, backup service health, and `zpool status -v zpool`. Reconnect clients +only after these checks pass. Record the last recoverable timestamp (actual +RPO) and elapsed service outage (actual RTO) in the incident log. + +## Scaled isolated rehearsal (2026-09-30) + +The lab setup, repeatable checks, and preserved VM state are recorded in +[`atlas-dr-lab.md`](atlas-dr-lab.md). + +On Ikaros, a local libvirt `atlas-dr-lab` VM used a 30 GiB Rocky 9.8 system +disk and four separate, disposable 4 GiB virtio data disks with stable +`/dev/disk/by-id` identities. The official Rocky cloud image matched its +published SHA-256. The lab inventory was separate from production, used a +fresh lab-only password hash and the operator's public SSH key, and disabled +sharing, Borg, USB backup, monitoring, media services, and the Prometheus pull. +No production disk, Vault secret, or production data was attached or copied. + +1. The existing `packages_rocky` and `profile_atlas` roles installed OpenZFS, + created a RAIDZ2 `zpool` through the explicit one-time pool gate, and built + all 12 declared datasets with a lab-sized 1 GiB backup reservation. The + pool creation gate was set false immediately afterward. +2. A 4 MiB canary file was written under the lab `archive` dataset and a ZFS + snapshot created. The pool was cleanly exported and the VM shut down. +3. Only the system-disk volume was replaced by a fresh Rocky cloud image; + the four virtio data volumes were retained. Ansible reinstalled OpenZFS. + Read-only `zpool import -d /dev/disk/by-id` showed the expected RAIDZ2 + topology and pool GUID `8880368391795119587` before an ordinary import + without `-f`, rewind, or rollback. +4. The imported pool was healthy. The canary SHA-256 matched its pre-rebuild + value. `profile_atlas` completed against the imported pool and a second + run reported `changed=0`. A file restored from the preserved snapshot into + `/var/tmp` matched SHA-256, owner, group, mode, size, and mtime; the temporary + copy was removed. Final checks found SELinux Enforcing, 12 datasets, the + snapshot, no failed units, and a healthy pool. The VM was shut down while + retaining its disposable volumes for a future rehearsal. + +This proves the **sequence** for a cleanly exported, small pool and the tested +Ansible subset, not recovery duration or capacity at 2 TB. The earlier +2026-09-25 independent production ZFS/USB file restores and the earlier Borg +temporary-directory restore remain separate evidence. The VM did not restore +production USB/Borg archives, exercise services with production data, test an +unclean import, or prove the provisional RPO/RTO. Before relying on 24h/72h, +measure a representative full restore and service cutover in a suitably sized +future change window. Never use the production Atlas pool for a rehearsal.