mirror of
https://github.com/fscotto/infra.git
synced 2026-10-03 13:29:58 +00:00
Document and rehearse Atlas disaster recovery
This commit is contained in:
141
docs/atlas-recovery.md
Normal file
141
docs/atlas-recovery.md
Normal file
@@ -0,0 +1,141 @@
|
||||
# Atlas recovery runbook
|
||||
|
||||
This runbook is for a **replacement Rocky Linux 9 installation**, not a normal
|
||||
playbook run. A scaled whole-OS rebuild with a disposable pool passed in an
|
||||
isolated VM on 2026-09-30, but no production-size whole-host recovery has been
|
||||
tested. The existing production pool must be imported, never created or
|
||||
rewritten. The provisional targets are **RPO 24 hours**
|
||||
and **RTO 72 hours**, for Archive and Atlas services alike. They are planning
|
||||
objectives, not demonstrated recovery times. The manual USB cadence may leave
|
||||
an older copy; a recent Borg archive is needed to meet the RPO after total
|
||||
pool loss.
|
||||
|
||||
## Before an incident
|
||||
|
||||
- Keep an offline copy of the encrypted Ansible Vault, its unlock material,
|
||||
the exported Borg repository key, and the Borg passphrase. Do not store
|
||||
unlock material in this repository or in a recovery command line.
|
||||
On 2026-09-30 the operator confirmed these are available independently of
|
||||
Atlas and the Ansible controller; their usability has not been tested here.
|
||||
- Keep the Atlas installation media and a reproducible checkout of this
|
||||
repository available independently of Atlas. Record the exact Git revision
|
||||
used for a successful deployment.
|
||||
- Record the pool's current disk identities with `zpool status -P zpool` and
|
||||
`lsblk -o NAME,SIZE,MODEL,SERIAL,FSTYPE,UUID`. Compare these with
|
||||
`atlas_zpool_disks` before touching a replacement host. The `host_vars`
|
||||
values are historical identifiers, not evidence that a newly attached disk
|
||||
is the same device.
|
||||
- Verify that the latest hourly/daily snapshots, Borg archive, and offline USB
|
||||
version exist and note their timestamps. A timer being enabled is not proof
|
||||
that a backup completed.
|
||||
|
||||
## Incident gate
|
||||
|
||||
1. Identify whether the fault is the OS disk, one or more pool disks, accidental
|
||||
deletion, or an unavailable host. Preserve failed media when possible.
|
||||
2. Stop writes to affected services and capture the last known good backup
|
||||
timestamps. Do not run `zpool create`, `zpool destroy`, `zfs rollback`,
|
||||
`zpool import -F`, `zpool import -X`, `zpool import -f`, or disk formatting
|
||||
as a diagnostic shortcut.
|
||||
3. Choose one recovery source below. Do not merge several sources into the
|
||||
production namespace without comparing their timestamps and content.
|
||||
|
||||
## Rebuild the OS and import the existing pool
|
||||
|
||||
1. Install Rocky Linux 9 on a **separate system disk**. Configure basic network,
|
||||
SSH, a temporary sudo administrator, SELinux enforcing, and the current
|
||||
OpenZFS kmod repository. Keep the pool drives untouched.
|
||||
2. Run read-only identification: `lsblk -f`, `zpool import`, and
|
||||
`zpool import -d /dev/disk/by-id`. Check the pool GUID, vdev layout, and
|
||||
stable drive identities against the incident record. If any differ, stop.
|
||||
3. Import only after matching the expected pool and host ownership. A pool
|
||||
cleanly exported from the old host can be imported with
|
||||
`zpool import -d /dev/disk/by-id zpool`. If it reports that the pool is
|
||||
active elsewhere or needs a rewind/force, stop and investigate rather than
|
||||
adding flags. Verify with `zpool status -v zpool`, `zfs list -r zpool`,
|
||||
`zfs get -r mountpoint,canmount zpool`, and `findmnt -R /zpool`.
|
||||
4. Leave `atlas_create_pool: false`. Ensure `host_vars/atlas.yml` reflects the
|
||||
replacement host's actual SSH address and disk identities before running
|
||||
Ansible. Apply `ansible/site.yml --limit atlas` with the bootstrap admin
|
||||
connection override as documented in the Atlas setup section of README.
|
||||
This may start shares/services, so keep clients disconnected or services
|
||||
gated until data and permissions are verified.
|
||||
5. Check `getenforce`, `zpool status -v zpool`, `systemctl --failed`, SSH,
|
||||
firewalld, Cockpit, NFS, SMB, and the backup/monitoring timers. Do not
|
||||
report recovery complete on the basis of Ansible success alone.
|
||||
|
||||
## Choose the data source
|
||||
|
||||
- **Local snapshot, pool intact:** inspect `zfs list -t snapshot -r zpool`.
|
||||
Mount/access the chosen snapshot read-only and copy selected files to an
|
||||
empty staging directory; compare content, owner, mode, mtime, and POSIX ACL.
|
||||
Move into the live namespace only after an operator-approved scope review.
|
||||
Do not use an automatic rollback: it can discard newer changes in the
|
||||
dataset and descendants.
|
||||
- **Offline USB:** verify the configured LUKS and ext4 UUIDs from
|
||||
`host_vars/atlas.yml` before unlocking. Mount ext4 read-only with `ro,noload`,
|
||||
use only a published `atlas/latest` version, and restore to an empty staging
|
||||
directory. Compare checksums and metadata. The USB copy intentionally omits
|
||||
generic xattrs and SELinux labels; relabel only the restored destination.
|
||||
Never run the backup service to perform a restore.
|
||||
- **Hetzner Borg:** use the dedicated pinned host key, repository path,
|
||||
offline exported recovery key, and Vault-backed passphrase. List archives
|
||||
and extract a selected archive into an empty staging directory, never the
|
||||
live `/zpool` tree. A repository check and sample restore were previously
|
||||
performed; that does not prove this incident's archive is complete. Compare
|
||||
content and metadata before publication. Avoid `borg break-lock` while any
|
||||
backup/check job may still be active.
|
||||
|
||||
After publishing restored files, run the explicit Ansible `restorecon` tag only
|
||||
for the paths actually restored, for example:
|
||||
|
||||
```bash
|
||||
ansible-playbook ansible/site.yml --limit atlas --tags restorecon \
|
||||
-e '{"atlas_restorecon_paths":["/zpool/archive"]}'
|
||||
```
|
||||
|
||||
Then check ownership/ACLs, application-specific integrity, SMB/NFS client
|
||||
access, backup service health, and `zpool status -v zpool`. Reconnect clients
|
||||
only after these checks pass. Record the last recoverable timestamp (actual
|
||||
RPO) and elapsed service outage (actual RTO) in the incident log.
|
||||
|
||||
## Scaled isolated rehearsal (2026-09-30)
|
||||
|
||||
The lab setup, repeatable checks, and preserved VM state are recorded in
|
||||
[`atlas-dr-lab.md`](atlas-dr-lab.md).
|
||||
|
||||
On Ikaros, a local libvirt `atlas-dr-lab` VM used a 30 GiB Rocky 9.8 system
|
||||
disk and four separate, disposable 4 GiB virtio data disks with stable
|
||||
`/dev/disk/by-id` identities. The official Rocky cloud image matched its
|
||||
published SHA-256. The lab inventory was separate from production, used a
|
||||
fresh lab-only password hash and the operator's public SSH key, and disabled
|
||||
sharing, Borg, USB backup, monitoring, media services, and the Prometheus pull.
|
||||
No production disk, Vault secret, or production data was attached or copied.
|
||||
|
||||
1. The existing `packages_rocky` and `profile_atlas` roles installed OpenZFS,
|
||||
created a RAIDZ2 `zpool` through the explicit one-time pool gate, and built
|
||||
all 12 declared datasets with a lab-sized 1 GiB backup reservation. The
|
||||
pool creation gate was set false immediately afterward.
|
||||
2. A 4 MiB canary file was written under the lab `archive` dataset and a ZFS
|
||||
snapshot created. The pool was cleanly exported and the VM shut down.
|
||||
3. Only the system-disk volume was replaced by a fresh Rocky cloud image;
|
||||
the four virtio data volumes were retained. Ansible reinstalled OpenZFS.
|
||||
Read-only `zpool import -d /dev/disk/by-id` showed the expected RAIDZ2
|
||||
topology and pool GUID `8880368391795119587` before an ordinary import
|
||||
without `-f`, rewind, or rollback.
|
||||
4. The imported pool was healthy. The canary SHA-256 matched its pre-rebuild
|
||||
value. `profile_atlas` completed against the imported pool and a second
|
||||
run reported `changed=0`. A file restored from the preserved snapshot into
|
||||
`/var/tmp` matched SHA-256, owner, group, mode, size, and mtime; the temporary
|
||||
copy was removed. Final checks found SELinux Enforcing, 12 datasets, the
|
||||
snapshot, no failed units, and a healthy pool. The VM was shut down while
|
||||
retaining its disposable volumes for a future rehearsal.
|
||||
|
||||
This proves the **sequence** for a cleanly exported, small pool and the tested
|
||||
Ansible subset, not recovery duration or capacity at 2 TB. The earlier
|
||||
2026-09-25 independent production ZFS/USB file restores and the earlier Borg
|
||||
temporary-directory restore remain separate evidence. The VM did not restore
|
||||
production USB/Borg archives, exercise services with production data, test an
|
||||
unclean import, or prove the provisional RPO/RTO. Before relying on 24h/72h,
|
||||
measure a representative full restore and service cutover in a suitably sized
|
||||
future change window. Never use the production Atlas pool for a rehearsal.
|
||||
Reference in New Issue
Block a user