mirror of
https://github.com/fscotto/infra.git
synced 2026-10-03 21:39:50 +00:00
142 lines
8.2 KiB
Markdown
142 lines
8.2 KiB
Markdown
# Atlas recovery runbook
|
|
|
|
This runbook is for a **replacement Rocky Linux 9 installation**, not a normal
|
|
playbook run. A scaled whole-OS rebuild with a disposable pool passed in an
|
|
isolated VM on 2026-09-30, but no production-size whole-host recovery has been
|
|
tested. The existing production pool must be imported, never created or
|
|
rewritten. The provisional targets are **RPO 24 hours**
|
|
and **RTO 72 hours**, for Archive and Atlas services alike. They are planning
|
|
objectives, not demonstrated recovery times. The manual USB cadence may leave
|
|
an older copy; a recent Borg archive is needed to meet the RPO after total
|
|
pool loss.
|
|
|
|
## Before an incident
|
|
|
|
- Keep an offline copy of the encrypted Ansible Vault, its unlock material,
|
|
the exported Borg repository key, and the Borg passphrase. Do not store
|
|
unlock material in this repository or in a recovery command line.
|
|
On 2026-09-30 the operator confirmed these are available independently of
|
|
Atlas and the Ansible controller; their usability has not been tested here.
|
|
- Keep the Atlas installation media and a reproducible checkout of this
|
|
repository available independently of Atlas. Record the exact Git revision
|
|
used for a successful deployment.
|
|
- Record the pool's current disk identities with `zpool status -P zpool` and
|
|
`lsblk -o NAME,SIZE,MODEL,SERIAL,FSTYPE,UUID`. Compare these with
|
|
`atlas_zpool_disks` before touching a replacement host. The `host_vars`
|
|
values are historical identifiers, not evidence that a newly attached disk
|
|
is the same device.
|
|
- Verify that the latest hourly/daily snapshots, Borg archive, and offline USB
|
|
version exist and note their timestamps. A timer being enabled is not proof
|
|
that a backup completed.
|
|
|
|
## Incident gate
|
|
|
|
1. Identify whether the fault is the OS disk, one or more pool disks, accidental
|
|
deletion, or an unavailable host. Preserve failed media when possible.
|
|
2. Stop writes to affected services and capture the last known good backup
|
|
timestamps. Do not run `zpool create`, `zpool destroy`, `zfs rollback`,
|
|
`zpool import -F`, `zpool import -X`, `zpool import -f`, or disk formatting
|
|
as a diagnostic shortcut.
|
|
3. Choose one recovery source below. Do not merge several sources into the
|
|
production namespace without comparing their timestamps and content.
|
|
|
|
## Rebuild the OS and import the existing pool
|
|
|
|
1. Install Rocky Linux 9 on a **separate system disk**. Configure basic network,
|
|
SSH, a temporary sudo administrator, SELinux enforcing, and the current
|
|
OpenZFS kmod repository. Keep the pool drives untouched.
|
|
2. Run read-only identification: `lsblk -f`, `zpool import`, and
|
|
`zpool import -d /dev/disk/by-id`. Check the pool GUID, vdev layout, and
|
|
stable drive identities against the incident record. If any differ, stop.
|
|
3. Import only after matching the expected pool and host ownership. A pool
|
|
cleanly exported from the old host can be imported with
|
|
`zpool import -d /dev/disk/by-id zpool`. If it reports that the pool is
|
|
active elsewhere or needs a rewind/force, stop and investigate rather than
|
|
adding flags. Verify with `zpool status -v zpool`, `zfs list -r zpool`,
|
|
`zfs get -r mountpoint,canmount zpool`, and `findmnt -R /zpool`.
|
|
4. Leave `atlas_create_pool: false`. Ensure `host_vars/atlas.yml` reflects the
|
|
replacement host's actual SSH address and disk identities before running
|
|
Ansible. Apply `ansible/site.yml --limit atlas` with the bootstrap admin
|
|
connection override as documented in the Atlas setup section of README.
|
|
This may start shares/services, so keep clients disconnected or services
|
|
gated until data and permissions are verified.
|
|
5. Check `getenforce`, `zpool status -v zpool`, `systemctl --failed`, SSH,
|
|
firewalld, Cockpit, NFS, SMB, and the backup/monitoring timers. Do not
|
|
report recovery complete on the basis of Ansible success alone.
|
|
|
|
## Choose the data source
|
|
|
|
- **Local snapshot, pool intact:** inspect `zfs list -t snapshot -r zpool`.
|
|
Mount/access the chosen snapshot read-only and copy selected files to an
|
|
empty staging directory; compare content, owner, mode, mtime, and POSIX ACL.
|
|
Move into the live namespace only after an operator-approved scope review.
|
|
Do not use an automatic rollback: it can discard newer changes in the
|
|
dataset and descendants.
|
|
- **Offline USB:** verify the configured LUKS and ext4 UUIDs from
|
|
`host_vars/atlas.yml` before unlocking. Mount ext4 read-only with `ro,noload`,
|
|
use only a published `atlas/latest` version, and restore to an empty staging
|
|
directory. Compare checksums and metadata. The USB copy intentionally omits
|
|
generic xattrs and SELinux labels; relabel only the restored destination.
|
|
Never run the backup service to perform a restore.
|
|
- **Hetzner Borg:** use the dedicated pinned host key, repository path,
|
|
offline exported recovery key, and Vault-backed passphrase. List archives
|
|
and extract a selected archive into an empty staging directory, never the
|
|
live `/zpool` tree. A repository check and sample restore were previously
|
|
performed; that does not prove this incident's archive is complete. Compare
|
|
content and metadata before publication. Avoid `borg break-lock` while any
|
|
backup/check job may still be active.
|
|
|
|
After publishing restored files, run the explicit Ansible `restorecon` tag only
|
|
for the paths actually restored, for example:
|
|
|
|
```bash
|
|
ansible-playbook ansible/site.yml --limit atlas --tags restorecon \
|
|
-e '{"atlas_restorecon_paths":["/zpool/archive"]}'
|
|
```
|
|
|
|
Then check ownership/ACLs, application-specific integrity, SMB/NFS client
|
|
access, backup service health, and `zpool status -v zpool`. Reconnect clients
|
|
only after these checks pass. Record the last recoverable timestamp (actual
|
|
RPO) and elapsed service outage (actual RTO) in the incident log.
|
|
|
|
## Scaled isolated rehearsal (2026-09-30)
|
|
|
|
The lab setup, repeatable checks, and preserved VM state are recorded in
|
|
[`atlas-dr-lab.md`](atlas-dr-lab.md).
|
|
|
|
On Ikaros, a local libvirt `atlas-dr-lab` VM used a 30 GiB Rocky 9.8 system
|
|
disk and four separate, disposable 4 GiB virtio data disks with stable
|
|
`/dev/disk/by-id` identities. The official Rocky cloud image matched its
|
|
published SHA-256. The lab inventory was separate from production, used a
|
|
fresh lab-only password hash and the operator's public SSH key, and disabled
|
|
sharing, Borg, USB backup, monitoring, media services, and the Prometheus pull.
|
|
No production disk, Vault secret, or production data was attached or copied.
|
|
|
|
1. The existing `packages_rocky` and `profile_atlas` roles installed OpenZFS,
|
|
created a RAIDZ2 `zpool` through the explicit one-time pool gate, and built
|
|
all 12 declared datasets with a lab-sized 1 GiB backup reservation. The
|
|
pool creation gate was set false immediately afterward.
|
|
2. A 4 MiB canary file was written under the lab `archive` dataset and a ZFS
|
|
snapshot created. The pool was cleanly exported and the VM shut down.
|
|
3. Only the system-disk volume was replaced by a fresh Rocky cloud image;
|
|
the four virtio data volumes were retained. Ansible reinstalled OpenZFS.
|
|
Read-only `zpool import -d /dev/disk/by-id` showed the expected RAIDZ2
|
|
topology and pool GUID `8880368391795119587` before an ordinary import
|
|
without `-f`, rewind, or rollback.
|
|
4. The imported pool was healthy. The canary SHA-256 matched its pre-rebuild
|
|
value. `profile_atlas` completed against the imported pool and a second
|
|
run reported `changed=0`. A file restored from the preserved snapshot into
|
|
`/var/tmp` matched SHA-256, owner, group, mode, size, and mtime; the temporary
|
|
copy was removed. Final checks found SELinux Enforcing, 12 datasets, the
|
|
snapshot, no failed units, and a healthy pool. The VM was shut down while
|
|
retaining its disposable volumes for a future rehearsal.
|
|
|
|
This proves the **sequence** for a cleanly exported, small pool and the tested
|
|
Ansible subset, not recovery duration or capacity at 2 TB. The earlier
|
|
2026-09-25 independent production ZFS/USB file restores and the earlier Borg
|
|
temporary-directory restore remain separate evidence. The VM did not restore
|
|
production USB/Borg archives, exercise services with production data, test an
|
|
unclean import, or prove the provisional RPO/RTO. Before relying on 24h/72h,
|
|
measure a representative full restore and service cutover in a suitably sized
|
|
future change window. Never use the production Atlas pool for a rehearsal.
|