Files
infra/docs/atlas-dr-lab.md
2026-09-30 21:19:28 +02:00

75 lines
4.6 KiB
Markdown

# Isolated Atlas DR lab
This is a **scaled rehearsal**, not a substitute for a full-data restore. The
`atlas-dr-lab` libvirt VM on Ikaros was left **shut off** on 2026-09-30. Its
persistent volumes are in the default libvirt pool: the current 30 GiB OS
volume `atlas-dr-lab-os-rebuild2.qcow2`, the pre-rebuild OS volume
`atlas-dr-lab-os.qcow2`, and four independent 4 GiB
`atlas-dr-lab-data{1,2,3,4}.qcow2` volumes. The VM uses libvirt's `default`
NAT network (last DHCP address `192.168.122.168`), 2 vCPU, and 4 GiB RAM.
The data disks have `virtio-atlasdrdata{1,2,3,4}` serials. No physical disk or
production Atlas storage is attached. The VM has no autostart.
## Rebuild inputs and isolation
- Use Rocky's **9.8 GenericCloud Base x86_64** image
`Rocky-9-GenericCloud-Base-9.8-20260525.0.x86_64.qcow2` from
`https://download.rockylinux.org/pub/rocky/9.8/images/x86_64/`.
Verify its `.CHECKSUM` file; the observed SHA-256 was
`92c206cc6f790c61583247eefe87890f8828420662c17cacf247cec78ab4eec8`.
- Use a dedicated lab-only inventory merged **after** the repository
inventory, and always `--limit atlas_dr_lab`. The temporary 2026-09-30
inventory/playbook and logs are in `/tmp/atlas-dr-lab-image/`; copy a
sanitized inventory to durable private storage before `/tmp` is cleared if
the lab will be repeated. Never reuse `host_vars/atlas.yml`, production
Vault secrets, or production disk by-id paths for the lab.
- The lab host belongs to `platform_rocky` and `atlas`. It uses `dradmin`
(UID/GID 1000) with the operator's **public** SSH key and a random,
unknown password hash, the libvirt DHCP address, pool `zpool`, mount root
`/zpool`, the four `virtio-atlasdrdata*` by-id paths, a 1 GiB backup
reservation, and `rocky_manage_openzfs_repo: true` with only `zfs` in
`host_packages`. The following gates remain false: sharing, firewall,
media stack, ZFS timers, Borg, USB, monitoring, and Prometheus pull.
`atlas_manage_storage` is true. Set `atlas_create_pool: true` **only for the
first disposable pool creation**, then set it false before any later run.
- A minimal lab playbook selects `atlas_dr_lab`, `become: true`, and the
existing `packages_rocky` and `profile_atlas` roles. Use a separate
`ANSIBLE_CONFIG` without the production Vault password script, and keep
host-key checking on with a lab-specific known-hosts file. The 2026-09-30
runs used `-i ansible/inventory/hosts.yml -i <lab-inventory.yml>` and
`--limit atlas_dr_lab` throughout.
## Rehearsal and narrow checks
1. Before any pool operation, compare `virsh -c qemu:///system domblklist
atlas-dr-lab` with the four intended qcow2 paths, and in the guest compare
`/dev/disk/by-id/virtio-atlasdrdata*` with `lsblk`. Do not proceed if a
physical disk or production identity appears.
2. For a first-time disposable build only, run the lab playbook with
`--tags pool` and `atlas_create_pool: true`, then immediately set the gate
false. Run the full lab playbook and check `zpool status -P zpool`,
`zfs list -r zpool`, SELinux, and failed systemd units.
3. Write a non-sensitive canary under the lab `/zpool/archive` and snapshot
it. Record the pool GUID and canary SHA-256. Export the lab pool cleanly,
shut down the VM, and replace **only the OS volume** with a fresh verified
Rocky image. Preserve all four data volumes. Reconfigure cloud-init for a
new instance; the seed CD-ROM must use **SATA**. The SCSI seed attachment
tried during this rehearsal was not detected by cloud-init and was
replaced with a SATA attachment before proceeding.
4. On the new OS, apply `packages_rocky` to reinstall OpenZFS. First run
`zpool import -d /dev/disk/by-id` **without importing**, compare GUID and
vdev membership, then use ordinary `zpool import -d /dev/disk/by-id zpool`.
Do not use `-f`, `-F`, `-X`, rollback, or pool creation.
5. Reapply `profile_atlas` with the lab gates and `atlas_create_pool: false`.
Verify the canary, restored snapshot file in an empty temporary directory,
dataset hierarchy, SELinux, and pool health. A second full playbook run
should report `changed=0`. Remove temporary restored files and shut down
the VM after testing.
The observed 2026-09-30 pool GUID was `8880368391795119587`; the canary
SHA-256 was `949701c7a95fadae1fddc21abe846c4312212dbfeb7477948f3188fc3ec34a78`.
The post-rebuild Ansible run succeeded, a repeat run reported `changed=0`,
12 datasets and the original snapshot were present, and the pool was healthy.
The snapshot-restored file matched content and basic metadata. See
[`atlas-recovery.md`](atlas-recovery.md) for the production runbook and limits.