4.6 KiB
Isolated Atlas DR lab
This is a scaled rehearsal, not a substitute for a full-data restore. The
atlas-dr-lab libvirt VM on Ikaros was left shut off on 2026-09-30. Its
persistent volumes are in the default libvirt pool: the current 30 GiB OS
volume atlas-dr-lab-os-rebuild2.qcow2, the pre-rebuild OS volume
atlas-dr-lab-os.qcow2, and four independent 4 GiB
atlas-dr-lab-data{1,2,3,4}.qcow2 volumes. The VM uses libvirt's default
NAT network (last DHCP address 192.168.122.168), 2 vCPU, and 4 GiB RAM.
The data disks have virtio-atlasdrdata{1,2,3,4} serials. No physical disk or
production Atlas storage is attached. The VM has no autostart.
Rebuild inputs and isolation
- Use Rocky's 9.8 GenericCloud Base x86_64 image
Rocky-9-GenericCloud-Base-9.8-20260525.0.x86_64.qcow2fromhttps://download.rockylinux.org/pub/rocky/9.8/images/x86_64/. Verify its.CHECKSUMfile; the observed SHA-256 was92c206cc6f790c61583247eefe87890f8828420662c17cacf247cec78ab4eec8. - Use a dedicated lab-only inventory merged after the repository
inventory, and always
--limit atlas_dr_lab. The temporary 2026-09-30 inventory/playbook and logs are in/tmp/atlas-dr-lab-image/; copy a sanitized inventory to durable private storage before/tmpis cleared if the lab will be repeated. Never reusehost_vars/atlas.yml, production Vault secrets, or production disk by-id paths for the lab. - The lab host belongs to
platform_rockyandatlas. It usesdradmin(UID/GID 1000) with the operator's public SSH key and a random, unknown password hash, the libvirt DHCP address, poolzpool, mount root/zpool, the fourvirtio-atlasdrdata*by-id paths, a 1 GiB backup reservation, androcky_manage_openzfs_repo: truewith onlyzfsinhost_packages. The following gates remain false: sharing, firewall, media stack, ZFS timers, Borg, USB, monitoring, and Prometheus pull.atlas_manage_storageis true. Setatlas_create_pool: trueonly for the first disposable pool creation, then set it false before any later run. - A minimal lab playbook selects
atlas_dr_lab,become: true, and the existingpackages_rockyandprofile_atlasroles. Use a separateANSIBLE_CONFIGwithout the production Vault password script, and keep host-key checking on with a lab-specific known-hosts file. The 2026-09-30 runs used-i ansible/inventory/hosts.yml -i <lab-inventory.yml>and--limit atlas_dr_labthroughout.
Rehearsal and narrow checks
- Before any pool operation, compare
virsh -c qemu:///system domblklist atlas-dr-labwith the four intended qcow2 paths, and in the guest compare/dev/disk/by-id/virtio-atlasdrdata*withlsblk. Do not proceed if a physical disk or production identity appears. - For a first-time disposable build only, run the lab playbook with
--tags poolandatlas_create_pool: true, then immediately set the gate false. Run the full lab playbook and checkzpool status -P zpool,zfs list -r zpool, SELinux, and failed systemd units. - Write a non-sensitive canary under the lab
/zpool/archiveand snapshot it. Record the pool GUID and canary SHA-256. Export the lab pool cleanly, shut down the VM, and replace only the OS volume with a fresh verified Rocky image. Preserve all four data volumes. Reconfigure cloud-init for a new instance; the seed CD-ROM must use SATA. The SCSI seed attachment tried during this rehearsal was not detected by cloud-init and was replaced with a SATA attachment before proceeding. - On the new OS, apply
packages_rockyto reinstall OpenZFS. First runzpool import -d /dev/disk/by-idwithout importing, compare GUID and vdev membership, then use ordinaryzpool import -d /dev/disk/by-id zpool. Do not use-f,-F,-X, rollback, or pool creation. - Reapply
profile_atlaswith the lab gates andatlas_create_pool: false. Verify the canary, restored snapshot file in an empty temporary directory, dataset hierarchy, SELinux, and pool health. A second full playbook run should reportchanged=0. Remove temporary restored files and shut down the VM after testing.
The observed 2026-09-30 pool GUID was 8880368391795119587; the canary
SHA-256 was 949701c7a95fadae1fddc21abe846c4312212dbfeb7477948f3188fc3ec34a78.
The post-rebuild Ansible run succeeded, a repeat run reported changed=0,
12 datasets and the original snapshot were present, and the pool was healthy.
The snapshot-restored file matched content and basic metadata. See
atlas-recovery.md for the production runbook and limits.