3.4 KiB
Atlas Rocky/OpenZFS update and reboot procedure (draft)
This is an operator-controlled maintenance procedure. The playbook does not reboot Atlas, replace a pool device, or perform a pool feature upgrade.
Preflight
- Schedule an outage and confirm no Borg, USB, snapshot, scrub, or resilver
job is active. A service in
activatingis still active; do not interrupt it. - Check
zpool status -v zpool(including scrub status),zfs list -r zpool,systemctl --failed, andsystemctl list-timers --all. Resolve pool errors first. Record currentuname -r,modinfo zfs | grep '^version:',rpm -q kernel-core kmod-zfs zfs, and the current boot entry. - Confirm a recent successful Borg archive and a usable snapshot. Confirm the latest published offline USB version and its physical availability; do not start a USB backup merely to satisfy a checklist without capacity, UUID, and operator checks. Record timestamps, not just timer state.
- Ensure console/KVM or another independent recovery route is available.
Check free space in
/bootand the root filesystem. Review proposed DNF transactions before consenting to package changes.
Change window
- Stop client writes and quiesce stateful applications deliberately. Record
which services were stopped; do not assume
ansible-playbook --checkdoes this. Avoid updating during a running scrub or backup. - Use
dnf upgrade --assumenofirst to review the kernel,kmod-zfs,zfs, and dependencies. Confirm a matching kmod will be available for the target kernel. If compatibility is uncertain, defer the update. - Apply the approved DNF transaction. Do not run
zpool upgradeor enable new pool feature flags as part of ordinary OS maintenance; that can remove downgrade options. Preserve at least one known-good boot entry. - Reboot manually during the agreed outage. Ansible must not trigger it.
Post-boot gate
- Verify
uname -r,modinfo zfs,rpm -q kernel-core kmod-zfs zfs,zpool status -v zpool,zfs list -r zpool, andfindmnt -R /zpool. - Verify SELinux remains enforcing; inspect
systemctl --failedand the journal for ZFS, mount, SSH, NFS, SMB, Cockpit, Podman, and backup errors. - Validate a read-only file listing through SMB and an NFS client access
check before reopening writes. Check the rootless temporary services and
all backup/monitoring timers. Run the Atlas health monitor in
--dry-runmode, then a real check after inspection. - Re-enable clients and record versions, downtime, anomalies, and next successful snapshot/Borg run. A green boot alone is not a completed update.
Failure response
If the new kernel cannot load ZFS, boot the previous known-good kernel from the console and inspect package/kmod matching before trying another reboot. Do not force-import, rewind, clear errors, or upgrade pool features to make a failed OS update appear successful. Preserve logs and stop for a recovery decision if the pool does not import cleanly.
The procedure-definition item is complete, but the procedure is not yet rehearsed on a replacement host or during a real Atlas update. Record the first controlled execution and its post-boot evidence separately.
Read-only preflight on 2026-09-30 observed kernel
5.14.0-687.52.1.el9_8.x86_64, ZFS module/package 2.2.11-1, a healthy
zpool, enforcing SELinux, and no failed systemd units. This did not review
an upgrade transaction, stop services, or reboot the host.