From 8844d00e24503febf504ff6d1b81ef5785680fda Mon Sep 17 00:00:00 2001 From: Fabio Scotto di Santolo Date: Wed, 30 Sep 2026 21:19:48 +0200 Subject: [PATCH] Define controlled Atlas kernel and OpenZFS updates --- AGENTS.md | 4 ++- docs/atlas-updates.md | 63 +++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 66 insertions(+), 1 deletion(-) create mode 100644 docs/atlas-updates.md diff --git a/AGENTS.md b/AGENTS.md index c14349c..9adc333 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -244,7 +244,9 @@ successfully. The first monthly scrub remains a runtime check. the second Ansible run was idempotent. Earlier independent production ZFS, USB, and Borg file restore tests remain separate evidence. A production-size full restore, unclean import, and measured 24h/72h compliance are not claimed. -- [ ] Define a controlled Rocky kernel/OpenZFS update and reboot procedure. +- [x] Define a controlled Rocky kernel/OpenZFS update and reboot procedure in + `docs/atlas-updates.md`. The first real change-window execution is not yet + validated; the procedure never reboots automatically or upgrades pool features. - [ ] Add the Atlas-initiated least-privilege Prometheus backup pull: Prometheus exposes only prepared read-only dumps through a dedicated account and Atlas retains the private SSH key, pinned host key, atomic pull, verification, retention and systemd service/timer. diff --git a/docs/atlas-updates.md b/docs/atlas-updates.md new file mode 100644 index 0000000..40e90eb --- /dev/null +++ b/docs/atlas-updates.md @@ -0,0 +1,63 @@ +# Atlas Rocky/OpenZFS update and reboot procedure (draft) + +This is an operator-controlled maintenance procedure. The playbook does not +reboot Atlas, replace a pool device, or perform a pool feature upgrade. + +## Preflight + +1. Schedule an outage and confirm no Borg, USB, snapshot, scrub, or resilver + job is active. A service in `activating` is still active; do not interrupt it. +2. Check `zpool status -v zpool` (including scrub status), `zfs list -r zpool`, + `systemctl --failed`, and `systemctl list-timers --all`. Resolve pool errors + first. Record current `uname -r`, `modinfo zfs | grep '^version:'`, + `rpm -q kernel-core kmod-zfs zfs`, and the current boot entry. +3. Confirm a recent successful Borg archive and a usable snapshot. Confirm + the latest published offline USB version and its physical availability; + do not start a USB backup merely to satisfy a checklist without capacity, + UUID, and operator checks. Record timestamps, not just timer state. +4. Ensure console/KVM or another independent recovery route is available. + Check free space in `/boot` and the root filesystem. Review proposed DNF + transactions before consenting to package changes. + +## Change window + +1. Stop client writes and quiesce stateful applications deliberately. Record + which services were stopped; do not assume `ansible-playbook --check` does + this. Avoid updating during a running scrub or backup. +2. Use `dnf upgrade --assumeno` first to review the kernel, `kmod-zfs`, `zfs`, + and dependencies. Confirm a matching kmod will be available for the target + kernel. If compatibility is uncertain, defer the update. +3. Apply the approved DNF transaction. Do not run `zpool upgrade` or enable + new pool feature flags as part of ordinary OS maintenance; that can remove + downgrade options. Preserve at least one known-good boot entry. +4. Reboot **manually** during the agreed outage. Ansible must not trigger it. + +## Post-boot gate + +1. Verify `uname -r`, `modinfo zfs`, `rpm -q kernel-core kmod-zfs zfs`, + `zpool status -v zpool`, `zfs list -r zpool`, and `findmnt -R /zpool`. +2. Verify SELinux remains enforcing; inspect `systemctl --failed` and the + journal for ZFS, mount, SSH, NFS, SMB, Cockpit, Podman, and backup errors. +3. Validate a read-only file listing through SMB and an NFS client access + check before reopening writes. Check the rootless temporary services and + all backup/monitoring timers. Run the Atlas health monitor in `--dry-run` + mode, then a real check after inspection. +4. Re-enable clients and record versions, downtime, anomalies, and next + successful snapshot/Borg run. A green boot alone is not a completed update. + +## Failure response + +If the new kernel cannot load ZFS, boot the previous known-good kernel from +the console and inspect package/kmod matching before trying another reboot. +Do not force-import, rewind, clear errors, or upgrade pool features to make a +failed OS update appear successful. Preserve logs and stop for a recovery +decision if the pool does not import cleanly. + +The procedure-definition item is complete, but the procedure is **not yet +rehearsed** on a replacement host or during a real Atlas update. Record the +first controlled execution and its post-boot evidence separately. + +Read-only preflight on 2026-09-30 observed kernel +`5.14.0-687.52.1.el9_8.x86_64`, ZFS module/package `2.2.11-1`, a healthy +`zpool`, enforcing SELinux, and no failed systemd units. This did not review +an upgrade transaction, stop services, or reboot the host.