Define controlled Atlas kernel and OpenZFS updates

This commit is contained in:
Fabio Scotto di Santolo
2026-09-30 21:19:48 +02:00
parent 9798fe3a12
commit 8844d00e24
2 changed files with 66 additions and 1 deletions

View File

@@ -244,7 +244,9 @@ successfully. The first monthly scrub remains a runtime check.
the second Ansible run was idempotent. Earlier independent production ZFS, USB, and Borg file the second Ansible run was idempotent. Earlier independent production ZFS, USB, and Borg file
restore tests remain separate evidence. A production-size full restore, unclean import, and restore tests remain separate evidence. A production-size full restore, unclean import, and
measured 24h/72h compliance are not claimed. measured 24h/72h compliance are not claimed.
- [ ] Define a controlled Rocky kernel/OpenZFS update and reboot procedure. - [x] Define a controlled Rocky kernel/OpenZFS update and reboot procedure in
`docs/atlas-updates.md`. The first real change-window execution is not yet
validated; the procedure never reboots automatically or upgrades pool features.
- [ ] Add the Atlas-initiated least-privilege Prometheus backup pull: Prometheus exposes only prepared - [ ] Add the Atlas-initiated least-privilege Prometheus backup pull: Prometheus exposes only prepared
read-only dumps through a dedicated account and Atlas retains the private SSH key, pinned host key, read-only dumps through a dedicated account and Atlas retains the private SSH key, pinned host key,
atomic pull, verification, retention and systemd service/timer. atomic pull, verification, retention and systemd service/timer.

63
docs/atlas-updates.md Normal file
View File

@@ -0,0 +1,63 @@
# Atlas Rocky/OpenZFS update and reboot procedure (draft)
This is an operator-controlled maintenance procedure. The playbook does not
reboot Atlas, replace a pool device, or perform a pool feature upgrade.
## Preflight
1. Schedule an outage and confirm no Borg, USB, snapshot, scrub, or resilver
job is active. A service in `activating` is still active; do not interrupt it.
2. Check `zpool status -v zpool` (including scrub status), `zfs list -r zpool`,
`systemctl --failed`, and `systemctl list-timers --all`. Resolve pool errors
first. Record current `uname -r`, `modinfo zfs | grep '^version:'`,
`rpm -q kernel-core kmod-zfs zfs`, and the current boot entry.
3. Confirm a recent successful Borg archive and a usable snapshot. Confirm
the latest published offline USB version and its physical availability;
do not start a USB backup merely to satisfy a checklist without capacity,
UUID, and operator checks. Record timestamps, not just timer state.
4. Ensure console/KVM or another independent recovery route is available.
Check free space in `/boot` and the root filesystem. Review proposed DNF
transactions before consenting to package changes.
## Change window
1. Stop client writes and quiesce stateful applications deliberately. Record
which services were stopped; do not assume `ansible-playbook --check` does
this. Avoid updating during a running scrub or backup.
2. Use `dnf upgrade --assumeno` first to review the kernel, `kmod-zfs`, `zfs`,
and dependencies. Confirm a matching kmod will be available for the target
kernel. If compatibility is uncertain, defer the update.
3. Apply the approved DNF transaction. Do not run `zpool upgrade` or enable
new pool feature flags as part of ordinary OS maintenance; that can remove
downgrade options. Preserve at least one known-good boot entry.
4. Reboot **manually** during the agreed outage. Ansible must not trigger it.
## Post-boot gate
1. Verify `uname -r`, `modinfo zfs`, `rpm -q kernel-core kmod-zfs zfs`,
`zpool status -v zpool`, `zfs list -r zpool`, and `findmnt -R /zpool`.
2. Verify SELinux remains enforcing; inspect `systemctl --failed` and the
journal for ZFS, mount, SSH, NFS, SMB, Cockpit, Podman, and backup errors.
3. Validate a read-only file listing through SMB and an NFS client access
check before reopening writes. Check the rootless temporary services and
all backup/monitoring timers. Run the Atlas health monitor in `--dry-run`
mode, then a real check after inspection.
4. Re-enable clients and record versions, downtime, anomalies, and next
successful snapshot/Borg run. A green boot alone is not a completed update.
## Failure response
If the new kernel cannot load ZFS, boot the previous known-good kernel from
the console and inspect package/kmod matching before trying another reboot.
Do not force-import, rewind, clear errors, or upgrade pool features to make a
failed OS update appear successful. Preserve logs and stop for a recovery
decision if the pool does not import cleanly.
The procedure-definition item is complete, but the procedure is **not yet
rehearsed** on a replacement host or during a real Atlas update. Record the
first controlled execution and its post-boot evidence separately.
Read-only preflight on 2026-09-30 observed kernel
`5.14.0-687.52.1.el9_8.x86_64`, ZFS module/package `2.2.11-1`, a healthy
`zpool`, enforcing SELinux, and no failed systemd units. This did not review
an upgrade transaction, stop services, or reboot the host.