Compare commits

...

5 Commits

Author SHA1 Message Date
Fabio Scotto di Santolo
802cb8c7ba Merge pull request #10 from fscotto/feature/atlas-priority2-recovery
Feature/atlas priority2 recovery
2026-09-30 21:25:36 +02:00
Fabio Scotto di Santolo
0144600a4a Add verified Prometheus backup pull to Atlas 2026-09-30 21:21:48 +02:00
Fabio Scotto di Santolo
3d2ef02c98 Keep Atlas SMB and NFS namespaces separate 2026-09-30 21:20:15 +02:00
Fabio Scotto di Santolo
8844d00e24 Define controlled Atlas kernel and OpenZFS updates 2026-09-30 21:19:48 +02:00
Fabio Scotto di Santolo
9798fe3a12 Document and rehearse Atlas disaster recovery 2026-09-30 21:19:28 +02:00
27 changed files with 1163 additions and 29 deletions

View File

@@ -226,7 +226,7 @@ successfully. The first monthly scrub remains a runtime check.
read-only ZFS snapshot test restored one file to `/var/tmp`, confirmed matching contents, ownership,
mode, mtime and ACL, then removed its temporary copy and on-demand mount. This is a file-level smoke
test, not full dataset recovery. An independent USB file restore passed on 2026-09-25 with matching
content and metadata; full disaster recovery remains a separate Priority 2 task.
content and metadata; the later scaled OS-rebuild rehearsal is documented under Priority 2.
- [x] Add monitoring and alerting for pool health, scrub/resilver, SMART data, temperatures, free space,
snapshot/local-backup growth, Hetzner Storage Box quota, and failed maintenance or backup timers.
The half-hourly Atlas health monitor and systemd final-failure hooks are deployed. A live probe
@@ -237,15 +237,26 @@ successfully. The first monthly scrub remains a runtime check.
on Atlas, but a new real failure notification has not been deliberately triggered.
### Priority 2 - NAS operability and recovery
- [ ] Document and test disaster recovery: rebuild Atlas with Ansible, import the existing pool, restore
from snapshot/USB/Hetzner, preserve Vault and Borg recovery material offline, and define RPO/RTO.
- [ ] Define a controlled Rocky kernel/OpenZFS update and reboot procedure.
- [ ] Add the Atlas-initiated least-privilege Prometheus backup pull: Prometheus exposes only prepared
- [x] Document and test disaster recovery in `docs/atlas-recovery.md`: the operator confirmed Vault
and Borg recovery material is available offline; provisional targets are RPO 24h/RTO 72h. On
2026-09-30 an isolated small Rocky VM was rebuilt with the Atlas Ansible roles, imported its
preserved RAIDZ2 pool without force/rewind, and restored a file from the preserved snapshot;
the second Ansible run was idempotent. Earlier independent production ZFS, USB, and Borg file
restore tests remain separate evidence. A production-size full restore, unclean import, and
measured 24h/72h compliance are not claimed.
- [x] Define a controlled Rocky kernel/OpenZFS update and reboot procedure in
`docs/atlas-updates.md`. The first real change-window execution is not yet
validated; the procedure never reboots automatically or upgrades pool features.
- [x] Add the Atlas-initiated least-privilege Prometheus backup pull: Prometheus exposes only prepared
read-only dumps through a dedicated account and Atlas retains the private SSH key, pinned host key,
atomic pull, verification, retention and systemd service/timer.
- [ ] Decide whether a common SMB/NFS namespace is required. `Archive` (SMB) and `photobook` (NFS) are
intentionally distinct today; only if a shared namespace is selected, finalize its UID/GID, group,
and POSIX ACL model and test the same files through both protocols.
atomic pull, verification, retention and systemd service/timer. The dedicated key/account and unit
files are deployed; live read-only SSH, shell denial, and write denial were verified. On 2026-09-30
a manual export, Atlas pull, checksum verification, and temporary restore passed; both SQLite
databases passed integrity checks and a restored Git repository passed `git fsck`. Both daily
timers are enabled for 02:00/03:00 Europe/Rome; their first scheduled results remain unverified.
- [x] Decide whether a common SMB/NFS namespace is required: no. `Archive` (SMB) and `photobook` (NFS)
remain intentionally distinct; `docs/atlas-sharing-decision.md` records the decision. No ACL or export
change is authorized by this decision.
### Priority 3 - Service expansion
- [ ] After data protection and recovery are validated, populate `/zpool/media/music` and validate Navidrome.

View File

@@ -485,7 +485,7 @@ etichettata di 45Drives Alerts usare
### Timer systemd di Atlas
Tutti i nove timer gestiti sono abilitati. Gli orari sono locali ad Atlas (`Europe/Rome`); Borg e
Tutti i dieci timer gestiti sono abilitati. Gli orari sono locali ad Atlas (`Europe/Rome`); Borg e
monitoraggio aggiungono il ritardo casuale indicato. Tutti hanno `Persistent=true`: un evento perso
viene recuperato quando il timer torna attivo.
@@ -500,10 +500,12 @@ viene recuperato quando il timer torna attivo.
| `atlas-borg-check.timer` | `*-*-15 06:00:00` — giorno 15 alle 06:00, più 0–30 min casuali | Controllo repository Borg |
| `atlas-usb-reminder.timer` | `Sat *-*-01..07 10:00:00 Europe/Rome` — primo sabato alle 10:00 | Solo promemoria 45Drives Alerts |
| `atlas-health-monitor.timer` | `*:0/30` — ogni mezz'ora, più 0–5 min casuali | Controlli di salute in sola lettura |
| `atlas-prometheus-pull.timer` | `*-*-* 03:00:00 Europe/Rome` — ogni giorno alle 03:00 | Pull e verifica del backup preparato su Prometheus |
`atlas-usb-backup.service` **non ha timer** e va avviato manualmente. Il timer del fornitore
`zfs-scrub-weekly@zpool.timer` è disabilitato a favore dello scrub mensile. Il futuro pull del backup
Prometheus non ha ancora un timer, perché non è implementato. Durante un backup Borg attivo,
`zfs-scrub-weekly@zpool.timer` è disabilitato a favore dello scrub mensile. Il timer di preparazione
su Prometheus è attivo alle 02:00 Europe/Rome; export, pull e ripristino temporaneo manuali sono
riusciti il 2026-09-30, ma il primo ciclo pianificato va ancora verificato. Durante un backup Borg attivo,
`systemctl list-timers` può mostrare `-` per il prossimo evento senza che il timer sia disabilitato.
Per vedere la pianificazione corrente: `systemctl list-timers --all` su Atlas.
@@ -519,7 +521,10 @@ cutover. L'attuale iCloudPD su Aegis e l'export NFS Photobook restano configurat
all'approvazione e alla verifica di questa migrazione separata. Anche il servizio Atlas sarà
temporaneo in attesa di Uranus.
Il pull dei backup di Prometheus e i test completi di disaster recovery restano da fare. Il backlog
Il primo ciclo pianificato del backup di Prometheus e una prova di disaster recovery a dimensione reale
restano da verificare. Il 2026-09-30 una VM Rocky isolata ha superato ricostruzione OS con Ansible,
import del pool RAIDZ2 fittizio e ripristino da snapshot; RPO 24 ore/RTO 72 ore restano obiettivi
provvisori, non tempi misurati. Dettagli e limiti sono in `docs/atlas-recovery.md`. Il backlog
prioritizzato è in `AGENTS.md`.
---

View File

@@ -500,7 +500,7 @@ monitoring. For a labelled 45Drives Alerts delivery test, use
### Atlas systemd timers
All nine managed timers below are enabled. Times are local to Atlas (`Europe/Rome`); Borg and monitoring
All ten managed timers below are enabled. Times are local to Atlas (`Europe/Rome`); Borg and monitoring
add the indicated randomized delay. Every timer has `Persistent=true`, so a missed calendar run is
scheduled after the timer becomes active again.
@@ -515,10 +515,12 @@ scheduled after the timer becomes active again.
| `atlas-borg-check.timer` | `*-*-15 06:00:00` — 15th of the month at 06:00, plus 0–30 min random delay | Borg repository check |
| `atlas-usb-reminder.timer` | `Sat *-*-01..07 10:00:00 Europe/Rome` — first Saturday at 10:00 | 45Drives Alerts reminder only |
| `atlas-health-monitor.timer` | `*:0/30` — every half-hour, plus 0–5 min random delay | Read-only health checks |
| `atlas-prometheus-pull.timer` | `*-*-* 03:00:00 Europe/Rome` — daily at 03:00 | Pull and verify the prepared Prometheus backup |
`atlas-usb-backup.service` has **no timer**: the encrypted USB backup must be started manually.
The vendor's `zfs-scrub-weekly@zpool.timer` is intentionally disabled in favor of the monthly scrub.
The future Prometheus backup pull has no timer yet because that workflow is not implemented. While a
The Prometheus export timer runs at 02:00 Europe/Rome; its first scheduled run and the Atlas pull
remain to be observed. A manual export, pull, and temporary restore passed. While a
Borg backup is still running, `systemctl list-timers` may show `-` for its next trigger; this does not
mean the timer has been disabled. Inspect the current schedule on Atlas with
`systemctl list-timers --all`.
@@ -534,9 +536,21 @@ state outside `Archive`, then test permissions, SELinux, backups and recovery be
The current Aegis iCloudPD service and Atlas Photobook NFS export remain configured until that
separate migration is approved and validated; the eventual Atlas service is temporary until Uranus.
Prometheus backup pulls and full disaster-recovery tests remain follow-up work. The prioritized
The first scheduled Prometheus backup runs and production-size disaster-recovery tests remain follow-up work. The prioritized
operational backlog is kept in `AGENTS.md`.
Priority 2 procedures and decisions are recorded in
[`docs/atlas-recovery.md`](docs/atlas-recovery.md),
[`docs/atlas-updates.md`](docs/atlas-updates.md), and
[`docs/atlas-sharing-decision.md`](docs/atlas-sharing-decision.md).
The provisional Atlas recovery objectives are RPO 24 hours and RTO 72 hours;
an isolated small-VM OS rebuild, pool import, Ansible reapplication, and
snapshot restore passed, but full-size recovery time is unmeasured. `Archive` (SMB) and
`photobook` (NFS) remain deliberately separate.
The Prometheus pull architecture and manual export/pull/restore evidence are in
[`docs/prometheus-backup.md`](docs/prometheus-backup.md). Both daily timers are
enabled; their first scheduled runs remain to be verified.
## How layering works
A host can intentionally belong to more than one inventory group. The final configuration is the combination of the host and its groups, not a one-host/one-play mapping.

View File

@@ -80,5 +80,30 @@ server_sshd_settings:
server_sshd_allow_users:
- "{{ server_username }}"
server_backup_export_enabled: false
server_backup_username: prometheus-backup
server_backup_public_key_name: atlas-pull
server_backup_export_root: /var/lib/prometheus-backup-export
server_backup_rrsync_path: /usr/share/doc/rsync/support/rrsync
server_backup_export_calendar: "*-*-* 02:00:00 Europe/Rome"
server_backup_export_start_timer: false
server_backup_export_source_keep: 3
server_backup_export_paths:
- opt/npm/data
- opt/npm/letsencrypt
- opt/gitea/data
- home/git/.ssh
- opt/docker/server/docker-compose.yml
- etc/systemd/system/podman-compose-server.service
- etc/ssh/sshd_config
- etc/ssh/sshd_config.d
- etc/firewalld
- etc/wireguard/wg0.conf
server_backup_export_excludes:
- opt/npm/data/logs
- opt/gitea/data/gitea/log
- opt/gitea/data/gitea/tmp
- opt/gitea/data/gitea/sessions
- opt/gitea/data/gitea/indexers
server_ssh_authorized_keys: []
server_ssh_authorized_key_directory: "{{ server_user_home }}/.ssh/authorized_keys.d"

View File

@@ -49,6 +49,7 @@ atlas_zfs_backup_reservation: 500G
atlas_zfs_dataset_photobook: media/photobook
atlas_mount_root: /zpool
atlas_manage_storage: true
atlas_prometheus_pull_start_timer: true
atlas_manage_zfs_snapshots: true
atlas_zfs_snapshot_prefix: atlas-auto
atlas_zfs_snapshot_policies:
@@ -91,6 +92,12 @@ atlas_usb_backup_mapper_name: zpool-backup
atlas_manage_usb_reminder: true
atlas_usb_reminder_calendar: "Sat *-*-01..07 10:00:00 Europe/Rome"
atlas_manage_monitoring: true
atlas_manage_prometheus_backup_pull: true
# Prometheus ED25519 host key read through the controller's strict SSH trust on 2026-09-30.
# Fingerprint: SHA256:rfedk7DHI9mLB3UHk/4F3HHlSIiswtCAFsAXvfh6iXk
atlas_prometheus_ssh_host_key: >-
179.237.102.172 ssh-ed25519
AAAAC3NzaC1lZDI1NTE5AAAAIC4b+QXlPupoEx71W9NKs9tTeYjBqTkVMqbGB97nMNWv
# Physical pool disks and the system NVMe; the disconnected USB disk is intentionally excluded.
atlas_monitor_smart_devices:
- { name: pool-1, path: "{{ atlas_zpool_disks[0] }}", warning_c: 50, critical_c: 55 }

View File

@@ -6,6 +6,8 @@ ansible_port: 22
ansible_ssh_private_key_file: /home/fscotto/.ssh/id_ed25519
server_username: rocky
server_backup_export_enabled: true
server_backup_export_start_timer: true
server_duckdns_domain: fscotto
server_ssh_authorized_keys:
- name: ikaros

View File

@@ -115,6 +115,15 @@ atlas_monitor_notifier: "{{ atlas_usb_reminder_notifier }}"
atlas_monitor_smart_devices: []
atlas_monitor_timers: []
atlas_monitor_failure_units: []
atlas_monitor_effective_timers: >-
{{ atlas_monitor_timers
+ ([{'name': 'atlas-prometheus-pull.timer', 'max_age_hours': 26}]
if atlas_manage_prometheus_backup_pull | bool and atlas_prometheus_pull_start_timer | bool
else []) }}
atlas_monitor_effective_failure_units: >-
{{ atlas_monitor_failure_units
+ (['atlas-prometheus-pull.service']
if atlas_manage_prometheus_backup_pull | bool else []) }}
atlas_monitor_remote_capacity: {}
atlas_monitor_pool_warning_percent: 80
atlas_monitor_pool_critical_percent: 90
@@ -141,6 +150,19 @@ atlas_music_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_music }}"
atlas_backup_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_backup }}"
atlas_host_backups_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_host_backups }}"
atlas_backup_prometheus_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_backup_prometheus }}"
atlas_manage_prometheus_backup_pull: false
atlas_prometheus_pull_ssh_dir: /etc/atlas-prometheus-pull
atlas_prometheus_pull_private_key_path: "{{ atlas_prometheus_pull_ssh_dir }}/id_ed25519"
atlas_prometheus_pull_known_hosts_path: "{{ atlas_prometheus_pull_ssh_dir }}/known_hosts"
atlas_prometheus_ssh_host_key: ""
atlas_prometheus_pull_source_user: prometheus-backup
atlas_prometheus_pull_source_port: 22
atlas_prometheus_pull_calendar: "*-*-* 03:00:00 Europe/Rome"
atlas_prometheus_pull_start_timer: false
atlas_prometheus_pull_keep_daily: 30
atlas_prometheus_pull_keep_weekly: 8
atlas_prometheus_pull_keep_monthly: 12
atlas_prometheus_pull_max_age_hours: 24
atlas_photobook_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_photobook }}"
atlas_45drives_repo_url: https://repo.45drives.com/repofiles/rocky/45drives-enterprise.repo

View File

@@ -0,0 +1,56 @@
#!/usr/bin/env python3
"""Prune only verified, named Prometheus backup versions after publication."""
import datetime as dt
import pathlib
import re
import shutil
import sys
def main() -> None:
if len(sys.argv) != 5:
raise SystemExit("Usage: atlas-prometheus-prune SNAPSHOTS DAILY WEEKLY MONTHLY")
root = pathlib.Path(sys.argv[1])
counts = [int(value) for value in sys.argv[2:]]
if not root.is_dir() or root.is_symlink() or min(counts) < 1:
raise SystemExit("Invalid backup directory or retention counts")
versions = []
for entry in root.iterdir():
if not entry.is_dir() or entry.is_symlink():
continue
if not re.fullmatch(r"[0-9]{8}T[0-9]{6}Z", entry.name):
continue
try:
when = dt.datetime.strptime(entry.name, "%Y%m%dT%H%M%SZ")
except ValueError:
continue
if not all((entry / name).is_file() for name in ("payload.tar", "payload.sha256", "metadata.json")):
continue
versions.append((when, entry))
versions.sort(reverse=True)
if not versions:
raise SystemExit("No published backup versions found; refusing to prune")
keep = {entry for _, entry in versions[: counts[0]]}
for count, key in (
(counts[1], lambda when: when.isocalendar()[:2]),
(counts[2], lambda when: (when.year, when.month)),
):
periods = set()
for when, entry in versions:
period = key(when)
if period in periods:
continue
periods.add(period)
keep.add(entry)
if len(periods) >= count:
break
for _, entry in versions:
if entry not in keep:
shutil.rmtree(entry)
if __name__ == "__main__":
main()

View File

@@ -23,6 +23,12 @@
- name: Import Atlas offline USB backup tasks
ansible.builtin.import_tasks: usb_backup.yml
- name: Import Atlas Prometheus backup pull identity tasks
ansible.builtin.import_tasks: prometheus_pull_identity.yml
- name: Import Atlas Prometheus backup pull job tasks
ansible.builtin.import_tasks: prometheus_pull_job.yml
- name: Import Atlas health monitoring tasks
ansible.builtin.import_tasks: monitoring.yml

View File

@@ -7,8 +7,8 @@
- atlas_zfs_pool != 'CHANGEME_ZFS_POOL'
- atlas_monitor_calendar | length > 0
- atlas_monitor_smart_devices | length > 0
- atlas_monitor_timers | length > 0
- atlas_monitor_failure_units | length > 0
- atlas_monitor_effective_timers | length > 0
- atlas_monitor_effective_failure_units | length > 0
- atlas_monitor_remote_capacity.user == atlas_borg_repository_user
- atlas_monitor_remote_capacity.host == atlas_borg_repository_host
- atlas_monitor_remote_capacity.run_as == atlas_borg_username
@@ -49,7 +49,7 @@
that:
- item.name is match('^[a-zA-Z0-9@_.-]+\\.timer$')
- item.max_age_hours | int >= 0
loop: "{{ atlas_monitor_timers }}"
loop: "{{ atlas_monitor_effective_timers }}"
loop_control:
label: "{{ item.name }}"
when: atlas_manage_monitoring | bool
@@ -59,7 +59,7 @@
ansible.builtin.assert:
that:
- item is match('^[a-zA-Z0-9@_.-]+\\.service$')
loop: "{{ atlas_monitor_failure_units }}"
loop: "{{ atlas_monitor_effective_failure_units }}"
when: atlas_manage_monitoring | bool
- name: Validate Atlas health monitor calendar
@@ -144,7 +144,7 @@
owner: root
group: root
mode: "0755"
loop: "{{ atlas_monitor_failure_units }}"
loop: "{{ atlas_monitor_effective_failure_units }}"
when: atlas_manage_monitoring | bool
- name: Notify 45Drives Alerts when an Atlas job fails
@@ -155,7 +155,7 @@
owner: root
group: root
mode: "0644"
loop: "{{ atlas_monitor_failure_units }}"
loop: "{{ atlas_monitor_effective_failure_units }}"
when: atlas_manage_monitoring | bool
- name: Reload systemd after installing Atlas monitoring

View File

@@ -0,0 +1,66 @@
---
- name: Validate Atlas Prometheus pull identity inputs
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
ansible.builtin.assert:
that:
- atlas_prometheus_pull_ssh_dir.startswith('/etc/')
- atlas_prometheus_pull_private_key_path.startswith(atlas_prometheus_pull_ssh_dir ~ '/')
- atlas_prometheus_pull_known_hosts_path.startswith(atlas_prometheus_pull_ssh_dir ~ '/')
- atlas_prometheus_ssh_host_key.startswith(
(hostvars['prometheus'].ansible_host | string) ~ ' ssh-ed25519 '
)
fail_msg: Pin the verified Prometheus ED25519 SSH host key before enabling the pull.
when: atlas_manage_prometheus_backup_pull | bool
- name: Create private Atlas Prometheus pull SSH directory
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
ansible.builtin.file:
path: "{{ atlas_prometheus_pull_ssh_dir }}"
state: directory
owner: root
group: root
mode: "0700"
when: atlas_manage_prometheus_backup_pull | bool
- name: Generate Atlas-only Prometheus pull SSH identity
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
ansible.builtin.command:
argv:
- ssh-keygen
- -q
- -t
- ed25519
- -N
- ""
- -C
- atlas-prometheus-pull@atlas
- -f
- "{{ atlas_prometheus_pull_private_key_path }}"
creates: "{{ atlas_prometheus_pull_private_key_path }}"
when: atlas_manage_prometheus_backup_pull | bool
- name: Protect Atlas-only Prometheus pull SSH identity
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
ansible.builtin.file:
path: "{{ item.path }}"
owner: root
group: root
mode: "{{ item.mode }}"
loop:
- { path: "{{ atlas_prometheus_pull_private_key_path }}", mode: "0600" }
- { path: "{{ atlas_prometheus_pull_private_key_path }}.pub", mode: "0644" }
loop_control:
label: "{{ item.path }}"
when:
- atlas_manage_prometheus_backup_pull | bool
- not ansible_check_mode
- name: Pin Prometheus SSH host key on Atlas
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
ansible.builtin.copy:
content: "{{ atlas_prometheus_ssh_host_key }}\n"
dest: "{{ atlas_prometheus_pull_known_hosts_path }}"
owner: root
group: root
mode: "0600"
when: atlas_manage_prometheus_backup_pull | bool

View File

@@ -0,0 +1,90 @@
---
- name: Validate Atlas Prometheus backup pull inputs
tags: [atlas, backup, prometheus_backup]
ansible.builtin.assert:
that:
- atlas_manage_storage | bool
- atlas_prometheus_pull_source_user is match('^[a-z_][a-z0-9_-]*$')
- atlas_prometheus_pull_source_port | int > 0
- atlas_prometheus_pull_source_port | int < 65536
- atlas_prometheus_pull_keep_daily | int > 0
- atlas_prometheus_pull_keep_weekly | int > 0
- atlas_prometheus_pull_keep_monthly | int > 0
- atlas_prometheus_pull_max_age_hours | int > 0
- atlas_backup_prometheus_mountpoint.startswith(atlas_mount_root ~ '/')
fail_msg: Define the Atlas backup destination, source account, and retention before enabling the pull.
when: atlas_manage_prometheus_backup_pull | bool
- name: Validate Atlas Prometheus backup pull calendar
tags: [atlas, backup, prometheus_backup]
ansible.builtin.command:
argv: [systemd-analyze, calendar, "{{ atlas_prometheus_pull_calendar }}"]
changed_when: false
check_mode: false
when: atlas_manage_prometheus_backup_pull | bool
- name: Create private Atlas Prometheus backup version directory
tags: [atlas, backup, prometheus_backup]
ansible.builtin.file:
path: "{{ atlas_backup_prometheus_mountpoint }}/snapshots"
state: directory
owner: root
group: root
mode: "0700"
when: atlas_manage_prometheus_backup_pull | bool
- name: Install Atlas Prometheus backup pull helper
tags: [atlas, backup, prometheus_backup]
ansible.builtin.template:
src: atlas-prometheus-pull.sh.j2
dest: /usr/local/sbin/atlas-prometheus-pull
owner: root
group: root
mode: "0750"
when: atlas_manage_prometheus_backup_pull | bool
- name: Install Atlas Prometheus backup retention helper
tags: [atlas, backup, prometheus_backup]
ansible.builtin.copy:
src: atlas-prometheus-prune.py
dest: /usr/local/libexec/atlas-prometheus-prune
owner: root
group: root
mode: "0750"
when: atlas_manage_prometheus_backup_pull | bool
- name: Install Atlas Prometheus backup pull systemd units
tags: [atlas, backup, prometheus_backup]
ansible.builtin.template:
src: "{{ item }}.j2"
dest: "/etc/systemd/system/{{ item }}"
owner: root
group: root
mode: "0644"
loop:
- atlas-prometheus-pull.service
- atlas-prometheus-pull.timer
loop_control:
label: "{{ item }}"
register: atlas_prometheus_pull_units
when: atlas_manage_prometheus_backup_pull | bool
- name: Reload systemd after Atlas Prometheus pull unit changes
tags: [atlas, backup, prometheus_backup]
ansible.builtin.systemd:
daemon_reload: true
when:
- atlas_manage_prometheus_backup_pull | bool
- atlas_prometheus_pull_units is changed
- not ansible_check_mode
- name: Enable Atlas Prometheus pull timer only after explicit activation
tags: [atlas, backup, prometheus_backup]
ansible.builtin.systemd:
name: atlas-prometheus-pull.timer
enabled: true
state: started
when:
- atlas_manage_prometheus_backup_pull | bool
- atlas_prometheus_pull_start_timer | bool
- not ansible_check_mode

View File

@@ -3,8 +3,8 @@
"backup_dataset": {{ (atlas_zfs_pool ~ '/' ~ atlas_zfs_dataset_backup) | to_json }},
"notifier": {{ atlas_monitor_notifier | to_json }},
"smart_devices": {{ atlas_monitor_smart_devices | to_json }},
"timers": {{ atlas_monitor_timers | to_json }},
"failure_units": {{ atlas_monitor_failure_units | to_json }},
"timers": {{ atlas_monitor_effective_timers | to_json }},
"failure_units": {{ atlas_monitor_effective_failure_units | to_json }},
"remote_capacity": {{ atlas_monitor_remote_capacity | to_json }},
"pool_warning_percent": {{ atlas_monitor_pool_warning_percent | int }},
"pool_critical_percent": {{ atlas_monitor_pool_critical_percent | int }},

View File

@@ -0,0 +1,19 @@
[Unit]
Description=Pull a prepared read-only Prometheus backup to Atlas
RequiresMountsFor={{ atlas_backup_prometheus_mountpoint }}
Wants=network-online.target
After=network-online.target zfs.target
ConditionFileIsExecutable=/usr/local/sbin/atlas-prometheus-pull
ConditionPathExists={{ atlas_prometheus_pull_private_key_path }}
ConditionPathExists={{ atlas_prometheus_pull_known_hosts_path }}
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/atlas-prometheus-pull
User=root
Group=root
UMask=0077
TimeoutStartSec=infinity
Nice=15
IOSchedulingClass=best-effort
IOSchedulingPriority=7

View File

@@ -0,0 +1,75 @@
#!/usr/bin/env bash
set -Eeuo pipefail
umask 077
backup_root={{ atlas_backup_prometheus_mountpoint | quote }}
snapshots="$backup_root/snapshots"
stage=''
exec 9>/run/lock/atlas-prometheus-pull.lock
flock -n 9 || { echo 'A Prometheus pull is already running' >&2; exit 1; }
cleanup() {
local rc=$?
trap - EXIT
if (( rc != 0 )) && [[ -n "$stage" && -d "$stage" ]]; then
rm -rf -- "$stage"
fi
exit "$rc"
}
trap cleanup EXIT
zpool list -H -o name {{ atlas_zfs_pool | quote }} >/dev/null
findmnt -rn --mountpoint "$backup_root" >/dev/null
stage=$(mktemp -d "$backup_root/.staging.XXXXXXXX")
ssh_cmd='/usr/bin/ssh -F /dev/null -o BatchMode=yes -o StrictHostKeyChecking=yes -o UserKnownHostsFile={{ atlas_prometheus_pull_known_hosts_path }} -o IdentitiesOnly=yes -i {{ atlas_prometheus_pull_private_key_path }} -p {{ atlas_prometheus_pull_source_port }}'
rsync -a --partial --delay-updates -e "$ssh_cmd" \
{{ (atlas_prometheus_pull_source_user ~ '@' ~ hostvars['prometheus'].ansible_host ~ ':current/') | quote }} \
"$stage/"
test -s "$stage/payload.tar"
test -s "$stage/payload.sha256"
test -s "$stage/metadata.json"
(cd "$stage" && sha256sum -c payload.sha256)
tar -tf "$stage/payload.tar" >/dev/null
stamp=$(python3 - "$stage/metadata.json" <<'PY'
import json
import datetime as dt
import re
import sys
with open(sys.argv[1], encoding="utf-8") as stream:
metadata = json.load(stream)
stamp = metadata.get("created_utc", "")
if metadata.get("schema") != 1 or metadata.get("host") != "prometheus":
raise SystemExit("Unexpected Prometheus backup metadata")
if not re.fullmatch(r"[0-9]{8}T[0-9]{6}Z", stamp):
raise SystemExit("Invalid Prometheus backup timestamp")
created = dt.datetime.strptime(stamp, "%Y%m%dT%H%M%SZ").replace(tzinfo=dt.timezone.utc)
age = dt.datetime.now(dt.timezone.utc) - created
if age.total_seconds() < -300 or age > dt.timedelta(hours={{ atlas_prometheus_pull_max_age_hours }}):
raise SystemExit("Prometheus backup is outside the configured freshness window")
print(stamp)
PY
)
if [[ -e "$snapshots/$stamp" ]]; then
cmp "$stage/payload.sha256" "$snapshots/$stamp/payload.sha256"
cmp "$stage/metadata.json" "$snapshots/$stamp/metadata.json"
(cd "$snapshots/$stamp" && sha256sum -c payload.sha256)
rm -rf -- "${stage:?}"
stage=''
else
chown -R root:root "$stage"
chmod 0700 "$stage"
chmod 0600 "$stage/payload.tar" "$stage/payload.sha256" "$stage/metadata.json"
mv -- "$stage" "$snapshots/$stamp"
stage=''
fi
latest_link=$(readlink "$backup_root/latest" 2>/dev/null || true)
latest_stamp=${latest_link##*/}
if [[ -z "$latest_stamp" || "$stamp" > "$latest_stamp" ]]; then
ln -s "snapshots/$stamp" "$backup_root/.latest.new"
mv -Tf -- "$backup_root/.latest.new" "$backup_root/latest"
fi
python3 /usr/local/libexec/atlas-prometheus-prune "$snapshots" \
{{ atlas_prometheus_pull_keep_daily }} {{ atlas_prometheus_pull_keep_weekly }} {{ atlas_prometheus_pull_keep_monthly }}
echo "Verified and published Prometheus backup $stamp"

View File

@@ -0,0 +1,10 @@
[Unit]
Description=Schedule Atlas pull of prepared Prometheus backups
[Timer]
OnCalendar={{ atlas_prometheus_pull_calendar }}
Persistent=true
Unit=atlas-prometheus-pull.service
[Install]
WantedBy=timers.target

View File

@@ -0,0 +1,103 @@
---
- name: Validate Prometheus backup export identity inputs
tags: [services, backup, prometheus_backup]
ansible.builtin.assert:
that:
- inventory_hostname == 'prometheus'
- server_backup_username is match('^[a-z_][a-z0-9_-]*$')
- server_backup_username not in ['root', server_username]
- server_backup_export_root.startswith('/var/lib/')
- server_backup_public_key_name is match('^[a-z0-9_-]+$')
- hostvars['atlas'].atlas_manage_prometheus_backup_pull | default(false) | bool
fail_msg: Enable Atlas and Prometheus backup roles together with dedicated identity settings.
when: server_backup_export_enabled | bool
- name: Create dedicated Prometheus backup export group
tags: [services, backup, prometheus_backup]
ansible.builtin.group:
name: "{{ server_backup_username }}"
system: true
state: present
when: server_backup_export_enabled | bool
- name: Create locked Prometheus backup export account
tags: [services, backup, prometheus_backup]
ansible.builtin.user:
name: "{{ server_backup_username }}"
group: "{{ server_backup_username }}"
groups: []
append: false
comment: Read-only prepared backup export for Atlas
home: "{{ server_backup_export_root }}"
create_home: false
shell: /bin/bash
password_lock: true
system: true
state: present
when: server_backup_export_enabled | bool
- name: Require restricted rrsync helper on Prometheus
tags: [services, backup, prometheus_backup]
ansible.builtin.stat:
path: "{{ server_backup_rrsync_path }}"
register: server_backup_rrsync_file
when: server_backup_export_enabled | bool
- name: Validate restricted rrsync helper
tags: [services, backup, prometheus_backup]
ansible.builtin.assert:
that:
- server_backup_rrsync_file.stat.exists
- server_backup_rrsync_file.stat.isreg
- server_backup_rrsync_file.stat.pw_name == 'root'
fail_msg: Rocky rsync must provide the root-owned rrsync support script.
when: server_backup_export_enabled | bool
- name: Create prepared backup export root
tags: [services, backup, prometheus_backup]
ansible.builtin.file:
path: "{{ server_backup_export_root }}"
state: directory
owner: root
group: "{{ server_backup_username }}"
mode: "0750"
when: server_backup_export_enabled | bool
- name: Create restricted Prometheus backup SSH directories
tags: [services, backup, prometheus_backup]
ansible.builtin.file:
path: "{{ item }}"
state: directory
owner: root
group: "{{ server_backup_username }}"
mode: "0750"
loop:
- "{{ server_backup_export_root }}/.ssh"
- "{{ server_backup_export_root }}/.ssh/authorized_keys.d"
when: server_backup_export_enabled | bool
- name: Read Atlas public key for Prometheus backup pull
tags: [services, backup, prometheus_backup]
ansible.builtin.slurp:
src: "{{ hostvars['atlas'].atlas_prometheus_pull_private_key_path | default('/etc/atlas-prometheus-pull/id_ed25519') }}.pub"
delegate_to: atlas
become: true
register: server_backup_atlas_public_key
when:
- server_backup_export_enabled | bool
- not ansible_check_mode
- name: Authorize only restricted read-only backup access from Atlas
tags: [services, backup, prometheus_backup]
ansible.builtin.copy:
content: >-
{{ 'command="/usr/bin/python3 ' ~ server_backup_rrsync_path ~ ' -ro '
~ server_backup_export_root ~ '/versions",restrict '
~ (server_backup_atlas_public_key.content | b64decode | trim) ~ '\n' }}
dest: "{{ server_backup_export_root }}/.ssh/authorized_keys.d/{{ server_backup_public_key_name }}"
owner: root
group: "{{ server_backup_username }}"
mode: "0640"
when:
- server_backup_export_enabled | bool
- not ansible_check_mode

View File

@@ -0,0 +1,90 @@
---
- name: Validate Prometheus backup export job inputs
tags: [services, backup, prometheus_backup]
ansible.builtin.assert:
that:
- server_backup_export_source_keep | int >= 2
- server_backup_export_paths | length > 0
- server_backup_export_paths | unique | length == server_backup_export_paths | length
- >-
server_backup_export_paths
| select('match', '^[a-zA-Z0-9][a-zA-Z0-9._/-]*$') | list | length
== server_backup_export_paths | length
- >-
server_backup_export_paths
| reject('search', '(^|/)\.\.(/|$)') | list | length
== server_backup_export_paths | length
- >-
server_backup_export_excludes
| select('match', '^[a-zA-Z0-9][a-zA-Z0-9._/-]*$') | list | length
== server_backup_export_excludes | length
- >-
server_backup_export_excludes
| reject('search', '(^|/)\.\.(/|$)') | list | length
== server_backup_export_excludes | length
fail_msg: Define safe relative paths and at least two prepared export versions.
when: server_backup_export_enabled | bool
- name: Validate Prometheus backup export calendar
tags: [services, backup, prometheus_backup]
ansible.builtin.command:
argv: [systemd-analyze, calendar, "{{ server_backup_export_calendar }}"]
changed_when: false
check_mode: false
when: server_backup_export_enabled | bool
- name: Ensure prepared Prometheus backup versions directory exists
tags: [services, backup, prometheus_backup]
ansible.builtin.file:
path: "{{ server_backup_export_root }}/versions"
state: directory
owner: root
group: "{{ server_backup_username }}"
mode: "0750"
when: server_backup_export_enabled | bool
- name: Install Prometheus backup export helper
tags: [services, backup, prometheus_backup]
ansible.builtin.template:
src: prometheus-backup-export.sh.j2
dest: /usr/local/sbin/prometheus-backup-export
owner: root
group: root
mode: "0750"
when: server_backup_export_enabled | bool
- name: Install Prometheus backup export systemd units
tags: [services, backup, prometheus_backup]
ansible.builtin.template:
src: "{{ item }}.j2"
dest: "/etc/systemd/system/{{ item }}"
owner: root
group: root
mode: "0644"
loop:
- prometheus-backup-export.service
- prometheus-backup-export.timer
loop_control:
label: "{{ item }}"
register: server_backup_export_units
when: server_backup_export_enabled | bool
- name: Reload systemd after Prometheus backup export unit changes
tags: [services, backup, prometheus_backup]
ansible.builtin.systemd:
daemon_reload: true
when:
- server_backup_export_enabled | bool
- server_backup_export_units is changed
- not ansible_check_mode
- name: Enable Prometheus backup export timer only after explicit activation
tags: [services, backup, prometheus_backup]
ansible.builtin.systemd:
name: prometheus-backup-export.timer
enabled: true
state: started
when:
- server_backup_export_enabled | bool
- server_backup_export_start_timer | bool
- not ansible_check_mode

View File

@@ -53,6 +53,12 @@
tags: [services, podman]
ansible.builtin.include_tasks: podman-compose.yml
- name: Import Prometheus backup export identity tasks
ansible.builtin.import_tasks: backup_export_identity.yml
- name: Import Prometheus backup export job tasks
ansible.builtin.import_tasks: backup_export_job.yml
- name: Ensure server SSH authorized key fragments directory exists
tags: [services, ssh]
ansible.builtin.file:
@@ -77,13 +83,17 @@
when: server_ssh_authorized_keys | length > 0
- name: Configure server SSH authorized key fragments
tags: [services, ssh]
tags: [services, ssh, prometheus_backup]
ansible.builtin.lineinfile:
path: /etc/ssh/sshd_config
regexp: '^\s*AuthorizedKeysFile\s+'
line: >-
AuthorizedKeysFile {{ server_ssh_authorized_keys | map(attribute='name')
| map('regex_replace', '^', '%h/.ssh/authorized_keys.d/') | join(' ') }}
AuthorizedKeysFile {{
((server_ssh_authorized_keys | map(attribute='name')
| map('regex_replace', '^', '%h/.ssh/authorized_keys.d/') | list)
+ (['%h/.ssh/authorized_keys.d/' ~ server_backup_public_key_name]
if server_backup_export_enabled | bool else [])) | join(' ')
}}
state: present
validate: "sshd -t -f %s"
notify: Reload SSH service
@@ -100,11 +110,13 @@
notify: Reload SSH service
- name: Restrict SSH login to allowed users on server
tags: [services]
tags: [services, prometheus_backup]
ansible.builtin.lineinfile:
path: /etc/ssh/sshd_config
regexp: '^\s*AllowUsers\s+'
line: "AllowUsers {{ server_sshd_allow_users | join(' ') }}"
line: >-
AllowUsers {{ (server_sshd_allow_users
+ ([server_backup_username] if server_backup_export_enabled | bool else [])) | join(' ') }}
state: present
validate: "sshd -t -f %s"
notify: Reload SSH service

View File

@@ -0,0 +1,15 @@
[Unit]
Description=Prepare a read-only Prometheus application backup for Atlas
RequiresMountsFor=/opt/npm /opt/gitea {{ server_backup_export_root }}
ConditionFileIsExecutable=/usr/local/sbin/prometheus-backup-export
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/prometheus-backup-export
User=root
Group=root
UMask=0077
TimeoutStartSec=infinity
Nice=10
IOSchedulingClass=best-effort
IOSchedulingPriority=7

View File

@@ -0,0 +1,92 @@
#!/usr/bin/env bash
set -Eeuo pipefail
umask 077
export_root={{ server_backup_export_root | quote }}
versions="$export_root/versions"
stack_unit=podman-compose-server.service
stamp=$(date -u +%Y%m%dT%H%M%SZ)
stage=''
stack_stopped=false
exec 9>/run/lock/prometheus-backup-export.lock
flock -n 9 || { echo 'A backup export is already running' >&2; exit 1; }
cleanup() {
local rc=$?
trap - EXIT
if "$stack_stopped"; then
if systemctl is-active --quiet "$stack_unit"; then
systemctl restart "$stack_unit" || rc=1
else
systemctl start "$stack_unit" || rc=1
fi
fi
if (( rc != 0 )) && [[ -n "$stage" && -d "$stage" ]]; then
rm -rf -- "$stage"
fi
exit "$rc"
}
trap cleanup EXIT
trap 'exit 129' HUP
trap 'exit 130' INT
trap 'exit 143' TERM
systemctl is-active --quiet "$stack_unit" || {
echo 'The managed Compose stack must be active before preparing a backup' >&2
exit 1
}
paths=(
{% for path in server_backup_export_paths %}
{{ path | quote }}
{% endfor %}
)
excludes=(
{% for path in server_backup_export_excludes %}
--exclude={{ path | quote }}
{% endfor %}
)
for path in "${paths[@]}"; do
[[ -e "/$path" ]] || { echo "Required backup path missing: /$path" >&2; exit 1; }
done
[[ ! -e "$versions/$stamp" ]] || { echo "Export version already exists: $stamp" >&2; exit 1; }
stage=$(mktemp -d "$export_root/.staging.XXXXXXXX")
# SQLite databases and their accompanying files are copied while both
# managed containers are stopped. The EXIT trap restarts the stack on error.
stack_stopped=true
systemctl stop "$stack_unit"
tar --acls --xattrs --selinux "${excludes[@]}" -C / -cf "$stage/payload.tar" "${paths[@]}"
systemctl start "$stack_unit"
for container in nginx-proxy-manager gitea; do
running=false
for _ in {1..30}; do
if [[ $(podman inspect --format '{{ '{{.State.Running}}' }}' "$container" 2>/dev/null) == true ]]; then
running=true
break
fi
sleep 2
done
"$running" || { echo "Container did not restart: $container" >&2; exit 1; }
done
stack_stopped=false
tar -tf "$stage/payload.tar" >/dev/null
(cd "$stage" && sha256sum payload.tar >payload.sha256)
printf '{"schema":1,"host":"prometheus","created_utc":"%s"}\n' "$stamp" >"$stage/metadata.json"
chown root:{{ server_backup_username }} "$stage" "$stage/payload.tar" "$stage/payload.sha256" "$stage/metadata.json"
chmod 0750 "$stage"
chmod 0640 "$stage/payload.tar" "$stage/payload.sha256" "$stage/metadata.json"
mv -- "$stage" "$versions/$stamp"
stage=''
ln -s "$stamp" "$versions/.current.new"
mv -Tf -- "$versions/.current.new" "$versions/current"
# Keep a small source-side safety window; Atlas owns long-term retention.
mapfile -t old_versions < <(find "$versions" -mindepth 1 -maxdepth 1 -type d \
-printf '%f\n' | grep -E '^[0-9]{8}T[0-9]{6}Z$' | sort -r | tail -n +{{ server_backup_export_source_keep + 1 }})
for old in "${old_versions[@]}"; do
rm -rf -- "${versions:?}/$old"
done
echo "Prepared Prometheus backup export $stamp"

View File

@@ -0,0 +1,10 @@
[Unit]
Description=Prepare daily Prometheus application backup for Atlas
[Timer]
OnCalendar={{ server_backup_export_calendar }}
Persistent=false
Unit=prometheus-backup-export.service
[Install]
WantedBy=timers.target

74
docs/atlas-dr-lab.md Normal file
View File

@@ -0,0 +1,74 @@
# Isolated Atlas DR lab
This is a **scaled rehearsal**, not a substitute for a full-data restore. The
`atlas-dr-lab` libvirt VM on Ikaros was left **shut off** on 2026-09-30. Its
persistent volumes are in the default libvirt pool: the current 30 GiB OS
volume `atlas-dr-lab-os-rebuild2.qcow2`, the pre-rebuild OS volume
`atlas-dr-lab-os.qcow2`, and four independent 4 GiB
`atlas-dr-lab-data{1,2,3,4}.qcow2` volumes. The VM uses libvirt's `default`
NAT network (last DHCP address `192.168.122.168`), 2 vCPU, and 4 GiB RAM.
The data disks have `virtio-atlasdrdata{1,2,3,4}` serials. No physical disk or
production Atlas storage is attached. The VM has no autostart.
## Rebuild inputs and isolation
- Use Rocky's **9.8 GenericCloud Base x86_64** image
`Rocky-9-GenericCloud-Base-9.8-20260525.0.x86_64.qcow2` from
`https://download.rockylinux.org/pub/rocky/9.8/images/x86_64/`.
Verify its `.CHECKSUM` file; the observed SHA-256 was
`92c206cc6f790c61583247eefe87890f8828420662c17cacf247cec78ab4eec8`.
- Use a dedicated lab-only inventory merged **after** the repository
inventory, and always `--limit atlas_dr_lab`. The temporary 2026-09-30
inventory/playbook and logs are in `/tmp/atlas-dr-lab-image/`; copy a
sanitized inventory to durable private storage before `/tmp` is cleared if
the lab will be repeated. Never reuse `host_vars/atlas.yml`, production
Vault secrets, or production disk by-id paths for the lab.
- The lab host belongs to `platform_rocky` and `atlas`. It uses `dradmin`
(UID/GID 1000) with the operator's **public** SSH key and a random,
unknown password hash, the libvirt DHCP address, pool `zpool`, mount root
`/zpool`, the four `virtio-atlasdrdata*` by-id paths, a 1 GiB backup
reservation, and `rocky_manage_openzfs_repo: true` with only `zfs` in
`host_packages`. The following gates remain false: sharing, firewall,
media stack, ZFS timers, Borg, USB, monitoring, and Prometheus pull.
`atlas_manage_storage` is true. Set `atlas_create_pool: true` **only for the
first disposable pool creation**, then set it false before any later run.
- A minimal lab playbook selects `atlas_dr_lab`, `become: true`, and the
existing `packages_rocky` and `profile_atlas` roles. Use a separate
`ANSIBLE_CONFIG` without the production Vault password script, and keep
host-key checking on with a lab-specific known-hosts file. The 2026-09-30
runs used `-i ansible/inventory/hosts.yml -i <lab-inventory.yml>` and
`--limit atlas_dr_lab` throughout.
## Rehearsal and narrow checks
1. Before any pool operation, compare `virsh -c qemu:///system domblklist
atlas-dr-lab` with the four intended qcow2 paths, and in the guest compare
`/dev/disk/by-id/virtio-atlasdrdata*` with `lsblk`. Do not proceed if a
physical disk or production identity appears.
2. For a first-time disposable build only, run the lab playbook with
`--tags pool` and `atlas_create_pool: true`, then immediately set the gate
false. Run the full lab playbook and check `zpool status -P zpool`,
`zfs list -r zpool`, SELinux, and failed systemd units.
3. Write a non-sensitive canary under the lab `/zpool/archive` and snapshot
it. Record the pool GUID and canary SHA-256. Export the lab pool cleanly,
shut down the VM, and replace **only the OS volume** with a fresh verified
Rocky image. Preserve all four data volumes. Reconfigure cloud-init for a
new instance; the seed CD-ROM must use **SATA**. The SCSI seed attachment
tried during this rehearsal was not detected by cloud-init and was
replaced with a SATA attachment before proceeding.
4. On the new OS, apply `packages_rocky` to reinstall OpenZFS. First run
`zpool import -d /dev/disk/by-id` **without importing**, compare GUID and
vdev membership, then use ordinary `zpool import -d /dev/disk/by-id zpool`.
Do not use `-f`, `-F`, `-X`, rollback, or pool creation.
5. Reapply `profile_atlas` with the lab gates and `atlas_create_pool: false`.
Verify the canary, restored snapshot file in an empty temporary directory,
dataset hierarchy, SELinux, and pool health. A second full playbook run
should report `changed=0`. Remove temporary restored files and shut down
the VM after testing.
The observed 2026-09-30 pool GUID was `8880368391795119587`; the canary
SHA-256 was `949701c7a95fadae1fddc21abe846c4312212dbfeb7477948f3188fc3ec34a78`.
The post-rebuild Ansible run succeeded, a repeat run reported `changed=0`,
12 datasets and the original snapshot were present, and the pool was healthy.
The snapshot-restored file matched content and basic metadata. See
[`atlas-recovery.md`](atlas-recovery.md) for the production runbook and limits.

141
docs/atlas-recovery.md Normal file
View File

@@ -0,0 +1,141 @@
# Atlas recovery runbook
This runbook is for a **replacement Rocky Linux 9 installation**, not a normal
playbook run. A scaled whole-OS rebuild with a disposable pool passed in an
isolated VM on 2026-09-30, but no production-size whole-host recovery has been
tested. The existing production pool must be imported, never created or
rewritten. The provisional targets are **RPO 24 hours**
and **RTO 72 hours**, for Archive and Atlas services alike. They are planning
objectives, not demonstrated recovery times. The manual USB cadence may leave
an older copy; a recent Borg archive is needed to meet the RPO after total
pool loss.
## Before an incident
- Keep an offline copy of the encrypted Ansible Vault, its unlock material,
the exported Borg repository key, and the Borg passphrase. Do not store
unlock material in this repository or in a recovery command line.
On 2026-09-30 the operator confirmed these are available independently of
Atlas and the Ansible controller; their usability has not been tested here.
- Keep the Atlas installation media and a reproducible checkout of this
repository available independently of Atlas. Record the exact Git revision
used for a successful deployment.
- Record the pool's current disk identities with `zpool status -P zpool` and
`lsblk -o NAME,SIZE,MODEL,SERIAL,FSTYPE,UUID`. Compare these with
`atlas_zpool_disks` before touching a replacement host. The `host_vars`
values are historical identifiers, not evidence that a newly attached disk
is the same device.
- Verify that the latest hourly/daily snapshots, Borg archive, and offline USB
version exist and note their timestamps. A timer being enabled is not proof
that a backup completed.
## Incident gate
1. Identify whether the fault is the OS disk, one or more pool disks, accidental
deletion, or an unavailable host. Preserve failed media when possible.
2. Stop writes to affected services and capture the last known good backup
timestamps. Do not run `zpool create`, `zpool destroy`, `zfs rollback`,
`zpool import -F`, `zpool import -X`, `zpool import -f`, or disk formatting
as a diagnostic shortcut.
3. Choose one recovery source below. Do not merge several sources into the
production namespace without comparing their timestamps and content.
## Rebuild the OS and import the existing pool
1. Install Rocky Linux 9 on a **separate system disk**. Configure basic network,
SSH, a temporary sudo administrator, SELinux enforcing, and the current
OpenZFS kmod repository. Keep the pool drives untouched.
2. Run read-only identification: `lsblk -f`, `zpool import`, and
`zpool import -d /dev/disk/by-id`. Check the pool GUID, vdev layout, and
stable drive identities against the incident record. If any differ, stop.
3. Import only after matching the expected pool and host ownership. A pool
cleanly exported from the old host can be imported with
`zpool import -d /dev/disk/by-id zpool`. If it reports that the pool is
active elsewhere or needs a rewind/force, stop and investigate rather than
adding flags. Verify with `zpool status -v zpool`, `zfs list -r zpool`,
`zfs get -r mountpoint,canmount zpool`, and `findmnt -R /zpool`.
4. Leave `atlas_create_pool: false`. Ensure `host_vars/atlas.yml` reflects the
replacement host's actual SSH address and disk identities before running
Ansible. Apply `ansible/site.yml --limit atlas` with the bootstrap admin
connection override as documented in the Atlas setup section of README.
This may start shares/services, so keep clients disconnected or services
gated until data and permissions are verified.
5. Check `getenforce`, `zpool status -v zpool`, `systemctl --failed`, SSH,
firewalld, Cockpit, NFS, SMB, and the backup/monitoring timers. Do not
report recovery complete on the basis of Ansible success alone.
## Choose the data source
- **Local snapshot, pool intact:** inspect `zfs list -t snapshot -r zpool`.
Mount/access the chosen snapshot read-only and copy selected files to an
empty staging directory; compare content, owner, mode, mtime, and POSIX ACL.
Move into the live namespace only after an operator-approved scope review.
Do not use an automatic rollback: it can discard newer changes in the
dataset and descendants.
- **Offline USB:** verify the configured LUKS and ext4 UUIDs from
`host_vars/atlas.yml` before unlocking. Mount ext4 read-only with `ro,noload`,
use only a published `atlas/latest` version, and restore to an empty staging
directory. Compare checksums and metadata. The USB copy intentionally omits
generic xattrs and SELinux labels; relabel only the restored destination.
Never run the backup service to perform a restore.
- **Hetzner Borg:** use the dedicated pinned host key, repository path,
offline exported recovery key, and Vault-backed passphrase. List archives
and extract a selected archive into an empty staging directory, never the
live `/zpool` tree. A repository check and sample restore were previously
performed; that does not prove this incident's archive is complete. Compare
content and metadata before publication. Avoid `borg break-lock` while any
backup/check job may still be active.
After publishing restored files, run the explicit Ansible `restorecon` tag only
for the paths actually restored, for example:
```bash
ansible-playbook ansible/site.yml --limit atlas --tags restorecon \
-e '{"atlas_restorecon_paths":["/zpool/archive"]}'
```
Then check ownership/ACLs, application-specific integrity, SMB/NFS client
access, backup service health, and `zpool status -v zpool`. Reconnect clients
only after these checks pass. Record the last recoverable timestamp (actual
RPO) and elapsed service outage (actual RTO) in the incident log.
## Scaled isolated rehearsal (2026-09-30)
The lab setup, repeatable checks, and preserved VM state are recorded in
[`atlas-dr-lab.md`](atlas-dr-lab.md).
On Ikaros, a local libvirt `atlas-dr-lab` VM used a 30 GiB Rocky 9.8 system
disk and four separate, disposable 4 GiB virtio data disks with stable
`/dev/disk/by-id` identities. The official Rocky cloud image matched its
published SHA-256. The lab inventory was separate from production, used a
fresh lab-only password hash and the operator's public SSH key, and disabled
sharing, Borg, USB backup, monitoring, media services, and the Prometheus pull.
No production disk, Vault secret, or production data was attached or copied.
1. The existing `packages_rocky` and `profile_atlas` roles installed OpenZFS,
created a RAIDZ2 `zpool` through the explicit one-time pool gate, and built
all 12 declared datasets with a lab-sized 1 GiB backup reservation. The
pool creation gate was set false immediately afterward.
2. A 4 MiB canary file was written under the lab `archive` dataset and a ZFS
snapshot created. The pool was cleanly exported and the VM shut down.
3. Only the system-disk volume was replaced by a fresh Rocky cloud image;
the four virtio data volumes were retained. Ansible reinstalled OpenZFS.
Read-only `zpool import -d /dev/disk/by-id` showed the expected RAIDZ2
topology and pool GUID `8880368391795119587` before an ordinary import
without `-f`, rewind, or rollback.
4. The imported pool was healthy. The canary SHA-256 matched its pre-rebuild
value. `profile_atlas` completed against the imported pool and a second
run reported `changed=0`. A file restored from the preserved snapshot into
`/var/tmp` matched SHA-256, owner, group, mode, size, and mtime; the temporary
copy was removed. Final checks found SELinux Enforcing, 12 datasets, the
snapshot, no failed units, and a healthy pool. The VM was shut down while
retaining its disposable volumes for a future rehearsal.
This proves the **sequence** for a cleanly exported, small pool and the tested
Ansible subset, not recovery duration or capacity at 2 TB. The earlier
2026-09-25 independent production ZFS/USB file restores and the earlier Borg
temporary-directory restore remain separate evidence. The VM did not restore
production USB/Borg archives, exercise services with production data, test an
unclean import, or prove the provisional RPO/RTO. Before relying on 24h/72h,
measure a representative full restore and service cutover in a suitably sized
future change window. Never use the production Atlas pool for a rehearsal.

View File

@@ -0,0 +1,20 @@
# Atlas SMB/NFS namespace decision
Decision date: 2026-09-30. Keep the current namespaces **separate**.
- `/zpool/archive` is the SMB3 `Archive` share for authorized Samba accounts.
- `/zpool/media/photobook` is the Aegis-only NFSv4 export, `all_squash`-mapped
to UID/GID `1100`.
- No new dual-protocol namespace, broad export, group, or ACL model is needed.
Existing permissions and client access remain unchanged.
The two paths serve different ownership and exposure needs. A common namespace
would expand the permissions design and require same-file SMB/NFS interoperability
testing without a present requirement. Revisit only when a specific workflow
needs both protocols on the same files; then decide UID/GID, group, POSIX ACL,
SELinux policy and client behavior before changing exports or permissions.
Read-only Atlas verification on 2026-09-30 confirmed that Samba `Archive` points
to `/zpool/archive`, NFS exports `/zpool/media/photobook` only to
`192.168.178.54` with `all_squash` and anonymous UID/GID `1100`, both datasets
are distinct, and `zpool` is healthy. No sharing configuration was changed.

63
docs/atlas-updates.md Normal file
View File

@@ -0,0 +1,63 @@
# Atlas Rocky/OpenZFS update and reboot procedure (draft)
This is an operator-controlled maintenance procedure. The playbook does not
reboot Atlas, replace a pool device, or perform a pool feature upgrade.
## Preflight
1. Schedule an outage and confirm no Borg, USB, snapshot, scrub, or resilver
job is active. A service in `activating` is still active; do not interrupt it.
2. Check `zpool status -v zpool` (including scrub status), `zfs list -r zpool`,
`systemctl --failed`, and `systemctl list-timers --all`. Resolve pool errors
first. Record current `uname -r`, `modinfo zfs | grep '^version:'`,
`rpm -q kernel-core kmod-zfs zfs`, and the current boot entry.
3. Confirm a recent successful Borg archive and a usable snapshot. Confirm
the latest published offline USB version and its physical availability;
do not start a USB backup merely to satisfy a checklist without capacity,
UUID, and operator checks. Record timestamps, not just timer state.
4. Ensure console/KVM or another independent recovery route is available.
Check free space in `/boot` and the root filesystem. Review proposed DNF
transactions before consenting to package changes.
## Change window
1. Stop client writes and quiesce stateful applications deliberately. Record
which services were stopped; do not assume `ansible-playbook --check` does
this. Avoid updating during a running scrub or backup.
2. Use `dnf upgrade --assumeno` first to review the kernel, `kmod-zfs`, `zfs`,
and dependencies. Confirm a matching kmod will be available for the target
kernel. If compatibility is uncertain, defer the update.
3. Apply the approved DNF transaction. Do not run `zpool upgrade` or enable
new pool feature flags as part of ordinary OS maintenance; that can remove
downgrade options. Preserve at least one known-good boot entry.
4. Reboot **manually** during the agreed outage. Ansible must not trigger it.
## Post-boot gate
1. Verify `uname -r`, `modinfo zfs`, `rpm -q kernel-core kmod-zfs zfs`,
`zpool status -v zpool`, `zfs list -r zpool`, and `findmnt -R /zpool`.
2. Verify SELinux remains enforcing; inspect `systemctl --failed` and the
journal for ZFS, mount, SSH, NFS, SMB, Cockpit, Podman, and backup errors.
3. Validate a read-only file listing through SMB and an NFS client access
check before reopening writes. Check the rootless temporary services and
all backup/monitoring timers. Run the Atlas health monitor in `--dry-run`
mode, then a real check after inspection.
4. Re-enable clients and record versions, downtime, anomalies, and next
successful snapshot/Borg run. A green boot alone is not a completed update.
## Failure response
If the new kernel cannot load ZFS, boot the previous known-good kernel from
the console and inspect package/kmod matching before trying another reboot.
Do not force-import, rewind, clear errors, or upgrade pool features to make a
failed OS update appear successful. Preserve logs and stop for a recovery
decision if the pool does not import cleanly.
The procedure-definition item is complete, but the procedure is **not yet
rehearsed** on a replacement host or during a real Atlas update. Record the
first controlled execution and its post-boot evidence separately.
Read-only preflight on 2026-09-30 observed kernel
`5.14.0-687.52.1.el9_8.x86_64`, ZFS module/package `2.2.11-1`, a healthy
`zpool`, enforcing SELinux, and no failed systemd units. This did not review
an upgrade transaction, stop services, or reboot the host.

106
docs/prometheus-backup.md Normal file
View File

@@ -0,0 +1,106 @@
# Prometheus to Atlas backup pull
The playbook and both hosts have the dedicated identity, restricted SSH
access, helpers, and systemd units. A manual export, pull, and temporary
restore passed on 2026-09-30. Both timers are enabled; their first scheduled
runs are pending, so daily operation is not yet verified.
## Declared design
- Prometheus prepares a tar archive of Nginx Proxy Manager and Gitea data,
their managed Compose configuration, SSH/firewalld/WireGuard configuration,
and the Gitea SSH path. NPM access logs and regenerable Gitea logs, sessions,
temporary files, and indexers are excluded. The archive contains credentials,
certificates, and the WireGuard private key: protect both copies accordingly.
- The approved consistency mode stops the managed Compose stack for the local
tar creation at 02:00 Europe/Rome, then restarts it even if archiving fails.
A manual test outside that window requires separate approval.
- Prometheus publishes the archive with its checksum as a versioned, read-only
source under `/var/lib/prometheus-backup-export`. A locked service account
has no sudo or supplementary groups. Its only authorized SSH key is forced
through Rocky's `rrsync -ro`; root owns the key file and export directories,
so the account cannot add an unrestricted key or change prepared data.
- Atlas generates and retains the private Ed25519 identity under
`/etc/atlas-prometheus-pull`. Its pinned Prometheus host key came through
the controller's already strict SSH trust; the observed fingerprint was
`SHA256:rfedk7DHI9mLB3UHk/4F3HHlSIiswtCAFsAXvfh6iXk` on 2026-09-30.
Atlas pulls only the prepared `current/` version, verifies SHA-256, tar
readability, metadata, and source freshness, then publishes atomically
below `/zpool/backup/hosts/prometheus/snapshots`. Long-term retention runs
only after publication. A local `rrsync` fixture verified the in-tree
`current` symlink. A live Atlas-to-Prometheus SSH test verified that the
account could list only the prepared versions directory,
cannot obtain a shell, and cannot write to the export. The key is restricted
to `/var/lib/prometheus-backup-export/versions`, not the account's `.ssh`.
- Approved source preparation is 02:00 Europe/Rome, pull 03:00, three source
versions, and 30 daily/8 weekly/12 monthly Atlas versions. The source
timer is non-persistent to avoid an unexpected outage after a missed run.
Atlas rejects a prepared source older than 24 hours.
- The Atlas pull joins the existing health monitor's timer/failure checks
only when enabled. Its failure hook uses 45Drives Alerts; email delivery
is not claimed. A failed source preparation should produce a stale-source
pull failure, not a silently successful reuse of an old archive.
## Activation and verification
1. The user confirmed downtime/consistency mode, schedule, retention, and
targeted configuration scope. Review the tar path list and exclusions
against the actual containers.
2. The identity and units are deployed. Re-run the targeted
check, confirm the Atlas public key remains only the restricted Prometheus
account's key, and verify `sshd -T -C user=prometheus-backup,...` plus
read-only SSH denial tests after any SSH configuration change.
3. During an agreed window, start the Prometheus export service manually.
Confirm Compose is healthy afterward, inspect the archive without exposing
file contents, and verify the checksum/metadata.
4. Start the Atlas pull service manually. Confirm the SSH host pin, source
freshness, checksum, tar listing, published `latest`, retention behavior,
clean temporary directories, and healthy pool.
5. Independently restore the selected archive to an empty staging directory
(never `/`) and compare the SQLite databases, Git repositories, NPM data,
Compose file, permissions, and representative files. Test application
startup only in an isolated environment or an approved restore window.
6. The two timers were enabled after the manual test. Verify their calendars
and the next actual run. A successful manual test is not proof of scheduled
operation.
Narrow static validation:
```bash
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
ansible-playbook ansible/site.yml --syntax-check
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
ansible-playbook ansible/site.yml --limit prometheus,atlas \
--tags prometheus_backup --check --diff
```
Do not run the export service as part of a routine playbook deployment. The
service restart and any restore/cutover require separate operator decisions.
On 2026-09-30 the initial targeted `--check --diff` run ended `changed=0`
with gates false. After enabling **implementation only**, a targeted real run
installed the identities and units; both timers were confirmed `disabled` and
`inactive`, the Compose stack stayed active, and the new account was locked
with no supplementary groups. No application was stopped.
The rendered shell helpers passed `bash -n` and ShellCheck; the retention
helper passed an isolated 400-version fixture. These static/isolated checks
were followed by live SSH, export, pull, and temporary restore checks.
Read-only preflight on 2026-09-30 found the Compose service active, all
declared source paths present, both timers inactive, and no prepared versions.
The source filesystem had about 6.0 GB free. Of the 2.1 GB NPM data tree,
2.1 GB was excluded access logs, so the expected archive is much smaller than
the raw tree size; capacity still needs verification after actual exports.
The manual export produced a 285,777,920-byte tar (273 MiB allocated at the
source), and Prometheus retained about 5.8 GB free. NPM and Gitea restarted;
both containers were running and their local HTTP endpoints returned 200.
Atlas pulled the same version, verified SHA-256, published `latest`, and kept
the pool healthy. A full extract to `/var/tmp` yielded 4,747 files; both
SQLite databases passed `PRAGMA integrity_check`, and one restored Gitea Git
repository passed `git fsck`. The temporary restore directory was removed.
This did not test application startup on an isolated host.
After these checks, Ansible enabled the Prometheus 02:00 Europe/Rome export
timer and Atlas 03:00 Europe/Rome pull timer. The next scheduled occurrences
were displayed for 2026-10-01. Atlas' health monitor now includes the pull
timer. Check both actual service results after the first scheduled run before
claiming unattended operation.