diff --git a/AGENTS.md b/AGENTS.md index 4148393..bb031b9 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -226,7 +226,7 @@ successfully. The first monthly scrub remains a runtime check. read-only ZFS snapshot test restored one file to `/var/tmp`, confirmed matching contents, ownership, mode, mtime and ACL, then removed its temporary copy and on-demand mount. This is a file-level smoke test, not full dataset recovery. An independent USB file restore passed on 2026-09-25 with matching - content and metadata; full disaster recovery remains a separate Priority 2 task. + content and metadata; the later scaled OS-rebuild rehearsal is documented under Priority 2. - [x] Add monitoring and alerting for pool health, scrub/resilver, SMART data, temperatures, free space, snapshot/local-backup growth, Hetzner Storage Box quota, and failed maintenance or backup timers. The half-hourly Atlas health monitor and systemd final-failure hooks are deployed. A live probe @@ -237,15 +237,26 @@ successfully. The first monthly scrub remains a runtime check. on Atlas, but a new real failure notification has not been deliberately triggered. ### Priority 2 - NAS operability and recovery -- [ ] Document and test disaster recovery: rebuild Atlas with Ansible, import the existing pool, restore - from snapshot/USB/Hetzner, preserve Vault and Borg recovery material offline, and define RPO/RTO. -- [ ] Define a controlled Rocky kernel/OpenZFS update and reboot procedure. -- [ ] Add the Atlas-initiated least-privilege Prometheus backup pull: Prometheus exposes only prepared +- [x] Document and test disaster recovery in `docs/atlas-recovery.md`: the operator confirmed Vault + and Borg recovery material is available offline; provisional targets are RPO 24h/RTO 72h. On + 2026-09-30 an isolated small Rocky VM was rebuilt with the Atlas Ansible roles, imported its + preserved RAIDZ2 pool without force/rewind, and restored a file from the preserved snapshot; + the second Ansible run was idempotent. Earlier independent production ZFS, USB, and Borg file + restore tests remain separate evidence. A production-size full restore, unclean import, and + measured 24h/72h compliance are not claimed. +- [x] Define a controlled Rocky kernel/OpenZFS update and reboot procedure in + `docs/atlas-updates.md`. The first real change-window execution is not yet + validated; the procedure never reboots automatically or upgrades pool features. +- [x] Add the Atlas-initiated least-privilege Prometheus backup pull: Prometheus exposes only prepared read-only dumps through a dedicated account and Atlas retains the private SSH key, pinned host key, - atomic pull, verification, retention and systemd service/timer. -- [ ] Decide whether a common SMB/NFS namespace is required. `Archive` (SMB) and `photobook` (NFS) are - intentionally distinct today; only if a shared namespace is selected, finalize its UID/GID, group, - and POSIX ACL model and test the same files through both protocols. + atomic pull, verification, retention and systemd service/timer. The dedicated key/account and unit + files are deployed; live read-only SSH, shell denial, and write denial were verified. On 2026-09-30 + a manual export, Atlas pull, checksum verification, and temporary restore passed; both SQLite + databases passed integrity checks and a restored Git repository passed `git fsck`. Both daily + timers are enabled for 02:00/03:00 Europe/Rome; their first scheduled results remain unverified. +- [x] Decide whether a common SMB/NFS namespace is required: no. `Archive` (SMB) and `photobook` (NFS) + remain intentionally distinct; `docs/atlas-sharing-decision.md` records the decision. No ACL or export + change is authorized by this decision. ### Priority 3 - Service expansion - [ ] After data protection and recovery are validated, populate `/zpool/media/music` and validate Navidrome. diff --git a/README.it.md b/README.it.md index 7675600..e8b4147 100644 --- a/README.it.md +++ b/README.it.md @@ -485,7 +485,7 @@ etichettata di 45Drives Alerts usare ### Timer systemd di Atlas -Tutti i nove timer gestiti sono abilitati. Gli orari sono locali ad Atlas (`Europe/Rome`); Borg e +Tutti i dieci timer gestiti sono abilitati. Gli orari sono locali ad Atlas (`Europe/Rome`); Borg e monitoraggio aggiungono il ritardo casuale indicato. Tutti hanno `Persistent=true`: un evento perso viene recuperato quando il timer torna attivo. @@ -500,10 +500,12 @@ viene recuperato quando il timer torna attivo. | `atlas-borg-check.timer` | `*-*-15 06:00:00` — giorno 15 alle 06:00, più 0–30 min casuali | Controllo repository Borg | | `atlas-usb-reminder.timer` | `Sat *-*-01..07 10:00:00 Europe/Rome` — primo sabato alle 10:00 | Solo promemoria 45Drives Alerts | | `atlas-health-monitor.timer` | `*:0/30` — ogni mezz'ora, più 0–5 min casuali | Controlli di salute in sola lettura | +| `atlas-prometheus-pull.timer` | `*-*-* 03:00:00 Europe/Rome` — ogni giorno alle 03:00 | Pull e verifica del backup preparato su Prometheus | `atlas-usb-backup.service` **non ha timer** e va avviato manualmente. Il timer del fornitore -`zfs-scrub-weekly@zpool.timer` è disabilitato a favore dello scrub mensile. Il futuro pull del backup -Prometheus non ha ancora un timer, perché non è implementato. Durante un backup Borg attivo, +`zfs-scrub-weekly@zpool.timer` è disabilitato a favore dello scrub mensile. Il timer di preparazione +su Prometheus è attivo alle 02:00 Europe/Rome; export, pull e ripristino temporaneo manuali sono +riusciti il 2026-09-30, ma il primo ciclo pianificato va ancora verificato. Durante un backup Borg attivo, `systemctl list-timers` può mostrare `-` per il prossimo evento senza che il timer sia disabilitato. Per vedere la pianificazione corrente: `systemctl list-timers --all` su Atlas. @@ -519,7 +521,10 @@ cutover. L'attuale iCloudPD su Aegis e l'export NFS Photobook restano configurat all'approvazione e alla verifica di questa migrazione separata. Anche il servizio Atlas sarà temporaneo in attesa di Uranus. -Il pull dei backup di Prometheus e i test completi di disaster recovery restano da fare. Il backlog +Il primo ciclo pianificato del backup di Prometheus e una prova di disaster recovery a dimensione reale +restano da verificare. Il 2026-09-30 una VM Rocky isolata ha superato ricostruzione OS con Ansible, +import del pool RAIDZ2 fittizio e ripristino da snapshot; RPO 24 ore/RTO 72 ore restano obiettivi +provvisori, non tempi misurati. Dettagli e limiti sono in `docs/atlas-recovery.md`. Il backlog prioritizzato è in `AGENTS.md`. --- diff --git a/README.md b/README.md index a950a08..f341f42 100644 --- a/README.md +++ b/README.md @@ -500,7 +500,7 @@ monitoring. For a labelled 45Drives Alerts delivery test, use ### Atlas systemd timers -All nine managed timers below are enabled. Times are local to Atlas (`Europe/Rome`); Borg and monitoring +All ten managed timers below are enabled. Times are local to Atlas (`Europe/Rome`); Borg and monitoring add the indicated randomized delay. Every timer has `Persistent=true`, so a missed calendar run is scheduled after the timer becomes active again. @@ -515,10 +515,12 @@ scheduled after the timer becomes active again. | `atlas-borg-check.timer` | `*-*-15 06:00:00` — 15th of the month at 06:00, plus 0–30 min random delay | Borg repository check | | `atlas-usb-reminder.timer` | `Sat *-*-01..07 10:00:00 Europe/Rome` — first Saturday at 10:00 | 45Drives Alerts reminder only | | `atlas-health-monitor.timer` | `*:0/30` — every half-hour, plus 0–5 min random delay | Read-only health checks | +| `atlas-prometheus-pull.timer` | `*-*-* 03:00:00 Europe/Rome` — daily at 03:00 | Pull and verify the prepared Prometheus backup | `atlas-usb-backup.service` has **no timer**: the encrypted USB backup must be started manually. The vendor's `zfs-scrub-weekly@zpool.timer` is intentionally disabled in favor of the monthly scrub. -The future Prometheus backup pull has no timer yet because that workflow is not implemented. While a +The Prometheus export timer runs at 02:00 Europe/Rome; its first scheduled run and the Atlas pull +remain to be observed. A manual export, pull, and temporary restore passed. While a Borg backup is still running, `systemctl list-timers` may show `-` for its next trigger; this does not mean the timer has been disabled. Inspect the current schedule on Atlas with `systemctl list-timers --all`. @@ -534,9 +536,21 @@ state outside `Archive`, then test permissions, SELinux, backups and recovery be The current Aegis iCloudPD service and Atlas Photobook NFS export remain configured until that separate migration is approved and validated; the eventual Atlas service is temporary until Uranus. -Prometheus backup pulls and full disaster-recovery tests remain follow-up work. The prioritized +The first scheduled Prometheus backup runs and production-size disaster-recovery tests remain follow-up work. The prioritized operational backlog is kept in `AGENTS.md`. +Priority 2 procedures and decisions are recorded in +[`docs/atlas-recovery.md`](docs/atlas-recovery.md), +[`docs/atlas-updates.md`](docs/atlas-updates.md), and +[`docs/atlas-sharing-decision.md`](docs/atlas-sharing-decision.md). +The provisional Atlas recovery objectives are RPO 24 hours and RTO 72 hours; +an isolated small-VM OS rebuild, pool import, Ansible reapplication, and +snapshot restore passed, but full-size recovery time is unmeasured. `Archive` (SMB) and +`photobook` (NFS) remain deliberately separate. +The Prometheus pull architecture and manual export/pull/restore evidence are in +[`docs/prometheus-backup.md`](docs/prometheus-backup.md). Both daily timers are +enabled; their first scheduled runs remain to be verified. + ## How layering works A host can intentionally belong to more than one inventory group. The final configuration is the combination of the host and its groups, not a one-host/one-play mapping. diff --git a/ansible/inventory/group_vars/server.yml b/ansible/inventory/group_vars/server.yml index bc026f1..d1b8e19 100644 --- a/ansible/inventory/group_vars/server.yml +++ b/ansible/inventory/group_vars/server.yml @@ -80,5 +80,30 @@ server_sshd_settings: server_sshd_allow_users: - "{{ server_username }}" +server_backup_export_enabled: false +server_backup_username: prometheus-backup +server_backup_public_key_name: atlas-pull +server_backup_export_root: /var/lib/prometheus-backup-export +server_backup_rrsync_path: /usr/share/doc/rsync/support/rrsync +server_backup_export_calendar: "*-*-* 02:00:00 Europe/Rome" +server_backup_export_start_timer: false +server_backup_export_source_keep: 3 +server_backup_export_paths: + - opt/npm/data + - opt/npm/letsencrypt + - opt/gitea/data + - home/git/.ssh + - opt/docker/server/docker-compose.yml + - etc/systemd/system/podman-compose-server.service + - etc/ssh/sshd_config + - etc/ssh/sshd_config.d + - etc/firewalld + - etc/wireguard/wg0.conf +server_backup_export_excludes: + - opt/npm/data/logs + - opt/gitea/data/gitea/log + - opt/gitea/data/gitea/tmp + - opt/gitea/data/gitea/sessions + - opt/gitea/data/gitea/indexers server_ssh_authorized_keys: [] server_ssh_authorized_key_directory: "{{ server_user_home }}/.ssh/authorized_keys.d" diff --git a/ansible/inventory/host_vars/atlas.yml b/ansible/inventory/host_vars/atlas.yml index fb7063c..1e06140 100644 --- a/ansible/inventory/host_vars/atlas.yml +++ b/ansible/inventory/host_vars/atlas.yml @@ -49,6 +49,7 @@ atlas_zfs_backup_reservation: 500G atlas_zfs_dataset_photobook: media/photobook atlas_mount_root: /zpool atlas_manage_storage: true +atlas_prometheus_pull_start_timer: true atlas_manage_zfs_snapshots: true atlas_zfs_snapshot_prefix: atlas-auto atlas_zfs_snapshot_policies: @@ -91,6 +92,12 @@ atlas_usb_backup_mapper_name: zpool-backup atlas_manage_usb_reminder: true atlas_usb_reminder_calendar: "Sat *-*-01..07 10:00:00 Europe/Rome" atlas_manage_monitoring: true +atlas_manage_prometheus_backup_pull: true +# Prometheus ED25519 host key read through the controller's strict SSH trust on 2026-09-30. +# Fingerprint: SHA256:rfedk7DHI9mLB3UHk/4F3HHlSIiswtCAFsAXvfh6iXk +atlas_prometheus_ssh_host_key: >- + 179.237.102.172 ssh-ed25519 + AAAAC3NzaC1lZDI1NTE5AAAAIC4b+QXlPupoEx71W9NKs9tTeYjBqTkVMqbGB97nMNWv # Physical pool disks and the system NVMe; the disconnected USB disk is intentionally excluded. atlas_monitor_smart_devices: - { name: pool-1, path: "{{ atlas_zpool_disks[0] }}", warning_c: 50, critical_c: 55 } diff --git a/ansible/inventory/host_vars/prometheus.yml b/ansible/inventory/host_vars/prometheus.yml index f03432b..bb41b77 100644 --- a/ansible/inventory/host_vars/prometheus.yml +++ b/ansible/inventory/host_vars/prometheus.yml @@ -6,6 +6,8 @@ ansible_port: 22 ansible_ssh_private_key_file: /home/fscotto/.ssh/id_ed25519 server_username: rocky +server_backup_export_enabled: true +server_backup_export_start_timer: true server_duckdns_domain: fscotto server_ssh_authorized_keys: - name: ikaros diff --git a/ansible/roles/profile_atlas/defaults/main.yml b/ansible/roles/profile_atlas/defaults/main.yml index 85d5591..15f0208 100644 --- a/ansible/roles/profile_atlas/defaults/main.yml +++ b/ansible/roles/profile_atlas/defaults/main.yml @@ -115,6 +115,15 @@ atlas_monitor_notifier: "{{ atlas_usb_reminder_notifier }}" atlas_monitor_smart_devices: [] atlas_monitor_timers: [] atlas_monitor_failure_units: [] +atlas_monitor_effective_timers: >- + {{ atlas_monitor_timers + + ([{'name': 'atlas-prometheus-pull.timer', 'max_age_hours': 26}] + if atlas_manage_prometheus_backup_pull | bool and atlas_prometheus_pull_start_timer | bool + else []) }} +atlas_monitor_effective_failure_units: >- + {{ atlas_monitor_failure_units + + (['atlas-prometheus-pull.service'] + if atlas_manage_prometheus_backup_pull | bool else []) }} atlas_monitor_remote_capacity: {} atlas_monitor_pool_warning_percent: 80 atlas_monitor_pool_critical_percent: 90 @@ -141,6 +150,19 @@ atlas_music_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_music }}" atlas_backup_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_backup }}" atlas_host_backups_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_host_backups }}" atlas_backup_prometheus_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_backup_prometheus }}" +atlas_manage_prometheus_backup_pull: false +atlas_prometheus_pull_ssh_dir: /etc/atlas-prometheus-pull +atlas_prometheus_pull_private_key_path: "{{ atlas_prometheus_pull_ssh_dir }}/id_ed25519" +atlas_prometheus_pull_known_hosts_path: "{{ atlas_prometheus_pull_ssh_dir }}/known_hosts" +atlas_prometheus_ssh_host_key: "" +atlas_prometheus_pull_source_user: prometheus-backup +atlas_prometheus_pull_source_port: 22 +atlas_prometheus_pull_calendar: "*-*-* 03:00:00 Europe/Rome" +atlas_prometheus_pull_start_timer: false +atlas_prometheus_pull_keep_daily: 30 +atlas_prometheus_pull_keep_weekly: 8 +atlas_prometheus_pull_keep_monthly: 12 +atlas_prometheus_pull_max_age_hours: 24 atlas_photobook_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_photobook }}" atlas_45drives_repo_url: https://repo.45drives.com/repofiles/rocky/45drives-enterprise.repo diff --git a/ansible/roles/profile_atlas/files/atlas-prometheus-prune.py b/ansible/roles/profile_atlas/files/atlas-prometheus-prune.py new file mode 100644 index 0000000..6c5a0ca --- /dev/null +++ b/ansible/roles/profile_atlas/files/atlas-prometheus-prune.py @@ -0,0 +1,56 @@ +#!/usr/bin/env python3 +"""Prune only verified, named Prometheus backup versions after publication.""" + +import datetime as dt +import pathlib +import re +import shutil +import sys + + +def main() -> None: + if len(sys.argv) != 5: + raise SystemExit("Usage: atlas-prometheus-prune SNAPSHOTS DAILY WEEKLY MONTHLY") + root = pathlib.Path(sys.argv[1]) + counts = [int(value) for value in sys.argv[2:]] + if not root.is_dir() or root.is_symlink() or min(counts) < 1: + raise SystemExit("Invalid backup directory or retention counts") + versions = [] + for entry in root.iterdir(): + if not entry.is_dir() or entry.is_symlink(): + continue + if not re.fullmatch(r"[0-9]{8}T[0-9]{6}Z", entry.name): + continue + try: + when = dt.datetime.strptime(entry.name, "%Y%m%dT%H%M%SZ") + except ValueError: + continue + if not all((entry / name).is_file() for name in ("payload.tar", "payload.sha256", "metadata.json")): + continue + versions.append((when, entry)) + versions.sort(reverse=True) + if not versions: + raise SystemExit("No published backup versions found; refusing to prune") + + keep = {entry for _, entry in versions[: counts[0]]} + for count, key in ( + (counts[1], lambda when: when.isocalendar()[:2]), + (counts[2], lambda when: (when.year, when.month)), + ): + periods = set() + for when, entry in versions: + period = key(when) + if period in periods: + continue + periods.add(period) + keep.add(entry) + if len(periods) >= count: + break + + for _, entry in versions: + if entry not in keep: + shutil.rmtree(entry) + + +if __name__ == "__main__": + main() diff --git a/ansible/roles/profile_atlas/tasks/main.yml b/ansible/roles/profile_atlas/tasks/main.yml index 4b87ecc..2403c66 100644 --- a/ansible/roles/profile_atlas/tasks/main.yml +++ b/ansible/roles/profile_atlas/tasks/main.yml @@ -23,6 +23,12 @@ - name: Import Atlas offline USB backup tasks ansible.builtin.import_tasks: usb_backup.yml +- name: Import Atlas Prometheus backup pull identity tasks + ansible.builtin.import_tasks: prometheus_pull_identity.yml + +- name: Import Atlas Prometheus backup pull job tasks + ansible.builtin.import_tasks: prometheus_pull_job.yml + - name: Import Atlas health monitoring tasks ansible.builtin.import_tasks: monitoring.yml diff --git a/ansible/roles/profile_atlas/tasks/monitoring.yml b/ansible/roles/profile_atlas/tasks/monitoring.yml index cacc460..2b34ae0 100644 --- a/ansible/roles/profile_atlas/tasks/monitoring.yml +++ b/ansible/roles/profile_atlas/tasks/monitoring.yml @@ -7,8 +7,8 @@ - atlas_zfs_pool != 'CHANGEME_ZFS_POOL' - atlas_monitor_calendar | length > 0 - atlas_monitor_smart_devices | length > 0 - - atlas_monitor_timers | length > 0 - - atlas_monitor_failure_units | length > 0 + - atlas_monitor_effective_timers | length > 0 + - atlas_monitor_effective_failure_units | length > 0 - atlas_monitor_remote_capacity.user == atlas_borg_repository_user - atlas_monitor_remote_capacity.host == atlas_borg_repository_host - atlas_monitor_remote_capacity.run_as == atlas_borg_username @@ -49,7 +49,7 @@ that: - item.name is match('^[a-zA-Z0-9@_.-]+\\.timer$') - item.max_age_hours | int >= 0 - loop: "{{ atlas_monitor_timers }}" + loop: "{{ atlas_monitor_effective_timers }}" loop_control: label: "{{ item.name }}" when: atlas_manage_monitoring | bool @@ -59,7 +59,7 @@ ansible.builtin.assert: that: - item is match('^[a-zA-Z0-9@_.-]+\\.service$') - loop: "{{ atlas_monitor_failure_units }}" + loop: "{{ atlas_monitor_effective_failure_units }}" when: atlas_manage_monitoring | bool - name: Validate Atlas health monitor calendar @@ -144,7 +144,7 @@ owner: root group: root mode: "0755" - loop: "{{ atlas_monitor_failure_units }}" + loop: "{{ atlas_monitor_effective_failure_units }}" when: atlas_manage_monitoring | bool - name: Notify 45Drives Alerts when an Atlas job fails @@ -155,7 +155,7 @@ owner: root group: root mode: "0644" - loop: "{{ atlas_monitor_failure_units }}" + loop: "{{ atlas_monitor_effective_failure_units }}" when: atlas_manage_monitoring | bool - name: Reload systemd after installing Atlas monitoring diff --git a/ansible/roles/profile_atlas/tasks/prometheus_pull_identity.yml b/ansible/roles/profile_atlas/tasks/prometheus_pull_identity.yml new file mode 100644 index 0000000..a96de9e --- /dev/null +++ b/ansible/roles/profile_atlas/tasks/prometheus_pull_identity.yml @@ -0,0 +1,66 @@ +--- +- name: Validate Atlas Prometheus pull identity inputs + tags: [atlas, backup, prometheus_backup, prometheus_backup_key] + ansible.builtin.assert: + that: + - atlas_prometheus_pull_ssh_dir.startswith('/etc/') + - atlas_prometheus_pull_private_key_path.startswith(atlas_prometheus_pull_ssh_dir ~ '/') + - atlas_prometheus_pull_known_hosts_path.startswith(atlas_prometheus_pull_ssh_dir ~ '/') + - atlas_prometheus_ssh_host_key.startswith( + (hostvars['prometheus'].ansible_host | string) ~ ' ssh-ed25519 ' + ) + fail_msg: Pin the verified Prometheus ED25519 SSH host key before enabling the pull. + when: atlas_manage_prometheus_backup_pull | bool + +- name: Create private Atlas Prometheus pull SSH directory + tags: [atlas, backup, prometheus_backup, prometheus_backup_key] + ansible.builtin.file: + path: "{{ atlas_prometheus_pull_ssh_dir }}" + state: directory + owner: root + group: root + mode: "0700" + when: atlas_manage_prometheus_backup_pull | bool + +- name: Generate Atlas-only Prometheus pull SSH identity + tags: [atlas, backup, prometheus_backup, prometheus_backup_key] + ansible.builtin.command: + argv: + - ssh-keygen + - -q + - -t + - ed25519 + - -N + - "" + - -C + - atlas-prometheus-pull@atlas + - -f + - "{{ atlas_prometheus_pull_private_key_path }}" + creates: "{{ atlas_prometheus_pull_private_key_path }}" + when: atlas_manage_prometheus_backup_pull | bool + +- name: Protect Atlas-only Prometheus pull SSH identity + tags: [atlas, backup, prometheus_backup, prometheus_backup_key] + ansible.builtin.file: + path: "{{ item.path }}" + owner: root + group: root + mode: "{{ item.mode }}" + loop: + - { path: "{{ atlas_prometheus_pull_private_key_path }}", mode: "0600" } + - { path: "{{ atlas_prometheus_pull_private_key_path }}.pub", mode: "0644" } + loop_control: + label: "{{ item.path }}" + when: + - atlas_manage_prometheus_backup_pull | bool + - not ansible_check_mode + +- name: Pin Prometheus SSH host key on Atlas + tags: [atlas, backup, prometheus_backup, prometheus_backup_key] + ansible.builtin.copy: + content: "{{ atlas_prometheus_ssh_host_key }}\n" + dest: "{{ atlas_prometheus_pull_known_hosts_path }}" + owner: root + group: root + mode: "0600" + when: atlas_manage_prometheus_backup_pull | bool diff --git a/ansible/roles/profile_atlas/tasks/prometheus_pull_job.yml b/ansible/roles/profile_atlas/tasks/prometheus_pull_job.yml new file mode 100644 index 0000000..5e12213 --- /dev/null +++ b/ansible/roles/profile_atlas/tasks/prometheus_pull_job.yml @@ -0,0 +1,90 @@ +--- +- name: Validate Atlas Prometheus backup pull inputs + tags: [atlas, backup, prometheus_backup] + ansible.builtin.assert: + that: + - atlas_manage_storage | bool + - atlas_prometheus_pull_source_user is match('^[a-z_][a-z0-9_-]*$') + - atlas_prometheus_pull_source_port | int > 0 + - atlas_prometheus_pull_source_port | int < 65536 + - atlas_prometheus_pull_keep_daily | int > 0 + - atlas_prometheus_pull_keep_weekly | int > 0 + - atlas_prometheus_pull_keep_monthly | int > 0 + - atlas_prometheus_pull_max_age_hours | int > 0 + - atlas_backup_prometheus_mountpoint.startswith(atlas_mount_root ~ '/') + fail_msg: Define the Atlas backup destination, source account, and retention before enabling the pull. + when: atlas_manage_prometheus_backup_pull | bool + +- name: Validate Atlas Prometheus backup pull calendar + tags: [atlas, backup, prometheus_backup] + ansible.builtin.command: + argv: [systemd-analyze, calendar, "{{ atlas_prometheus_pull_calendar }}"] + changed_when: false + check_mode: false + when: atlas_manage_prometheus_backup_pull | bool + +- name: Create private Atlas Prometheus backup version directory + tags: [atlas, backup, prometheus_backup] + ansible.builtin.file: + path: "{{ atlas_backup_prometheus_mountpoint }}/snapshots" + state: directory + owner: root + group: root + mode: "0700" + when: atlas_manage_prometheus_backup_pull | bool + +- name: Install Atlas Prometheus backup pull helper + tags: [atlas, backup, prometheus_backup] + ansible.builtin.template: + src: atlas-prometheus-pull.sh.j2 + dest: /usr/local/sbin/atlas-prometheus-pull + owner: root + group: root + mode: "0750" + when: atlas_manage_prometheus_backup_pull | bool + +- name: Install Atlas Prometheus backup retention helper + tags: [atlas, backup, prometheus_backup] + ansible.builtin.copy: + src: atlas-prometheus-prune.py + dest: /usr/local/libexec/atlas-prometheus-prune + owner: root + group: root + mode: "0750" + when: atlas_manage_prometheus_backup_pull | bool + +- name: Install Atlas Prometheus backup pull systemd units + tags: [atlas, backup, prometheus_backup] + ansible.builtin.template: + src: "{{ item }}.j2" + dest: "/etc/systemd/system/{{ item }}" + owner: root + group: root + mode: "0644" + loop: + - atlas-prometheus-pull.service + - atlas-prometheus-pull.timer + loop_control: + label: "{{ item }}" + register: atlas_prometheus_pull_units + when: atlas_manage_prometheus_backup_pull | bool + +- name: Reload systemd after Atlas Prometheus pull unit changes + tags: [atlas, backup, prometheus_backup] + ansible.builtin.systemd: + daemon_reload: true + when: + - atlas_manage_prometheus_backup_pull | bool + - atlas_prometheus_pull_units is changed + - not ansible_check_mode + +- name: Enable Atlas Prometheus pull timer only after explicit activation + tags: [atlas, backup, prometheus_backup] + ansible.builtin.systemd: + name: atlas-prometheus-pull.timer + enabled: true + state: started + when: + - atlas_manage_prometheus_backup_pull | bool + - atlas_prometheus_pull_start_timer | bool + - not ansible_check_mode diff --git a/ansible/roles/profile_atlas/templates/atlas-health-monitor.json.j2 b/ansible/roles/profile_atlas/templates/atlas-health-monitor.json.j2 index 9191fa9..ef9b780 100644 --- a/ansible/roles/profile_atlas/templates/atlas-health-monitor.json.j2 +++ b/ansible/roles/profile_atlas/templates/atlas-health-monitor.json.j2 @@ -3,8 +3,8 @@ "backup_dataset": {{ (atlas_zfs_pool ~ '/' ~ atlas_zfs_dataset_backup) | to_json }}, "notifier": {{ atlas_monitor_notifier | to_json }}, "smart_devices": {{ atlas_monitor_smart_devices | to_json }}, - "timers": {{ atlas_monitor_timers | to_json }}, - "failure_units": {{ atlas_monitor_failure_units | to_json }}, + "timers": {{ atlas_monitor_effective_timers | to_json }}, + "failure_units": {{ atlas_monitor_effective_failure_units | to_json }}, "remote_capacity": {{ atlas_monitor_remote_capacity | to_json }}, "pool_warning_percent": {{ atlas_monitor_pool_warning_percent | int }}, "pool_critical_percent": {{ atlas_monitor_pool_critical_percent | int }}, diff --git a/ansible/roles/profile_atlas/templates/atlas-prometheus-pull.service.j2 b/ansible/roles/profile_atlas/templates/atlas-prometheus-pull.service.j2 new file mode 100644 index 0000000..3432d96 --- /dev/null +++ b/ansible/roles/profile_atlas/templates/atlas-prometheus-pull.service.j2 @@ -0,0 +1,19 @@ +[Unit] +Description=Pull a prepared read-only Prometheus backup to Atlas +RequiresMountsFor={{ atlas_backup_prometheus_mountpoint }} +Wants=network-online.target +After=network-online.target zfs.target +ConditionFileIsExecutable=/usr/local/sbin/atlas-prometheus-pull +ConditionPathExists={{ atlas_prometheus_pull_private_key_path }} +ConditionPathExists={{ atlas_prometheus_pull_known_hosts_path }} + +[Service] +Type=oneshot +ExecStart=/usr/local/sbin/atlas-prometheus-pull +User=root +Group=root +UMask=0077 +TimeoutStartSec=infinity +Nice=15 +IOSchedulingClass=best-effort +IOSchedulingPriority=7 diff --git a/ansible/roles/profile_atlas/templates/atlas-prometheus-pull.sh.j2 b/ansible/roles/profile_atlas/templates/atlas-prometheus-pull.sh.j2 new file mode 100644 index 0000000..8857839 --- /dev/null +++ b/ansible/roles/profile_atlas/templates/atlas-prometheus-pull.sh.j2 @@ -0,0 +1,75 @@ +#!/usr/bin/env bash +set -Eeuo pipefail +umask 077 + +backup_root={{ atlas_backup_prometheus_mountpoint | quote }} +snapshots="$backup_root/snapshots" +stage='' +exec 9>/run/lock/atlas-prometheus-pull.lock +flock -n 9 || { echo 'A Prometheus pull is already running' >&2; exit 1; } + +cleanup() { + local rc=$? + trap - EXIT + if (( rc != 0 )) && [[ -n "$stage" && -d "$stage" ]]; then + rm -rf -- "$stage" + fi + exit "$rc" +} +trap cleanup EXIT + +zpool list -H -o name {{ atlas_zfs_pool | quote }} >/dev/null +findmnt -rn --mountpoint "$backup_root" >/dev/null +stage=$(mktemp -d "$backup_root/.staging.XXXXXXXX") +ssh_cmd='/usr/bin/ssh -F /dev/null -o BatchMode=yes -o StrictHostKeyChecking=yes -o UserKnownHostsFile={{ atlas_prometheus_pull_known_hosts_path }} -o IdentitiesOnly=yes -i {{ atlas_prometheus_pull_private_key_path }} -p {{ atlas_prometheus_pull_source_port }}' +rsync -a --partial --delay-updates -e "$ssh_cmd" \ + {{ (atlas_prometheus_pull_source_user ~ '@' ~ hostvars['prometheus'].ansible_host ~ ':current/') | quote }} \ + "$stage/" + +test -s "$stage/payload.tar" +test -s "$stage/payload.sha256" +test -s "$stage/metadata.json" +(cd "$stage" && sha256sum -c payload.sha256) +tar -tf "$stage/payload.tar" >/dev/null +stamp=$(python3 - "$stage/metadata.json" <<'PY' +import json +import datetime as dt +import re +import sys + +with open(sys.argv[1], encoding="utf-8") as stream: + metadata = json.load(stream) +stamp = metadata.get("created_utc", "") +if metadata.get("schema") != 1 or metadata.get("host") != "prometheus": + raise SystemExit("Unexpected Prometheus backup metadata") +if not re.fullmatch(r"[0-9]{8}T[0-9]{6}Z", stamp): + raise SystemExit("Invalid Prometheus backup timestamp") +created = dt.datetime.strptime(stamp, "%Y%m%dT%H%M%SZ").replace(tzinfo=dt.timezone.utc) +age = dt.datetime.now(dt.timezone.utc) - created +if age.total_seconds() < -300 or age > dt.timedelta(hours={{ atlas_prometheus_pull_max_age_hours }}): + raise SystemExit("Prometheus backup is outside the configured freshness window") +print(stamp) +PY +) +if [[ -e "$snapshots/$stamp" ]]; then + cmp "$stage/payload.sha256" "$snapshots/$stamp/payload.sha256" + cmp "$stage/metadata.json" "$snapshots/$stamp/metadata.json" + (cd "$snapshots/$stamp" && sha256sum -c payload.sha256) + rm -rf -- "${stage:?}" + stage='' +else + chown -R root:root "$stage" + chmod 0700 "$stage" + chmod 0600 "$stage/payload.tar" "$stage/payload.sha256" "$stage/metadata.json" + mv -- "$stage" "$snapshots/$stamp" + stage='' +fi +latest_link=$(readlink "$backup_root/latest" 2>/dev/null || true) +latest_stamp=${latest_link##*/} +if [[ -z "$latest_stamp" || "$stamp" > "$latest_stamp" ]]; then + ln -s "snapshots/$stamp" "$backup_root/.latest.new" + mv -Tf -- "$backup_root/.latest.new" "$backup_root/latest" +fi +python3 /usr/local/libexec/atlas-prometheus-prune "$snapshots" \ + {{ atlas_prometheus_pull_keep_daily }} {{ atlas_prometheus_pull_keep_weekly }} {{ atlas_prometheus_pull_keep_monthly }} +echo "Verified and published Prometheus backup $stamp" diff --git a/ansible/roles/profile_atlas/templates/atlas-prometheus-pull.timer.j2 b/ansible/roles/profile_atlas/templates/atlas-prometheus-pull.timer.j2 new file mode 100644 index 0000000..3564706 --- /dev/null +++ b/ansible/roles/profile_atlas/templates/atlas-prometheus-pull.timer.j2 @@ -0,0 +1,10 @@ +[Unit] +Description=Schedule Atlas pull of prepared Prometheus backups + +[Timer] +OnCalendar={{ atlas_prometheus_pull_calendar }} +Persistent=true +Unit=atlas-prometheus-pull.service + +[Install] +WantedBy=timers.target diff --git a/ansible/roles/profile_server/tasks/backup_export_identity.yml b/ansible/roles/profile_server/tasks/backup_export_identity.yml new file mode 100644 index 0000000..7ac5664 --- /dev/null +++ b/ansible/roles/profile_server/tasks/backup_export_identity.yml @@ -0,0 +1,103 @@ +--- +- name: Validate Prometheus backup export identity inputs + tags: [services, backup, prometheus_backup] + ansible.builtin.assert: + that: + - inventory_hostname == 'prometheus' + - server_backup_username is match('^[a-z_][a-z0-9_-]*$') + - server_backup_username not in ['root', server_username] + - server_backup_export_root.startswith('/var/lib/') + - server_backup_public_key_name is match('^[a-z0-9_-]+$') + - hostvars['atlas'].atlas_manage_prometheus_backup_pull | default(false) | bool + fail_msg: Enable Atlas and Prometheus backup roles together with dedicated identity settings. + when: server_backup_export_enabled | bool + +- name: Create dedicated Prometheus backup export group + tags: [services, backup, prometheus_backup] + ansible.builtin.group: + name: "{{ server_backup_username }}" + system: true + state: present + when: server_backup_export_enabled | bool + +- name: Create locked Prometheus backup export account + tags: [services, backup, prometheus_backup] + ansible.builtin.user: + name: "{{ server_backup_username }}" + group: "{{ server_backup_username }}" + groups: [] + append: false + comment: Read-only prepared backup export for Atlas + home: "{{ server_backup_export_root }}" + create_home: false + shell: /bin/bash + password_lock: true + system: true + state: present + when: server_backup_export_enabled | bool + +- name: Require restricted rrsync helper on Prometheus + tags: [services, backup, prometheus_backup] + ansible.builtin.stat: + path: "{{ server_backup_rrsync_path }}" + register: server_backup_rrsync_file + when: server_backup_export_enabled | bool + +- name: Validate restricted rrsync helper + tags: [services, backup, prometheus_backup] + ansible.builtin.assert: + that: + - server_backup_rrsync_file.stat.exists + - server_backup_rrsync_file.stat.isreg + - server_backup_rrsync_file.stat.pw_name == 'root' + fail_msg: Rocky rsync must provide the root-owned rrsync support script. + when: server_backup_export_enabled | bool + +- name: Create prepared backup export root + tags: [services, backup, prometheus_backup] + ansible.builtin.file: + path: "{{ server_backup_export_root }}" + state: directory + owner: root + group: "{{ server_backup_username }}" + mode: "0750" + when: server_backup_export_enabled | bool + +- name: Create restricted Prometheus backup SSH directories + tags: [services, backup, prometheus_backup] + ansible.builtin.file: + path: "{{ item }}" + state: directory + owner: root + group: "{{ server_backup_username }}" + mode: "0750" + loop: + - "{{ server_backup_export_root }}/.ssh" + - "{{ server_backup_export_root }}/.ssh/authorized_keys.d" + when: server_backup_export_enabled | bool + +- name: Read Atlas public key for Prometheus backup pull + tags: [services, backup, prometheus_backup] + ansible.builtin.slurp: + src: "{{ hostvars['atlas'].atlas_prometheus_pull_private_key_path | default('/etc/atlas-prometheus-pull/id_ed25519') }}.pub" + delegate_to: atlas + become: true + register: server_backup_atlas_public_key + when: + - server_backup_export_enabled | bool + - not ansible_check_mode + +- name: Authorize only restricted read-only backup access from Atlas + tags: [services, backup, prometheus_backup] + ansible.builtin.copy: + content: >- + {{ 'command="/usr/bin/python3 ' ~ server_backup_rrsync_path ~ ' -ro ' + ~ server_backup_export_root ~ '/versions",restrict ' + ~ (server_backup_atlas_public_key.content | b64decode | trim) ~ '\n' }} + dest: "{{ server_backup_export_root }}/.ssh/authorized_keys.d/{{ server_backup_public_key_name }}" + owner: root + group: "{{ server_backup_username }}" + mode: "0640" + when: + - server_backup_export_enabled | bool + - not ansible_check_mode diff --git a/ansible/roles/profile_server/tasks/backup_export_job.yml b/ansible/roles/profile_server/tasks/backup_export_job.yml new file mode 100644 index 0000000..48cff77 --- /dev/null +++ b/ansible/roles/profile_server/tasks/backup_export_job.yml @@ -0,0 +1,90 @@ +--- +- name: Validate Prometheus backup export job inputs + tags: [services, backup, prometheus_backup] + ansible.builtin.assert: + that: + - server_backup_export_source_keep | int >= 2 + - server_backup_export_paths | length > 0 + - server_backup_export_paths | unique | length == server_backup_export_paths | length + - >- + server_backup_export_paths + | select('match', '^[a-zA-Z0-9][a-zA-Z0-9._/-]*$') | list | length + == server_backup_export_paths | length + - >- + server_backup_export_paths + | reject('search', '(^|/)\.\.(/|$)') | list | length + == server_backup_export_paths | length + - >- + server_backup_export_excludes + | select('match', '^[a-zA-Z0-9][a-zA-Z0-9._/-]*$') | list | length + == server_backup_export_excludes | length + - >- + server_backup_export_excludes + | reject('search', '(^|/)\.\.(/|$)') | list | length + == server_backup_export_excludes | length + fail_msg: Define safe relative paths and at least two prepared export versions. + when: server_backup_export_enabled | bool + +- name: Validate Prometheus backup export calendar + tags: [services, backup, prometheus_backup] + ansible.builtin.command: + argv: [systemd-analyze, calendar, "{{ server_backup_export_calendar }}"] + changed_when: false + check_mode: false + when: server_backup_export_enabled | bool + +- name: Ensure prepared Prometheus backup versions directory exists + tags: [services, backup, prometheus_backup] + ansible.builtin.file: + path: "{{ server_backup_export_root }}/versions" + state: directory + owner: root + group: "{{ server_backup_username }}" + mode: "0750" + when: server_backup_export_enabled | bool + +- name: Install Prometheus backup export helper + tags: [services, backup, prometheus_backup] + ansible.builtin.template: + src: prometheus-backup-export.sh.j2 + dest: /usr/local/sbin/prometheus-backup-export + owner: root + group: root + mode: "0750" + when: server_backup_export_enabled | bool + +- name: Install Prometheus backup export systemd units + tags: [services, backup, prometheus_backup] + ansible.builtin.template: + src: "{{ item }}.j2" + dest: "/etc/systemd/system/{{ item }}" + owner: root + group: root + mode: "0644" + loop: + - prometheus-backup-export.service + - prometheus-backup-export.timer + loop_control: + label: "{{ item }}" + register: server_backup_export_units + when: server_backup_export_enabled | bool + +- name: Reload systemd after Prometheus backup export unit changes + tags: [services, backup, prometheus_backup] + ansible.builtin.systemd: + daemon_reload: true + when: + - server_backup_export_enabled | bool + - server_backup_export_units is changed + - not ansible_check_mode + +- name: Enable Prometheus backup export timer only after explicit activation + tags: [services, backup, prometheus_backup] + ansible.builtin.systemd: + name: prometheus-backup-export.timer + enabled: true + state: started + when: + - server_backup_export_enabled | bool + - server_backup_export_start_timer | bool + - not ansible_check_mode diff --git a/ansible/roles/profile_server/tasks/main.yml b/ansible/roles/profile_server/tasks/main.yml index 395a9b1..73bda6f 100644 --- a/ansible/roles/profile_server/tasks/main.yml +++ b/ansible/roles/profile_server/tasks/main.yml @@ -53,6 +53,12 @@ tags: [services, podman] ansible.builtin.include_tasks: podman-compose.yml +- name: Import Prometheus backup export identity tasks + ansible.builtin.import_tasks: backup_export_identity.yml + +- name: Import Prometheus backup export job tasks + ansible.builtin.import_tasks: backup_export_job.yml + - name: Ensure server SSH authorized key fragments directory exists tags: [services, ssh] ansible.builtin.file: @@ -77,13 +83,17 @@ when: server_ssh_authorized_keys | length > 0 - name: Configure server SSH authorized key fragments - tags: [services, ssh] + tags: [services, ssh, prometheus_backup] ansible.builtin.lineinfile: path: /etc/ssh/sshd_config regexp: '^\s*AuthorizedKeysFile\s+' line: >- - AuthorizedKeysFile {{ server_ssh_authorized_keys | map(attribute='name') - | map('regex_replace', '^', '%h/.ssh/authorized_keys.d/') | join(' ') }} + AuthorizedKeysFile {{ + ((server_ssh_authorized_keys | map(attribute='name') + | map('regex_replace', '^', '%h/.ssh/authorized_keys.d/') | list) + + (['%h/.ssh/authorized_keys.d/' ~ server_backup_public_key_name] + if server_backup_export_enabled | bool else [])) | join(' ') + }} state: present validate: "sshd -t -f %s" notify: Reload SSH service @@ -100,11 +110,13 @@ notify: Reload SSH service - name: Restrict SSH login to allowed users on server - tags: [services] + tags: [services, prometheus_backup] ansible.builtin.lineinfile: path: /etc/ssh/sshd_config regexp: '^\s*AllowUsers\s+' - line: "AllowUsers {{ server_sshd_allow_users | join(' ') }}" + line: >- + AllowUsers {{ (server_sshd_allow_users + + ([server_backup_username] if server_backup_export_enabled | bool else [])) | join(' ') }} state: present validate: "sshd -t -f %s" notify: Reload SSH service diff --git a/ansible/roles/profile_server/templates/prometheus-backup-export.service.j2 b/ansible/roles/profile_server/templates/prometheus-backup-export.service.j2 new file mode 100644 index 0000000..a12224f --- /dev/null +++ b/ansible/roles/profile_server/templates/prometheus-backup-export.service.j2 @@ -0,0 +1,15 @@ +[Unit] +Description=Prepare a read-only Prometheus application backup for Atlas +RequiresMountsFor=/opt/npm /opt/gitea {{ server_backup_export_root }} +ConditionFileIsExecutable=/usr/local/sbin/prometheus-backup-export + +[Service] +Type=oneshot +ExecStart=/usr/local/sbin/prometheus-backup-export +User=root +Group=root +UMask=0077 +TimeoutStartSec=infinity +Nice=10 +IOSchedulingClass=best-effort +IOSchedulingPriority=7 diff --git a/ansible/roles/profile_server/templates/prometheus-backup-export.sh.j2 b/ansible/roles/profile_server/templates/prometheus-backup-export.sh.j2 new file mode 100644 index 0000000..9be174b --- /dev/null +++ b/ansible/roles/profile_server/templates/prometheus-backup-export.sh.j2 @@ -0,0 +1,92 @@ +#!/usr/bin/env bash +set -Eeuo pipefail +umask 077 + +export_root={{ server_backup_export_root | quote }} +versions="$export_root/versions" +stack_unit=podman-compose-server.service +stamp=$(date -u +%Y%m%dT%H%M%SZ) +stage='' +stack_stopped=false + +exec 9>/run/lock/prometheus-backup-export.lock +flock -n 9 || { echo 'A backup export is already running' >&2; exit 1; } + +cleanup() { + local rc=$? + trap - EXIT + if "$stack_stopped"; then + if systemctl is-active --quiet "$stack_unit"; then + systemctl restart "$stack_unit" || rc=1 + else + systemctl start "$stack_unit" || rc=1 + fi + fi + if (( rc != 0 )) && [[ -n "$stage" && -d "$stage" ]]; then + rm -rf -- "$stage" + fi + exit "$rc" +} +trap cleanup EXIT +trap 'exit 129' HUP +trap 'exit 130' INT +trap 'exit 143' TERM + +systemctl is-active --quiet "$stack_unit" || { + echo 'The managed Compose stack must be active before preparing a backup' >&2 + exit 1 +} + +paths=( +{% for path in server_backup_export_paths %} + {{ path | quote }} +{% endfor %} +) +excludes=( +{% for path in server_backup_export_excludes %} + --exclude={{ path | quote }} +{% endfor %} +) +for path in "${paths[@]}"; do + [[ -e "/$path" ]] || { echo "Required backup path missing: /$path" >&2; exit 1; } +done +[[ ! -e "$versions/$stamp" ]] || { echo "Export version already exists: $stamp" >&2; exit 1; } +stage=$(mktemp -d "$export_root/.staging.XXXXXXXX") + +# SQLite databases and their accompanying files are copied while both +# managed containers are stopped. The EXIT trap restarts the stack on error. +stack_stopped=true +systemctl stop "$stack_unit" +tar --acls --xattrs --selinux "${excludes[@]}" -C / -cf "$stage/payload.tar" "${paths[@]}" +systemctl start "$stack_unit" +for container in nginx-proxy-manager gitea; do + running=false + for _ in {1..30}; do + if [[ $(podman inspect --format '{{ '{{.State.Running}}' }}' "$container" 2>/dev/null) == true ]]; then + running=true + break + fi + sleep 2 + done + "$running" || { echo "Container did not restart: $container" >&2; exit 1; } +done +stack_stopped=false + +tar -tf "$stage/payload.tar" >/dev/null +(cd "$stage" && sha256sum payload.tar >payload.sha256) +printf '{"schema":1,"host":"prometheus","created_utc":"%s"}\n' "$stamp" >"$stage/metadata.json" +chown root:{{ server_backup_username }} "$stage" "$stage/payload.tar" "$stage/payload.sha256" "$stage/metadata.json" +chmod 0750 "$stage" +chmod 0640 "$stage/payload.tar" "$stage/payload.sha256" "$stage/metadata.json" +mv -- "$stage" "$versions/$stamp" +stage='' +ln -s "$stamp" "$versions/.current.new" +mv -Tf -- "$versions/.current.new" "$versions/current" + +# Keep a small source-side safety window; Atlas owns long-term retention. +mapfile -t old_versions < <(find "$versions" -mindepth 1 -maxdepth 1 -type d \ + -printf '%f\n' | grep -E '^[0-9]{8}T[0-9]{6}Z$' | sort -r | tail -n +{{ server_backup_export_source_keep + 1 }}) +for old in "${old_versions[@]}"; do + rm -rf -- "${versions:?}/$old" +done +echo "Prepared Prometheus backup export $stamp" diff --git a/ansible/roles/profile_server/templates/prometheus-backup-export.timer.j2 b/ansible/roles/profile_server/templates/prometheus-backup-export.timer.j2 new file mode 100644 index 0000000..e0585b7 --- /dev/null +++ b/ansible/roles/profile_server/templates/prometheus-backup-export.timer.j2 @@ -0,0 +1,10 @@ +[Unit] +Description=Prepare daily Prometheus application backup for Atlas + +[Timer] +OnCalendar={{ server_backup_export_calendar }} +Persistent=false +Unit=prometheus-backup-export.service + +[Install] +WantedBy=timers.target diff --git a/docs/atlas-dr-lab.md b/docs/atlas-dr-lab.md new file mode 100644 index 0000000..93ed96b --- /dev/null +++ b/docs/atlas-dr-lab.md @@ -0,0 +1,74 @@ +# Isolated Atlas DR lab + +This is a **scaled rehearsal**, not a substitute for a full-data restore. The +`atlas-dr-lab` libvirt VM on Ikaros was left **shut off** on 2026-09-30. Its +persistent volumes are in the default libvirt pool: the current 30 GiB OS +volume `atlas-dr-lab-os-rebuild2.qcow2`, the pre-rebuild OS volume +`atlas-dr-lab-os.qcow2`, and four independent 4 GiB +`atlas-dr-lab-data{1,2,3,4}.qcow2` volumes. The VM uses libvirt's `default` +NAT network (last DHCP address `192.168.122.168`), 2 vCPU, and 4 GiB RAM. +The data disks have `virtio-atlasdrdata{1,2,3,4}` serials. No physical disk or +production Atlas storage is attached. The VM has no autostart. + +## Rebuild inputs and isolation + +- Use Rocky's **9.8 GenericCloud Base x86_64** image + `Rocky-9-GenericCloud-Base-9.8-20260525.0.x86_64.qcow2` from + `https://download.rockylinux.org/pub/rocky/9.8/images/x86_64/`. + Verify its `.CHECKSUM` file; the observed SHA-256 was + `92c206cc6f790c61583247eefe87890f8828420662c17cacf247cec78ab4eec8`. +- Use a dedicated lab-only inventory merged **after** the repository + inventory, and always `--limit atlas_dr_lab`. The temporary 2026-09-30 + inventory/playbook and logs are in `/tmp/atlas-dr-lab-image/`; copy a + sanitized inventory to durable private storage before `/tmp` is cleared if + the lab will be repeated. Never reuse `host_vars/atlas.yml`, production + Vault secrets, or production disk by-id paths for the lab. +- The lab host belongs to `platform_rocky` and `atlas`. It uses `dradmin` + (UID/GID 1000) with the operator's **public** SSH key and a random, + unknown password hash, the libvirt DHCP address, pool `zpool`, mount root + `/zpool`, the four `virtio-atlasdrdata*` by-id paths, a 1 GiB backup + reservation, and `rocky_manage_openzfs_repo: true` with only `zfs` in + `host_packages`. The following gates remain false: sharing, firewall, + media stack, ZFS timers, Borg, USB, monitoring, and Prometheus pull. + `atlas_manage_storage` is true. Set `atlas_create_pool: true` **only for the + first disposable pool creation**, then set it false before any later run. +- A minimal lab playbook selects `atlas_dr_lab`, `become: true`, and the + existing `packages_rocky` and `profile_atlas` roles. Use a separate + `ANSIBLE_CONFIG` without the production Vault password script, and keep + host-key checking on with a lab-specific known-hosts file. The 2026-09-30 + runs used `-i ansible/inventory/hosts.yml -i ` and + `--limit atlas_dr_lab` throughout. + +## Rehearsal and narrow checks + +1. Before any pool operation, compare `virsh -c qemu:///system domblklist + atlas-dr-lab` with the four intended qcow2 paths, and in the guest compare + `/dev/disk/by-id/virtio-atlasdrdata*` with `lsblk`. Do not proceed if a + physical disk or production identity appears. +2. For a first-time disposable build only, run the lab playbook with + `--tags pool` and `atlas_create_pool: true`, then immediately set the gate + false. Run the full lab playbook and check `zpool status -P zpool`, + `zfs list -r zpool`, SELinux, and failed systemd units. +3. Write a non-sensitive canary under the lab `/zpool/archive` and snapshot + it. Record the pool GUID and canary SHA-256. Export the lab pool cleanly, + shut down the VM, and replace **only the OS volume** with a fresh verified + Rocky image. Preserve all four data volumes. Reconfigure cloud-init for a + new instance; the seed CD-ROM must use **SATA**. The SCSI seed attachment + tried during this rehearsal was not detected by cloud-init and was + replaced with a SATA attachment before proceeding. +4. On the new OS, apply `packages_rocky` to reinstall OpenZFS. First run + `zpool import -d /dev/disk/by-id` **without importing**, compare GUID and + vdev membership, then use ordinary `zpool import -d /dev/disk/by-id zpool`. + Do not use `-f`, `-F`, `-X`, rollback, or pool creation. +5. Reapply `profile_atlas` with the lab gates and `atlas_create_pool: false`. + Verify the canary, restored snapshot file in an empty temporary directory, + dataset hierarchy, SELinux, and pool health. A second full playbook run + should report `changed=0`. Remove temporary restored files and shut down + the VM after testing. + +The observed 2026-09-30 pool GUID was `8880368391795119587`; the canary +SHA-256 was `949701c7a95fadae1fddc21abe846c4312212dbfeb7477948f3188fc3ec34a78`. +The post-rebuild Ansible run succeeded, a repeat run reported `changed=0`, +12 datasets and the original snapshot were present, and the pool was healthy. +The snapshot-restored file matched content and basic metadata. See +[`atlas-recovery.md`](atlas-recovery.md) for the production runbook and limits. diff --git a/docs/atlas-recovery.md b/docs/atlas-recovery.md new file mode 100644 index 0000000..4b70ff5 --- /dev/null +++ b/docs/atlas-recovery.md @@ -0,0 +1,141 @@ +# Atlas recovery runbook + +This runbook is for a **replacement Rocky Linux 9 installation**, not a normal +playbook run. A scaled whole-OS rebuild with a disposable pool passed in an +isolated VM on 2026-09-30, but no production-size whole-host recovery has been +tested. The existing production pool must be imported, never created or +rewritten. The provisional targets are **RPO 24 hours** +and **RTO 72 hours**, for Archive and Atlas services alike. They are planning +objectives, not demonstrated recovery times. The manual USB cadence may leave +an older copy; a recent Borg archive is needed to meet the RPO after total +pool loss. + +## Before an incident + +- Keep an offline copy of the encrypted Ansible Vault, its unlock material, + the exported Borg repository key, and the Borg passphrase. Do not store + unlock material in this repository or in a recovery command line. + On 2026-09-30 the operator confirmed these are available independently of + Atlas and the Ansible controller; their usability has not been tested here. +- Keep the Atlas installation media and a reproducible checkout of this + repository available independently of Atlas. Record the exact Git revision + used for a successful deployment. +- Record the pool's current disk identities with `zpool status -P zpool` and + `lsblk -o NAME,SIZE,MODEL,SERIAL,FSTYPE,UUID`. Compare these with + `atlas_zpool_disks` before touching a replacement host. The `host_vars` + values are historical identifiers, not evidence that a newly attached disk + is the same device. +- Verify that the latest hourly/daily snapshots, Borg archive, and offline USB + version exist and note their timestamps. A timer being enabled is not proof + that a backup completed. + +## Incident gate + +1. Identify whether the fault is the OS disk, one or more pool disks, accidental + deletion, or an unavailable host. Preserve failed media when possible. +2. Stop writes to affected services and capture the last known good backup + timestamps. Do not run `zpool create`, `zpool destroy`, `zfs rollback`, + `zpool import -F`, `zpool import -X`, `zpool import -f`, or disk formatting + as a diagnostic shortcut. +3. Choose one recovery source below. Do not merge several sources into the + production namespace without comparing their timestamps and content. + +## Rebuild the OS and import the existing pool + +1. Install Rocky Linux 9 on a **separate system disk**. Configure basic network, + SSH, a temporary sudo administrator, SELinux enforcing, and the current + OpenZFS kmod repository. Keep the pool drives untouched. +2. Run read-only identification: `lsblk -f`, `zpool import`, and + `zpool import -d /dev/disk/by-id`. Check the pool GUID, vdev layout, and + stable drive identities against the incident record. If any differ, stop. +3. Import only after matching the expected pool and host ownership. A pool + cleanly exported from the old host can be imported with + `zpool import -d /dev/disk/by-id zpool`. If it reports that the pool is + active elsewhere or needs a rewind/force, stop and investigate rather than + adding flags. Verify with `zpool status -v zpool`, `zfs list -r zpool`, + `zfs get -r mountpoint,canmount zpool`, and `findmnt -R /zpool`. +4. Leave `atlas_create_pool: false`. Ensure `host_vars/atlas.yml` reflects the + replacement host's actual SSH address and disk identities before running + Ansible. Apply `ansible/site.yml --limit atlas` with the bootstrap admin + connection override as documented in the Atlas setup section of README. + This may start shares/services, so keep clients disconnected or services + gated until data and permissions are verified. +5. Check `getenforce`, `zpool status -v zpool`, `systemctl --failed`, SSH, + firewalld, Cockpit, NFS, SMB, and the backup/monitoring timers. Do not + report recovery complete on the basis of Ansible success alone. + +## Choose the data source + +- **Local snapshot, pool intact:** inspect `zfs list -t snapshot -r zpool`. + Mount/access the chosen snapshot read-only and copy selected files to an + empty staging directory; compare content, owner, mode, mtime, and POSIX ACL. + Move into the live namespace only after an operator-approved scope review. + Do not use an automatic rollback: it can discard newer changes in the + dataset and descendants. +- **Offline USB:** verify the configured LUKS and ext4 UUIDs from + `host_vars/atlas.yml` before unlocking. Mount ext4 read-only with `ro,noload`, + use only a published `atlas/latest` version, and restore to an empty staging + directory. Compare checksums and metadata. The USB copy intentionally omits + generic xattrs and SELinux labels; relabel only the restored destination. + Never run the backup service to perform a restore. +- **Hetzner Borg:** use the dedicated pinned host key, repository path, + offline exported recovery key, and Vault-backed passphrase. List archives + and extract a selected archive into an empty staging directory, never the + live `/zpool` tree. A repository check and sample restore were previously + performed; that does not prove this incident's archive is complete. Compare + content and metadata before publication. Avoid `borg break-lock` while any + backup/check job may still be active. + +After publishing restored files, run the explicit Ansible `restorecon` tag only +for the paths actually restored, for example: + +```bash +ansible-playbook ansible/site.yml --limit atlas --tags restorecon \ + -e '{"atlas_restorecon_paths":["/zpool/archive"]}' +``` + +Then check ownership/ACLs, application-specific integrity, SMB/NFS client +access, backup service health, and `zpool status -v zpool`. Reconnect clients +only after these checks pass. Record the last recoverable timestamp (actual +RPO) and elapsed service outage (actual RTO) in the incident log. + +## Scaled isolated rehearsal (2026-09-30) + +The lab setup, repeatable checks, and preserved VM state are recorded in +[`atlas-dr-lab.md`](atlas-dr-lab.md). + +On Ikaros, a local libvirt `atlas-dr-lab` VM used a 30 GiB Rocky 9.8 system +disk and four separate, disposable 4 GiB virtio data disks with stable +`/dev/disk/by-id` identities. The official Rocky cloud image matched its +published SHA-256. The lab inventory was separate from production, used a +fresh lab-only password hash and the operator's public SSH key, and disabled +sharing, Borg, USB backup, monitoring, media services, and the Prometheus pull. +No production disk, Vault secret, or production data was attached or copied. + +1. The existing `packages_rocky` and `profile_atlas` roles installed OpenZFS, + created a RAIDZ2 `zpool` through the explicit one-time pool gate, and built + all 12 declared datasets with a lab-sized 1 GiB backup reservation. The + pool creation gate was set false immediately afterward. +2. A 4 MiB canary file was written under the lab `archive` dataset and a ZFS + snapshot created. The pool was cleanly exported and the VM shut down. +3. Only the system-disk volume was replaced by a fresh Rocky cloud image; + the four virtio data volumes were retained. Ansible reinstalled OpenZFS. + Read-only `zpool import -d /dev/disk/by-id` showed the expected RAIDZ2 + topology and pool GUID `8880368391795119587` before an ordinary import + without `-f`, rewind, or rollback. +4. The imported pool was healthy. The canary SHA-256 matched its pre-rebuild + value. `profile_atlas` completed against the imported pool and a second + run reported `changed=0`. A file restored from the preserved snapshot into + `/var/tmp` matched SHA-256, owner, group, mode, size, and mtime; the temporary + copy was removed. Final checks found SELinux Enforcing, 12 datasets, the + snapshot, no failed units, and a healthy pool. The VM was shut down while + retaining its disposable volumes for a future rehearsal. + +This proves the **sequence** for a cleanly exported, small pool and the tested +Ansible subset, not recovery duration or capacity at 2 TB. The earlier +2026-09-25 independent production ZFS/USB file restores and the earlier Borg +temporary-directory restore remain separate evidence. The VM did not restore +production USB/Borg archives, exercise services with production data, test an +unclean import, or prove the provisional RPO/RTO. Before relying on 24h/72h, +measure a representative full restore and service cutover in a suitably sized +future change window. Never use the production Atlas pool for a rehearsal. diff --git a/docs/atlas-sharing-decision.md b/docs/atlas-sharing-decision.md new file mode 100644 index 0000000..81bb9ea --- /dev/null +++ b/docs/atlas-sharing-decision.md @@ -0,0 +1,20 @@ +# Atlas SMB/NFS namespace decision + +Decision date: 2026-09-30. Keep the current namespaces **separate**. + +- `/zpool/archive` is the SMB3 `Archive` share for authorized Samba accounts. +- `/zpool/media/photobook` is the Aegis-only NFSv4 export, `all_squash`-mapped + to UID/GID `1100`. +- No new dual-protocol namespace, broad export, group, or ACL model is needed. + Existing permissions and client access remain unchanged. + +The two paths serve different ownership and exposure needs. A common namespace +would expand the permissions design and require same-file SMB/NFS interoperability +testing without a present requirement. Revisit only when a specific workflow +needs both protocols on the same files; then decide UID/GID, group, POSIX ACL, +SELinux policy and client behavior before changing exports or permissions. + +Read-only Atlas verification on 2026-09-30 confirmed that Samba `Archive` points +to `/zpool/archive`, NFS exports `/zpool/media/photobook` only to +`192.168.178.54` with `all_squash` and anonymous UID/GID `1100`, both datasets +are distinct, and `zpool` is healthy. No sharing configuration was changed. diff --git a/docs/atlas-updates.md b/docs/atlas-updates.md new file mode 100644 index 0000000..40e90eb --- /dev/null +++ b/docs/atlas-updates.md @@ -0,0 +1,63 @@ +# Atlas Rocky/OpenZFS update and reboot procedure (draft) + +This is an operator-controlled maintenance procedure. The playbook does not +reboot Atlas, replace a pool device, or perform a pool feature upgrade. + +## Preflight + +1. Schedule an outage and confirm no Borg, USB, snapshot, scrub, or resilver + job is active. A service in `activating` is still active; do not interrupt it. +2. Check `zpool status -v zpool` (including scrub status), `zfs list -r zpool`, + `systemctl --failed`, and `systemctl list-timers --all`. Resolve pool errors + first. Record current `uname -r`, `modinfo zfs | grep '^version:'`, + `rpm -q kernel-core kmod-zfs zfs`, and the current boot entry. +3. Confirm a recent successful Borg archive and a usable snapshot. Confirm + the latest published offline USB version and its physical availability; + do not start a USB backup merely to satisfy a checklist without capacity, + UUID, and operator checks. Record timestamps, not just timer state. +4. Ensure console/KVM or another independent recovery route is available. + Check free space in `/boot` and the root filesystem. Review proposed DNF + transactions before consenting to package changes. + +## Change window + +1. Stop client writes and quiesce stateful applications deliberately. Record + which services were stopped; do not assume `ansible-playbook --check` does + this. Avoid updating during a running scrub or backup. +2. Use `dnf upgrade --assumeno` first to review the kernel, `kmod-zfs`, `zfs`, + and dependencies. Confirm a matching kmod will be available for the target + kernel. If compatibility is uncertain, defer the update. +3. Apply the approved DNF transaction. Do not run `zpool upgrade` or enable + new pool feature flags as part of ordinary OS maintenance; that can remove + downgrade options. Preserve at least one known-good boot entry. +4. Reboot **manually** during the agreed outage. Ansible must not trigger it. + +## Post-boot gate + +1. Verify `uname -r`, `modinfo zfs`, `rpm -q kernel-core kmod-zfs zfs`, + `zpool status -v zpool`, `zfs list -r zpool`, and `findmnt -R /zpool`. +2. Verify SELinux remains enforcing; inspect `systemctl --failed` and the + journal for ZFS, mount, SSH, NFS, SMB, Cockpit, Podman, and backup errors. +3. Validate a read-only file listing through SMB and an NFS client access + check before reopening writes. Check the rootless temporary services and + all backup/monitoring timers. Run the Atlas health monitor in `--dry-run` + mode, then a real check after inspection. +4. Re-enable clients and record versions, downtime, anomalies, and next + successful snapshot/Borg run. A green boot alone is not a completed update. + +## Failure response + +If the new kernel cannot load ZFS, boot the previous known-good kernel from +the console and inspect package/kmod matching before trying another reboot. +Do not force-import, rewind, clear errors, or upgrade pool features to make a +failed OS update appear successful. Preserve logs and stop for a recovery +decision if the pool does not import cleanly. + +The procedure-definition item is complete, but the procedure is **not yet +rehearsed** on a replacement host or during a real Atlas update. Record the +first controlled execution and its post-boot evidence separately. + +Read-only preflight on 2026-09-30 observed kernel +`5.14.0-687.52.1.el9_8.x86_64`, ZFS module/package `2.2.11-1`, a healthy +`zpool`, enforcing SELinux, and no failed systemd units. This did not review +an upgrade transaction, stop services, or reboot the host. diff --git a/docs/prometheus-backup.md b/docs/prometheus-backup.md new file mode 100644 index 0000000..37468cc --- /dev/null +++ b/docs/prometheus-backup.md @@ -0,0 +1,106 @@ +# Prometheus to Atlas backup pull + +The playbook and both hosts have the dedicated identity, restricted SSH +access, helpers, and systemd units. A manual export, pull, and temporary +restore passed on 2026-09-30. Both timers are enabled; their first scheduled +runs are pending, so daily operation is not yet verified. + +## Declared design + +- Prometheus prepares a tar archive of Nginx Proxy Manager and Gitea data, + their managed Compose configuration, SSH/firewalld/WireGuard configuration, + and the Gitea SSH path. NPM access logs and regenerable Gitea logs, sessions, + temporary files, and indexers are excluded. The archive contains credentials, + certificates, and the WireGuard private key: protect both copies accordingly. +- The approved consistency mode stops the managed Compose stack for the local + tar creation at 02:00 Europe/Rome, then restarts it even if archiving fails. + A manual test outside that window requires separate approval. +- Prometheus publishes the archive with its checksum as a versioned, read-only + source under `/var/lib/prometheus-backup-export`. A locked service account + has no sudo or supplementary groups. Its only authorized SSH key is forced + through Rocky's `rrsync -ro`; root owns the key file and export directories, + so the account cannot add an unrestricted key or change prepared data. +- Atlas generates and retains the private Ed25519 identity under + `/etc/atlas-prometheus-pull`. Its pinned Prometheus host key came through + the controller's already strict SSH trust; the observed fingerprint was + `SHA256:rfedk7DHI9mLB3UHk/4F3HHlSIiswtCAFsAXvfh6iXk` on 2026-09-30. + Atlas pulls only the prepared `current/` version, verifies SHA-256, tar + readability, metadata, and source freshness, then publishes atomically + below `/zpool/backup/hosts/prometheus/snapshots`. Long-term retention runs + only after publication. A local `rrsync` fixture verified the in-tree + `current` symlink. A live Atlas-to-Prometheus SSH test verified that the + account could list only the prepared versions directory, + cannot obtain a shell, and cannot write to the export. The key is restricted + to `/var/lib/prometheus-backup-export/versions`, not the account's `.ssh`. +- Approved source preparation is 02:00 Europe/Rome, pull 03:00, three source + versions, and 30 daily/8 weekly/12 monthly Atlas versions. The source + timer is non-persistent to avoid an unexpected outage after a missed run. + Atlas rejects a prepared source older than 24 hours. +- The Atlas pull joins the existing health monitor's timer/failure checks + only when enabled. Its failure hook uses 45Drives Alerts; email delivery + is not claimed. A failed source preparation should produce a stale-source + pull failure, not a silently successful reuse of an old archive. + +## Activation and verification + +1. The user confirmed downtime/consistency mode, schedule, retention, and + targeted configuration scope. Review the tar path list and exclusions + against the actual containers. +2. The identity and units are deployed. Re-run the targeted + check, confirm the Atlas public key remains only the restricted Prometheus + account's key, and verify `sshd -T -C user=prometheus-backup,...` plus + read-only SSH denial tests after any SSH configuration change. +3. During an agreed window, start the Prometheus export service manually. + Confirm Compose is healthy afterward, inspect the archive without exposing + file contents, and verify the checksum/metadata. +4. Start the Atlas pull service manually. Confirm the SSH host pin, source + freshness, checksum, tar listing, published `latest`, retention behavior, + clean temporary directories, and healthy pool. +5. Independently restore the selected archive to an empty staging directory + (never `/`) and compare the SQLite databases, Git repositories, NPM data, + Compose file, permissions, and representative files. Test application + startup only in an isolated environment or an approved restore window. +6. The two timers were enabled after the manual test. Verify their calendars + and the next actual run. A successful manual test is not proof of scheduled + operation. + +Narrow static validation: + +```bash +ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \ +ansible-playbook ansible/site.yml --syntax-check +ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \ +ansible-playbook ansible/site.yml --limit prometheus,atlas \ + --tags prometheus_backup --check --diff +``` + +Do not run the export service as part of a routine playbook deployment. The +service restart and any restore/cutover require separate operator decisions. + +On 2026-09-30 the initial targeted `--check --diff` run ended `changed=0` +with gates false. After enabling **implementation only**, a targeted real run +installed the identities and units; both timers were confirmed `disabled` and +`inactive`, the Compose stack stayed active, and the new account was locked +with no supplementary groups. No application was stopped. +The rendered shell helpers passed `bash -n` and ShellCheck; the retention +helper passed an isolated 400-version fixture. These static/isolated checks +were followed by live SSH, export, pull, and temporary restore checks. +Read-only preflight on 2026-09-30 found the Compose service active, all +declared source paths present, both timers inactive, and no prepared versions. +The source filesystem had about 6.0 GB free. Of the 2.1 GB NPM data tree, +2.1 GB was excluded access logs, so the expected archive is much smaller than +the raw tree size; capacity still needs verification after actual exports. +The manual export produced a 285,777,920-byte tar (273 MiB allocated at the +source), and Prometheus retained about 5.8 GB free. NPM and Gitea restarted; +both containers were running and their local HTTP endpoints returned 200. +Atlas pulled the same version, verified SHA-256, published `latest`, and kept +the pool healthy. A full extract to `/var/tmp` yielded 4,747 files; both +SQLite databases passed `PRAGMA integrity_check`, and one restored Gitea Git +repository passed `git fsck`. The temporary restore directory was removed. +This did not test application startup on an isolated host. + +After these checks, Ansible enabled the Prometheus 02:00 Europe/Rome export +timer and Atlas 03:00 Europe/Rome pull timer. The next scheduled occurrences +were displayed for 2026-10-01. Atlas' health monitor now includes the pull +timer. Check both actual service results after the first scheduled run before +claiming unattended operation.