Add verified Prometheus backup pull to Atlas

This commit is contained in:
Fabio Scotto di Santolo
2026-09-30 21:21:48 +02:00
parent 3d2ef02c98
commit 0144600a4a
23 changed files with 851 additions and 22 deletions

View File

@@ -247,9 +247,13 @@ successfully. The first monthly scrub remains a runtime check.
- [x] Define a controlled Rocky kernel/OpenZFS update and reboot procedure in - [x] Define a controlled Rocky kernel/OpenZFS update and reboot procedure in
`docs/atlas-updates.md`. The first real change-window execution is not yet `docs/atlas-updates.md`. The first real change-window execution is not yet
validated; the procedure never reboots automatically or upgrades pool features. validated; the procedure never reboots automatically or upgrades pool features.
- [ ] Add the Atlas-initiated least-privilege Prometheus backup pull: Prometheus exposes only prepared - [x] Add the Atlas-initiated least-privilege Prometheus backup pull: Prometheus exposes only prepared
read-only dumps through a dedicated account and Atlas retains the private SSH key, pinned host key, read-only dumps through a dedicated account and Atlas retains the private SSH key, pinned host key,
atomic pull, verification, retention and systemd service/timer. atomic pull, verification, retention and systemd service/timer. The dedicated key/account and unit
files are deployed; live read-only SSH, shell denial, and write denial were verified. On 2026-09-30
a manual export, Atlas pull, checksum verification, and temporary restore passed; both SQLite
databases passed integrity checks and a restored Git repository passed `git fsck`. Both daily
timers are enabled for 02:00/03:00 Europe/Rome; their first scheduled results remain unverified.
- [x] Decide whether a common SMB/NFS namespace is required: no. `Archive` (SMB) and `photobook` (NFS) - [x] Decide whether a common SMB/NFS namespace is required: no. `Archive` (SMB) and `photobook` (NFS)
remain intentionally distinct; `docs/atlas-sharing-decision.md` records the decision. No ACL or export remain intentionally distinct; `docs/atlas-sharing-decision.md` records the decision. No ACL or export
change is authorized by this decision. change is authorized by this decision.

View File

@@ -485,7 +485,7 @@ etichettata di 45Drives Alerts usare
### Timer systemd di Atlas ### Timer systemd di Atlas
Tutti i nove timer gestiti sono abilitati. Gli orari sono locali ad Atlas (`Europe/Rome`); Borg e Tutti i dieci timer gestiti sono abilitati. Gli orari sono locali ad Atlas (`Europe/Rome`); Borg e
monitoraggio aggiungono il ritardo casuale indicato. Tutti hanno `Persistent=true`: un evento perso monitoraggio aggiungono il ritardo casuale indicato. Tutti hanno `Persistent=true`: un evento perso
viene recuperato quando il timer torna attivo. viene recuperato quando il timer torna attivo.
@@ -500,10 +500,12 @@ viene recuperato quando il timer torna attivo.
| `atlas-borg-check.timer` | `*-*-15 06:00:00` — giorno 15 alle 06:00, più 0–30 min casuali | Controllo repository Borg | | `atlas-borg-check.timer` | `*-*-15 06:00:00` — giorno 15 alle 06:00, più 0–30 min casuali | Controllo repository Borg |
| `atlas-usb-reminder.timer` | `Sat *-*-01..07 10:00:00 Europe/Rome` — primo sabato alle 10:00 | Solo promemoria 45Drives Alerts | | `atlas-usb-reminder.timer` | `Sat *-*-01..07 10:00:00 Europe/Rome` — primo sabato alle 10:00 | Solo promemoria 45Drives Alerts |
| `atlas-health-monitor.timer` | `*:0/30` — ogni mezz'ora, più 0–5 min casuali | Controlli di salute in sola lettura | | `atlas-health-monitor.timer` | `*:0/30` — ogni mezz'ora, più 0–5 min casuali | Controlli di salute in sola lettura |
| `atlas-prometheus-pull.timer` | `*-*-* 03:00:00 Europe/Rome` — ogni giorno alle 03:00 | Pull e verifica del backup preparato su Prometheus |
`atlas-usb-backup.service` **non ha timer** e va avviato manualmente. Il timer del fornitore `atlas-usb-backup.service` **non ha timer** e va avviato manualmente. Il timer del fornitore
`zfs-scrub-weekly@zpool.timer` è disabilitato a favore dello scrub mensile. Il futuro pull del backup `zfs-scrub-weekly@zpool.timer` è disabilitato a favore dello scrub mensile. Il timer di preparazione
Prometheus non ha ancora un timer, perché non è implementato. Durante un backup Borg attivo, su Prometheus è attivo alle 02:00 Europe/Rome; export, pull e ripristino temporaneo manuali sono
riusciti il 2026-09-30, ma il primo ciclo pianificato va ancora verificato. Durante un backup Borg attivo,
`systemctl list-timers` può mostrare `-` per il prossimo evento senza che il timer sia disabilitato. `systemctl list-timers` può mostrare `-` per il prossimo evento senza che il timer sia disabilitato.
Per vedere la pianificazione corrente: `systemctl list-timers --all` su Atlas. Per vedere la pianificazione corrente: `systemctl list-timers --all` su Atlas.
@@ -519,7 +521,10 @@ cutover. L'attuale iCloudPD su Aegis e l'export NFS Photobook restano configurat
all'approvazione e alla verifica di questa migrazione separata. Anche il servizio Atlas sarà all'approvazione e alla verifica di questa migrazione separata. Anche il servizio Atlas sarà
temporaneo in attesa di Uranus. temporaneo in attesa di Uranus.
Il pull dei backup di Prometheus e i test completi di disaster recovery restano da fare. Il backlog Il primo ciclo pianificato del backup di Prometheus e una prova di disaster recovery a dimensione reale
restano da verificare. Il 2026-09-30 una VM Rocky isolata ha superato ricostruzione OS con Ansible,
import del pool RAIDZ2 fittizio e ripristino da snapshot; RPO 24 ore/RTO 72 ore restano obiettivi
provvisori, non tempi misurati. Dettagli e limiti sono in `docs/atlas-recovery.md`. Il backlog
prioritizzato è in `AGENTS.md`. prioritizzato è in `AGENTS.md`.
--- ---

View File

@@ -500,7 +500,7 @@ monitoring. For a labelled 45Drives Alerts delivery test, use
### Atlas systemd timers ### Atlas systemd timers
All nine managed timers below are enabled. Times are local to Atlas (`Europe/Rome`); Borg and monitoring All ten managed timers below are enabled. Times are local to Atlas (`Europe/Rome`); Borg and monitoring
add the indicated randomized delay. Every timer has `Persistent=true`, so a missed calendar run is add the indicated randomized delay. Every timer has `Persistent=true`, so a missed calendar run is
scheduled after the timer becomes active again. scheduled after the timer becomes active again.
@@ -515,10 +515,12 @@ scheduled after the timer becomes active again.
| `atlas-borg-check.timer` | `*-*-15 06:00:00` — 15th of the month at 06:00, plus 0–30 min random delay | Borg repository check | | `atlas-borg-check.timer` | `*-*-15 06:00:00` — 15th of the month at 06:00, plus 0–30 min random delay | Borg repository check |
| `atlas-usb-reminder.timer` | `Sat *-*-01..07 10:00:00 Europe/Rome` — first Saturday at 10:00 | 45Drives Alerts reminder only | | `atlas-usb-reminder.timer` | `Sat *-*-01..07 10:00:00 Europe/Rome` — first Saturday at 10:00 | 45Drives Alerts reminder only |
| `atlas-health-monitor.timer` | `*:0/30` — every half-hour, plus 0–5 min random delay | Read-only health checks | | `atlas-health-monitor.timer` | `*:0/30` — every half-hour, plus 0–5 min random delay | Read-only health checks |
| `atlas-prometheus-pull.timer` | `*-*-* 03:00:00 Europe/Rome` — daily at 03:00 | Pull and verify the prepared Prometheus backup |
`atlas-usb-backup.service` has **no timer**: the encrypted USB backup must be started manually. `atlas-usb-backup.service` has **no timer**: the encrypted USB backup must be started manually.
The vendor's `zfs-scrub-weekly@zpool.timer` is intentionally disabled in favor of the monthly scrub. The vendor's `zfs-scrub-weekly@zpool.timer` is intentionally disabled in favor of the monthly scrub.
The future Prometheus backup pull has no timer yet because that workflow is not implemented. While a The Prometheus export timer runs at 02:00 Europe/Rome; its first scheduled run and the Atlas pull
remain to be observed. A manual export, pull, and temporary restore passed. While a
Borg backup is still running, `systemctl list-timers` may show `-` for its next trigger; this does not Borg backup is still running, `systemctl list-timers` may show `-` for its next trigger; this does not
mean the timer has been disabled. Inspect the current schedule on Atlas with mean the timer has been disabled. Inspect the current schedule on Atlas with
`systemctl list-timers --all`. `systemctl list-timers --all`.
@@ -534,9 +536,21 @@ state outside `Archive`, then test permissions, SELinux, backups and recovery be
The current Aegis iCloudPD service and Atlas Photobook NFS export remain configured until that The current Aegis iCloudPD service and Atlas Photobook NFS export remain configured until that
separate migration is approved and validated; the eventual Atlas service is temporary until Uranus. separate migration is approved and validated; the eventual Atlas service is temporary until Uranus.
Prometheus backup pulls and full disaster-recovery tests remain follow-up work. The prioritized The first scheduled Prometheus backup runs and production-size disaster-recovery tests remain follow-up work. The prioritized
operational backlog is kept in `AGENTS.md`. operational backlog is kept in `AGENTS.md`.
Priority 2 procedures and decisions are recorded in
[`docs/atlas-recovery.md`](docs/atlas-recovery.md),
[`docs/atlas-updates.md`](docs/atlas-updates.md), and
[`docs/atlas-sharing-decision.md`](docs/atlas-sharing-decision.md).
The provisional Atlas recovery objectives are RPO 24 hours and RTO 72 hours;
an isolated small-VM OS rebuild, pool import, Ansible reapplication, and
snapshot restore passed, but full-size recovery time is unmeasured. `Archive` (SMB) and
`photobook` (NFS) remain deliberately separate.
The Prometheus pull architecture and manual export/pull/restore evidence are in
[`docs/prometheus-backup.md`](docs/prometheus-backup.md). Both daily timers are
enabled; their first scheduled runs remain to be verified.
## How layering works ## How layering works
A host can intentionally belong to more than one inventory group. The final configuration is the combination of the host and its groups, not a one-host/one-play mapping. A host can intentionally belong to more than one inventory group. The final configuration is the combination of the host and its groups, not a one-host/one-play mapping.

View File

@@ -80,5 +80,30 @@ server_sshd_settings:
server_sshd_allow_users: server_sshd_allow_users:
- "{{ server_username }}" - "{{ server_username }}"
server_backup_export_enabled: false
server_backup_username: prometheus-backup
server_backup_public_key_name: atlas-pull
server_backup_export_root: /var/lib/prometheus-backup-export
server_backup_rrsync_path: /usr/share/doc/rsync/support/rrsync
server_backup_export_calendar: "*-*-* 02:00:00 Europe/Rome"
server_backup_export_start_timer: false
server_backup_export_source_keep: 3
server_backup_export_paths:
- opt/npm/data
- opt/npm/letsencrypt
- opt/gitea/data
- home/git/.ssh
- opt/docker/server/docker-compose.yml
- etc/systemd/system/podman-compose-server.service
- etc/ssh/sshd_config
- etc/ssh/sshd_config.d
- etc/firewalld
- etc/wireguard/wg0.conf
server_backup_export_excludes:
- opt/npm/data/logs
- opt/gitea/data/gitea/log
- opt/gitea/data/gitea/tmp
- opt/gitea/data/gitea/sessions
- opt/gitea/data/gitea/indexers
server_ssh_authorized_keys: [] server_ssh_authorized_keys: []
server_ssh_authorized_key_directory: "{{ server_user_home }}/.ssh/authorized_keys.d" server_ssh_authorized_key_directory: "{{ server_user_home }}/.ssh/authorized_keys.d"

View File

@@ -49,6 +49,7 @@ atlas_zfs_backup_reservation: 500G
atlas_zfs_dataset_photobook: media/photobook atlas_zfs_dataset_photobook: media/photobook
atlas_mount_root: /zpool atlas_mount_root: /zpool
atlas_manage_storage: true atlas_manage_storage: true
atlas_prometheus_pull_start_timer: true
atlas_manage_zfs_snapshots: true atlas_manage_zfs_snapshots: true
atlas_zfs_snapshot_prefix: atlas-auto atlas_zfs_snapshot_prefix: atlas-auto
atlas_zfs_snapshot_policies: atlas_zfs_snapshot_policies:
@@ -91,6 +92,12 @@ atlas_usb_backup_mapper_name: zpool-backup
atlas_manage_usb_reminder: true atlas_manage_usb_reminder: true
atlas_usb_reminder_calendar: "Sat *-*-01..07 10:00:00 Europe/Rome" atlas_usb_reminder_calendar: "Sat *-*-01..07 10:00:00 Europe/Rome"
atlas_manage_monitoring: true atlas_manage_monitoring: true
atlas_manage_prometheus_backup_pull: true
# Prometheus ED25519 host key read through the controller's strict SSH trust on 2026-09-30.
# Fingerprint: SHA256:rfedk7DHI9mLB3UHk/4F3HHlSIiswtCAFsAXvfh6iXk
atlas_prometheus_ssh_host_key: >-
179.237.102.172 ssh-ed25519
AAAAC3NzaC1lZDI1NTE5AAAAIC4b+QXlPupoEx71W9NKs9tTeYjBqTkVMqbGB97nMNWv
# Physical pool disks and the system NVMe; the disconnected USB disk is intentionally excluded. # Physical pool disks and the system NVMe; the disconnected USB disk is intentionally excluded.
atlas_monitor_smart_devices: atlas_monitor_smart_devices:
- { name: pool-1, path: "{{ atlas_zpool_disks[0] }}", warning_c: 50, critical_c: 55 } - { name: pool-1, path: "{{ atlas_zpool_disks[0] }}", warning_c: 50, critical_c: 55 }

View File

@@ -6,6 +6,8 @@ ansible_port: 22
ansible_ssh_private_key_file: /home/fscotto/.ssh/id_ed25519 ansible_ssh_private_key_file: /home/fscotto/.ssh/id_ed25519
server_username: rocky server_username: rocky
server_backup_export_enabled: true
server_backup_export_start_timer: true
server_duckdns_domain: fscotto server_duckdns_domain: fscotto
server_ssh_authorized_keys: server_ssh_authorized_keys:
- name: ikaros - name: ikaros

View File

@@ -115,6 +115,15 @@ atlas_monitor_notifier: "{{ atlas_usb_reminder_notifier }}"
atlas_monitor_smart_devices: [] atlas_monitor_smart_devices: []
atlas_monitor_timers: [] atlas_monitor_timers: []
atlas_monitor_failure_units: [] atlas_monitor_failure_units: []
atlas_monitor_effective_timers: >-
{{ atlas_monitor_timers
+ ([{'name': 'atlas-prometheus-pull.timer', 'max_age_hours': 26}]
if atlas_manage_prometheus_backup_pull | bool and atlas_prometheus_pull_start_timer | bool
else []) }}
atlas_monitor_effective_failure_units: >-
{{ atlas_monitor_failure_units
+ (['atlas-prometheus-pull.service']
if atlas_manage_prometheus_backup_pull | bool else []) }}
atlas_monitor_remote_capacity: {} atlas_monitor_remote_capacity: {}
atlas_monitor_pool_warning_percent: 80 atlas_monitor_pool_warning_percent: 80
atlas_monitor_pool_critical_percent: 90 atlas_monitor_pool_critical_percent: 90
@@ -141,6 +150,19 @@ atlas_music_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_music }}"
atlas_backup_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_backup }}" atlas_backup_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_backup }}"
atlas_host_backups_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_host_backups }}" atlas_host_backups_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_host_backups }}"
atlas_backup_prometheus_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_backup_prometheus }}" atlas_backup_prometheus_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_backup_prometheus }}"
atlas_manage_prometheus_backup_pull: false
atlas_prometheus_pull_ssh_dir: /etc/atlas-prometheus-pull
atlas_prometheus_pull_private_key_path: "{{ atlas_prometheus_pull_ssh_dir }}/id_ed25519"
atlas_prometheus_pull_known_hosts_path: "{{ atlas_prometheus_pull_ssh_dir }}/known_hosts"
atlas_prometheus_ssh_host_key: ""
atlas_prometheus_pull_source_user: prometheus-backup
atlas_prometheus_pull_source_port: 22
atlas_prometheus_pull_calendar: "*-*-* 03:00:00 Europe/Rome"
atlas_prometheus_pull_start_timer: false
atlas_prometheus_pull_keep_daily: 30
atlas_prometheus_pull_keep_weekly: 8
atlas_prometheus_pull_keep_monthly: 12
atlas_prometheus_pull_max_age_hours: 24
atlas_photobook_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_photobook }}" atlas_photobook_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_photobook }}"
atlas_45drives_repo_url: https://repo.45drives.com/repofiles/rocky/45drives-enterprise.repo atlas_45drives_repo_url: https://repo.45drives.com/repofiles/rocky/45drives-enterprise.repo

View File

@@ -0,0 +1,56 @@
#!/usr/bin/env python3
"""Prune only verified, named Prometheus backup versions after publication."""
import datetime as dt
import pathlib
import re
import shutil
import sys
def main() -> None:
if len(sys.argv) != 5:
raise SystemExit("Usage: atlas-prometheus-prune SNAPSHOTS DAILY WEEKLY MONTHLY")
root = pathlib.Path(sys.argv[1])
counts = [int(value) for value in sys.argv[2:]]
if not root.is_dir() or root.is_symlink() or min(counts) < 1:
raise SystemExit("Invalid backup directory or retention counts")
versions = []
for entry in root.iterdir():
if not entry.is_dir() or entry.is_symlink():
continue
if not re.fullmatch(r"[0-9]{8}T[0-9]{6}Z", entry.name):
continue
try:
when = dt.datetime.strptime(entry.name, "%Y%m%dT%H%M%SZ")
except ValueError:
continue
if not all((entry / name).is_file() for name in ("payload.tar", "payload.sha256", "metadata.json")):
continue
versions.append((when, entry))
versions.sort(reverse=True)
if not versions:
raise SystemExit("No published backup versions found; refusing to prune")
keep = {entry for _, entry in versions[: counts[0]]}
for count, key in (
(counts[1], lambda when: when.isocalendar()[:2]),
(counts[2], lambda when: (when.year, when.month)),
):
periods = set()
for when, entry in versions:
period = key(when)
if period in periods:
continue
periods.add(period)
keep.add(entry)
if len(periods) >= count:
break
for _, entry in versions:
if entry not in keep:
shutil.rmtree(entry)
if __name__ == "__main__":
main()

View File

@@ -23,6 +23,12 @@
- name: Import Atlas offline USB backup tasks - name: Import Atlas offline USB backup tasks
ansible.builtin.import_tasks: usb_backup.yml ansible.builtin.import_tasks: usb_backup.yml
- name: Import Atlas Prometheus backup pull identity tasks
ansible.builtin.import_tasks: prometheus_pull_identity.yml
- name: Import Atlas Prometheus backup pull job tasks
ansible.builtin.import_tasks: prometheus_pull_job.yml
- name: Import Atlas health monitoring tasks - name: Import Atlas health monitoring tasks
ansible.builtin.import_tasks: monitoring.yml ansible.builtin.import_tasks: monitoring.yml

View File

@@ -7,8 +7,8 @@
- atlas_zfs_pool != 'CHANGEME_ZFS_POOL' - atlas_zfs_pool != 'CHANGEME_ZFS_POOL'
- atlas_monitor_calendar | length > 0 - atlas_monitor_calendar | length > 0
- atlas_monitor_smart_devices | length > 0 - atlas_monitor_smart_devices | length > 0
- atlas_monitor_timers | length > 0 - atlas_monitor_effective_timers | length > 0
- atlas_monitor_failure_units | length > 0 - atlas_monitor_effective_failure_units | length > 0
- atlas_monitor_remote_capacity.user == atlas_borg_repository_user - atlas_monitor_remote_capacity.user == atlas_borg_repository_user
- atlas_monitor_remote_capacity.host == atlas_borg_repository_host - atlas_monitor_remote_capacity.host == atlas_borg_repository_host
- atlas_monitor_remote_capacity.run_as == atlas_borg_username - atlas_monitor_remote_capacity.run_as == atlas_borg_username
@@ -49,7 +49,7 @@
that: that:
- item.name is match('^[a-zA-Z0-9@_.-]+\\.timer$') - item.name is match('^[a-zA-Z0-9@_.-]+\\.timer$')
- item.max_age_hours | int >= 0 - item.max_age_hours | int >= 0
loop: "{{ atlas_monitor_timers }}" loop: "{{ atlas_monitor_effective_timers }}"
loop_control: loop_control:
label: "{{ item.name }}" label: "{{ item.name }}"
when: atlas_manage_monitoring | bool when: atlas_manage_monitoring | bool
@@ -59,7 +59,7 @@
ansible.builtin.assert: ansible.builtin.assert:
that: that:
- item is match('^[a-zA-Z0-9@_.-]+\\.service$') - item is match('^[a-zA-Z0-9@_.-]+\\.service$')
loop: "{{ atlas_monitor_failure_units }}" loop: "{{ atlas_monitor_effective_failure_units }}"
when: atlas_manage_monitoring | bool when: atlas_manage_monitoring | bool
- name: Validate Atlas health monitor calendar - name: Validate Atlas health monitor calendar
@@ -144,7 +144,7 @@
owner: root owner: root
group: root group: root
mode: "0755" mode: "0755"
loop: "{{ atlas_monitor_failure_units }}" loop: "{{ atlas_monitor_effective_failure_units }}"
when: atlas_manage_monitoring | bool when: atlas_manage_monitoring | bool
- name: Notify 45Drives Alerts when an Atlas job fails - name: Notify 45Drives Alerts when an Atlas job fails
@@ -155,7 +155,7 @@
owner: root owner: root
group: root group: root
mode: "0644" mode: "0644"
loop: "{{ atlas_monitor_failure_units }}" loop: "{{ atlas_monitor_effective_failure_units }}"
when: atlas_manage_monitoring | bool when: atlas_manage_monitoring | bool
- name: Reload systemd after installing Atlas monitoring - name: Reload systemd after installing Atlas monitoring

View File

@@ -0,0 +1,66 @@
---
- name: Validate Atlas Prometheus pull identity inputs
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
ansible.builtin.assert:
that:
- atlas_prometheus_pull_ssh_dir.startswith('/etc/')
- atlas_prometheus_pull_private_key_path.startswith(atlas_prometheus_pull_ssh_dir ~ '/')
- atlas_prometheus_pull_known_hosts_path.startswith(atlas_prometheus_pull_ssh_dir ~ '/')
- atlas_prometheus_ssh_host_key.startswith(
(hostvars['prometheus'].ansible_host | string) ~ ' ssh-ed25519 '
)
fail_msg: Pin the verified Prometheus ED25519 SSH host key before enabling the pull.
when: atlas_manage_prometheus_backup_pull | bool
- name: Create private Atlas Prometheus pull SSH directory
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
ansible.builtin.file:
path: "{{ atlas_prometheus_pull_ssh_dir }}"
state: directory
owner: root
group: root
mode: "0700"
when: atlas_manage_prometheus_backup_pull | bool
- name: Generate Atlas-only Prometheus pull SSH identity
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
ansible.builtin.command:
argv:
- ssh-keygen
- -q
- -t
- ed25519
- -N
- ""
- -C
- atlas-prometheus-pull@atlas
- -f
- "{{ atlas_prometheus_pull_private_key_path }}"
creates: "{{ atlas_prometheus_pull_private_key_path }}"
when: atlas_manage_prometheus_backup_pull | bool
- name: Protect Atlas-only Prometheus pull SSH identity
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
ansible.builtin.file:
path: "{{ item.path }}"
owner: root
group: root
mode: "{{ item.mode }}"
loop:
- { path: "{{ atlas_prometheus_pull_private_key_path }}", mode: "0600" }
- { path: "{{ atlas_prometheus_pull_private_key_path }}.pub", mode: "0644" }
loop_control:
label: "{{ item.path }}"
when:
- atlas_manage_prometheus_backup_pull | bool
- not ansible_check_mode
- name: Pin Prometheus SSH host key on Atlas
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
ansible.builtin.copy:
content: "{{ atlas_prometheus_ssh_host_key }}\n"
dest: "{{ atlas_prometheus_pull_known_hosts_path }}"
owner: root
group: root
mode: "0600"
when: atlas_manage_prometheus_backup_pull | bool

View File

@@ -0,0 +1,90 @@
---
- name: Validate Atlas Prometheus backup pull inputs
tags: [atlas, backup, prometheus_backup]
ansible.builtin.assert:
that:
- atlas_manage_storage | bool
- atlas_prometheus_pull_source_user is match('^[a-z_][a-z0-9_-]*$')
- atlas_prometheus_pull_source_port | int > 0
- atlas_prometheus_pull_source_port | int < 65536
- atlas_prometheus_pull_keep_daily | int > 0
- atlas_prometheus_pull_keep_weekly | int > 0
- atlas_prometheus_pull_keep_monthly | int > 0
- atlas_prometheus_pull_max_age_hours | int > 0
- atlas_backup_prometheus_mountpoint.startswith(atlas_mount_root ~ '/')
fail_msg: Define the Atlas backup destination, source account, and retention before enabling the pull.
when: atlas_manage_prometheus_backup_pull | bool
- name: Validate Atlas Prometheus backup pull calendar
tags: [atlas, backup, prometheus_backup]
ansible.builtin.command:
argv: [systemd-analyze, calendar, "{{ atlas_prometheus_pull_calendar }}"]
changed_when: false
check_mode: false
when: atlas_manage_prometheus_backup_pull | bool
- name: Create private Atlas Prometheus backup version directory
tags: [atlas, backup, prometheus_backup]
ansible.builtin.file:
path: "{{ atlas_backup_prometheus_mountpoint }}/snapshots"
state: directory
owner: root
group: root
mode: "0700"
when: atlas_manage_prometheus_backup_pull | bool
- name: Install Atlas Prometheus backup pull helper
tags: [atlas, backup, prometheus_backup]
ansible.builtin.template:
src: atlas-prometheus-pull.sh.j2
dest: /usr/local/sbin/atlas-prometheus-pull
owner: root
group: root
mode: "0750"
when: atlas_manage_prometheus_backup_pull | bool
- name: Install Atlas Prometheus backup retention helper
tags: [atlas, backup, prometheus_backup]
ansible.builtin.copy:
src: atlas-prometheus-prune.py
dest: /usr/local/libexec/atlas-prometheus-prune
owner: root
group: root
mode: "0750"
when: atlas_manage_prometheus_backup_pull | bool
- name: Install Atlas Prometheus backup pull systemd units
tags: [atlas, backup, prometheus_backup]
ansible.builtin.template:
src: "{{ item }}.j2"
dest: "/etc/systemd/system/{{ item }}"
owner: root
group: root
mode: "0644"
loop:
- atlas-prometheus-pull.service
- atlas-prometheus-pull.timer
loop_control:
label: "{{ item }}"
register: atlas_prometheus_pull_units
when: atlas_manage_prometheus_backup_pull | bool
- name: Reload systemd after Atlas Prometheus pull unit changes
tags: [atlas, backup, prometheus_backup]
ansible.builtin.systemd:
daemon_reload: true
when:
- atlas_manage_prometheus_backup_pull | bool
- atlas_prometheus_pull_units is changed
- not ansible_check_mode
- name: Enable Atlas Prometheus pull timer only after explicit activation
tags: [atlas, backup, prometheus_backup]
ansible.builtin.systemd:
name: atlas-prometheus-pull.timer
enabled: true
state: started
when:
- atlas_manage_prometheus_backup_pull | bool
- atlas_prometheus_pull_start_timer | bool
- not ansible_check_mode

View File

@@ -3,8 +3,8 @@
"backup_dataset": {{ (atlas_zfs_pool ~ '/' ~ atlas_zfs_dataset_backup) | to_json }}, "backup_dataset": {{ (atlas_zfs_pool ~ '/' ~ atlas_zfs_dataset_backup) | to_json }},
"notifier": {{ atlas_monitor_notifier | to_json }}, "notifier": {{ atlas_monitor_notifier | to_json }},
"smart_devices": {{ atlas_monitor_smart_devices | to_json }}, "smart_devices": {{ atlas_monitor_smart_devices | to_json }},
"timers": {{ atlas_monitor_timers | to_json }}, "timers": {{ atlas_monitor_effective_timers | to_json }},
"failure_units": {{ atlas_monitor_failure_units | to_json }}, "failure_units": {{ atlas_monitor_effective_failure_units | to_json }},
"remote_capacity": {{ atlas_monitor_remote_capacity | to_json }}, "remote_capacity": {{ atlas_monitor_remote_capacity | to_json }},
"pool_warning_percent": {{ atlas_monitor_pool_warning_percent | int }}, "pool_warning_percent": {{ atlas_monitor_pool_warning_percent | int }},
"pool_critical_percent": {{ atlas_monitor_pool_critical_percent | int }}, "pool_critical_percent": {{ atlas_monitor_pool_critical_percent | int }},

View File

@@ -0,0 +1,19 @@
[Unit]
Description=Pull a prepared read-only Prometheus backup to Atlas
RequiresMountsFor={{ atlas_backup_prometheus_mountpoint }}
Wants=network-online.target
After=network-online.target zfs.target
ConditionFileIsExecutable=/usr/local/sbin/atlas-prometheus-pull
ConditionPathExists={{ atlas_prometheus_pull_private_key_path }}
ConditionPathExists={{ atlas_prometheus_pull_known_hosts_path }}
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/atlas-prometheus-pull
User=root
Group=root
UMask=0077
TimeoutStartSec=infinity
Nice=15
IOSchedulingClass=best-effort
IOSchedulingPriority=7

View File

@@ -0,0 +1,75 @@
#!/usr/bin/env bash
set -Eeuo pipefail
umask 077
backup_root={{ atlas_backup_prometheus_mountpoint | quote }}
snapshots="$backup_root/snapshots"
stage=''
exec 9>/run/lock/atlas-prometheus-pull.lock
flock -n 9 || { echo 'A Prometheus pull is already running' >&2; exit 1; }
cleanup() {
local rc=$?
trap - EXIT
if (( rc != 0 )) && [[ -n "$stage" && -d "$stage" ]]; then
rm -rf -- "$stage"
fi
exit "$rc"
}
trap cleanup EXIT
zpool list -H -o name {{ atlas_zfs_pool | quote }} >/dev/null
findmnt -rn --mountpoint "$backup_root" >/dev/null
stage=$(mktemp -d "$backup_root/.staging.XXXXXXXX")
ssh_cmd='/usr/bin/ssh -F /dev/null -o BatchMode=yes -o StrictHostKeyChecking=yes -o UserKnownHostsFile={{ atlas_prometheus_pull_known_hosts_path }} -o IdentitiesOnly=yes -i {{ atlas_prometheus_pull_private_key_path }} -p {{ atlas_prometheus_pull_source_port }}'
rsync -a --partial --delay-updates -e "$ssh_cmd" \
{{ (atlas_prometheus_pull_source_user ~ '@' ~ hostvars['prometheus'].ansible_host ~ ':current/') | quote }} \
"$stage/"
test -s "$stage/payload.tar"
test -s "$stage/payload.sha256"
test -s "$stage/metadata.json"
(cd "$stage" && sha256sum -c payload.sha256)
tar -tf "$stage/payload.tar" >/dev/null
stamp=$(python3 - "$stage/metadata.json" <<'PY'
import json
import datetime as dt
import re
import sys
with open(sys.argv[1], encoding="utf-8") as stream:
metadata = json.load(stream)
stamp = metadata.get("created_utc", "")
if metadata.get("schema") != 1 or metadata.get("host") != "prometheus":
raise SystemExit("Unexpected Prometheus backup metadata")
if not re.fullmatch(r"[0-9]{8}T[0-9]{6}Z", stamp):
raise SystemExit("Invalid Prometheus backup timestamp")
created = dt.datetime.strptime(stamp, "%Y%m%dT%H%M%SZ").replace(tzinfo=dt.timezone.utc)
age = dt.datetime.now(dt.timezone.utc) - created
if age.total_seconds() < -300 or age > dt.timedelta(hours={{ atlas_prometheus_pull_max_age_hours }}):
raise SystemExit("Prometheus backup is outside the configured freshness window")
print(stamp)
PY
)
if [[ -e "$snapshots/$stamp" ]]; then
cmp "$stage/payload.sha256" "$snapshots/$stamp/payload.sha256"
cmp "$stage/metadata.json" "$snapshots/$stamp/metadata.json"
(cd "$snapshots/$stamp" && sha256sum -c payload.sha256)
rm -rf -- "${stage:?}"
stage=''
else
chown -R root:root "$stage"
chmod 0700 "$stage"
chmod 0600 "$stage/payload.tar" "$stage/payload.sha256" "$stage/metadata.json"
mv -- "$stage" "$snapshots/$stamp"
stage=''
fi
latest_link=$(readlink "$backup_root/latest" 2>/dev/null || true)
latest_stamp=${latest_link##*/}
if [[ -z "$latest_stamp" || "$stamp" > "$latest_stamp" ]]; then
ln -s "snapshots/$stamp" "$backup_root/.latest.new"
mv -Tf -- "$backup_root/.latest.new" "$backup_root/latest"
fi
python3 /usr/local/libexec/atlas-prometheus-prune "$snapshots" \
{{ atlas_prometheus_pull_keep_daily }} {{ atlas_prometheus_pull_keep_weekly }} {{ atlas_prometheus_pull_keep_monthly }}
echo "Verified and published Prometheus backup $stamp"

View File

@@ -0,0 +1,10 @@
[Unit]
Description=Schedule Atlas pull of prepared Prometheus backups
[Timer]
OnCalendar={{ atlas_prometheus_pull_calendar }}
Persistent=true
Unit=atlas-prometheus-pull.service
[Install]
WantedBy=timers.target

View File

@@ -0,0 +1,103 @@
---
- name: Validate Prometheus backup export identity inputs
tags: [services, backup, prometheus_backup]
ansible.builtin.assert:
that:
- inventory_hostname == 'prometheus'
- server_backup_username is match('^[a-z_][a-z0-9_-]*$')
- server_backup_username not in ['root', server_username]
- server_backup_export_root.startswith('/var/lib/')
- server_backup_public_key_name is match('^[a-z0-9_-]+$')
- hostvars['atlas'].atlas_manage_prometheus_backup_pull | default(false) | bool
fail_msg: Enable Atlas and Prometheus backup roles together with dedicated identity settings.
when: server_backup_export_enabled | bool
- name: Create dedicated Prometheus backup export group
tags: [services, backup, prometheus_backup]
ansible.builtin.group:
name: "{{ server_backup_username }}"
system: true
state: present
when: server_backup_export_enabled | bool
- name: Create locked Prometheus backup export account
tags: [services, backup, prometheus_backup]
ansible.builtin.user:
name: "{{ server_backup_username }}"
group: "{{ server_backup_username }}"
groups: []
append: false
comment: Read-only prepared backup export for Atlas
home: "{{ server_backup_export_root }}"
create_home: false
shell: /bin/bash
password_lock: true
system: true
state: present
when: server_backup_export_enabled | bool
- name: Require restricted rrsync helper on Prometheus
tags: [services, backup, prometheus_backup]
ansible.builtin.stat:
path: "{{ server_backup_rrsync_path }}"
register: server_backup_rrsync_file
when: server_backup_export_enabled | bool
- name: Validate restricted rrsync helper
tags: [services, backup, prometheus_backup]
ansible.builtin.assert:
that:
- server_backup_rrsync_file.stat.exists
- server_backup_rrsync_file.stat.isreg
- server_backup_rrsync_file.stat.pw_name == 'root'
fail_msg: Rocky rsync must provide the root-owned rrsync support script.
when: server_backup_export_enabled | bool
- name: Create prepared backup export root
tags: [services, backup, prometheus_backup]
ansible.builtin.file:
path: "{{ server_backup_export_root }}"
state: directory
owner: root
group: "{{ server_backup_username }}"
mode: "0750"
when: server_backup_export_enabled | bool
- name: Create restricted Prometheus backup SSH directories
tags: [services, backup, prometheus_backup]
ansible.builtin.file:
path: "{{ item }}"
state: directory
owner: root
group: "{{ server_backup_username }}"
mode: "0750"
loop:
- "{{ server_backup_export_root }}/.ssh"
- "{{ server_backup_export_root }}/.ssh/authorized_keys.d"
when: server_backup_export_enabled | bool
- name: Read Atlas public key for Prometheus backup pull
tags: [services, backup, prometheus_backup]
ansible.builtin.slurp:
src: "{{ hostvars['atlas'].atlas_prometheus_pull_private_key_path | default('/etc/atlas-prometheus-pull/id_ed25519') }}.pub"
delegate_to: atlas
become: true
register: server_backup_atlas_public_key
when:
- server_backup_export_enabled | bool
- not ansible_check_mode
- name: Authorize only restricted read-only backup access from Atlas
tags: [services, backup, prometheus_backup]
ansible.builtin.copy:
content: >-
{{ 'command="/usr/bin/python3 ' ~ server_backup_rrsync_path ~ ' -ro '
~ server_backup_export_root ~ '/versions",restrict '
~ (server_backup_atlas_public_key.content | b64decode | trim) ~ '\n' }}
dest: "{{ server_backup_export_root }}/.ssh/authorized_keys.d/{{ server_backup_public_key_name }}"
owner: root
group: "{{ server_backup_username }}"
mode: "0640"
when:
- server_backup_export_enabled | bool
- not ansible_check_mode

View File

@@ -0,0 +1,90 @@
---
- name: Validate Prometheus backup export job inputs
tags: [services, backup, prometheus_backup]
ansible.builtin.assert:
that:
- server_backup_export_source_keep | int >= 2
- server_backup_export_paths | length > 0
- server_backup_export_paths | unique | length == server_backup_export_paths | length
- >-
server_backup_export_paths
| select('match', '^[a-zA-Z0-9][a-zA-Z0-9._/-]*$') | list | length
== server_backup_export_paths | length
- >-
server_backup_export_paths
| reject('search', '(^|/)\.\.(/|$)') | list | length
== server_backup_export_paths | length
- >-
server_backup_export_excludes
| select('match', '^[a-zA-Z0-9][a-zA-Z0-9._/-]*$') | list | length
== server_backup_export_excludes | length
- >-
server_backup_export_excludes
| reject('search', '(^|/)\.\.(/|$)') | list | length
== server_backup_export_excludes | length
fail_msg: Define safe relative paths and at least two prepared export versions.
when: server_backup_export_enabled | bool
- name: Validate Prometheus backup export calendar
tags: [services, backup, prometheus_backup]
ansible.builtin.command:
argv: [systemd-analyze, calendar, "{{ server_backup_export_calendar }}"]
changed_when: false
check_mode: false
when: server_backup_export_enabled | bool
- name: Ensure prepared Prometheus backup versions directory exists
tags: [services, backup, prometheus_backup]
ansible.builtin.file:
path: "{{ server_backup_export_root }}/versions"
state: directory
owner: root
group: "{{ server_backup_username }}"
mode: "0750"
when: server_backup_export_enabled | bool
- name: Install Prometheus backup export helper
tags: [services, backup, prometheus_backup]
ansible.builtin.template:
src: prometheus-backup-export.sh.j2
dest: /usr/local/sbin/prometheus-backup-export
owner: root
group: root
mode: "0750"
when: server_backup_export_enabled | bool
- name: Install Prometheus backup export systemd units
tags: [services, backup, prometheus_backup]
ansible.builtin.template:
src: "{{ item }}.j2"
dest: "/etc/systemd/system/{{ item }}"
owner: root
group: root
mode: "0644"
loop:
- prometheus-backup-export.service
- prometheus-backup-export.timer
loop_control:
label: "{{ item }}"
register: server_backup_export_units
when: server_backup_export_enabled | bool
- name: Reload systemd after Prometheus backup export unit changes
tags: [services, backup, prometheus_backup]
ansible.builtin.systemd:
daemon_reload: true
when:
- server_backup_export_enabled | bool
- server_backup_export_units is changed
- not ansible_check_mode
- name: Enable Prometheus backup export timer only after explicit activation
tags: [services, backup, prometheus_backup]
ansible.builtin.systemd:
name: prometheus-backup-export.timer
enabled: true
state: started
when:
- server_backup_export_enabled | bool
- server_backup_export_start_timer | bool
- not ansible_check_mode

View File

@@ -53,6 +53,12 @@
tags: [services, podman] tags: [services, podman]
ansible.builtin.include_tasks: podman-compose.yml ansible.builtin.include_tasks: podman-compose.yml
- name: Import Prometheus backup export identity tasks
ansible.builtin.import_tasks: backup_export_identity.yml
- name: Import Prometheus backup export job tasks
ansible.builtin.import_tasks: backup_export_job.yml
- name: Ensure server SSH authorized key fragments directory exists - name: Ensure server SSH authorized key fragments directory exists
tags: [services, ssh] tags: [services, ssh]
ansible.builtin.file: ansible.builtin.file:
@@ -77,13 +83,17 @@
when: server_ssh_authorized_keys | length > 0 when: server_ssh_authorized_keys | length > 0
- name: Configure server SSH authorized key fragments - name: Configure server SSH authorized key fragments
tags: [services, ssh] tags: [services, ssh, prometheus_backup]
ansible.builtin.lineinfile: ansible.builtin.lineinfile:
path: /etc/ssh/sshd_config path: /etc/ssh/sshd_config
regexp: '^\s*AuthorizedKeysFile\s+' regexp: '^\s*AuthorizedKeysFile\s+'
line: >- line: >-
AuthorizedKeysFile {{ server_ssh_authorized_keys | map(attribute='name') AuthorizedKeysFile {{
| map('regex_replace', '^', '%h/.ssh/authorized_keys.d/') | join(' ') }} ((server_ssh_authorized_keys | map(attribute='name')
| map('regex_replace', '^', '%h/.ssh/authorized_keys.d/') | list)
+ (['%h/.ssh/authorized_keys.d/' ~ server_backup_public_key_name]
if server_backup_export_enabled | bool else [])) | join(' ')
}}
state: present state: present
validate: "sshd -t -f %s" validate: "sshd -t -f %s"
notify: Reload SSH service notify: Reload SSH service
@@ -100,11 +110,13 @@
notify: Reload SSH service notify: Reload SSH service
- name: Restrict SSH login to allowed users on server - name: Restrict SSH login to allowed users on server
tags: [services] tags: [services, prometheus_backup]
ansible.builtin.lineinfile: ansible.builtin.lineinfile:
path: /etc/ssh/sshd_config path: /etc/ssh/sshd_config
regexp: '^\s*AllowUsers\s+' regexp: '^\s*AllowUsers\s+'
line: "AllowUsers {{ server_sshd_allow_users | join(' ') }}" line: >-
AllowUsers {{ (server_sshd_allow_users
+ ([server_backup_username] if server_backup_export_enabled | bool else [])) | join(' ') }}
state: present state: present
validate: "sshd -t -f %s" validate: "sshd -t -f %s"
notify: Reload SSH service notify: Reload SSH service

View File

@@ -0,0 +1,15 @@
[Unit]
Description=Prepare a read-only Prometheus application backup for Atlas
RequiresMountsFor=/opt/npm /opt/gitea {{ server_backup_export_root }}
ConditionFileIsExecutable=/usr/local/sbin/prometheus-backup-export
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/prometheus-backup-export
User=root
Group=root
UMask=0077
TimeoutStartSec=infinity
Nice=10
IOSchedulingClass=best-effort
IOSchedulingPriority=7

View File

@@ -0,0 +1,92 @@
#!/usr/bin/env bash
set -Eeuo pipefail
umask 077
export_root={{ server_backup_export_root | quote }}
versions="$export_root/versions"
stack_unit=podman-compose-server.service
stamp=$(date -u +%Y%m%dT%H%M%SZ)
stage=''
stack_stopped=false
exec 9>/run/lock/prometheus-backup-export.lock
flock -n 9 || { echo 'A backup export is already running' >&2; exit 1; }
cleanup() {
local rc=$?
trap - EXIT
if "$stack_stopped"; then
if systemctl is-active --quiet "$stack_unit"; then
systemctl restart "$stack_unit" || rc=1
else
systemctl start "$stack_unit" || rc=1
fi
fi
if (( rc != 0 )) && [[ -n "$stage" && -d "$stage" ]]; then
rm -rf -- "$stage"
fi
exit "$rc"
}
trap cleanup EXIT
trap 'exit 129' HUP
trap 'exit 130' INT
trap 'exit 143' TERM
systemctl is-active --quiet "$stack_unit" || {
echo 'The managed Compose stack must be active before preparing a backup' >&2
exit 1
}
paths=(
{% for path in server_backup_export_paths %}
{{ path | quote }}
{% endfor %}
)
excludes=(
{% for path in server_backup_export_excludes %}
--exclude={{ path | quote }}
{% endfor %}
)
for path in "${paths[@]}"; do
[[ -e "/$path" ]] || { echo "Required backup path missing: /$path" >&2; exit 1; }
done
[[ ! -e "$versions/$stamp" ]] || { echo "Export version already exists: $stamp" >&2; exit 1; }
stage=$(mktemp -d "$export_root/.staging.XXXXXXXX")
# SQLite databases and their accompanying files are copied while both
# managed containers are stopped. The EXIT trap restarts the stack on error.
stack_stopped=true
systemctl stop "$stack_unit"
tar --acls --xattrs --selinux "${excludes[@]}" -C / -cf "$stage/payload.tar" "${paths[@]}"
systemctl start "$stack_unit"
for container in nginx-proxy-manager gitea; do
running=false
for _ in {1..30}; do
if [[ $(podman inspect --format '{{ '{{.State.Running}}' }}' "$container" 2>/dev/null) == true ]]; then
running=true
break
fi
sleep 2
done
"$running" || { echo "Container did not restart: $container" >&2; exit 1; }
done
stack_stopped=false
tar -tf "$stage/payload.tar" >/dev/null
(cd "$stage" && sha256sum payload.tar >payload.sha256)
printf '{"schema":1,"host":"prometheus","created_utc":"%s"}\n' "$stamp" >"$stage/metadata.json"
chown root:{{ server_backup_username }} "$stage" "$stage/payload.tar" "$stage/payload.sha256" "$stage/metadata.json"
chmod 0750 "$stage"
chmod 0640 "$stage/payload.tar" "$stage/payload.sha256" "$stage/metadata.json"
mv -- "$stage" "$versions/$stamp"
stage=''
ln -s "$stamp" "$versions/.current.new"
mv -Tf -- "$versions/.current.new" "$versions/current"
# Keep a small source-side safety window; Atlas owns long-term retention.
mapfile -t old_versions < <(find "$versions" -mindepth 1 -maxdepth 1 -type d \
-printf '%f\n' | grep -E '^[0-9]{8}T[0-9]{6}Z$' | sort -r | tail -n +{{ server_backup_export_source_keep + 1 }})
for old in "${old_versions[@]}"; do
rm -rf -- "${versions:?}/$old"
done
echo "Prepared Prometheus backup export $stamp"

View File

@@ -0,0 +1,10 @@
[Unit]
Description=Prepare daily Prometheus application backup for Atlas
[Timer]
OnCalendar={{ server_backup_export_calendar }}
Persistent=false
Unit=prometheus-backup-export.service
[Install]
WantedBy=timers.target

106
docs/prometheus-backup.md Normal file
View File

@@ -0,0 +1,106 @@
# Prometheus to Atlas backup pull
The playbook and both hosts have the dedicated identity, restricted SSH
access, helpers, and systemd units. A manual export, pull, and temporary
restore passed on 2026-09-30. Both timers are enabled; their first scheduled
runs are pending, so daily operation is not yet verified.
## Declared design
- Prometheus prepares a tar archive of Nginx Proxy Manager and Gitea data,
their managed Compose configuration, SSH/firewalld/WireGuard configuration,
and the Gitea SSH path. NPM access logs and regenerable Gitea logs, sessions,
temporary files, and indexers are excluded. The archive contains credentials,
certificates, and the WireGuard private key: protect both copies accordingly.
- The approved consistency mode stops the managed Compose stack for the local
tar creation at 02:00 Europe/Rome, then restarts it even if archiving fails.
A manual test outside that window requires separate approval.
- Prometheus publishes the archive with its checksum as a versioned, read-only
source under `/var/lib/prometheus-backup-export`. A locked service account
has no sudo or supplementary groups. Its only authorized SSH key is forced
through Rocky's `rrsync -ro`; root owns the key file and export directories,
so the account cannot add an unrestricted key or change prepared data.
- Atlas generates and retains the private Ed25519 identity under
`/etc/atlas-prometheus-pull`. Its pinned Prometheus host key came through
the controller's already strict SSH trust; the observed fingerprint was
`SHA256:rfedk7DHI9mLB3UHk/4F3HHlSIiswtCAFsAXvfh6iXk` on 2026-09-30.
Atlas pulls only the prepared `current/` version, verifies SHA-256, tar
readability, metadata, and source freshness, then publishes atomically
below `/zpool/backup/hosts/prometheus/snapshots`. Long-term retention runs
only after publication. A local `rrsync` fixture verified the in-tree
`current` symlink. A live Atlas-to-Prometheus SSH test verified that the
account could list only the prepared versions directory,
cannot obtain a shell, and cannot write to the export. The key is restricted
to `/var/lib/prometheus-backup-export/versions`, not the account's `.ssh`.
- Approved source preparation is 02:00 Europe/Rome, pull 03:00, three source
versions, and 30 daily/8 weekly/12 monthly Atlas versions. The source
timer is non-persistent to avoid an unexpected outage after a missed run.
Atlas rejects a prepared source older than 24 hours.
- The Atlas pull joins the existing health monitor's timer/failure checks
only when enabled. Its failure hook uses 45Drives Alerts; email delivery
is not claimed. A failed source preparation should produce a stale-source
pull failure, not a silently successful reuse of an old archive.
## Activation and verification
1. The user confirmed downtime/consistency mode, schedule, retention, and
targeted configuration scope. Review the tar path list and exclusions
against the actual containers.
2. The identity and units are deployed. Re-run the targeted
check, confirm the Atlas public key remains only the restricted Prometheus
account's key, and verify `sshd -T -C user=prometheus-backup,...` plus
read-only SSH denial tests after any SSH configuration change.
3. During an agreed window, start the Prometheus export service manually.
Confirm Compose is healthy afterward, inspect the archive without exposing
file contents, and verify the checksum/metadata.
4. Start the Atlas pull service manually. Confirm the SSH host pin, source
freshness, checksum, tar listing, published `latest`, retention behavior,
clean temporary directories, and healthy pool.
5. Independently restore the selected archive to an empty staging directory
(never `/`) and compare the SQLite databases, Git repositories, NPM data,
Compose file, permissions, and representative files. Test application
startup only in an isolated environment or an approved restore window.
6. The two timers were enabled after the manual test. Verify their calendars
and the next actual run. A successful manual test is not proof of scheduled
operation.
Narrow static validation:
```bash
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
ansible-playbook ansible/site.yml --syntax-check
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
ansible-playbook ansible/site.yml --limit prometheus,atlas \
--tags prometheus_backup --check --diff
```
Do not run the export service as part of a routine playbook deployment. The
service restart and any restore/cutover require separate operator decisions.
On 2026-09-30 the initial targeted `--check --diff` run ended `changed=0`
with gates false. After enabling **implementation only**, a targeted real run
installed the identities and units; both timers were confirmed `disabled` and
`inactive`, the Compose stack stayed active, and the new account was locked
with no supplementary groups. No application was stopped.
The rendered shell helpers passed `bash -n` and ShellCheck; the retention
helper passed an isolated 400-version fixture. These static/isolated checks
were followed by live SSH, export, pull, and temporary restore checks.
Read-only preflight on 2026-09-30 found the Compose service active, all
declared source paths present, both timers inactive, and no prepared versions.
The source filesystem had about 6.0 GB free. Of the 2.1 GB NPM data tree,
2.1 GB was excluded access logs, so the expected archive is much smaller than
the raw tree size; capacity still needs verification after actual exports.
The manual export produced a 285,777,920-byte tar (273 MiB allocated at the
source), and Prometheus retained about 5.8 GB free. NPM and Gitea restarted;
both containers were running and their local HTTP endpoints returned 200.
Atlas pulled the same version, verified SHA-256, published `latest`, and kept
the pool healthy. A full extract to `/var/tmp` yielded 4,747 files; both
SQLite databases passed `PRAGMA integrity_check`, and one restored Gitea Git
repository passed `git fsck`. The temporary restore directory was removed.
This did not test application startup on an isolated host.
After these checks, Ansible enabled the Prometheus 02:00 Europe/Rome export
timer and Atlas 03:00 Europe/Rome pull timer. The next scheduled occurrences
were displayed for 2026-10-01. Atlas' health monitor now includes the pull
timer. Check both actual service results after the first scheduled run before
claiming unattended operation.