mirror of
https://github.com/fscotto/infra.git
synced 2026-10-03 13:29:58 +00:00
Merge pull request #10 from fscotto/feature/atlas-priority2-recovery
Feature/atlas priority2 recovery
This commit is contained in:
29
AGENTS.md
29
AGENTS.md
@@ -226,7 +226,7 @@ successfully. The first monthly scrub remains a runtime check.
|
|||||||
read-only ZFS snapshot test restored one file to `/var/tmp`, confirmed matching contents, ownership,
|
read-only ZFS snapshot test restored one file to `/var/tmp`, confirmed matching contents, ownership,
|
||||||
mode, mtime and ACL, then removed its temporary copy and on-demand mount. This is a file-level smoke
|
mode, mtime and ACL, then removed its temporary copy and on-demand mount. This is a file-level smoke
|
||||||
test, not full dataset recovery. An independent USB file restore passed on 2026-09-25 with matching
|
test, not full dataset recovery. An independent USB file restore passed on 2026-09-25 with matching
|
||||||
content and metadata; full disaster recovery remains a separate Priority 2 task.
|
content and metadata; the later scaled OS-rebuild rehearsal is documented under Priority 2.
|
||||||
- [x] Add monitoring and alerting for pool health, scrub/resilver, SMART data, temperatures, free space,
|
- [x] Add monitoring and alerting for pool health, scrub/resilver, SMART data, temperatures, free space,
|
||||||
snapshot/local-backup growth, Hetzner Storage Box quota, and failed maintenance or backup timers.
|
snapshot/local-backup growth, Hetzner Storage Box quota, and failed maintenance or backup timers.
|
||||||
The half-hourly Atlas health monitor and systemd final-failure hooks are deployed. A live probe
|
The half-hourly Atlas health monitor and systemd final-failure hooks are deployed. A live probe
|
||||||
@@ -237,15 +237,26 @@ successfully. The first monthly scrub remains a runtime check.
|
|||||||
on Atlas, but a new real failure notification has not been deliberately triggered.
|
on Atlas, but a new real failure notification has not been deliberately triggered.
|
||||||
|
|
||||||
### Priority 2 - NAS operability and recovery
|
### Priority 2 - NAS operability and recovery
|
||||||
- [ ] Document and test disaster recovery: rebuild Atlas with Ansible, import the existing pool, restore
|
- [x] Document and test disaster recovery in `docs/atlas-recovery.md`: the operator confirmed Vault
|
||||||
from snapshot/USB/Hetzner, preserve Vault and Borg recovery material offline, and define RPO/RTO.
|
and Borg recovery material is available offline; provisional targets are RPO 24h/RTO 72h. On
|
||||||
- [ ] Define a controlled Rocky kernel/OpenZFS update and reboot procedure.
|
2026-09-30 an isolated small Rocky VM was rebuilt with the Atlas Ansible roles, imported its
|
||||||
- [ ] Add the Atlas-initiated least-privilege Prometheus backup pull: Prometheus exposes only prepared
|
preserved RAIDZ2 pool without force/rewind, and restored a file from the preserved snapshot;
|
||||||
|
the second Ansible run was idempotent. Earlier independent production ZFS, USB, and Borg file
|
||||||
|
restore tests remain separate evidence. A production-size full restore, unclean import, and
|
||||||
|
measured 24h/72h compliance are not claimed.
|
||||||
|
- [x] Define a controlled Rocky kernel/OpenZFS update and reboot procedure in
|
||||||
|
`docs/atlas-updates.md`. The first real change-window execution is not yet
|
||||||
|
validated; the procedure never reboots automatically or upgrades pool features.
|
||||||
|
- [x] Add the Atlas-initiated least-privilege Prometheus backup pull: Prometheus exposes only prepared
|
||||||
read-only dumps through a dedicated account and Atlas retains the private SSH key, pinned host key,
|
read-only dumps through a dedicated account and Atlas retains the private SSH key, pinned host key,
|
||||||
atomic pull, verification, retention and systemd service/timer.
|
atomic pull, verification, retention and systemd service/timer. The dedicated key/account and unit
|
||||||
- [ ] Decide whether a common SMB/NFS namespace is required. `Archive` (SMB) and `photobook` (NFS) are
|
files are deployed; live read-only SSH, shell denial, and write denial were verified. On 2026-09-30
|
||||||
intentionally distinct today; only if a shared namespace is selected, finalize its UID/GID, group,
|
a manual export, Atlas pull, checksum verification, and temporary restore passed; both SQLite
|
||||||
and POSIX ACL model and test the same files through both protocols.
|
databases passed integrity checks and a restored Git repository passed `git fsck`. Both daily
|
||||||
|
timers are enabled for 02:00/03:00 Europe/Rome; their first scheduled results remain unverified.
|
||||||
|
- [x] Decide whether a common SMB/NFS namespace is required: no. `Archive` (SMB) and `photobook` (NFS)
|
||||||
|
remain intentionally distinct; `docs/atlas-sharing-decision.md` records the decision. No ACL or export
|
||||||
|
change is authorized by this decision.
|
||||||
|
|
||||||
### Priority 3 - Service expansion
|
### Priority 3 - Service expansion
|
||||||
- [ ] After data protection and recovery are validated, populate `/zpool/media/music` and validate Navidrome.
|
- [ ] After data protection and recovery are validated, populate `/zpool/media/music` and validate Navidrome.
|
||||||
|
|||||||
13
README.it.md
13
README.it.md
@@ -485,7 +485,7 @@ etichettata di 45Drives Alerts usare
|
|||||||
|
|
||||||
### Timer systemd di Atlas
|
### Timer systemd di Atlas
|
||||||
|
|
||||||
Tutti i nove timer gestiti sono abilitati. Gli orari sono locali ad Atlas (`Europe/Rome`); Borg e
|
Tutti i dieci timer gestiti sono abilitati. Gli orari sono locali ad Atlas (`Europe/Rome`); Borg e
|
||||||
monitoraggio aggiungono il ritardo casuale indicato. Tutti hanno `Persistent=true`: un evento perso
|
monitoraggio aggiungono il ritardo casuale indicato. Tutti hanno `Persistent=true`: un evento perso
|
||||||
viene recuperato quando il timer torna attivo.
|
viene recuperato quando il timer torna attivo.
|
||||||
|
|
||||||
@@ -500,10 +500,12 @@ viene recuperato quando il timer torna attivo.
|
|||||||
| `atlas-borg-check.timer` | `*-*-15 06:00:00` — giorno 15 alle 06:00, più 0–30 min casuali | Controllo repository Borg |
|
| `atlas-borg-check.timer` | `*-*-15 06:00:00` — giorno 15 alle 06:00, più 0–30 min casuali | Controllo repository Borg |
|
||||||
| `atlas-usb-reminder.timer` | `Sat *-*-01..07 10:00:00 Europe/Rome` — primo sabato alle 10:00 | Solo promemoria 45Drives Alerts |
|
| `atlas-usb-reminder.timer` | `Sat *-*-01..07 10:00:00 Europe/Rome` — primo sabato alle 10:00 | Solo promemoria 45Drives Alerts |
|
||||||
| `atlas-health-monitor.timer` | `*:0/30` — ogni mezz'ora, più 0–5 min casuali | Controlli di salute in sola lettura |
|
| `atlas-health-monitor.timer` | `*:0/30` — ogni mezz'ora, più 0–5 min casuali | Controlli di salute in sola lettura |
|
||||||
|
| `atlas-prometheus-pull.timer` | `*-*-* 03:00:00 Europe/Rome` — ogni giorno alle 03:00 | Pull e verifica del backup preparato su Prometheus |
|
||||||
|
|
||||||
`atlas-usb-backup.service` **non ha timer** e va avviato manualmente. Il timer del fornitore
|
`atlas-usb-backup.service` **non ha timer** e va avviato manualmente. Il timer del fornitore
|
||||||
`zfs-scrub-weekly@zpool.timer` è disabilitato a favore dello scrub mensile. Il futuro pull del backup
|
`zfs-scrub-weekly@zpool.timer` è disabilitato a favore dello scrub mensile. Il timer di preparazione
|
||||||
Prometheus non ha ancora un timer, perché non è implementato. Durante un backup Borg attivo,
|
su Prometheus è attivo alle 02:00 Europe/Rome; export, pull e ripristino temporaneo manuali sono
|
||||||
|
riusciti il 2026-09-30, ma il primo ciclo pianificato va ancora verificato. Durante un backup Borg attivo,
|
||||||
`systemctl list-timers` può mostrare `-` per il prossimo evento senza che il timer sia disabilitato.
|
`systemctl list-timers` può mostrare `-` per il prossimo evento senza che il timer sia disabilitato.
|
||||||
Per vedere la pianificazione corrente: `systemctl list-timers --all` su Atlas.
|
Per vedere la pianificazione corrente: `systemctl list-timers --all` su Atlas.
|
||||||
|
|
||||||
@@ -519,7 +521,10 @@ cutover. L'attuale iCloudPD su Aegis e l'export NFS Photobook restano configurat
|
|||||||
all'approvazione e alla verifica di questa migrazione separata. Anche il servizio Atlas sarà
|
all'approvazione e alla verifica di questa migrazione separata. Anche il servizio Atlas sarà
|
||||||
temporaneo in attesa di Uranus.
|
temporaneo in attesa di Uranus.
|
||||||
|
|
||||||
Il pull dei backup di Prometheus e i test completi di disaster recovery restano da fare. Il backlog
|
Il primo ciclo pianificato del backup di Prometheus e una prova di disaster recovery a dimensione reale
|
||||||
|
restano da verificare. Il 2026-09-30 una VM Rocky isolata ha superato ricostruzione OS con Ansible,
|
||||||
|
import del pool RAIDZ2 fittizio e ripristino da snapshot; RPO 24 ore/RTO 72 ore restano obiettivi
|
||||||
|
provvisori, non tempi misurati. Dettagli e limiti sono in `docs/atlas-recovery.md`. Il backlog
|
||||||
prioritizzato è in `AGENTS.md`.
|
prioritizzato è in `AGENTS.md`.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|||||||
20
README.md
20
README.md
@@ -500,7 +500,7 @@ monitoring. For a labelled 45Drives Alerts delivery test, use
|
|||||||
|
|
||||||
### Atlas systemd timers
|
### Atlas systemd timers
|
||||||
|
|
||||||
All nine managed timers below are enabled. Times are local to Atlas (`Europe/Rome`); Borg and monitoring
|
All ten managed timers below are enabled. Times are local to Atlas (`Europe/Rome`); Borg and monitoring
|
||||||
add the indicated randomized delay. Every timer has `Persistent=true`, so a missed calendar run is
|
add the indicated randomized delay. Every timer has `Persistent=true`, so a missed calendar run is
|
||||||
scheduled after the timer becomes active again.
|
scheduled after the timer becomes active again.
|
||||||
|
|
||||||
@@ -515,10 +515,12 @@ scheduled after the timer becomes active again.
|
|||||||
| `atlas-borg-check.timer` | `*-*-15 06:00:00` — 15th of the month at 06:00, plus 0–30 min random delay | Borg repository check |
|
| `atlas-borg-check.timer` | `*-*-15 06:00:00` — 15th of the month at 06:00, plus 0–30 min random delay | Borg repository check |
|
||||||
| `atlas-usb-reminder.timer` | `Sat *-*-01..07 10:00:00 Europe/Rome` — first Saturday at 10:00 | 45Drives Alerts reminder only |
|
| `atlas-usb-reminder.timer` | `Sat *-*-01..07 10:00:00 Europe/Rome` — first Saturday at 10:00 | 45Drives Alerts reminder only |
|
||||||
| `atlas-health-monitor.timer` | `*:0/30` — every half-hour, plus 0–5 min random delay | Read-only health checks |
|
| `atlas-health-monitor.timer` | `*:0/30` — every half-hour, plus 0–5 min random delay | Read-only health checks |
|
||||||
|
| `atlas-prometheus-pull.timer` | `*-*-* 03:00:00 Europe/Rome` — daily at 03:00 | Pull and verify the prepared Prometheus backup |
|
||||||
|
|
||||||
`atlas-usb-backup.service` has **no timer**: the encrypted USB backup must be started manually.
|
`atlas-usb-backup.service` has **no timer**: the encrypted USB backup must be started manually.
|
||||||
The vendor's `zfs-scrub-weekly@zpool.timer` is intentionally disabled in favor of the monthly scrub.
|
The vendor's `zfs-scrub-weekly@zpool.timer` is intentionally disabled in favor of the monthly scrub.
|
||||||
The future Prometheus backup pull has no timer yet because that workflow is not implemented. While a
|
The Prometheus export timer runs at 02:00 Europe/Rome; its first scheduled run and the Atlas pull
|
||||||
|
remain to be observed. A manual export, pull, and temporary restore passed. While a
|
||||||
Borg backup is still running, `systemctl list-timers` may show `-` for its next trigger; this does not
|
Borg backup is still running, `systemctl list-timers` may show `-` for its next trigger; this does not
|
||||||
mean the timer has been disabled. Inspect the current schedule on Atlas with
|
mean the timer has been disabled. Inspect the current schedule on Atlas with
|
||||||
`systemctl list-timers --all`.
|
`systemctl list-timers --all`.
|
||||||
@@ -534,9 +536,21 @@ state outside `Archive`, then test permissions, SELinux, backups and recovery be
|
|||||||
The current Aegis iCloudPD service and Atlas Photobook NFS export remain configured until that
|
The current Aegis iCloudPD service and Atlas Photobook NFS export remain configured until that
|
||||||
separate migration is approved and validated; the eventual Atlas service is temporary until Uranus.
|
separate migration is approved and validated; the eventual Atlas service is temporary until Uranus.
|
||||||
|
|
||||||
Prometheus backup pulls and full disaster-recovery tests remain follow-up work. The prioritized
|
The first scheduled Prometheus backup runs and production-size disaster-recovery tests remain follow-up work. The prioritized
|
||||||
operational backlog is kept in `AGENTS.md`.
|
operational backlog is kept in `AGENTS.md`.
|
||||||
|
|
||||||
|
Priority 2 procedures and decisions are recorded in
|
||||||
|
[`docs/atlas-recovery.md`](docs/atlas-recovery.md),
|
||||||
|
[`docs/atlas-updates.md`](docs/atlas-updates.md), and
|
||||||
|
[`docs/atlas-sharing-decision.md`](docs/atlas-sharing-decision.md).
|
||||||
|
The provisional Atlas recovery objectives are RPO 24 hours and RTO 72 hours;
|
||||||
|
an isolated small-VM OS rebuild, pool import, Ansible reapplication, and
|
||||||
|
snapshot restore passed, but full-size recovery time is unmeasured. `Archive` (SMB) and
|
||||||
|
`photobook` (NFS) remain deliberately separate.
|
||||||
|
The Prometheus pull architecture and manual export/pull/restore evidence are in
|
||||||
|
[`docs/prometheus-backup.md`](docs/prometheus-backup.md). Both daily timers are
|
||||||
|
enabled; their first scheduled runs remain to be verified.
|
||||||
|
|
||||||
## How layering works
|
## How layering works
|
||||||
|
|
||||||
A host can intentionally belong to more than one inventory group. The final configuration is the combination of the host and its groups, not a one-host/one-play mapping.
|
A host can intentionally belong to more than one inventory group. The final configuration is the combination of the host and its groups, not a one-host/one-play mapping.
|
||||||
|
|||||||
@@ -80,5 +80,30 @@ server_sshd_settings:
|
|||||||
|
|
||||||
server_sshd_allow_users:
|
server_sshd_allow_users:
|
||||||
- "{{ server_username }}"
|
- "{{ server_username }}"
|
||||||
|
server_backup_export_enabled: false
|
||||||
|
server_backup_username: prometheus-backup
|
||||||
|
server_backup_public_key_name: atlas-pull
|
||||||
|
server_backup_export_root: /var/lib/prometheus-backup-export
|
||||||
|
server_backup_rrsync_path: /usr/share/doc/rsync/support/rrsync
|
||||||
|
server_backup_export_calendar: "*-*-* 02:00:00 Europe/Rome"
|
||||||
|
server_backup_export_start_timer: false
|
||||||
|
server_backup_export_source_keep: 3
|
||||||
|
server_backup_export_paths:
|
||||||
|
- opt/npm/data
|
||||||
|
- opt/npm/letsencrypt
|
||||||
|
- opt/gitea/data
|
||||||
|
- home/git/.ssh
|
||||||
|
- opt/docker/server/docker-compose.yml
|
||||||
|
- etc/systemd/system/podman-compose-server.service
|
||||||
|
- etc/ssh/sshd_config
|
||||||
|
- etc/ssh/sshd_config.d
|
||||||
|
- etc/firewalld
|
||||||
|
- etc/wireguard/wg0.conf
|
||||||
|
server_backup_export_excludes:
|
||||||
|
- opt/npm/data/logs
|
||||||
|
- opt/gitea/data/gitea/log
|
||||||
|
- opt/gitea/data/gitea/tmp
|
||||||
|
- opt/gitea/data/gitea/sessions
|
||||||
|
- opt/gitea/data/gitea/indexers
|
||||||
server_ssh_authorized_keys: []
|
server_ssh_authorized_keys: []
|
||||||
server_ssh_authorized_key_directory: "{{ server_user_home }}/.ssh/authorized_keys.d"
|
server_ssh_authorized_key_directory: "{{ server_user_home }}/.ssh/authorized_keys.d"
|
||||||
|
|||||||
@@ -49,6 +49,7 @@ atlas_zfs_backup_reservation: 500G
|
|||||||
atlas_zfs_dataset_photobook: media/photobook
|
atlas_zfs_dataset_photobook: media/photobook
|
||||||
atlas_mount_root: /zpool
|
atlas_mount_root: /zpool
|
||||||
atlas_manage_storage: true
|
atlas_manage_storage: true
|
||||||
|
atlas_prometheus_pull_start_timer: true
|
||||||
atlas_manage_zfs_snapshots: true
|
atlas_manage_zfs_snapshots: true
|
||||||
atlas_zfs_snapshot_prefix: atlas-auto
|
atlas_zfs_snapshot_prefix: atlas-auto
|
||||||
atlas_zfs_snapshot_policies:
|
atlas_zfs_snapshot_policies:
|
||||||
@@ -91,6 +92,12 @@ atlas_usb_backup_mapper_name: zpool-backup
|
|||||||
atlas_manage_usb_reminder: true
|
atlas_manage_usb_reminder: true
|
||||||
atlas_usb_reminder_calendar: "Sat *-*-01..07 10:00:00 Europe/Rome"
|
atlas_usb_reminder_calendar: "Sat *-*-01..07 10:00:00 Europe/Rome"
|
||||||
atlas_manage_monitoring: true
|
atlas_manage_monitoring: true
|
||||||
|
atlas_manage_prometheus_backup_pull: true
|
||||||
|
# Prometheus ED25519 host key read through the controller's strict SSH trust on 2026-09-30.
|
||||||
|
# Fingerprint: SHA256:rfedk7DHI9mLB3UHk/4F3HHlSIiswtCAFsAXvfh6iXk
|
||||||
|
atlas_prometheus_ssh_host_key: >-
|
||||||
|
179.237.102.172 ssh-ed25519
|
||||||
|
AAAAC3NzaC1lZDI1NTE5AAAAIC4b+QXlPupoEx71W9NKs9tTeYjBqTkVMqbGB97nMNWv
|
||||||
# Physical pool disks and the system NVMe; the disconnected USB disk is intentionally excluded.
|
# Physical pool disks and the system NVMe; the disconnected USB disk is intentionally excluded.
|
||||||
atlas_monitor_smart_devices:
|
atlas_monitor_smart_devices:
|
||||||
- { name: pool-1, path: "{{ atlas_zpool_disks[0] }}", warning_c: 50, critical_c: 55 }
|
- { name: pool-1, path: "{{ atlas_zpool_disks[0] }}", warning_c: 50, critical_c: 55 }
|
||||||
|
|||||||
@@ -6,6 +6,8 @@ ansible_port: 22
|
|||||||
ansible_ssh_private_key_file: /home/fscotto/.ssh/id_ed25519
|
ansible_ssh_private_key_file: /home/fscotto/.ssh/id_ed25519
|
||||||
|
|
||||||
server_username: rocky
|
server_username: rocky
|
||||||
|
server_backup_export_enabled: true
|
||||||
|
server_backup_export_start_timer: true
|
||||||
server_duckdns_domain: fscotto
|
server_duckdns_domain: fscotto
|
||||||
server_ssh_authorized_keys:
|
server_ssh_authorized_keys:
|
||||||
- name: ikaros
|
- name: ikaros
|
||||||
|
|||||||
@@ -115,6 +115,15 @@ atlas_monitor_notifier: "{{ atlas_usb_reminder_notifier }}"
|
|||||||
atlas_monitor_smart_devices: []
|
atlas_monitor_smart_devices: []
|
||||||
atlas_monitor_timers: []
|
atlas_monitor_timers: []
|
||||||
atlas_monitor_failure_units: []
|
atlas_monitor_failure_units: []
|
||||||
|
atlas_monitor_effective_timers: >-
|
||||||
|
{{ atlas_monitor_timers
|
||||||
|
+ ([{'name': 'atlas-prometheus-pull.timer', 'max_age_hours': 26}]
|
||||||
|
if atlas_manage_prometheus_backup_pull | bool and atlas_prometheus_pull_start_timer | bool
|
||||||
|
else []) }}
|
||||||
|
atlas_monitor_effective_failure_units: >-
|
||||||
|
{{ atlas_monitor_failure_units
|
||||||
|
+ (['atlas-prometheus-pull.service']
|
||||||
|
if atlas_manage_prometheus_backup_pull | bool else []) }}
|
||||||
atlas_monitor_remote_capacity: {}
|
atlas_monitor_remote_capacity: {}
|
||||||
atlas_monitor_pool_warning_percent: 80
|
atlas_monitor_pool_warning_percent: 80
|
||||||
atlas_monitor_pool_critical_percent: 90
|
atlas_monitor_pool_critical_percent: 90
|
||||||
@@ -141,6 +150,19 @@ atlas_music_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_music }}"
|
|||||||
atlas_backup_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_backup }}"
|
atlas_backup_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_backup }}"
|
||||||
atlas_host_backups_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_host_backups }}"
|
atlas_host_backups_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_host_backups }}"
|
||||||
atlas_backup_prometheus_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_backup_prometheus }}"
|
atlas_backup_prometheus_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_backup_prometheus }}"
|
||||||
|
atlas_manage_prometheus_backup_pull: false
|
||||||
|
atlas_prometheus_pull_ssh_dir: /etc/atlas-prometheus-pull
|
||||||
|
atlas_prometheus_pull_private_key_path: "{{ atlas_prometheus_pull_ssh_dir }}/id_ed25519"
|
||||||
|
atlas_prometheus_pull_known_hosts_path: "{{ atlas_prometheus_pull_ssh_dir }}/known_hosts"
|
||||||
|
atlas_prometheus_ssh_host_key: ""
|
||||||
|
atlas_prometheus_pull_source_user: prometheus-backup
|
||||||
|
atlas_prometheus_pull_source_port: 22
|
||||||
|
atlas_prometheus_pull_calendar: "*-*-* 03:00:00 Europe/Rome"
|
||||||
|
atlas_prometheus_pull_start_timer: false
|
||||||
|
atlas_prometheus_pull_keep_daily: 30
|
||||||
|
atlas_prometheus_pull_keep_weekly: 8
|
||||||
|
atlas_prometheus_pull_keep_monthly: 12
|
||||||
|
atlas_prometheus_pull_max_age_hours: 24
|
||||||
atlas_photobook_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_photobook }}"
|
atlas_photobook_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_photobook }}"
|
||||||
|
|
||||||
atlas_45drives_repo_url: https://repo.45drives.com/repofiles/rocky/45drives-enterprise.repo
|
atlas_45drives_repo_url: https://repo.45drives.com/repofiles/rocky/45drives-enterprise.repo
|
||||||
|
|||||||
56
ansible/roles/profile_atlas/files/atlas-prometheus-prune.py
Normal file
56
ansible/roles/profile_atlas/files/atlas-prometheus-prune.py
Normal file
@@ -0,0 +1,56 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Prune only verified, named Prometheus backup versions after publication."""
|
||||||
|
|
||||||
|
import datetime as dt
|
||||||
|
import pathlib
|
||||||
|
import re
|
||||||
|
import shutil
|
||||||
|
import sys
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
if len(sys.argv) != 5:
|
||||||
|
raise SystemExit("Usage: atlas-prometheus-prune SNAPSHOTS DAILY WEEKLY MONTHLY")
|
||||||
|
root = pathlib.Path(sys.argv[1])
|
||||||
|
counts = [int(value) for value in sys.argv[2:]]
|
||||||
|
if not root.is_dir() or root.is_symlink() or min(counts) < 1:
|
||||||
|
raise SystemExit("Invalid backup directory or retention counts")
|
||||||
|
versions = []
|
||||||
|
for entry in root.iterdir():
|
||||||
|
if not entry.is_dir() or entry.is_symlink():
|
||||||
|
continue
|
||||||
|
if not re.fullmatch(r"[0-9]{8}T[0-9]{6}Z", entry.name):
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
when = dt.datetime.strptime(entry.name, "%Y%m%dT%H%M%SZ")
|
||||||
|
except ValueError:
|
||||||
|
continue
|
||||||
|
if not all((entry / name).is_file() for name in ("payload.tar", "payload.sha256", "metadata.json")):
|
||||||
|
continue
|
||||||
|
versions.append((when, entry))
|
||||||
|
versions.sort(reverse=True)
|
||||||
|
if not versions:
|
||||||
|
raise SystemExit("No published backup versions found; refusing to prune")
|
||||||
|
|
||||||
|
keep = {entry for _, entry in versions[: counts[0]]}
|
||||||
|
for count, key in (
|
||||||
|
(counts[1], lambda when: when.isocalendar()[:2]),
|
||||||
|
(counts[2], lambda when: (when.year, when.month)),
|
||||||
|
):
|
||||||
|
periods = set()
|
||||||
|
for when, entry in versions:
|
||||||
|
period = key(when)
|
||||||
|
if period in periods:
|
||||||
|
continue
|
||||||
|
periods.add(period)
|
||||||
|
keep.add(entry)
|
||||||
|
if len(periods) >= count:
|
||||||
|
break
|
||||||
|
|
||||||
|
for _, entry in versions:
|
||||||
|
if entry not in keep:
|
||||||
|
shutil.rmtree(entry)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -23,6 +23,12 @@
|
|||||||
- name: Import Atlas offline USB backup tasks
|
- name: Import Atlas offline USB backup tasks
|
||||||
ansible.builtin.import_tasks: usb_backup.yml
|
ansible.builtin.import_tasks: usb_backup.yml
|
||||||
|
|
||||||
|
- name: Import Atlas Prometheus backup pull identity tasks
|
||||||
|
ansible.builtin.import_tasks: prometheus_pull_identity.yml
|
||||||
|
|
||||||
|
- name: Import Atlas Prometheus backup pull job tasks
|
||||||
|
ansible.builtin.import_tasks: prometheus_pull_job.yml
|
||||||
|
|
||||||
- name: Import Atlas health monitoring tasks
|
- name: Import Atlas health monitoring tasks
|
||||||
ansible.builtin.import_tasks: monitoring.yml
|
ansible.builtin.import_tasks: monitoring.yml
|
||||||
|
|
||||||
|
|||||||
@@ -7,8 +7,8 @@
|
|||||||
- atlas_zfs_pool != 'CHANGEME_ZFS_POOL'
|
- atlas_zfs_pool != 'CHANGEME_ZFS_POOL'
|
||||||
- atlas_monitor_calendar | length > 0
|
- atlas_monitor_calendar | length > 0
|
||||||
- atlas_monitor_smart_devices | length > 0
|
- atlas_monitor_smart_devices | length > 0
|
||||||
- atlas_monitor_timers | length > 0
|
- atlas_monitor_effective_timers | length > 0
|
||||||
- atlas_monitor_failure_units | length > 0
|
- atlas_monitor_effective_failure_units | length > 0
|
||||||
- atlas_monitor_remote_capacity.user == atlas_borg_repository_user
|
- atlas_monitor_remote_capacity.user == atlas_borg_repository_user
|
||||||
- atlas_monitor_remote_capacity.host == atlas_borg_repository_host
|
- atlas_monitor_remote_capacity.host == atlas_borg_repository_host
|
||||||
- atlas_monitor_remote_capacity.run_as == atlas_borg_username
|
- atlas_monitor_remote_capacity.run_as == atlas_borg_username
|
||||||
@@ -49,7 +49,7 @@
|
|||||||
that:
|
that:
|
||||||
- item.name is match('^[a-zA-Z0-9@_.-]+\\.timer$')
|
- item.name is match('^[a-zA-Z0-9@_.-]+\\.timer$')
|
||||||
- item.max_age_hours | int >= 0
|
- item.max_age_hours | int >= 0
|
||||||
loop: "{{ atlas_monitor_timers }}"
|
loop: "{{ atlas_monitor_effective_timers }}"
|
||||||
loop_control:
|
loop_control:
|
||||||
label: "{{ item.name }}"
|
label: "{{ item.name }}"
|
||||||
when: atlas_manage_monitoring | bool
|
when: atlas_manage_monitoring | bool
|
||||||
@@ -59,7 +59,7 @@
|
|||||||
ansible.builtin.assert:
|
ansible.builtin.assert:
|
||||||
that:
|
that:
|
||||||
- item is match('^[a-zA-Z0-9@_.-]+\\.service$')
|
- item is match('^[a-zA-Z0-9@_.-]+\\.service$')
|
||||||
loop: "{{ atlas_monitor_failure_units }}"
|
loop: "{{ atlas_monitor_effective_failure_units }}"
|
||||||
when: atlas_manage_monitoring | bool
|
when: atlas_manage_monitoring | bool
|
||||||
|
|
||||||
- name: Validate Atlas health monitor calendar
|
- name: Validate Atlas health monitor calendar
|
||||||
@@ -144,7 +144,7 @@
|
|||||||
owner: root
|
owner: root
|
||||||
group: root
|
group: root
|
||||||
mode: "0755"
|
mode: "0755"
|
||||||
loop: "{{ atlas_monitor_failure_units }}"
|
loop: "{{ atlas_monitor_effective_failure_units }}"
|
||||||
when: atlas_manage_monitoring | bool
|
when: atlas_manage_monitoring | bool
|
||||||
|
|
||||||
- name: Notify 45Drives Alerts when an Atlas job fails
|
- name: Notify 45Drives Alerts when an Atlas job fails
|
||||||
@@ -155,7 +155,7 @@
|
|||||||
owner: root
|
owner: root
|
||||||
group: root
|
group: root
|
||||||
mode: "0644"
|
mode: "0644"
|
||||||
loop: "{{ atlas_monitor_failure_units }}"
|
loop: "{{ atlas_monitor_effective_failure_units }}"
|
||||||
when: atlas_manage_monitoring | bool
|
when: atlas_manage_monitoring | bool
|
||||||
|
|
||||||
- name: Reload systemd after installing Atlas monitoring
|
- name: Reload systemd after installing Atlas monitoring
|
||||||
|
|||||||
@@ -0,0 +1,66 @@
|
|||||||
|
---
|
||||||
|
- name: Validate Atlas Prometheus pull identity inputs
|
||||||
|
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
|
||||||
|
ansible.builtin.assert:
|
||||||
|
that:
|
||||||
|
- atlas_prometheus_pull_ssh_dir.startswith('/etc/')
|
||||||
|
- atlas_prometheus_pull_private_key_path.startswith(atlas_prometheus_pull_ssh_dir ~ '/')
|
||||||
|
- atlas_prometheus_pull_known_hosts_path.startswith(atlas_prometheus_pull_ssh_dir ~ '/')
|
||||||
|
- atlas_prometheus_ssh_host_key.startswith(
|
||||||
|
(hostvars['prometheus'].ansible_host | string) ~ ' ssh-ed25519 '
|
||||||
|
)
|
||||||
|
fail_msg: Pin the verified Prometheus ED25519 SSH host key before enabling the pull.
|
||||||
|
when: atlas_manage_prometheus_backup_pull | bool
|
||||||
|
|
||||||
|
- name: Create private Atlas Prometheus pull SSH directory
|
||||||
|
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
|
||||||
|
ansible.builtin.file:
|
||||||
|
path: "{{ atlas_prometheus_pull_ssh_dir }}"
|
||||||
|
state: directory
|
||||||
|
owner: root
|
||||||
|
group: root
|
||||||
|
mode: "0700"
|
||||||
|
when: atlas_manage_prometheus_backup_pull | bool
|
||||||
|
|
||||||
|
- name: Generate Atlas-only Prometheus pull SSH identity
|
||||||
|
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
|
||||||
|
ansible.builtin.command:
|
||||||
|
argv:
|
||||||
|
- ssh-keygen
|
||||||
|
- -q
|
||||||
|
- -t
|
||||||
|
- ed25519
|
||||||
|
- -N
|
||||||
|
- ""
|
||||||
|
- -C
|
||||||
|
- atlas-prometheus-pull@atlas
|
||||||
|
- -f
|
||||||
|
- "{{ atlas_prometheus_pull_private_key_path }}"
|
||||||
|
creates: "{{ atlas_prometheus_pull_private_key_path }}"
|
||||||
|
when: atlas_manage_prometheus_backup_pull | bool
|
||||||
|
|
||||||
|
- name: Protect Atlas-only Prometheus pull SSH identity
|
||||||
|
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
|
||||||
|
ansible.builtin.file:
|
||||||
|
path: "{{ item.path }}"
|
||||||
|
owner: root
|
||||||
|
group: root
|
||||||
|
mode: "{{ item.mode }}"
|
||||||
|
loop:
|
||||||
|
- { path: "{{ atlas_prometheus_pull_private_key_path }}", mode: "0600" }
|
||||||
|
- { path: "{{ atlas_prometheus_pull_private_key_path }}.pub", mode: "0644" }
|
||||||
|
loop_control:
|
||||||
|
label: "{{ item.path }}"
|
||||||
|
when:
|
||||||
|
- atlas_manage_prometheus_backup_pull | bool
|
||||||
|
- not ansible_check_mode
|
||||||
|
|
||||||
|
- name: Pin Prometheus SSH host key on Atlas
|
||||||
|
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
|
||||||
|
ansible.builtin.copy:
|
||||||
|
content: "{{ atlas_prometheus_ssh_host_key }}\n"
|
||||||
|
dest: "{{ atlas_prometheus_pull_known_hosts_path }}"
|
||||||
|
owner: root
|
||||||
|
group: root
|
||||||
|
mode: "0600"
|
||||||
|
when: atlas_manage_prometheus_backup_pull | bool
|
||||||
90
ansible/roles/profile_atlas/tasks/prometheus_pull_job.yml
Normal file
90
ansible/roles/profile_atlas/tasks/prometheus_pull_job.yml
Normal file
@@ -0,0 +1,90 @@
|
|||||||
|
---
|
||||||
|
- name: Validate Atlas Prometheus backup pull inputs
|
||||||
|
tags: [atlas, backup, prometheus_backup]
|
||||||
|
ansible.builtin.assert:
|
||||||
|
that:
|
||||||
|
- atlas_manage_storage | bool
|
||||||
|
- atlas_prometheus_pull_source_user is match('^[a-z_][a-z0-9_-]*$')
|
||||||
|
- atlas_prometheus_pull_source_port | int > 0
|
||||||
|
- atlas_prometheus_pull_source_port | int < 65536
|
||||||
|
- atlas_prometheus_pull_keep_daily | int > 0
|
||||||
|
- atlas_prometheus_pull_keep_weekly | int > 0
|
||||||
|
- atlas_prometheus_pull_keep_monthly | int > 0
|
||||||
|
- atlas_prometheus_pull_max_age_hours | int > 0
|
||||||
|
- atlas_backup_prometheus_mountpoint.startswith(atlas_mount_root ~ '/')
|
||||||
|
fail_msg: Define the Atlas backup destination, source account, and retention before enabling the pull.
|
||||||
|
when: atlas_manage_prometheus_backup_pull | bool
|
||||||
|
|
||||||
|
- name: Validate Atlas Prometheus backup pull calendar
|
||||||
|
tags: [atlas, backup, prometheus_backup]
|
||||||
|
ansible.builtin.command:
|
||||||
|
argv: [systemd-analyze, calendar, "{{ atlas_prometheus_pull_calendar }}"]
|
||||||
|
changed_when: false
|
||||||
|
check_mode: false
|
||||||
|
when: atlas_manage_prometheus_backup_pull | bool
|
||||||
|
|
||||||
|
- name: Create private Atlas Prometheus backup version directory
|
||||||
|
tags: [atlas, backup, prometheus_backup]
|
||||||
|
ansible.builtin.file:
|
||||||
|
path: "{{ atlas_backup_prometheus_mountpoint }}/snapshots"
|
||||||
|
state: directory
|
||||||
|
owner: root
|
||||||
|
group: root
|
||||||
|
mode: "0700"
|
||||||
|
when: atlas_manage_prometheus_backup_pull | bool
|
||||||
|
|
||||||
|
- name: Install Atlas Prometheus backup pull helper
|
||||||
|
tags: [atlas, backup, prometheus_backup]
|
||||||
|
ansible.builtin.template:
|
||||||
|
src: atlas-prometheus-pull.sh.j2
|
||||||
|
dest: /usr/local/sbin/atlas-prometheus-pull
|
||||||
|
owner: root
|
||||||
|
group: root
|
||||||
|
mode: "0750"
|
||||||
|
when: atlas_manage_prometheus_backup_pull | bool
|
||||||
|
|
||||||
|
- name: Install Atlas Prometheus backup retention helper
|
||||||
|
tags: [atlas, backup, prometheus_backup]
|
||||||
|
ansible.builtin.copy:
|
||||||
|
src: atlas-prometheus-prune.py
|
||||||
|
dest: /usr/local/libexec/atlas-prometheus-prune
|
||||||
|
owner: root
|
||||||
|
group: root
|
||||||
|
mode: "0750"
|
||||||
|
when: atlas_manage_prometheus_backup_pull | bool
|
||||||
|
|
||||||
|
- name: Install Atlas Prometheus backup pull systemd units
|
||||||
|
tags: [atlas, backup, prometheus_backup]
|
||||||
|
ansible.builtin.template:
|
||||||
|
src: "{{ item }}.j2"
|
||||||
|
dest: "/etc/systemd/system/{{ item }}"
|
||||||
|
owner: root
|
||||||
|
group: root
|
||||||
|
mode: "0644"
|
||||||
|
loop:
|
||||||
|
- atlas-prometheus-pull.service
|
||||||
|
- atlas-prometheus-pull.timer
|
||||||
|
loop_control:
|
||||||
|
label: "{{ item }}"
|
||||||
|
register: atlas_prometheus_pull_units
|
||||||
|
when: atlas_manage_prometheus_backup_pull | bool
|
||||||
|
|
||||||
|
- name: Reload systemd after Atlas Prometheus pull unit changes
|
||||||
|
tags: [atlas, backup, prometheus_backup]
|
||||||
|
ansible.builtin.systemd:
|
||||||
|
daemon_reload: true
|
||||||
|
when:
|
||||||
|
- atlas_manage_prometheus_backup_pull | bool
|
||||||
|
- atlas_prometheus_pull_units is changed
|
||||||
|
- not ansible_check_mode
|
||||||
|
|
||||||
|
- name: Enable Atlas Prometheus pull timer only after explicit activation
|
||||||
|
tags: [atlas, backup, prometheus_backup]
|
||||||
|
ansible.builtin.systemd:
|
||||||
|
name: atlas-prometheus-pull.timer
|
||||||
|
enabled: true
|
||||||
|
state: started
|
||||||
|
when:
|
||||||
|
- atlas_manage_prometheus_backup_pull | bool
|
||||||
|
- atlas_prometheus_pull_start_timer | bool
|
||||||
|
- not ansible_check_mode
|
||||||
@@ -3,8 +3,8 @@
|
|||||||
"backup_dataset": {{ (atlas_zfs_pool ~ '/' ~ atlas_zfs_dataset_backup) | to_json }},
|
"backup_dataset": {{ (atlas_zfs_pool ~ '/' ~ atlas_zfs_dataset_backup) | to_json }},
|
||||||
"notifier": {{ atlas_monitor_notifier | to_json }},
|
"notifier": {{ atlas_monitor_notifier | to_json }},
|
||||||
"smart_devices": {{ atlas_monitor_smart_devices | to_json }},
|
"smart_devices": {{ atlas_monitor_smart_devices | to_json }},
|
||||||
"timers": {{ atlas_monitor_timers | to_json }},
|
"timers": {{ atlas_monitor_effective_timers | to_json }},
|
||||||
"failure_units": {{ atlas_monitor_failure_units | to_json }},
|
"failure_units": {{ atlas_monitor_effective_failure_units | to_json }},
|
||||||
"remote_capacity": {{ atlas_monitor_remote_capacity | to_json }},
|
"remote_capacity": {{ atlas_monitor_remote_capacity | to_json }},
|
||||||
"pool_warning_percent": {{ atlas_monitor_pool_warning_percent | int }},
|
"pool_warning_percent": {{ atlas_monitor_pool_warning_percent | int }},
|
||||||
"pool_critical_percent": {{ atlas_monitor_pool_critical_percent | int }},
|
"pool_critical_percent": {{ atlas_monitor_pool_critical_percent | int }},
|
||||||
|
|||||||
@@ -0,0 +1,19 @@
|
|||||||
|
[Unit]
|
||||||
|
Description=Pull a prepared read-only Prometheus backup to Atlas
|
||||||
|
RequiresMountsFor={{ atlas_backup_prometheus_mountpoint }}
|
||||||
|
Wants=network-online.target
|
||||||
|
After=network-online.target zfs.target
|
||||||
|
ConditionFileIsExecutable=/usr/local/sbin/atlas-prometheus-pull
|
||||||
|
ConditionPathExists={{ atlas_prometheus_pull_private_key_path }}
|
||||||
|
ConditionPathExists={{ atlas_prometheus_pull_known_hosts_path }}
|
||||||
|
|
||||||
|
[Service]
|
||||||
|
Type=oneshot
|
||||||
|
ExecStart=/usr/local/sbin/atlas-prometheus-pull
|
||||||
|
User=root
|
||||||
|
Group=root
|
||||||
|
UMask=0077
|
||||||
|
TimeoutStartSec=infinity
|
||||||
|
Nice=15
|
||||||
|
IOSchedulingClass=best-effort
|
||||||
|
IOSchedulingPriority=7
|
||||||
@@ -0,0 +1,75 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
set -Eeuo pipefail
|
||||||
|
umask 077
|
||||||
|
|
||||||
|
backup_root={{ atlas_backup_prometheus_mountpoint | quote }}
|
||||||
|
snapshots="$backup_root/snapshots"
|
||||||
|
stage=''
|
||||||
|
exec 9>/run/lock/atlas-prometheus-pull.lock
|
||||||
|
flock -n 9 || { echo 'A Prometheus pull is already running' >&2; exit 1; }
|
||||||
|
|
||||||
|
cleanup() {
|
||||||
|
local rc=$?
|
||||||
|
trap - EXIT
|
||||||
|
if (( rc != 0 )) && [[ -n "$stage" && -d "$stage" ]]; then
|
||||||
|
rm -rf -- "$stage"
|
||||||
|
fi
|
||||||
|
exit "$rc"
|
||||||
|
}
|
||||||
|
trap cleanup EXIT
|
||||||
|
|
||||||
|
zpool list -H -o name {{ atlas_zfs_pool | quote }} >/dev/null
|
||||||
|
findmnt -rn --mountpoint "$backup_root" >/dev/null
|
||||||
|
stage=$(mktemp -d "$backup_root/.staging.XXXXXXXX")
|
||||||
|
ssh_cmd='/usr/bin/ssh -F /dev/null -o BatchMode=yes -o StrictHostKeyChecking=yes -o UserKnownHostsFile={{ atlas_prometheus_pull_known_hosts_path }} -o IdentitiesOnly=yes -i {{ atlas_prometheus_pull_private_key_path }} -p {{ atlas_prometheus_pull_source_port }}'
|
||||||
|
rsync -a --partial --delay-updates -e "$ssh_cmd" \
|
||||||
|
{{ (atlas_prometheus_pull_source_user ~ '@' ~ hostvars['prometheus'].ansible_host ~ ':current/') | quote }} \
|
||||||
|
"$stage/"
|
||||||
|
|
||||||
|
test -s "$stage/payload.tar"
|
||||||
|
test -s "$stage/payload.sha256"
|
||||||
|
test -s "$stage/metadata.json"
|
||||||
|
(cd "$stage" && sha256sum -c payload.sha256)
|
||||||
|
tar -tf "$stage/payload.tar" >/dev/null
|
||||||
|
stamp=$(python3 - "$stage/metadata.json" <<'PY'
|
||||||
|
import json
|
||||||
|
import datetime as dt
|
||||||
|
import re
|
||||||
|
import sys
|
||||||
|
|
||||||
|
with open(sys.argv[1], encoding="utf-8") as stream:
|
||||||
|
metadata = json.load(stream)
|
||||||
|
stamp = metadata.get("created_utc", "")
|
||||||
|
if metadata.get("schema") != 1 or metadata.get("host") != "prometheus":
|
||||||
|
raise SystemExit("Unexpected Prometheus backup metadata")
|
||||||
|
if not re.fullmatch(r"[0-9]{8}T[0-9]{6}Z", stamp):
|
||||||
|
raise SystemExit("Invalid Prometheus backup timestamp")
|
||||||
|
created = dt.datetime.strptime(stamp, "%Y%m%dT%H%M%SZ").replace(tzinfo=dt.timezone.utc)
|
||||||
|
age = dt.datetime.now(dt.timezone.utc) - created
|
||||||
|
if age.total_seconds() < -300 or age > dt.timedelta(hours={{ atlas_prometheus_pull_max_age_hours }}):
|
||||||
|
raise SystemExit("Prometheus backup is outside the configured freshness window")
|
||||||
|
print(stamp)
|
||||||
|
PY
|
||||||
|
)
|
||||||
|
if [[ -e "$snapshots/$stamp" ]]; then
|
||||||
|
cmp "$stage/payload.sha256" "$snapshots/$stamp/payload.sha256"
|
||||||
|
cmp "$stage/metadata.json" "$snapshots/$stamp/metadata.json"
|
||||||
|
(cd "$snapshots/$stamp" && sha256sum -c payload.sha256)
|
||||||
|
rm -rf -- "${stage:?}"
|
||||||
|
stage=''
|
||||||
|
else
|
||||||
|
chown -R root:root "$stage"
|
||||||
|
chmod 0700 "$stage"
|
||||||
|
chmod 0600 "$stage/payload.tar" "$stage/payload.sha256" "$stage/metadata.json"
|
||||||
|
mv -- "$stage" "$snapshots/$stamp"
|
||||||
|
stage=''
|
||||||
|
fi
|
||||||
|
latest_link=$(readlink "$backup_root/latest" 2>/dev/null || true)
|
||||||
|
latest_stamp=${latest_link##*/}
|
||||||
|
if [[ -z "$latest_stamp" || "$stamp" > "$latest_stamp" ]]; then
|
||||||
|
ln -s "snapshots/$stamp" "$backup_root/.latest.new"
|
||||||
|
mv -Tf -- "$backup_root/.latest.new" "$backup_root/latest"
|
||||||
|
fi
|
||||||
|
python3 /usr/local/libexec/atlas-prometheus-prune "$snapshots" \
|
||||||
|
{{ atlas_prometheus_pull_keep_daily }} {{ atlas_prometheus_pull_keep_weekly }} {{ atlas_prometheus_pull_keep_monthly }}
|
||||||
|
echo "Verified and published Prometheus backup $stamp"
|
||||||
@@ -0,0 +1,10 @@
|
|||||||
|
[Unit]
|
||||||
|
Description=Schedule Atlas pull of prepared Prometheus backups
|
||||||
|
|
||||||
|
[Timer]
|
||||||
|
OnCalendar={{ atlas_prometheus_pull_calendar }}
|
||||||
|
Persistent=true
|
||||||
|
Unit=atlas-prometheus-pull.service
|
||||||
|
|
||||||
|
[Install]
|
||||||
|
WantedBy=timers.target
|
||||||
103
ansible/roles/profile_server/tasks/backup_export_identity.yml
Normal file
103
ansible/roles/profile_server/tasks/backup_export_identity.yml
Normal file
@@ -0,0 +1,103 @@
|
|||||||
|
---
|
||||||
|
- name: Validate Prometheus backup export identity inputs
|
||||||
|
tags: [services, backup, prometheus_backup]
|
||||||
|
ansible.builtin.assert:
|
||||||
|
that:
|
||||||
|
- inventory_hostname == 'prometheus'
|
||||||
|
- server_backup_username is match('^[a-z_][a-z0-9_-]*$')
|
||||||
|
- server_backup_username not in ['root', server_username]
|
||||||
|
- server_backup_export_root.startswith('/var/lib/')
|
||||||
|
- server_backup_public_key_name is match('^[a-z0-9_-]+$')
|
||||||
|
- hostvars['atlas'].atlas_manage_prometheus_backup_pull | default(false) | bool
|
||||||
|
fail_msg: Enable Atlas and Prometheus backup roles together with dedicated identity settings.
|
||||||
|
when: server_backup_export_enabled | bool
|
||||||
|
|
||||||
|
- name: Create dedicated Prometheus backup export group
|
||||||
|
tags: [services, backup, prometheus_backup]
|
||||||
|
ansible.builtin.group:
|
||||||
|
name: "{{ server_backup_username }}"
|
||||||
|
system: true
|
||||||
|
state: present
|
||||||
|
when: server_backup_export_enabled | bool
|
||||||
|
|
||||||
|
- name: Create locked Prometheus backup export account
|
||||||
|
tags: [services, backup, prometheus_backup]
|
||||||
|
ansible.builtin.user:
|
||||||
|
name: "{{ server_backup_username }}"
|
||||||
|
group: "{{ server_backup_username }}"
|
||||||
|
groups: []
|
||||||
|
append: false
|
||||||
|
comment: Read-only prepared backup export for Atlas
|
||||||
|
home: "{{ server_backup_export_root }}"
|
||||||
|
create_home: false
|
||||||
|
shell: /bin/bash
|
||||||
|
password_lock: true
|
||||||
|
system: true
|
||||||
|
state: present
|
||||||
|
when: server_backup_export_enabled | bool
|
||||||
|
|
||||||
|
- name: Require restricted rrsync helper on Prometheus
|
||||||
|
tags: [services, backup, prometheus_backup]
|
||||||
|
ansible.builtin.stat:
|
||||||
|
path: "{{ server_backup_rrsync_path }}"
|
||||||
|
register: server_backup_rrsync_file
|
||||||
|
when: server_backup_export_enabled | bool
|
||||||
|
|
||||||
|
- name: Validate restricted rrsync helper
|
||||||
|
tags: [services, backup, prometheus_backup]
|
||||||
|
ansible.builtin.assert:
|
||||||
|
that:
|
||||||
|
- server_backup_rrsync_file.stat.exists
|
||||||
|
- server_backup_rrsync_file.stat.isreg
|
||||||
|
- server_backup_rrsync_file.stat.pw_name == 'root'
|
||||||
|
fail_msg: Rocky rsync must provide the root-owned rrsync support script.
|
||||||
|
when: server_backup_export_enabled | bool
|
||||||
|
|
||||||
|
- name: Create prepared backup export root
|
||||||
|
tags: [services, backup, prometheus_backup]
|
||||||
|
ansible.builtin.file:
|
||||||
|
path: "{{ server_backup_export_root }}"
|
||||||
|
state: directory
|
||||||
|
owner: root
|
||||||
|
group: "{{ server_backup_username }}"
|
||||||
|
mode: "0750"
|
||||||
|
when: server_backup_export_enabled | bool
|
||||||
|
|
||||||
|
- name: Create restricted Prometheus backup SSH directories
|
||||||
|
tags: [services, backup, prometheus_backup]
|
||||||
|
ansible.builtin.file:
|
||||||
|
path: "{{ item }}"
|
||||||
|
state: directory
|
||||||
|
owner: root
|
||||||
|
group: "{{ server_backup_username }}"
|
||||||
|
mode: "0750"
|
||||||
|
loop:
|
||||||
|
- "{{ server_backup_export_root }}/.ssh"
|
||||||
|
- "{{ server_backup_export_root }}/.ssh/authorized_keys.d"
|
||||||
|
when: server_backup_export_enabled | bool
|
||||||
|
|
||||||
|
- name: Read Atlas public key for Prometheus backup pull
|
||||||
|
tags: [services, backup, prometheus_backup]
|
||||||
|
ansible.builtin.slurp:
|
||||||
|
src: "{{ hostvars['atlas'].atlas_prometheus_pull_private_key_path | default('/etc/atlas-prometheus-pull/id_ed25519') }}.pub"
|
||||||
|
delegate_to: atlas
|
||||||
|
become: true
|
||||||
|
register: server_backup_atlas_public_key
|
||||||
|
when:
|
||||||
|
- server_backup_export_enabled | bool
|
||||||
|
- not ansible_check_mode
|
||||||
|
|
||||||
|
- name: Authorize only restricted read-only backup access from Atlas
|
||||||
|
tags: [services, backup, prometheus_backup]
|
||||||
|
ansible.builtin.copy:
|
||||||
|
content: >-
|
||||||
|
{{ 'command="/usr/bin/python3 ' ~ server_backup_rrsync_path ~ ' -ro '
|
||||||
|
~ server_backup_export_root ~ '/versions",restrict '
|
||||||
|
~ (server_backup_atlas_public_key.content | b64decode | trim) ~ '\n' }}
|
||||||
|
dest: "{{ server_backup_export_root }}/.ssh/authorized_keys.d/{{ server_backup_public_key_name }}"
|
||||||
|
owner: root
|
||||||
|
group: "{{ server_backup_username }}"
|
||||||
|
mode: "0640"
|
||||||
|
when:
|
||||||
|
- server_backup_export_enabled | bool
|
||||||
|
- not ansible_check_mode
|
||||||
90
ansible/roles/profile_server/tasks/backup_export_job.yml
Normal file
90
ansible/roles/profile_server/tasks/backup_export_job.yml
Normal file
@@ -0,0 +1,90 @@
|
|||||||
|
---
|
||||||
|
- name: Validate Prometheus backup export job inputs
|
||||||
|
tags: [services, backup, prometheus_backup]
|
||||||
|
ansible.builtin.assert:
|
||||||
|
that:
|
||||||
|
- server_backup_export_source_keep | int >= 2
|
||||||
|
- server_backup_export_paths | length > 0
|
||||||
|
- server_backup_export_paths | unique | length == server_backup_export_paths | length
|
||||||
|
- >-
|
||||||
|
server_backup_export_paths
|
||||||
|
| select('match', '^[a-zA-Z0-9][a-zA-Z0-9._/-]*$') | list | length
|
||||||
|
== server_backup_export_paths | length
|
||||||
|
- >-
|
||||||
|
server_backup_export_paths
|
||||||
|
| reject('search', '(^|/)\.\.(/|$)') | list | length
|
||||||
|
== server_backup_export_paths | length
|
||||||
|
- >-
|
||||||
|
server_backup_export_excludes
|
||||||
|
| select('match', '^[a-zA-Z0-9][a-zA-Z0-9._/-]*$') | list | length
|
||||||
|
== server_backup_export_excludes | length
|
||||||
|
- >-
|
||||||
|
server_backup_export_excludes
|
||||||
|
| reject('search', '(^|/)\.\.(/|$)') | list | length
|
||||||
|
== server_backup_export_excludes | length
|
||||||
|
fail_msg: Define safe relative paths and at least two prepared export versions.
|
||||||
|
when: server_backup_export_enabled | bool
|
||||||
|
|
||||||
|
- name: Validate Prometheus backup export calendar
|
||||||
|
tags: [services, backup, prometheus_backup]
|
||||||
|
ansible.builtin.command:
|
||||||
|
argv: [systemd-analyze, calendar, "{{ server_backup_export_calendar }}"]
|
||||||
|
changed_when: false
|
||||||
|
check_mode: false
|
||||||
|
when: server_backup_export_enabled | bool
|
||||||
|
|
||||||
|
- name: Ensure prepared Prometheus backup versions directory exists
|
||||||
|
tags: [services, backup, prometheus_backup]
|
||||||
|
ansible.builtin.file:
|
||||||
|
path: "{{ server_backup_export_root }}/versions"
|
||||||
|
state: directory
|
||||||
|
owner: root
|
||||||
|
group: "{{ server_backup_username }}"
|
||||||
|
mode: "0750"
|
||||||
|
when: server_backup_export_enabled | bool
|
||||||
|
|
||||||
|
- name: Install Prometheus backup export helper
|
||||||
|
tags: [services, backup, prometheus_backup]
|
||||||
|
ansible.builtin.template:
|
||||||
|
src: prometheus-backup-export.sh.j2
|
||||||
|
dest: /usr/local/sbin/prometheus-backup-export
|
||||||
|
owner: root
|
||||||
|
group: root
|
||||||
|
mode: "0750"
|
||||||
|
when: server_backup_export_enabled | bool
|
||||||
|
|
||||||
|
- name: Install Prometheus backup export systemd units
|
||||||
|
tags: [services, backup, prometheus_backup]
|
||||||
|
ansible.builtin.template:
|
||||||
|
src: "{{ item }}.j2"
|
||||||
|
dest: "/etc/systemd/system/{{ item }}"
|
||||||
|
owner: root
|
||||||
|
group: root
|
||||||
|
mode: "0644"
|
||||||
|
loop:
|
||||||
|
- prometheus-backup-export.service
|
||||||
|
- prometheus-backup-export.timer
|
||||||
|
loop_control:
|
||||||
|
label: "{{ item }}"
|
||||||
|
register: server_backup_export_units
|
||||||
|
when: server_backup_export_enabled | bool
|
||||||
|
|
||||||
|
- name: Reload systemd after Prometheus backup export unit changes
|
||||||
|
tags: [services, backup, prometheus_backup]
|
||||||
|
ansible.builtin.systemd:
|
||||||
|
daemon_reload: true
|
||||||
|
when:
|
||||||
|
- server_backup_export_enabled | bool
|
||||||
|
- server_backup_export_units is changed
|
||||||
|
- not ansible_check_mode
|
||||||
|
|
||||||
|
- name: Enable Prometheus backup export timer only after explicit activation
|
||||||
|
tags: [services, backup, prometheus_backup]
|
||||||
|
ansible.builtin.systemd:
|
||||||
|
name: prometheus-backup-export.timer
|
||||||
|
enabled: true
|
||||||
|
state: started
|
||||||
|
when:
|
||||||
|
- server_backup_export_enabled | bool
|
||||||
|
- server_backup_export_start_timer | bool
|
||||||
|
- not ansible_check_mode
|
||||||
@@ -53,6 +53,12 @@
|
|||||||
tags: [services, podman]
|
tags: [services, podman]
|
||||||
ansible.builtin.include_tasks: podman-compose.yml
|
ansible.builtin.include_tasks: podman-compose.yml
|
||||||
|
|
||||||
|
- name: Import Prometheus backup export identity tasks
|
||||||
|
ansible.builtin.import_tasks: backup_export_identity.yml
|
||||||
|
|
||||||
|
- name: Import Prometheus backup export job tasks
|
||||||
|
ansible.builtin.import_tasks: backup_export_job.yml
|
||||||
|
|
||||||
- name: Ensure server SSH authorized key fragments directory exists
|
- name: Ensure server SSH authorized key fragments directory exists
|
||||||
tags: [services, ssh]
|
tags: [services, ssh]
|
||||||
ansible.builtin.file:
|
ansible.builtin.file:
|
||||||
@@ -77,13 +83,17 @@
|
|||||||
when: server_ssh_authorized_keys | length > 0
|
when: server_ssh_authorized_keys | length > 0
|
||||||
|
|
||||||
- name: Configure server SSH authorized key fragments
|
- name: Configure server SSH authorized key fragments
|
||||||
tags: [services, ssh]
|
tags: [services, ssh, prometheus_backup]
|
||||||
ansible.builtin.lineinfile:
|
ansible.builtin.lineinfile:
|
||||||
path: /etc/ssh/sshd_config
|
path: /etc/ssh/sshd_config
|
||||||
regexp: '^\s*AuthorizedKeysFile\s+'
|
regexp: '^\s*AuthorizedKeysFile\s+'
|
||||||
line: >-
|
line: >-
|
||||||
AuthorizedKeysFile {{ server_ssh_authorized_keys | map(attribute='name')
|
AuthorizedKeysFile {{
|
||||||
| map('regex_replace', '^', '%h/.ssh/authorized_keys.d/') | join(' ') }}
|
((server_ssh_authorized_keys | map(attribute='name')
|
||||||
|
| map('regex_replace', '^', '%h/.ssh/authorized_keys.d/') | list)
|
||||||
|
+ (['%h/.ssh/authorized_keys.d/' ~ server_backup_public_key_name]
|
||||||
|
if server_backup_export_enabled | bool else [])) | join(' ')
|
||||||
|
}}
|
||||||
state: present
|
state: present
|
||||||
validate: "sshd -t -f %s"
|
validate: "sshd -t -f %s"
|
||||||
notify: Reload SSH service
|
notify: Reload SSH service
|
||||||
@@ -100,11 +110,13 @@
|
|||||||
notify: Reload SSH service
|
notify: Reload SSH service
|
||||||
|
|
||||||
- name: Restrict SSH login to allowed users on server
|
- name: Restrict SSH login to allowed users on server
|
||||||
tags: [services]
|
tags: [services, prometheus_backup]
|
||||||
ansible.builtin.lineinfile:
|
ansible.builtin.lineinfile:
|
||||||
path: /etc/ssh/sshd_config
|
path: /etc/ssh/sshd_config
|
||||||
regexp: '^\s*AllowUsers\s+'
|
regexp: '^\s*AllowUsers\s+'
|
||||||
line: "AllowUsers {{ server_sshd_allow_users | join(' ') }}"
|
line: >-
|
||||||
|
AllowUsers {{ (server_sshd_allow_users
|
||||||
|
+ ([server_backup_username] if server_backup_export_enabled | bool else [])) | join(' ') }}
|
||||||
state: present
|
state: present
|
||||||
validate: "sshd -t -f %s"
|
validate: "sshd -t -f %s"
|
||||||
notify: Reload SSH service
|
notify: Reload SSH service
|
||||||
|
|||||||
@@ -0,0 +1,15 @@
|
|||||||
|
[Unit]
|
||||||
|
Description=Prepare a read-only Prometheus application backup for Atlas
|
||||||
|
RequiresMountsFor=/opt/npm /opt/gitea {{ server_backup_export_root }}
|
||||||
|
ConditionFileIsExecutable=/usr/local/sbin/prometheus-backup-export
|
||||||
|
|
||||||
|
[Service]
|
||||||
|
Type=oneshot
|
||||||
|
ExecStart=/usr/local/sbin/prometheus-backup-export
|
||||||
|
User=root
|
||||||
|
Group=root
|
||||||
|
UMask=0077
|
||||||
|
TimeoutStartSec=infinity
|
||||||
|
Nice=10
|
||||||
|
IOSchedulingClass=best-effort
|
||||||
|
IOSchedulingPriority=7
|
||||||
@@ -0,0 +1,92 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
set -Eeuo pipefail
|
||||||
|
umask 077
|
||||||
|
|
||||||
|
export_root={{ server_backup_export_root | quote }}
|
||||||
|
versions="$export_root/versions"
|
||||||
|
stack_unit=podman-compose-server.service
|
||||||
|
stamp=$(date -u +%Y%m%dT%H%M%SZ)
|
||||||
|
stage=''
|
||||||
|
stack_stopped=false
|
||||||
|
|
||||||
|
exec 9>/run/lock/prometheus-backup-export.lock
|
||||||
|
flock -n 9 || { echo 'A backup export is already running' >&2; exit 1; }
|
||||||
|
|
||||||
|
cleanup() {
|
||||||
|
local rc=$?
|
||||||
|
trap - EXIT
|
||||||
|
if "$stack_stopped"; then
|
||||||
|
if systemctl is-active --quiet "$stack_unit"; then
|
||||||
|
systemctl restart "$stack_unit" || rc=1
|
||||||
|
else
|
||||||
|
systemctl start "$stack_unit" || rc=1
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
if (( rc != 0 )) && [[ -n "$stage" && -d "$stage" ]]; then
|
||||||
|
rm -rf -- "$stage"
|
||||||
|
fi
|
||||||
|
exit "$rc"
|
||||||
|
}
|
||||||
|
trap cleanup EXIT
|
||||||
|
trap 'exit 129' HUP
|
||||||
|
trap 'exit 130' INT
|
||||||
|
trap 'exit 143' TERM
|
||||||
|
|
||||||
|
systemctl is-active --quiet "$stack_unit" || {
|
||||||
|
echo 'The managed Compose stack must be active before preparing a backup' >&2
|
||||||
|
exit 1
|
||||||
|
}
|
||||||
|
|
||||||
|
paths=(
|
||||||
|
{% for path in server_backup_export_paths %}
|
||||||
|
{{ path | quote }}
|
||||||
|
{% endfor %}
|
||||||
|
)
|
||||||
|
excludes=(
|
||||||
|
{% for path in server_backup_export_excludes %}
|
||||||
|
--exclude={{ path | quote }}
|
||||||
|
{% endfor %}
|
||||||
|
)
|
||||||
|
for path in "${paths[@]}"; do
|
||||||
|
[[ -e "/$path" ]] || { echo "Required backup path missing: /$path" >&2; exit 1; }
|
||||||
|
done
|
||||||
|
[[ ! -e "$versions/$stamp" ]] || { echo "Export version already exists: $stamp" >&2; exit 1; }
|
||||||
|
stage=$(mktemp -d "$export_root/.staging.XXXXXXXX")
|
||||||
|
|
||||||
|
# SQLite databases and their accompanying files are copied while both
|
||||||
|
# managed containers are stopped. The EXIT trap restarts the stack on error.
|
||||||
|
stack_stopped=true
|
||||||
|
systemctl stop "$stack_unit"
|
||||||
|
tar --acls --xattrs --selinux "${excludes[@]}" -C / -cf "$stage/payload.tar" "${paths[@]}"
|
||||||
|
systemctl start "$stack_unit"
|
||||||
|
for container in nginx-proxy-manager gitea; do
|
||||||
|
running=false
|
||||||
|
for _ in {1..30}; do
|
||||||
|
if [[ $(podman inspect --format '{{ '{{.State.Running}}' }}' "$container" 2>/dev/null) == true ]]; then
|
||||||
|
running=true
|
||||||
|
break
|
||||||
|
fi
|
||||||
|
sleep 2
|
||||||
|
done
|
||||||
|
"$running" || { echo "Container did not restart: $container" >&2; exit 1; }
|
||||||
|
done
|
||||||
|
stack_stopped=false
|
||||||
|
|
||||||
|
tar -tf "$stage/payload.tar" >/dev/null
|
||||||
|
(cd "$stage" && sha256sum payload.tar >payload.sha256)
|
||||||
|
printf '{"schema":1,"host":"prometheus","created_utc":"%s"}\n' "$stamp" >"$stage/metadata.json"
|
||||||
|
chown root:{{ server_backup_username }} "$stage" "$stage/payload.tar" "$stage/payload.sha256" "$stage/metadata.json"
|
||||||
|
chmod 0750 "$stage"
|
||||||
|
chmod 0640 "$stage/payload.tar" "$stage/payload.sha256" "$stage/metadata.json"
|
||||||
|
mv -- "$stage" "$versions/$stamp"
|
||||||
|
stage=''
|
||||||
|
ln -s "$stamp" "$versions/.current.new"
|
||||||
|
mv -Tf -- "$versions/.current.new" "$versions/current"
|
||||||
|
|
||||||
|
# Keep a small source-side safety window; Atlas owns long-term retention.
|
||||||
|
mapfile -t old_versions < <(find "$versions" -mindepth 1 -maxdepth 1 -type d \
|
||||||
|
-printf '%f\n' | grep -E '^[0-9]{8}T[0-9]{6}Z$' | sort -r | tail -n +{{ server_backup_export_source_keep + 1 }})
|
||||||
|
for old in "${old_versions[@]}"; do
|
||||||
|
rm -rf -- "${versions:?}/$old"
|
||||||
|
done
|
||||||
|
echo "Prepared Prometheus backup export $stamp"
|
||||||
@@ -0,0 +1,10 @@
|
|||||||
|
[Unit]
|
||||||
|
Description=Prepare daily Prometheus application backup for Atlas
|
||||||
|
|
||||||
|
[Timer]
|
||||||
|
OnCalendar={{ server_backup_export_calendar }}
|
||||||
|
Persistent=false
|
||||||
|
Unit=prometheus-backup-export.service
|
||||||
|
|
||||||
|
[Install]
|
||||||
|
WantedBy=timers.target
|
||||||
74
docs/atlas-dr-lab.md
Normal file
74
docs/atlas-dr-lab.md
Normal file
@@ -0,0 +1,74 @@
|
|||||||
|
# Isolated Atlas DR lab
|
||||||
|
|
||||||
|
This is a **scaled rehearsal**, not a substitute for a full-data restore. The
|
||||||
|
`atlas-dr-lab` libvirt VM on Ikaros was left **shut off** on 2026-09-30. Its
|
||||||
|
persistent volumes are in the default libvirt pool: the current 30 GiB OS
|
||||||
|
volume `atlas-dr-lab-os-rebuild2.qcow2`, the pre-rebuild OS volume
|
||||||
|
`atlas-dr-lab-os.qcow2`, and four independent 4 GiB
|
||||||
|
`atlas-dr-lab-data{1,2,3,4}.qcow2` volumes. The VM uses libvirt's `default`
|
||||||
|
NAT network (last DHCP address `192.168.122.168`), 2 vCPU, and 4 GiB RAM.
|
||||||
|
The data disks have `virtio-atlasdrdata{1,2,3,4}` serials. No physical disk or
|
||||||
|
production Atlas storage is attached. The VM has no autostart.
|
||||||
|
|
||||||
|
## Rebuild inputs and isolation
|
||||||
|
|
||||||
|
- Use Rocky's **9.8 GenericCloud Base x86_64** image
|
||||||
|
`Rocky-9-GenericCloud-Base-9.8-20260525.0.x86_64.qcow2` from
|
||||||
|
`https://download.rockylinux.org/pub/rocky/9.8/images/x86_64/`.
|
||||||
|
Verify its `.CHECKSUM` file; the observed SHA-256 was
|
||||||
|
`92c206cc6f790c61583247eefe87890f8828420662c17cacf247cec78ab4eec8`.
|
||||||
|
- Use a dedicated lab-only inventory merged **after** the repository
|
||||||
|
inventory, and always `--limit atlas_dr_lab`. The temporary 2026-09-30
|
||||||
|
inventory/playbook and logs are in `/tmp/atlas-dr-lab-image/`; copy a
|
||||||
|
sanitized inventory to durable private storage before `/tmp` is cleared if
|
||||||
|
the lab will be repeated. Never reuse `host_vars/atlas.yml`, production
|
||||||
|
Vault secrets, or production disk by-id paths for the lab.
|
||||||
|
- The lab host belongs to `platform_rocky` and `atlas`. It uses `dradmin`
|
||||||
|
(UID/GID 1000) with the operator's **public** SSH key and a random,
|
||||||
|
unknown password hash, the libvirt DHCP address, pool `zpool`, mount root
|
||||||
|
`/zpool`, the four `virtio-atlasdrdata*` by-id paths, a 1 GiB backup
|
||||||
|
reservation, and `rocky_manage_openzfs_repo: true` with only `zfs` in
|
||||||
|
`host_packages`. The following gates remain false: sharing, firewall,
|
||||||
|
media stack, ZFS timers, Borg, USB, monitoring, and Prometheus pull.
|
||||||
|
`atlas_manage_storage` is true. Set `atlas_create_pool: true` **only for the
|
||||||
|
first disposable pool creation**, then set it false before any later run.
|
||||||
|
- A minimal lab playbook selects `atlas_dr_lab`, `become: true`, and the
|
||||||
|
existing `packages_rocky` and `profile_atlas` roles. Use a separate
|
||||||
|
`ANSIBLE_CONFIG` without the production Vault password script, and keep
|
||||||
|
host-key checking on with a lab-specific known-hosts file. The 2026-09-30
|
||||||
|
runs used `-i ansible/inventory/hosts.yml -i <lab-inventory.yml>` and
|
||||||
|
`--limit atlas_dr_lab` throughout.
|
||||||
|
|
||||||
|
## Rehearsal and narrow checks
|
||||||
|
|
||||||
|
1. Before any pool operation, compare `virsh -c qemu:///system domblklist
|
||||||
|
atlas-dr-lab` with the four intended qcow2 paths, and in the guest compare
|
||||||
|
`/dev/disk/by-id/virtio-atlasdrdata*` with `lsblk`. Do not proceed if a
|
||||||
|
physical disk or production identity appears.
|
||||||
|
2. For a first-time disposable build only, run the lab playbook with
|
||||||
|
`--tags pool` and `atlas_create_pool: true`, then immediately set the gate
|
||||||
|
false. Run the full lab playbook and check `zpool status -P zpool`,
|
||||||
|
`zfs list -r zpool`, SELinux, and failed systemd units.
|
||||||
|
3. Write a non-sensitive canary under the lab `/zpool/archive` and snapshot
|
||||||
|
it. Record the pool GUID and canary SHA-256. Export the lab pool cleanly,
|
||||||
|
shut down the VM, and replace **only the OS volume** with a fresh verified
|
||||||
|
Rocky image. Preserve all four data volumes. Reconfigure cloud-init for a
|
||||||
|
new instance; the seed CD-ROM must use **SATA**. The SCSI seed attachment
|
||||||
|
tried during this rehearsal was not detected by cloud-init and was
|
||||||
|
replaced with a SATA attachment before proceeding.
|
||||||
|
4. On the new OS, apply `packages_rocky` to reinstall OpenZFS. First run
|
||||||
|
`zpool import -d /dev/disk/by-id` **without importing**, compare GUID and
|
||||||
|
vdev membership, then use ordinary `zpool import -d /dev/disk/by-id zpool`.
|
||||||
|
Do not use `-f`, `-F`, `-X`, rollback, or pool creation.
|
||||||
|
5. Reapply `profile_atlas` with the lab gates and `atlas_create_pool: false`.
|
||||||
|
Verify the canary, restored snapshot file in an empty temporary directory,
|
||||||
|
dataset hierarchy, SELinux, and pool health. A second full playbook run
|
||||||
|
should report `changed=0`. Remove temporary restored files and shut down
|
||||||
|
the VM after testing.
|
||||||
|
|
||||||
|
The observed 2026-09-30 pool GUID was `8880368391795119587`; the canary
|
||||||
|
SHA-256 was `949701c7a95fadae1fddc21abe846c4312212dbfeb7477948f3188fc3ec34a78`.
|
||||||
|
The post-rebuild Ansible run succeeded, a repeat run reported `changed=0`,
|
||||||
|
12 datasets and the original snapshot were present, and the pool was healthy.
|
||||||
|
The snapshot-restored file matched content and basic metadata. See
|
||||||
|
[`atlas-recovery.md`](atlas-recovery.md) for the production runbook and limits.
|
||||||
141
docs/atlas-recovery.md
Normal file
141
docs/atlas-recovery.md
Normal file
@@ -0,0 +1,141 @@
|
|||||||
|
# Atlas recovery runbook
|
||||||
|
|
||||||
|
This runbook is for a **replacement Rocky Linux 9 installation**, not a normal
|
||||||
|
playbook run. A scaled whole-OS rebuild with a disposable pool passed in an
|
||||||
|
isolated VM on 2026-09-30, but no production-size whole-host recovery has been
|
||||||
|
tested. The existing production pool must be imported, never created or
|
||||||
|
rewritten. The provisional targets are **RPO 24 hours**
|
||||||
|
and **RTO 72 hours**, for Archive and Atlas services alike. They are planning
|
||||||
|
objectives, not demonstrated recovery times. The manual USB cadence may leave
|
||||||
|
an older copy; a recent Borg archive is needed to meet the RPO after total
|
||||||
|
pool loss.
|
||||||
|
|
||||||
|
## Before an incident
|
||||||
|
|
||||||
|
- Keep an offline copy of the encrypted Ansible Vault, its unlock material,
|
||||||
|
the exported Borg repository key, and the Borg passphrase. Do not store
|
||||||
|
unlock material in this repository or in a recovery command line.
|
||||||
|
On 2026-09-30 the operator confirmed these are available independently of
|
||||||
|
Atlas and the Ansible controller; their usability has not been tested here.
|
||||||
|
- Keep the Atlas installation media and a reproducible checkout of this
|
||||||
|
repository available independently of Atlas. Record the exact Git revision
|
||||||
|
used for a successful deployment.
|
||||||
|
- Record the pool's current disk identities with `zpool status -P zpool` and
|
||||||
|
`lsblk -o NAME,SIZE,MODEL,SERIAL,FSTYPE,UUID`. Compare these with
|
||||||
|
`atlas_zpool_disks` before touching a replacement host. The `host_vars`
|
||||||
|
values are historical identifiers, not evidence that a newly attached disk
|
||||||
|
is the same device.
|
||||||
|
- Verify that the latest hourly/daily snapshots, Borg archive, and offline USB
|
||||||
|
version exist and note their timestamps. A timer being enabled is not proof
|
||||||
|
that a backup completed.
|
||||||
|
|
||||||
|
## Incident gate
|
||||||
|
|
||||||
|
1. Identify whether the fault is the OS disk, one or more pool disks, accidental
|
||||||
|
deletion, or an unavailable host. Preserve failed media when possible.
|
||||||
|
2. Stop writes to affected services and capture the last known good backup
|
||||||
|
timestamps. Do not run `zpool create`, `zpool destroy`, `zfs rollback`,
|
||||||
|
`zpool import -F`, `zpool import -X`, `zpool import -f`, or disk formatting
|
||||||
|
as a diagnostic shortcut.
|
||||||
|
3. Choose one recovery source below. Do not merge several sources into the
|
||||||
|
production namespace without comparing their timestamps and content.
|
||||||
|
|
||||||
|
## Rebuild the OS and import the existing pool
|
||||||
|
|
||||||
|
1. Install Rocky Linux 9 on a **separate system disk**. Configure basic network,
|
||||||
|
SSH, a temporary sudo administrator, SELinux enforcing, and the current
|
||||||
|
OpenZFS kmod repository. Keep the pool drives untouched.
|
||||||
|
2. Run read-only identification: `lsblk -f`, `zpool import`, and
|
||||||
|
`zpool import -d /dev/disk/by-id`. Check the pool GUID, vdev layout, and
|
||||||
|
stable drive identities against the incident record. If any differ, stop.
|
||||||
|
3. Import only after matching the expected pool and host ownership. A pool
|
||||||
|
cleanly exported from the old host can be imported with
|
||||||
|
`zpool import -d /dev/disk/by-id zpool`. If it reports that the pool is
|
||||||
|
active elsewhere or needs a rewind/force, stop and investigate rather than
|
||||||
|
adding flags. Verify with `zpool status -v zpool`, `zfs list -r zpool`,
|
||||||
|
`zfs get -r mountpoint,canmount zpool`, and `findmnt -R /zpool`.
|
||||||
|
4. Leave `atlas_create_pool: false`. Ensure `host_vars/atlas.yml` reflects the
|
||||||
|
replacement host's actual SSH address and disk identities before running
|
||||||
|
Ansible. Apply `ansible/site.yml --limit atlas` with the bootstrap admin
|
||||||
|
connection override as documented in the Atlas setup section of README.
|
||||||
|
This may start shares/services, so keep clients disconnected or services
|
||||||
|
gated until data and permissions are verified.
|
||||||
|
5. Check `getenforce`, `zpool status -v zpool`, `systemctl --failed`, SSH,
|
||||||
|
firewalld, Cockpit, NFS, SMB, and the backup/monitoring timers. Do not
|
||||||
|
report recovery complete on the basis of Ansible success alone.
|
||||||
|
|
||||||
|
## Choose the data source
|
||||||
|
|
||||||
|
- **Local snapshot, pool intact:** inspect `zfs list -t snapshot -r zpool`.
|
||||||
|
Mount/access the chosen snapshot read-only and copy selected files to an
|
||||||
|
empty staging directory; compare content, owner, mode, mtime, and POSIX ACL.
|
||||||
|
Move into the live namespace only after an operator-approved scope review.
|
||||||
|
Do not use an automatic rollback: it can discard newer changes in the
|
||||||
|
dataset and descendants.
|
||||||
|
- **Offline USB:** verify the configured LUKS and ext4 UUIDs from
|
||||||
|
`host_vars/atlas.yml` before unlocking. Mount ext4 read-only with `ro,noload`,
|
||||||
|
use only a published `atlas/latest` version, and restore to an empty staging
|
||||||
|
directory. Compare checksums and metadata. The USB copy intentionally omits
|
||||||
|
generic xattrs and SELinux labels; relabel only the restored destination.
|
||||||
|
Never run the backup service to perform a restore.
|
||||||
|
- **Hetzner Borg:** use the dedicated pinned host key, repository path,
|
||||||
|
offline exported recovery key, and Vault-backed passphrase. List archives
|
||||||
|
and extract a selected archive into an empty staging directory, never the
|
||||||
|
live `/zpool` tree. A repository check and sample restore were previously
|
||||||
|
performed; that does not prove this incident's archive is complete. Compare
|
||||||
|
content and metadata before publication. Avoid `borg break-lock` while any
|
||||||
|
backup/check job may still be active.
|
||||||
|
|
||||||
|
After publishing restored files, run the explicit Ansible `restorecon` tag only
|
||||||
|
for the paths actually restored, for example:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
ansible-playbook ansible/site.yml --limit atlas --tags restorecon \
|
||||||
|
-e '{"atlas_restorecon_paths":["/zpool/archive"]}'
|
||||||
|
```
|
||||||
|
|
||||||
|
Then check ownership/ACLs, application-specific integrity, SMB/NFS client
|
||||||
|
access, backup service health, and `zpool status -v zpool`. Reconnect clients
|
||||||
|
only after these checks pass. Record the last recoverable timestamp (actual
|
||||||
|
RPO) and elapsed service outage (actual RTO) in the incident log.
|
||||||
|
|
||||||
|
## Scaled isolated rehearsal (2026-09-30)
|
||||||
|
|
||||||
|
The lab setup, repeatable checks, and preserved VM state are recorded in
|
||||||
|
[`atlas-dr-lab.md`](atlas-dr-lab.md).
|
||||||
|
|
||||||
|
On Ikaros, a local libvirt `atlas-dr-lab` VM used a 30 GiB Rocky 9.8 system
|
||||||
|
disk and four separate, disposable 4 GiB virtio data disks with stable
|
||||||
|
`/dev/disk/by-id` identities. The official Rocky cloud image matched its
|
||||||
|
published SHA-256. The lab inventory was separate from production, used a
|
||||||
|
fresh lab-only password hash and the operator's public SSH key, and disabled
|
||||||
|
sharing, Borg, USB backup, monitoring, media services, and the Prometheus pull.
|
||||||
|
No production disk, Vault secret, or production data was attached or copied.
|
||||||
|
|
||||||
|
1. The existing `packages_rocky` and `profile_atlas` roles installed OpenZFS,
|
||||||
|
created a RAIDZ2 `zpool` through the explicit one-time pool gate, and built
|
||||||
|
all 12 declared datasets with a lab-sized 1 GiB backup reservation. The
|
||||||
|
pool creation gate was set false immediately afterward.
|
||||||
|
2. A 4 MiB canary file was written under the lab `archive` dataset and a ZFS
|
||||||
|
snapshot created. The pool was cleanly exported and the VM shut down.
|
||||||
|
3. Only the system-disk volume was replaced by a fresh Rocky cloud image;
|
||||||
|
the four virtio data volumes were retained. Ansible reinstalled OpenZFS.
|
||||||
|
Read-only `zpool import -d /dev/disk/by-id` showed the expected RAIDZ2
|
||||||
|
topology and pool GUID `8880368391795119587` before an ordinary import
|
||||||
|
without `-f`, rewind, or rollback.
|
||||||
|
4. The imported pool was healthy. The canary SHA-256 matched its pre-rebuild
|
||||||
|
value. `profile_atlas` completed against the imported pool and a second
|
||||||
|
run reported `changed=0`. A file restored from the preserved snapshot into
|
||||||
|
`/var/tmp` matched SHA-256, owner, group, mode, size, and mtime; the temporary
|
||||||
|
copy was removed. Final checks found SELinux Enforcing, 12 datasets, the
|
||||||
|
snapshot, no failed units, and a healthy pool. The VM was shut down while
|
||||||
|
retaining its disposable volumes for a future rehearsal.
|
||||||
|
|
||||||
|
This proves the **sequence** for a cleanly exported, small pool and the tested
|
||||||
|
Ansible subset, not recovery duration or capacity at 2 TB. The earlier
|
||||||
|
2026-09-25 independent production ZFS/USB file restores and the earlier Borg
|
||||||
|
temporary-directory restore remain separate evidence. The VM did not restore
|
||||||
|
production USB/Borg archives, exercise services with production data, test an
|
||||||
|
unclean import, or prove the provisional RPO/RTO. Before relying on 24h/72h,
|
||||||
|
measure a representative full restore and service cutover in a suitably sized
|
||||||
|
future change window. Never use the production Atlas pool for a rehearsal.
|
||||||
20
docs/atlas-sharing-decision.md
Normal file
20
docs/atlas-sharing-decision.md
Normal file
@@ -0,0 +1,20 @@
|
|||||||
|
# Atlas SMB/NFS namespace decision
|
||||||
|
|
||||||
|
Decision date: 2026-09-30. Keep the current namespaces **separate**.
|
||||||
|
|
||||||
|
- `/zpool/archive` is the SMB3 `Archive` share for authorized Samba accounts.
|
||||||
|
- `/zpool/media/photobook` is the Aegis-only NFSv4 export, `all_squash`-mapped
|
||||||
|
to UID/GID `1100`.
|
||||||
|
- No new dual-protocol namespace, broad export, group, or ACL model is needed.
|
||||||
|
Existing permissions and client access remain unchanged.
|
||||||
|
|
||||||
|
The two paths serve different ownership and exposure needs. A common namespace
|
||||||
|
would expand the permissions design and require same-file SMB/NFS interoperability
|
||||||
|
testing without a present requirement. Revisit only when a specific workflow
|
||||||
|
needs both protocols on the same files; then decide UID/GID, group, POSIX ACL,
|
||||||
|
SELinux policy and client behavior before changing exports or permissions.
|
||||||
|
|
||||||
|
Read-only Atlas verification on 2026-09-30 confirmed that Samba `Archive` points
|
||||||
|
to `/zpool/archive`, NFS exports `/zpool/media/photobook` only to
|
||||||
|
`192.168.178.54` with `all_squash` and anonymous UID/GID `1100`, both datasets
|
||||||
|
are distinct, and `zpool` is healthy. No sharing configuration was changed.
|
||||||
63
docs/atlas-updates.md
Normal file
63
docs/atlas-updates.md
Normal file
@@ -0,0 +1,63 @@
|
|||||||
|
# Atlas Rocky/OpenZFS update and reboot procedure (draft)
|
||||||
|
|
||||||
|
This is an operator-controlled maintenance procedure. The playbook does not
|
||||||
|
reboot Atlas, replace a pool device, or perform a pool feature upgrade.
|
||||||
|
|
||||||
|
## Preflight
|
||||||
|
|
||||||
|
1. Schedule an outage and confirm no Borg, USB, snapshot, scrub, or resilver
|
||||||
|
job is active. A service in `activating` is still active; do not interrupt it.
|
||||||
|
2. Check `zpool status -v zpool` (including scrub status), `zfs list -r zpool`,
|
||||||
|
`systemctl --failed`, and `systemctl list-timers --all`. Resolve pool errors
|
||||||
|
first. Record current `uname -r`, `modinfo zfs | grep '^version:'`,
|
||||||
|
`rpm -q kernel-core kmod-zfs zfs`, and the current boot entry.
|
||||||
|
3. Confirm a recent successful Borg archive and a usable snapshot. Confirm
|
||||||
|
the latest published offline USB version and its physical availability;
|
||||||
|
do not start a USB backup merely to satisfy a checklist without capacity,
|
||||||
|
UUID, and operator checks. Record timestamps, not just timer state.
|
||||||
|
4. Ensure console/KVM or another independent recovery route is available.
|
||||||
|
Check free space in `/boot` and the root filesystem. Review proposed DNF
|
||||||
|
transactions before consenting to package changes.
|
||||||
|
|
||||||
|
## Change window
|
||||||
|
|
||||||
|
1. Stop client writes and quiesce stateful applications deliberately. Record
|
||||||
|
which services were stopped; do not assume `ansible-playbook --check` does
|
||||||
|
this. Avoid updating during a running scrub or backup.
|
||||||
|
2. Use `dnf upgrade --assumeno` first to review the kernel, `kmod-zfs`, `zfs`,
|
||||||
|
and dependencies. Confirm a matching kmod will be available for the target
|
||||||
|
kernel. If compatibility is uncertain, defer the update.
|
||||||
|
3. Apply the approved DNF transaction. Do not run `zpool upgrade` or enable
|
||||||
|
new pool feature flags as part of ordinary OS maintenance; that can remove
|
||||||
|
downgrade options. Preserve at least one known-good boot entry.
|
||||||
|
4. Reboot **manually** during the agreed outage. Ansible must not trigger it.
|
||||||
|
|
||||||
|
## Post-boot gate
|
||||||
|
|
||||||
|
1. Verify `uname -r`, `modinfo zfs`, `rpm -q kernel-core kmod-zfs zfs`,
|
||||||
|
`zpool status -v zpool`, `zfs list -r zpool`, and `findmnt -R /zpool`.
|
||||||
|
2. Verify SELinux remains enforcing; inspect `systemctl --failed` and the
|
||||||
|
journal for ZFS, mount, SSH, NFS, SMB, Cockpit, Podman, and backup errors.
|
||||||
|
3. Validate a read-only file listing through SMB and an NFS client access
|
||||||
|
check before reopening writes. Check the rootless temporary services and
|
||||||
|
all backup/monitoring timers. Run the Atlas health monitor in `--dry-run`
|
||||||
|
mode, then a real check after inspection.
|
||||||
|
4. Re-enable clients and record versions, downtime, anomalies, and next
|
||||||
|
successful snapshot/Borg run. A green boot alone is not a completed update.
|
||||||
|
|
||||||
|
## Failure response
|
||||||
|
|
||||||
|
If the new kernel cannot load ZFS, boot the previous known-good kernel from
|
||||||
|
the console and inspect package/kmod matching before trying another reboot.
|
||||||
|
Do not force-import, rewind, clear errors, or upgrade pool features to make a
|
||||||
|
failed OS update appear successful. Preserve logs and stop for a recovery
|
||||||
|
decision if the pool does not import cleanly.
|
||||||
|
|
||||||
|
The procedure-definition item is complete, but the procedure is **not yet
|
||||||
|
rehearsed** on a replacement host or during a real Atlas update. Record the
|
||||||
|
first controlled execution and its post-boot evidence separately.
|
||||||
|
|
||||||
|
Read-only preflight on 2026-09-30 observed kernel
|
||||||
|
`5.14.0-687.52.1.el9_8.x86_64`, ZFS module/package `2.2.11-1`, a healthy
|
||||||
|
`zpool`, enforcing SELinux, and no failed systemd units. This did not review
|
||||||
|
an upgrade transaction, stop services, or reboot the host.
|
||||||
106
docs/prometheus-backup.md
Normal file
106
docs/prometheus-backup.md
Normal file
@@ -0,0 +1,106 @@
|
|||||||
|
# Prometheus to Atlas backup pull
|
||||||
|
|
||||||
|
The playbook and both hosts have the dedicated identity, restricted SSH
|
||||||
|
access, helpers, and systemd units. A manual export, pull, and temporary
|
||||||
|
restore passed on 2026-09-30. Both timers are enabled; their first scheduled
|
||||||
|
runs are pending, so daily operation is not yet verified.
|
||||||
|
|
||||||
|
## Declared design
|
||||||
|
|
||||||
|
- Prometheus prepares a tar archive of Nginx Proxy Manager and Gitea data,
|
||||||
|
their managed Compose configuration, SSH/firewalld/WireGuard configuration,
|
||||||
|
and the Gitea SSH path. NPM access logs and regenerable Gitea logs, sessions,
|
||||||
|
temporary files, and indexers are excluded. The archive contains credentials,
|
||||||
|
certificates, and the WireGuard private key: protect both copies accordingly.
|
||||||
|
- The approved consistency mode stops the managed Compose stack for the local
|
||||||
|
tar creation at 02:00 Europe/Rome, then restarts it even if archiving fails.
|
||||||
|
A manual test outside that window requires separate approval.
|
||||||
|
- Prometheus publishes the archive with its checksum as a versioned, read-only
|
||||||
|
source under `/var/lib/prometheus-backup-export`. A locked service account
|
||||||
|
has no sudo or supplementary groups. Its only authorized SSH key is forced
|
||||||
|
through Rocky's `rrsync -ro`; root owns the key file and export directories,
|
||||||
|
so the account cannot add an unrestricted key or change prepared data.
|
||||||
|
- Atlas generates and retains the private Ed25519 identity under
|
||||||
|
`/etc/atlas-prometheus-pull`. Its pinned Prometheus host key came through
|
||||||
|
the controller's already strict SSH trust; the observed fingerprint was
|
||||||
|
`SHA256:rfedk7DHI9mLB3UHk/4F3HHlSIiswtCAFsAXvfh6iXk` on 2026-09-30.
|
||||||
|
Atlas pulls only the prepared `current/` version, verifies SHA-256, tar
|
||||||
|
readability, metadata, and source freshness, then publishes atomically
|
||||||
|
below `/zpool/backup/hosts/prometheus/snapshots`. Long-term retention runs
|
||||||
|
only after publication. A local `rrsync` fixture verified the in-tree
|
||||||
|
`current` symlink. A live Atlas-to-Prometheus SSH test verified that the
|
||||||
|
account could list only the prepared versions directory,
|
||||||
|
cannot obtain a shell, and cannot write to the export. The key is restricted
|
||||||
|
to `/var/lib/prometheus-backup-export/versions`, not the account's `.ssh`.
|
||||||
|
- Approved source preparation is 02:00 Europe/Rome, pull 03:00, three source
|
||||||
|
versions, and 30 daily/8 weekly/12 monthly Atlas versions. The source
|
||||||
|
timer is non-persistent to avoid an unexpected outage after a missed run.
|
||||||
|
Atlas rejects a prepared source older than 24 hours.
|
||||||
|
- The Atlas pull joins the existing health monitor's timer/failure checks
|
||||||
|
only when enabled. Its failure hook uses 45Drives Alerts; email delivery
|
||||||
|
is not claimed. A failed source preparation should produce a stale-source
|
||||||
|
pull failure, not a silently successful reuse of an old archive.
|
||||||
|
|
||||||
|
## Activation and verification
|
||||||
|
|
||||||
|
1. The user confirmed downtime/consistency mode, schedule, retention, and
|
||||||
|
targeted configuration scope. Review the tar path list and exclusions
|
||||||
|
against the actual containers.
|
||||||
|
2. The identity and units are deployed. Re-run the targeted
|
||||||
|
check, confirm the Atlas public key remains only the restricted Prometheus
|
||||||
|
account's key, and verify `sshd -T -C user=prometheus-backup,...` plus
|
||||||
|
read-only SSH denial tests after any SSH configuration change.
|
||||||
|
3. During an agreed window, start the Prometheus export service manually.
|
||||||
|
Confirm Compose is healthy afterward, inspect the archive without exposing
|
||||||
|
file contents, and verify the checksum/metadata.
|
||||||
|
4. Start the Atlas pull service manually. Confirm the SSH host pin, source
|
||||||
|
freshness, checksum, tar listing, published `latest`, retention behavior,
|
||||||
|
clean temporary directories, and healthy pool.
|
||||||
|
5. Independently restore the selected archive to an empty staging directory
|
||||||
|
(never `/`) and compare the SQLite databases, Git repositories, NPM data,
|
||||||
|
Compose file, permissions, and representative files. Test application
|
||||||
|
startup only in an isolated environment or an approved restore window.
|
||||||
|
6. The two timers were enabled after the manual test. Verify their calendars
|
||||||
|
and the next actual run. A successful manual test is not proof of scheduled
|
||||||
|
operation.
|
||||||
|
|
||||||
|
Narrow static validation:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
|
||||||
|
ansible-playbook ansible/site.yml --syntax-check
|
||||||
|
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
|
||||||
|
ansible-playbook ansible/site.yml --limit prometheus,atlas \
|
||||||
|
--tags prometheus_backup --check --diff
|
||||||
|
```
|
||||||
|
|
||||||
|
Do not run the export service as part of a routine playbook deployment. The
|
||||||
|
service restart and any restore/cutover require separate operator decisions.
|
||||||
|
|
||||||
|
On 2026-09-30 the initial targeted `--check --diff` run ended `changed=0`
|
||||||
|
with gates false. After enabling **implementation only**, a targeted real run
|
||||||
|
installed the identities and units; both timers were confirmed `disabled` and
|
||||||
|
`inactive`, the Compose stack stayed active, and the new account was locked
|
||||||
|
with no supplementary groups. No application was stopped.
|
||||||
|
The rendered shell helpers passed `bash -n` and ShellCheck; the retention
|
||||||
|
helper passed an isolated 400-version fixture. These static/isolated checks
|
||||||
|
were followed by live SSH, export, pull, and temporary restore checks.
|
||||||
|
Read-only preflight on 2026-09-30 found the Compose service active, all
|
||||||
|
declared source paths present, both timers inactive, and no prepared versions.
|
||||||
|
The source filesystem had about 6.0 GB free. Of the 2.1 GB NPM data tree,
|
||||||
|
2.1 GB was excluded access logs, so the expected archive is much smaller than
|
||||||
|
the raw tree size; capacity still needs verification after actual exports.
|
||||||
|
The manual export produced a 285,777,920-byte tar (273 MiB allocated at the
|
||||||
|
source), and Prometheus retained about 5.8 GB free. NPM and Gitea restarted;
|
||||||
|
both containers were running and their local HTTP endpoints returned 200.
|
||||||
|
Atlas pulled the same version, verified SHA-256, published `latest`, and kept
|
||||||
|
the pool healthy. A full extract to `/var/tmp` yielded 4,747 files; both
|
||||||
|
SQLite databases passed `PRAGMA integrity_check`, and one restored Gitea Git
|
||||||
|
repository passed `git fsck`. The temporary restore directory was removed.
|
||||||
|
This did not test application startup on an isolated host.
|
||||||
|
|
||||||
|
After these checks, Ansible enabled the Prometheus 02:00 Europe/Rome export
|
||||||
|
timer and Atlas 03:00 Europe/Rome pull timer. The next scheduled occurrences
|
||||||
|
were displayed for 2026-10-01. Atlas' health monitor now includes the pull
|
||||||
|
timer. Check both actual service results after the first scheduled run before
|
||||||
|
claiming unattended operation.
|
||||||
Reference in New Issue
Block a user