Compare commits

..

2 Commits

Author SHA1 Message Date
Fabio Scotto di Santolo
702283b430 Document Atlas backup and retention validation 2026-09-30 09:13:55 +02:00
Fabio Scotto di Santolo
797087c66f Fix Atlas Borg runtime permissions 2026-09-30 08:56:01 +02:00
5 changed files with 52 additions and 28 deletions

View File

@@ -180,27 +180,32 @@ Prometheus--Aegis WireGuard gateway are operational. The gateway handshake, forw
and TCP reachability to Atlas were verified. Temporary Navidrome and Syncthing are available through their
manual NPM Proxy Hosts; Syncthing uses `/data/Org` backed by the SMB-shared Archive dataset. Aegis has also
validated NFSv4.2 read, write, delete, and `all_squash` mapping to UID/GID `1100` end-to-end. The ZFS
snapshot timers are active and the first recursive hourly snapshot completed successfully; the first
scheduled retention prune and monthly scrub remain runtime checks.
snapshot timers are active; a recursive hourly snapshot and scheduled retention prune completed
successfully. The first monthly scrub remains a runtime check.
### Priority 1 - Data protection
- [x] Deploy Ansible-managed recursive ZFS snapshots with 24 hourly, 30 daily, 8 weekly, and 12 monthly
generations, plus a monthly scrub on the first Sunday at 03:00. The timers and first hourly snapshot were
verified on Atlas. Still observe the first scheduled retention prune and scrub; Cockpit Scheduler is for
visibility or manual operations only, and snapshot rollback is never automated.
verified on Atlas. Cockpit Scheduler is for visibility or manual operations only, and snapshot
rollback is never automated.
- [ ] Verify the first monthly ZFS scrub from its actual service result. Scheduled retention pruning
was observed on 2026-09-30; timer activation alone does not establish a successful scrub.
- [x] Activate and validate the encrypted offsite Borg backup to the Hetzner Storage Box. Atlas uses the
dedicated SSH identity, pinned ED25519 host key, Vault-backed `repokey` encryption, and a locked
non-login `borg` account with no sudo or supplementary groups. The initial snapshot-consistent backup,
Borg repository check, and temporary-directory restore completed successfully; the restored `Archive`
tree matched the live data, and temporary snapshots and mounts were removed. The exported recovery key
was copied offline. Daily backup retries and logging, 30 daily, 8 weekly and 12 monthly archives,
compaction, and monthly repository checks are enabled. Future runs report a ZFS-based estimated
percentage, and a post-exit helper handles host-namespace temporary snapshot cleanup. The active
run predates the new progress logging and still requires an observed final cleanup result.
compaction, and monthly repository checks are enabled. Runs report a ZFS-based estimated percentage.
On 2026-09-30 a successful incremental run also removed the stale 2026-09-29 snapshot and its own
temporary snapshot after exit; the earlier `RuntimeDirectory` cleanup failure is resolved.
- [x] Populate `/zpool/archive` with the currently available data so offsite and offline backup tests run
against a representative load.
- [ ] Run and evaluate Borg against the populated pool: duration, repository capacity, deduplication, and
a subsequent incremental archive must be observed before considering the offsite path fully validated.
- [x] Evaluate Borg against the populated pool. The 2026-09-29 archive took 1 h 32 min for 2.18 TB
original / 2.04 TB compressed data, with 13.49 GB deduplicated size; retention and compaction
succeeded. On 2026-09-30 a subsequent incremental archive completed in about 22 seconds with
successful cleanup. The monitor reported 37% Storage Box quota used. These are observed runs, not
a guarantee of future duration or compression ratio.
- [x] Add the UUID-bound offline USB backup with versioned rsync, locking, capacity checks, verification,
safe unmounting and a tested restore procedure; never trigger it for an arbitrary USB disk. The
LUKS/ext4 identities were read-only verified; the manual service and 45Drives Alerts reminder timer were
@@ -228,7 +233,8 @@ scheduled retention prune and monthly scrub remain runtime checks.
found no issues; the service and timer succeeded, and a labelled 45Drives Alerts test notification
was submitted. Alerts are deduplicated; email delivery is not claimed. The Storage Box quota probe
runs `df -m` over the dedicated pinned-key SSH identity and does not open the Borg repository.
Detailed archive size and deduplication remain part of the pending Borg evaluation.
The failed-job hook was corrected to pass the literal systemd unit name; its expansion was verified
on Atlas, but a new real failure notification has not been deliberately triggered.
### Priority 2 - NAS operability and recovery
- [ ] Document and test disaster recovery: rebuild Atlas with Ansible, import the existing pool, restore

View File

@@ -337,8 +337,8 @@ Gli snapshot ZFS ricorsivi coprono l'intero pool: 24 orari al minuto 05, 30 gior
8 settimanali la domenica alle 01:00 e 12 mensili il primo giorno alle 02:00. La retention elimina
solo gli snapshot con prefisso gestito `atlas-auto` e non esegue rollback. Lo scrub OpenZFS mensile è
previsto la prima domenica alle 03:00; il timer settimanale incompatibile è disabilitato. Il primo
snapshot orario ricorsivo è riuscito; la prima pulizia pianificata e il primo scrub schedulato
richiedono ancora una verifica a runtime.
snapshot orario ricorsivo è riuscito e la pulizia pianificata della retention è stata osservata il
2026-09-30. Il primo scrub mensile richiede ancora una verifica a runtime.
```bash
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
@@ -391,6 +391,13 @@ copiato un file di `/zpool/archive` in `/var/tmp`, verificando contenuto, propri
mtime e ACL POSIX; copia e mount temporanei sono stati rimossi senza interrompere Borg. Non è un test
di ripristino dell'intero dataset.
L'archivio del pool popolato del 2026-09-29 ha richiesto 1 h 32 min per 2,18 TB originali / 2,04 TB
compressi, con 13,49 GB di dimensione deduplicata. Retention e compattazione sono riuscite, ma un
errore di permessi su `RuntimeDirectory` ha impedito la pulizia dello snapshot dopo il job. Dopo la
correzione, l'archivio incrementale del 2026-09-30 è terminato in circa 22 secondi, ha rimosso lo
snapshot residuo e quello corrente ed è terminato con stato 0. Il monitor ha rilevato il 37% della
quota Storage Box utilizzata. Questi risultati non predicono durata o compressione dei prossimi run.
Il backup USB offline è distribuito come **servizio solo manuale** (`atlas_manage_usb_backup: true`):
Ansible non formatta, sblocca, monta né avvia automaticamente il disco. Il disco esistente è stato
verificato in sola lettura il 2026-09-23: UUID LUKS `577b3c43-ea37-4611-81a9-39d555cdfbd4`,
@@ -459,9 +466,10 @@ baseline di circa 24 ore. Sono controllati anche attivazione e freschezza dei ti
`OnFailure` segnalano errori di snapshot, scrub, Borg, USB, promemoria e monitoraggio. Il monitor non
riavvia Borg; avvisa solo se un run supera 14 giorni. Soglie e percorsi stabili dei dischi sono nelle
variabili host. Gli avvisi usano 45Drives Houston con deduplicazione; **la consegna email non è stata
verificata**. Il controllo live del 2026-09-25 non ha trovato problemi; la notifica di prova è stata
inviata e lo Storage Box risultava occupato al 22%. Dimensione dell'archivio Borg e deduplicazione
dettagliata richiedono ancora la fine del backup in corso.
verificata**. Il controllo live del 2026-09-25 non ha trovato problemi e ha inviato una notifica di
prova. Il 2026-09-30 il monitor ha rilevato zero problemi e una quota Storage Box occupata al 37%.
L'hook per i job falliti ora passa il nome letterale della unità systemd; l'espansione è stata
verificata senza inviare un falso allarme.
```bash
ansible-playbook ansible/site.yml --limit atlas --tags monitoring --check --diff
@@ -511,8 +519,8 @@ cutover. L'attuale iCloudPD su Aegis e l'export NFS Photobook restano configurat
all'approvazione e alla verifica di questa migrazione separata. Anche il servizio Atlas sarà
temporaneo in attesa di Uranus.
Il pull dei backup di Prometheus, la valutazione delle dimensioni degli archivi Borg e i test completi
di disaster recovery restano da fare. Il backlog prioritizzato è in `AGENTS.md`.
Il pull dei backup di Prometheus e i test completi di disaster recovery restano da fare. Il backlog
prioritizzato è in `AGENTS.md`.
---

View File

@@ -325,8 +325,8 @@ snapshots at minute 05, 30 daily snapshots at 00:15, 8 weekly snapshots on Sunda
monthly snapshots on the first day at 02:00. The retention helper prunes only snapshots carrying its
managed `atlas-auto` prefix and never rolls back a dataset. The OpenZFS monthly scrub timer is scheduled
for the first Sunday at 03:00; the conflicting weekly scrub timer is disabled explicitly. The first recursive
hourly snapshot completed successfully on Atlas; retention pruning and the first scheduled scrub still await
live runtime evidence. Validate this layer independently with:
hourly snapshot completed successfully on Atlas, and scheduled retention pruning was observed on
2026-09-30. The first monthly scrub still awaits runtime evidence. Validate this layer independently with:
```bash
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
@@ -381,6 +381,12 @@ ansible-playbook ansible/site.yml --limit atlas --tags packages,borg --check --d
Atlas runtime activation is complete: the initial backup and repository check succeeded, a full restore
to a temporary directory was validated against the live `Archive` tree, the recovery-key export was copied
to offline storage, and the temporary snapshot and bind mounts were cleaned up.
The populated-pool archive on 2026-09-29 took 1 h 32 min for 2.18 TB original / 2.04 TB compressed
data, with a 13.49 GB deduplicated archive size. Retention and compaction succeeded, but a
`RuntimeDirectory` permission error prevented post-exit snapshot cleanup. After correction, the
2026-09-30 incremental archive completed in about 22 seconds, removed the stale and current temporary
snapshots, and ended with service status 0. The monitor reported 37% Storage Box quota used. These
observations do not predict the duration or compression ratio of future runs.
On 2026-09-25 a separate ZFS restore smoke test copied a small file from an automatic daily
`zpool/archive` snapshot to `/var/tmp`, then confirmed matching contents, ownership, mode, mtime and
POSIX ACL. The temporary copy and on-demand snapshot mount were removed; Borg kept running. This
@@ -472,12 +478,13 @@ Storage Box quota via `df -m` over the dedicated `borg` account's pinned-key SSH
query never opens the Borg repository or its lock. Growth alerts compare against a roughly 24-hour
baseline and therefore begin only after enough samples exist. The monitor also checks
maintenance/backup timer activation and freshness; systemd `OnFailure` hooks report snapshot,
scrub, Borg, USB, reminder, and monitoring services when they enter the failed state. The ongoing
initial Borg run is never restarted by the monitor; only a run exceeding 14 days raises a warning.
scrub, Borg, USB, reminder, and monitoring services when they enter the failed state. An ongoing
Borg run is never restarted by the monitor; only a run exceeding 14 days raises a warning.
Thresholds and stable disk paths are declared in Atlas host variables. Alerts use the existing 45Drives
Houston notifier and repeated issues are deduplicated; **email delivery is not verified**. The
2026-09-25 live probe found no issues and a labelled test notification was submitted. The Storage Box
reported 22% used. Detailed Borg archive size and deduplication still require the active run to finish.
2026-09-25 live probe found no issues and a labelled test notification was submitted. On 2026-09-30
the monitor reported zero issues and 37% Storage Box quota used. The failed-job hook now passes the
literal systemd unit name; its expansion was verified without sending a false failure notification.
```bash
ansible-playbook ansible/site.yml --limit atlas --tags monitoring --check --diff
@@ -527,8 +534,8 @@ state outside `Archive`, then test permissions, SELinux, backups and recovery be
The current Aegis iCloudPD service and Atlas Photobook NFS export remain configured until that
separate migration is approved and validated; the eventual Atlas service is temporary until Uranus.
Prometheus backup pulls, Borg archive-size evaluation, and full disaster-recovery tests remain follow-up
work. The prioritized operational backlog is kept in `AGENTS.md`.
Prometheus backup pulls and full disaster-recovery tests remain follow-up work. The prioritized
operational backlog is kept in `AGENTS.md`.
## How layering works

View File

@@ -87,7 +87,10 @@ exec 9>/run/lock/atlas-zfs-snapshot.lock
zpool list -H -o name "$pool" >/dev/null
rm -rf "$stage"
mkdir -p "$stage"
chown root:"$borg_group" /run/atlas-borg "$stage"
# Keep systemd's root:root ownership of RuntimeDirectory: changing it makes
# ExecStopPost re-chown its contents, which SELinux denies for the marker.
setfacl -m "u:${borg_user}:rx" /run/atlas-borg
chown root:"$borg_group" "$stage"
chmod 0750 /run/atlas-borg "$stage"
flock 9

View File

@@ -1,12 +1,12 @@
[Unit]
Description=Submit a 45Drives Alert for failed Atlas job %I
Description=Submit a 45Drives Alert for failed Atlas job %i
Requires=houston-dbus.service
After=houston-dbus.service
ConditionFileIsExecutable=/usr/local/libexec/atlas-health-monitor
[Service]
Type=oneshot
ExecStart=/usr/local/libexec/atlas-health-monitor --job-failed %I
ExecStart=/usr/local/libexec/atlas-health-monitor --job-failed %i
User=root
Group=root
UMask=0077