diff --git a/AGENTS.md b/AGENTS.md index 4328dda..4148393 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -180,27 +180,32 @@ Prometheus--Aegis WireGuard gateway are operational. The gateway handshake, forw and TCP reachability to Atlas were verified. Temporary Navidrome and Syncthing are available through their manual NPM Proxy Hosts; Syncthing uses `/data/Org` backed by the SMB-shared Archive dataset. Aegis has also validated NFSv4.2 read, write, delete, and `all_squash` mapping to UID/GID `1100` end-to-end. The ZFS -snapshot timers are active and the first recursive hourly snapshot completed successfully; the first -scheduled retention prune and monthly scrub remain runtime checks. +snapshot timers are active; a recursive hourly snapshot and scheduled retention prune completed +successfully. The first monthly scrub remains a runtime check. ### Priority 1 - Data protection - [x] Deploy Ansible-managed recursive ZFS snapshots with 24 hourly, 30 daily, 8 weekly, and 12 monthly generations, plus a monthly scrub on the first Sunday at 03:00. The timers and first hourly snapshot were - verified on Atlas. Still observe the first scheduled retention prune and scrub; Cockpit Scheduler is for - visibility or manual operations only, and snapshot rollback is never automated. + verified on Atlas. Cockpit Scheduler is for visibility or manual operations only, and snapshot + rollback is never automated. +- [ ] Verify the first monthly ZFS scrub from its actual service result. Scheduled retention pruning + was observed on 2026-09-30; timer activation alone does not establish a successful scrub. - [x] Activate and validate the encrypted offsite Borg backup to the Hetzner Storage Box. Atlas uses the dedicated SSH identity, pinned ED25519 host key, Vault-backed `repokey` encryption, and a locked non-login `borg` account with no sudo or supplementary groups. The initial snapshot-consistent backup, Borg repository check, and temporary-directory restore completed successfully; the restored `Archive` tree matched the live data, and temporary snapshots and mounts were removed. The exported recovery key was copied offline. Daily backup retries and logging, 30 daily, 8 weekly and 12 monthly archives, - compaction, and monthly repository checks are enabled. Future runs report a ZFS-based estimated - percentage, and a post-exit helper handles host-namespace temporary snapshot cleanup. The active - run predates the new progress logging and still requires an observed final cleanup result. + compaction, and monthly repository checks are enabled. Runs report a ZFS-based estimated percentage. + On 2026-09-30 a successful incremental run also removed the stale 2026-09-29 snapshot and its own + temporary snapshot after exit; the earlier `RuntimeDirectory` cleanup failure is resolved. - [x] Populate `/zpool/archive` with the currently available data so offsite and offline backup tests run against a representative load. -- [ ] Run and evaluate Borg against the populated pool: duration, repository capacity, deduplication, and - a subsequent incremental archive must be observed before considering the offsite path fully validated. +- [x] Evaluate Borg against the populated pool. The 2026-09-29 archive took 1 h 32 min for 2.18 TB + original / 2.04 TB compressed data, with 13.49 GB deduplicated size; retention and compaction + succeeded. On 2026-09-30 a subsequent incremental archive completed in about 22 seconds with + successful cleanup. The monitor reported 37% Storage Box quota used. These are observed runs, not + a guarantee of future duration or compression ratio. - [x] Add the UUID-bound offline USB backup with versioned rsync, locking, capacity checks, verification, safe unmounting and a tested restore procedure; never trigger it for an arbitrary USB disk. The LUKS/ext4 identities were read-only verified; the manual service and 45Drives Alerts reminder timer were @@ -228,7 +233,8 @@ scheduled retention prune and monthly scrub remain runtime checks. found no issues; the service and timer succeeded, and a labelled 45Drives Alerts test notification was submitted. Alerts are deduplicated; email delivery is not claimed. The Storage Box quota probe runs `df -m` over the dedicated pinned-key SSH identity and does not open the Borg repository. - Detailed archive size and deduplication remain part of the pending Borg evaluation. + The failed-job hook was corrected to pass the literal systemd unit name; its expansion was verified + on Atlas, but a new real failure notification has not been deliberately triggered. ### Priority 2 - NAS operability and recovery - [ ] Document and test disaster recovery: rebuild Atlas with Ansible, import the existing pool, restore diff --git a/README.it.md b/README.it.md index f98615e..7675600 100644 --- a/README.it.md +++ b/README.it.md @@ -337,8 +337,8 @@ Gli snapshot ZFS ricorsivi coprono l'intero pool: 24 orari al minuto 05, 30 gior 8 settimanali la domenica alle 01:00 e 12 mensili il primo giorno alle 02:00. La retention elimina solo gli snapshot con prefisso gestito `atlas-auto` e non esegue rollback. Lo scrub OpenZFS mensile è previsto la prima domenica alle 03:00; il timer settimanale incompatibile è disabilitato. Il primo -snapshot orario ricorsivo è riuscito; la prima pulizia pianificata e il primo scrub schedulato -richiedono ancora una verifica a runtime. +snapshot orario ricorsivo è riuscito e la pulizia pianificata della retention è stata osservata il +2026-09-30. Il primo scrub mensile richiede ancora una verifica a runtime. ```bash ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \ @@ -391,6 +391,13 @@ copiato un file di `/zpool/archive` in `/var/tmp`, verificando contenuto, propri mtime e ACL POSIX; copia e mount temporanei sono stati rimossi senza interrompere Borg. Non è un test di ripristino dell'intero dataset. +L'archivio del pool popolato del 2026-09-29 ha richiesto 1 h 32 min per 2,18 TB originali / 2,04 TB +compressi, con 13,49 GB di dimensione deduplicata. Retention e compattazione sono riuscite, ma un +errore di permessi su `RuntimeDirectory` ha impedito la pulizia dello snapshot dopo il job. Dopo la +correzione, l'archivio incrementale del 2026-09-30 è terminato in circa 22 secondi, ha rimosso lo +snapshot residuo e quello corrente ed è terminato con stato 0. Il monitor ha rilevato il 37% della +quota Storage Box utilizzata. Questi risultati non predicono durata o compressione dei prossimi run. + Il backup USB offline è distribuito come **servizio solo manuale** (`atlas_manage_usb_backup: true`): Ansible non formatta, sblocca, monta né avvia automaticamente il disco. Il disco esistente è stato verificato in sola lettura il 2026-09-23: UUID LUKS `577b3c43-ea37-4611-81a9-39d555cdfbd4`, @@ -459,9 +466,10 @@ baseline di circa 24 ore. Sono controllati anche attivazione e freschezza dei ti `OnFailure` segnalano errori di snapshot, scrub, Borg, USB, promemoria e monitoraggio. Il monitor non riavvia Borg; avvisa solo se un run supera 14 giorni. Soglie e percorsi stabili dei dischi sono nelle variabili host. Gli avvisi usano 45Drives Houston con deduplicazione; **la consegna email non è stata -verificata**. Il controllo live del 2026-09-25 non ha trovato problemi; la notifica di prova è stata -inviata e lo Storage Box risultava occupato al 22%. Dimensione dell'archivio Borg e deduplicazione -dettagliata richiedono ancora la fine del backup in corso. +verificata**. Il controllo live del 2026-09-25 non ha trovato problemi e ha inviato una notifica di +prova. Il 2026-09-30 il monitor ha rilevato zero problemi e una quota Storage Box occupata al 37%. +L'hook per i job falliti ora passa il nome letterale della unità systemd; l'espansione è stata +verificata senza inviare un falso allarme. ```bash ansible-playbook ansible/site.yml --limit atlas --tags monitoring --check --diff @@ -511,8 +519,8 @@ cutover. L'attuale iCloudPD su Aegis e l'export NFS Photobook restano configurat all'approvazione e alla verifica di questa migrazione separata. Anche il servizio Atlas sarà temporaneo in attesa di Uranus. -Il pull dei backup di Prometheus, la valutazione delle dimensioni degli archivi Borg e i test completi -di disaster recovery restano da fare. Il backlog prioritizzato è in `AGENTS.md`. +Il pull dei backup di Prometheus e i test completi di disaster recovery restano da fare. Il backlog +prioritizzato è in `AGENTS.md`. --- diff --git a/README.md b/README.md index 1d05151..a950a08 100644 --- a/README.md +++ b/README.md @@ -325,8 +325,8 @@ snapshots at minute 05, 30 daily snapshots at 00:15, 8 weekly snapshots on Sunda monthly snapshots on the first day at 02:00. The retention helper prunes only snapshots carrying its managed `atlas-auto` prefix and never rolls back a dataset. The OpenZFS monthly scrub timer is scheduled for the first Sunday at 03:00; the conflicting weekly scrub timer is disabled explicitly. The first recursive -hourly snapshot completed successfully on Atlas; retention pruning and the first scheduled scrub still await -live runtime evidence. Validate this layer independently with: +hourly snapshot completed successfully on Atlas, and scheduled retention pruning was observed on +2026-09-30. The first monthly scrub still awaits runtime evidence. Validate this layer independently with: ```bash ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \ @@ -381,6 +381,12 @@ ansible-playbook ansible/site.yml --limit atlas --tags packages,borg --check --d Atlas runtime activation is complete: the initial backup and repository check succeeded, a full restore to a temporary directory was validated against the live `Archive` tree, the recovery-key export was copied to offline storage, and the temporary snapshot and bind mounts were cleaned up. +The populated-pool archive on 2026-09-29 took 1 h 32 min for 2.18 TB original / 2.04 TB compressed +data, with a 13.49 GB deduplicated archive size. Retention and compaction succeeded, but a +`RuntimeDirectory` permission error prevented post-exit snapshot cleanup. After correction, the +2026-09-30 incremental archive completed in about 22 seconds, removed the stale and current temporary +snapshots, and ended with service status 0. The monitor reported 37% Storage Box quota used. These +observations do not predict the duration or compression ratio of future runs. On 2026-09-25 a separate ZFS restore smoke test copied a small file from an automatic daily `zpool/archive` snapshot to `/var/tmp`, then confirmed matching contents, ownership, mode, mtime and POSIX ACL. The temporary copy and on-demand snapshot mount were removed; Borg kept running. This @@ -472,12 +478,13 @@ Storage Box quota via `df -m` over the dedicated `borg` account's pinned-key SSH query never opens the Borg repository or its lock. Growth alerts compare against a roughly 24-hour baseline and therefore begin only after enough samples exist. The monitor also checks maintenance/backup timer activation and freshness; systemd `OnFailure` hooks report snapshot, -scrub, Borg, USB, reminder, and monitoring services when they enter the failed state. The ongoing -initial Borg run is never restarted by the monitor; only a run exceeding 14 days raises a warning. +scrub, Borg, USB, reminder, and monitoring services when they enter the failed state. An ongoing +Borg run is never restarted by the monitor; only a run exceeding 14 days raises a warning. Thresholds and stable disk paths are declared in Atlas host variables. Alerts use the existing 45Drives Houston notifier and repeated issues are deduplicated; **email delivery is not verified**. The -2026-09-25 live probe found no issues and a labelled test notification was submitted. The Storage Box -reported 22% used. Detailed Borg archive size and deduplication still require the active run to finish. +2026-09-25 live probe found no issues and a labelled test notification was submitted. On 2026-09-30 +the monitor reported zero issues and 37% Storage Box quota used. The failed-job hook now passes the +literal systemd unit name; its expansion was verified without sending a false failure notification. ```bash ansible-playbook ansible/site.yml --limit atlas --tags monitoring --check --diff @@ -527,8 +534,8 @@ state outside `Archive`, then test permissions, SELinux, backups and recovery be The current Aegis iCloudPD service and Atlas Photobook NFS export remain configured until that separate migration is approved and validated; the eventual Atlas service is temporary until Uranus. -Prometheus backup pulls, Borg archive-size evaluation, and full disaster-recovery tests remain follow-up -work. The prioritized operational backlog is kept in `AGENTS.md`. +Prometheus backup pulls and full disaster-recovery tests remain follow-up work. The prioritized +operational backlog is kept in `AGENTS.md`. ## How layering works