mirror of
https://github.com/fscotto/infra.git
synced 2026-10-03 21:39:50 +00:00
Document Atlas backup and retention validation
This commit is contained in:
26
AGENTS.md
26
AGENTS.md
@@ -180,27 +180,32 @@ Prometheus--Aegis WireGuard gateway are operational. The gateway handshake, forw
|
|||||||
and TCP reachability to Atlas were verified. Temporary Navidrome and Syncthing are available through their
|
and TCP reachability to Atlas were verified. Temporary Navidrome and Syncthing are available through their
|
||||||
manual NPM Proxy Hosts; Syncthing uses `/data/Org` backed by the SMB-shared Archive dataset. Aegis has also
|
manual NPM Proxy Hosts; Syncthing uses `/data/Org` backed by the SMB-shared Archive dataset. Aegis has also
|
||||||
validated NFSv4.2 read, write, delete, and `all_squash` mapping to UID/GID `1100` end-to-end. The ZFS
|
validated NFSv4.2 read, write, delete, and `all_squash` mapping to UID/GID `1100` end-to-end. The ZFS
|
||||||
snapshot timers are active and the first recursive hourly snapshot completed successfully; the first
|
snapshot timers are active; a recursive hourly snapshot and scheduled retention prune completed
|
||||||
scheduled retention prune and monthly scrub remain runtime checks.
|
successfully. The first monthly scrub remains a runtime check.
|
||||||
|
|
||||||
### Priority 1 - Data protection
|
### Priority 1 - Data protection
|
||||||
- [x] Deploy Ansible-managed recursive ZFS snapshots with 24 hourly, 30 daily, 8 weekly, and 12 monthly
|
- [x] Deploy Ansible-managed recursive ZFS snapshots with 24 hourly, 30 daily, 8 weekly, and 12 monthly
|
||||||
generations, plus a monthly scrub on the first Sunday at 03:00. The timers and first hourly snapshot were
|
generations, plus a monthly scrub on the first Sunday at 03:00. The timers and first hourly snapshot were
|
||||||
verified on Atlas. Still observe the first scheduled retention prune and scrub; Cockpit Scheduler is for
|
verified on Atlas. Cockpit Scheduler is for visibility or manual operations only, and snapshot
|
||||||
visibility or manual operations only, and snapshot rollback is never automated.
|
rollback is never automated.
|
||||||
|
- [ ] Verify the first monthly ZFS scrub from its actual service result. Scheduled retention pruning
|
||||||
|
was observed on 2026-09-30; timer activation alone does not establish a successful scrub.
|
||||||
- [x] Activate and validate the encrypted offsite Borg backup to the Hetzner Storage Box. Atlas uses the
|
- [x] Activate and validate the encrypted offsite Borg backup to the Hetzner Storage Box. Atlas uses the
|
||||||
dedicated SSH identity, pinned ED25519 host key, Vault-backed `repokey` encryption, and a locked
|
dedicated SSH identity, pinned ED25519 host key, Vault-backed `repokey` encryption, and a locked
|
||||||
non-login `borg` account with no sudo or supplementary groups. The initial snapshot-consistent backup,
|
non-login `borg` account with no sudo or supplementary groups. The initial snapshot-consistent backup,
|
||||||
Borg repository check, and temporary-directory restore completed successfully; the restored `Archive`
|
Borg repository check, and temporary-directory restore completed successfully; the restored `Archive`
|
||||||
tree matched the live data, and temporary snapshots and mounts were removed. The exported recovery key
|
tree matched the live data, and temporary snapshots and mounts were removed. The exported recovery key
|
||||||
was copied offline. Daily backup retries and logging, 30 daily, 8 weekly and 12 monthly archives,
|
was copied offline. Daily backup retries and logging, 30 daily, 8 weekly and 12 monthly archives,
|
||||||
compaction, and monthly repository checks are enabled. Future runs report a ZFS-based estimated
|
compaction, and monthly repository checks are enabled. Runs report a ZFS-based estimated percentage.
|
||||||
percentage, and a post-exit helper handles host-namespace temporary snapshot cleanup. The active
|
On 2026-09-30 a successful incremental run also removed the stale 2026-09-29 snapshot and its own
|
||||||
run predates the new progress logging and still requires an observed final cleanup result.
|
temporary snapshot after exit; the earlier `RuntimeDirectory` cleanup failure is resolved.
|
||||||
- [x] Populate `/zpool/archive` with the currently available data so offsite and offline backup tests run
|
- [x] Populate `/zpool/archive` with the currently available data so offsite and offline backup tests run
|
||||||
against a representative load.
|
against a representative load.
|
||||||
- [ ] Run and evaluate Borg against the populated pool: duration, repository capacity, deduplication, and
|
- [x] Evaluate Borg against the populated pool. The 2026-09-29 archive took 1 h 32 min for 2.18 TB
|
||||||
a subsequent incremental archive must be observed before considering the offsite path fully validated.
|
original / 2.04 TB compressed data, with 13.49 GB deduplicated size; retention and compaction
|
||||||
|
succeeded. On 2026-09-30 a subsequent incremental archive completed in about 22 seconds with
|
||||||
|
successful cleanup. The monitor reported 37% Storage Box quota used. These are observed runs, not
|
||||||
|
a guarantee of future duration or compression ratio.
|
||||||
- [x] Add the UUID-bound offline USB backup with versioned rsync, locking, capacity checks, verification,
|
- [x] Add the UUID-bound offline USB backup with versioned rsync, locking, capacity checks, verification,
|
||||||
safe unmounting and a tested restore procedure; never trigger it for an arbitrary USB disk. The
|
safe unmounting and a tested restore procedure; never trigger it for an arbitrary USB disk. The
|
||||||
LUKS/ext4 identities were read-only verified; the manual service and 45Drives Alerts reminder timer were
|
LUKS/ext4 identities were read-only verified; the manual service and 45Drives Alerts reminder timer were
|
||||||
@@ -228,7 +233,8 @@ scheduled retention prune and monthly scrub remain runtime checks.
|
|||||||
found no issues; the service and timer succeeded, and a labelled 45Drives Alerts test notification
|
found no issues; the service and timer succeeded, and a labelled 45Drives Alerts test notification
|
||||||
was submitted. Alerts are deduplicated; email delivery is not claimed. The Storage Box quota probe
|
was submitted. Alerts are deduplicated; email delivery is not claimed. The Storage Box quota probe
|
||||||
runs `df -m` over the dedicated pinned-key SSH identity and does not open the Borg repository.
|
runs `df -m` over the dedicated pinned-key SSH identity and does not open the Borg repository.
|
||||||
Detailed archive size and deduplication remain part of the pending Borg evaluation.
|
The failed-job hook was corrected to pass the literal systemd unit name; its expansion was verified
|
||||||
|
on Atlas, but a new real failure notification has not been deliberately triggered.
|
||||||
|
|
||||||
### Priority 2 - NAS operability and recovery
|
### Priority 2 - NAS operability and recovery
|
||||||
- [ ] Document and test disaster recovery: rebuild Atlas with Ansible, import the existing pool, restore
|
- [ ] Document and test disaster recovery: rebuild Atlas with Ansible, import the existing pool, restore
|
||||||
|
|||||||
22
README.it.md
22
README.it.md
@@ -337,8 +337,8 @@ Gli snapshot ZFS ricorsivi coprono l'intero pool: 24 orari al minuto 05, 30 gior
|
|||||||
8 settimanali la domenica alle 01:00 e 12 mensili il primo giorno alle 02:00. La retention elimina
|
8 settimanali la domenica alle 01:00 e 12 mensili il primo giorno alle 02:00. La retention elimina
|
||||||
solo gli snapshot con prefisso gestito `atlas-auto` e non esegue rollback. Lo scrub OpenZFS mensile è
|
solo gli snapshot con prefisso gestito `atlas-auto` e non esegue rollback. Lo scrub OpenZFS mensile è
|
||||||
previsto la prima domenica alle 03:00; il timer settimanale incompatibile è disabilitato. Il primo
|
previsto la prima domenica alle 03:00; il timer settimanale incompatibile è disabilitato. Il primo
|
||||||
snapshot orario ricorsivo è riuscito; la prima pulizia pianificata e il primo scrub schedulato
|
snapshot orario ricorsivo è riuscito e la pulizia pianificata della retention è stata osservata il
|
||||||
richiedono ancora una verifica a runtime.
|
2026-09-30. Il primo scrub mensile richiede ancora una verifica a runtime.
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
|
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
|
||||||
@@ -391,6 +391,13 @@ copiato un file di `/zpool/archive` in `/var/tmp`, verificando contenuto, propri
|
|||||||
mtime e ACL POSIX; copia e mount temporanei sono stati rimossi senza interrompere Borg. Non è un test
|
mtime e ACL POSIX; copia e mount temporanei sono stati rimossi senza interrompere Borg. Non è un test
|
||||||
di ripristino dell'intero dataset.
|
di ripristino dell'intero dataset.
|
||||||
|
|
||||||
|
L'archivio del pool popolato del 2026-09-29 ha richiesto 1 h 32 min per 2,18 TB originali / 2,04 TB
|
||||||
|
compressi, con 13,49 GB di dimensione deduplicata. Retention e compattazione sono riuscite, ma un
|
||||||
|
errore di permessi su `RuntimeDirectory` ha impedito la pulizia dello snapshot dopo il job. Dopo la
|
||||||
|
correzione, l'archivio incrementale del 2026-09-30 è terminato in circa 22 secondi, ha rimosso lo
|
||||||
|
snapshot residuo e quello corrente ed è terminato con stato 0. Il monitor ha rilevato il 37% della
|
||||||
|
quota Storage Box utilizzata. Questi risultati non predicono durata o compressione dei prossimi run.
|
||||||
|
|
||||||
Il backup USB offline è distribuito come **servizio solo manuale** (`atlas_manage_usb_backup: true`):
|
Il backup USB offline è distribuito come **servizio solo manuale** (`atlas_manage_usb_backup: true`):
|
||||||
Ansible non formatta, sblocca, monta né avvia automaticamente il disco. Il disco esistente è stato
|
Ansible non formatta, sblocca, monta né avvia automaticamente il disco. Il disco esistente è stato
|
||||||
verificato in sola lettura il 2026-09-23: UUID LUKS `577b3c43-ea37-4611-81a9-39d555cdfbd4`,
|
verificato in sola lettura il 2026-09-23: UUID LUKS `577b3c43-ea37-4611-81a9-39d555cdfbd4`,
|
||||||
@@ -459,9 +466,10 @@ baseline di circa 24 ore. Sono controllati anche attivazione e freschezza dei ti
|
|||||||
`OnFailure` segnalano errori di snapshot, scrub, Borg, USB, promemoria e monitoraggio. Il monitor non
|
`OnFailure` segnalano errori di snapshot, scrub, Borg, USB, promemoria e monitoraggio. Il monitor non
|
||||||
riavvia Borg; avvisa solo se un run supera 14 giorni. Soglie e percorsi stabili dei dischi sono nelle
|
riavvia Borg; avvisa solo se un run supera 14 giorni. Soglie e percorsi stabili dei dischi sono nelle
|
||||||
variabili host. Gli avvisi usano 45Drives Houston con deduplicazione; **la consegna email non è stata
|
variabili host. Gli avvisi usano 45Drives Houston con deduplicazione; **la consegna email non è stata
|
||||||
verificata**. Il controllo live del 2026-09-25 non ha trovato problemi; la notifica di prova è stata
|
verificata**. Il controllo live del 2026-09-25 non ha trovato problemi e ha inviato una notifica di
|
||||||
inviata e lo Storage Box risultava occupato al 22%. Dimensione dell'archivio Borg e deduplicazione
|
prova. Il 2026-09-30 il monitor ha rilevato zero problemi e una quota Storage Box occupata al 37%.
|
||||||
dettagliata richiedono ancora la fine del backup in corso.
|
L'hook per i job falliti ora passa il nome letterale della unità systemd; l'espansione è stata
|
||||||
|
verificata senza inviare un falso allarme.
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
ansible-playbook ansible/site.yml --limit atlas --tags monitoring --check --diff
|
ansible-playbook ansible/site.yml --limit atlas --tags monitoring --check --diff
|
||||||
@@ -511,8 +519,8 @@ cutover. L'attuale iCloudPD su Aegis e l'export NFS Photobook restano configurat
|
|||||||
all'approvazione e alla verifica di questa migrazione separata. Anche il servizio Atlas sarà
|
all'approvazione e alla verifica di questa migrazione separata. Anche il servizio Atlas sarà
|
||||||
temporaneo in attesa di Uranus.
|
temporaneo in attesa di Uranus.
|
||||||
|
|
||||||
Il pull dei backup di Prometheus, la valutazione delle dimensioni degli archivi Borg e i test completi
|
Il pull dei backup di Prometheus e i test completi di disaster recovery restano da fare. Il backlog
|
||||||
di disaster recovery restano da fare. Il backlog prioritizzato è in `AGENTS.md`.
|
prioritizzato è in `AGENTS.md`.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
23
README.md
23
README.md
@@ -325,8 +325,8 @@ snapshots at minute 05, 30 daily snapshots at 00:15, 8 weekly snapshots on Sunda
|
|||||||
monthly snapshots on the first day at 02:00. The retention helper prunes only snapshots carrying its
|
monthly snapshots on the first day at 02:00. The retention helper prunes only snapshots carrying its
|
||||||
managed `atlas-auto` prefix and never rolls back a dataset. The OpenZFS monthly scrub timer is scheduled
|
managed `atlas-auto` prefix and never rolls back a dataset. The OpenZFS monthly scrub timer is scheduled
|
||||||
for the first Sunday at 03:00; the conflicting weekly scrub timer is disabled explicitly. The first recursive
|
for the first Sunday at 03:00; the conflicting weekly scrub timer is disabled explicitly. The first recursive
|
||||||
hourly snapshot completed successfully on Atlas; retention pruning and the first scheduled scrub still await
|
hourly snapshot completed successfully on Atlas, and scheduled retention pruning was observed on
|
||||||
live runtime evidence. Validate this layer independently with:
|
2026-09-30. The first monthly scrub still awaits runtime evidence. Validate this layer independently with:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
|
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
|
||||||
@@ -381,6 +381,12 @@ ansible-playbook ansible/site.yml --limit atlas --tags packages,borg --check --d
|
|||||||
Atlas runtime activation is complete: the initial backup and repository check succeeded, a full restore
|
Atlas runtime activation is complete: the initial backup and repository check succeeded, a full restore
|
||||||
to a temporary directory was validated against the live `Archive` tree, the recovery-key export was copied
|
to a temporary directory was validated against the live `Archive` tree, the recovery-key export was copied
|
||||||
to offline storage, and the temporary snapshot and bind mounts were cleaned up.
|
to offline storage, and the temporary snapshot and bind mounts were cleaned up.
|
||||||
|
The populated-pool archive on 2026-09-29 took 1 h 32 min for 2.18 TB original / 2.04 TB compressed
|
||||||
|
data, with a 13.49 GB deduplicated archive size. Retention and compaction succeeded, but a
|
||||||
|
`RuntimeDirectory` permission error prevented post-exit snapshot cleanup. After correction, the
|
||||||
|
2026-09-30 incremental archive completed in about 22 seconds, removed the stale and current temporary
|
||||||
|
snapshots, and ended with service status 0. The monitor reported 37% Storage Box quota used. These
|
||||||
|
observations do not predict the duration or compression ratio of future runs.
|
||||||
On 2026-09-25 a separate ZFS restore smoke test copied a small file from an automatic daily
|
On 2026-09-25 a separate ZFS restore smoke test copied a small file from an automatic daily
|
||||||
`zpool/archive` snapshot to `/var/tmp`, then confirmed matching contents, ownership, mode, mtime and
|
`zpool/archive` snapshot to `/var/tmp`, then confirmed matching contents, ownership, mode, mtime and
|
||||||
POSIX ACL. The temporary copy and on-demand snapshot mount were removed; Borg kept running. This
|
POSIX ACL. The temporary copy and on-demand snapshot mount were removed; Borg kept running. This
|
||||||
@@ -472,12 +478,13 @@ Storage Box quota via `df -m` over the dedicated `borg` account's pinned-key SSH
|
|||||||
query never opens the Borg repository or its lock. Growth alerts compare against a roughly 24-hour
|
query never opens the Borg repository or its lock. Growth alerts compare against a roughly 24-hour
|
||||||
baseline and therefore begin only after enough samples exist. The monitor also checks
|
baseline and therefore begin only after enough samples exist. The monitor also checks
|
||||||
maintenance/backup timer activation and freshness; systemd `OnFailure` hooks report snapshot,
|
maintenance/backup timer activation and freshness; systemd `OnFailure` hooks report snapshot,
|
||||||
scrub, Borg, USB, reminder, and monitoring services when they enter the failed state. The ongoing
|
scrub, Borg, USB, reminder, and monitoring services when they enter the failed state. An ongoing
|
||||||
initial Borg run is never restarted by the monitor; only a run exceeding 14 days raises a warning.
|
Borg run is never restarted by the monitor; only a run exceeding 14 days raises a warning.
|
||||||
Thresholds and stable disk paths are declared in Atlas host variables. Alerts use the existing 45Drives
|
Thresholds and stable disk paths are declared in Atlas host variables. Alerts use the existing 45Drives
|
||||||
Houston notifier and repeated issues are deduplicated; **email delivery is not verified**. The
|
Houston notifier and repeated issues are deduplicated; **email delivery is not verified**. The
|
||||||
2026-09-25 live probe found no issues and a labelled test notification was submitted. The Storage Box
|
2026-09-25 live probe found no issues and a labelled test notification was submitted. On 2026-09-30
|
||||||
reported 22% used. Detailed Borg archive size and deduplication still require the active run to finish.
|
the monitor reported zero issues and 37% Storage Box quota used. The failed-job hook now passes the
|
||||||
|
literal systemd unit name; its expansion was verified without sending a false failure notification.
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
ansible-playbook ansible/site.yml --limit atlas --tags monitoring --check --diff
|
ansible-playbook ansible/site.yml --limit atlas --tags monitoring --check --diff
|
||||||
@@ -527,8 +534,8 @@ state outside `Archive`, then test permissions, SELinux, backups and recovery be
|
|||||||
The current Aegis iCloudPD service and Atlas Photobook NFS export remain configured until that
|
The current Aegis iCloudPD service and Atlas Photobook NFS export remain configured until that
|
||||||
separate migration is approved and validated; the eventual Atlas service is temporary until Uranus.
|
separate migration is approved and validated; the eventual Atlas service is temporary until Uranus.
|
||||||
|
|
||||||
Prometheus backup pulls, Borg archive-size evaluation, and full disaster-recovery tests remain follow-up
|
Prometheus backup pulls and full disaster-recovery tests remain follow-up work. The prioritized
|
||||||
work. The prioritized operational backlog is kept in `AGENTS.md`.
|
operational backlog is kept in `AGENTS.md`.
|
||||||
|
|
||||||
## How layering works
|
## How layering works
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user