From a609e68f42cc7007affd64e2fed5ae30b906f883 Mon Sep 17 00:00:00 2001 From: Fabio Scotto di Santolo Date: Thu, 1 Oct 2026 21:37:49 +0200 Subject: [PATCH] Record ZFS and Borg coverage for staged Gitea --- AGENTS.md | 10 ++++++++-- docs/atlas-gitea-migration.md | 20 ++++++++++++++++---- 2 files changed, 24 insertions(+), 6 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index b060986..8cd76dc 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -278,8 +278,14 @@ successfully. The first monthly scrub remains a runtime check. matched. The pinned rootless image answered HTTP and listened on internal SSH/2222 with `--network none`; the temporary container was removed and the Quadlet stayed inactive. A second restore run made no changes. This is a rehearsal copy, not the final consistent cutover copy. -- [ ] Verify target Gitea backup coverage with a new recursive ZFS snapshot, Borg archive, and an - independent restore of the staged dataset before accepting production writes. +- [x] Verify ZFS and Borg coverage of the staged Gitea dataset. On 2026-10-01 the managed recursive + hourly snapshot `atlas-auto-hourly-20261001T193401Z` included it, and the managed incremental + Borg archive `atlas-20261001T193420Z` included its database. A private one-file restore from + each independently matched the staged database and passed SQLite `quick_check`; temporary files + and snapshot mounts were removed, the Borg service ended successfully, and the pool was healthy. +- [ ] Include the new Gitea dataset in the next UUID-bound offline USB version and test a file restore + from that version before accepting production writes; the UUID-bound disk is connected but its + LUKS mapper is closed, so the manual backup still requires interactive unlock. - [ ] After an explicit outage approval, perform the final consistent copy and HTTPS/SSH cutover, then remove Gitea from Prometheus' desired stack and backup export without deleting source data. - [ ] Design and deploy Nextcloud as another explicitly temporary Atlas service before Uranus. Give it diff --git a/docs/atlas-gitea-migration.md b/docs/atlas-gitea-migration.md index 9b599ce..934c67d 100644 --- a/docs/atlas-gitea-migration.md +++ b/docs/atlas-gitea-migration.md @@ -61,8 +61,19 @@ with `--network none` answered HTTP internally and listened on internal SSH/2222. The container was removed; the user Quadlet remains inactive, with no staging listener. The second restore run changed nothing. This copy is deliberately stale once new source writes occur and **must not** be used as the -final cutover copy. Target snapshot/Borg inclusion and an independent restore -are still pending. +final cutover copy. + +Target backup checks on 2026-10-01: the managed recursive hourly ZFS snapshot +`atlas-auto-hourly-20261001T193401Z` contains the new dataset. The managed +Borg service completed archive `atlas-20261001T193420Z`, whose contents list +includes the staged Gitea database. A separate one-file restore from each +source into private `/var/tmp` directories matched the live staged database +and passed SQLite `quick_check`. Temporary files and the on-demand snapshot +mount were removed; the Borg temporary snapshot was cleaned up and the pool +remained healthy. This is file-level proof, **not** a full Gitea recovery. +The UUID-bound offline USB disk is connected but its LUKS mapper is closed; +its manual backup requires interactive unlock. It has not yet captured or +restored this new dataset. 1. Provision a dedicated target dataset and non-login service identity via Ansible, keeping UID/GID distinct from Atlas' reserved Immich `1100`. @@ -87,8 +98,9 @@ are still pending. container with no production ingress or outbound network. Because the source stays active, this is a rehearsal copy, not the final cutover copy. Regenerate Git hooks if the changed installation path requires it. -4. Confirm that Atlas snapshots, Borg, and offline USB include the new dataset; - test at least one independent restore before user traffic is accepted. +4. ZFS and Borg inclusion and one-file restores have passed. Complete a + UUID-bound offline USB version and a one-file restore for the new dataset + before accepting user traffic. ## Phase 2: explicit final cutover