Compare commits

...

33 Commits

Author SHA1 Message Date
Fabio Scotto di Santolo
18eb2d2eb2 Document confirmed Gitea domain transition completion 2026-10-03 15:15:21 +02:00
Fabio Scotto di Santolo
269fb13665 Document Gitea domain and retire Prometheus DuckDNS 2026-10-03 15:09:34 +02:00
Fabio Scotto di Santolo
2dfe766b7b Enable boot startup for Atlas iCloudPD 2026-10-03 13:59:05 +02:00
Fabio Scotto di Santolo
755f24bc72 Retire Prometheus Compose stack and document cleanup 2026-10-03 13:44:57 +02:00
Fabio Scotto di Santolo
7bc7f0e645 Feature/prometheus npm quadlet (#15)
* Stage Prometheus NPM Quadlet with backup-safe cutover

* Complete Prometheus NPM Quadlet cutover
2026-10-03 11:53:12 +02:00
Fabio Scotto di Santolo
1577eec19d Merge branch 'feature/gitea-https-validation' 2026-10-03 10:19:09 +02:00
Fabio Scotto di Santolo
e30683c3d1 Record Gitea HTTPS validation 2026-10-03 10:18:25 +02:00
Fabio Scotto di Santolo
4bd6aafb53 Feature/atlas icloudpd migration (#14)
* Design gated Atlas iCloudPD migration target

* Target Atlas iCloudPD photos to Photobook

* Record isolated iCloudPD Photobook ACL validation

* Record Aegis iCloudPD source audit gap

* Verify iCloudPD backup source scope and Borg access

* Record Atlas iCloudPD deployment gate checks

* Pin iCloudPD photo file and directory modes

* Validate inactive iCloudPD Quadlet on Atlas generator

* Keep iCloudPD in Archive and reserve Photobook for Immich

* Prepare guarded Aegis iCloudPD retirement

* Declare inactive Atlas iCloudPD storage and Quadlet

* Retire Aegis iCloudPD from desired state

* Clear retired Aegis iCloudPD failed-unit state

* Remove completed iCloudPD retirement tasks from Aegis

* Record initial Atlas iCloudPD service start

* Manage Atlas iCloudPD config from Vault

* Fix Atlas iCloudPD traceroute startup and config drift

* Use Atlas Vault key for iCloudPD Apple ID

* Add HEIC decoding to Fedora desktops

* Record completed iCloudPD ingestion and remaining recovery checks
2026-10-03 09:59:59 +02:00
Fabio Scotto di Santolo
9e76309833 Merge main and reconcile Atlas checklist 2026-10-02 17:54:33 +02:00
Fabio Scotto di Santolo
ed3fee06e8 Record operator-validated Gitea SSH pull and push 2026-10-02 17:47:40 +02:00
Fabio Scotto di Santolo
309d64b4ed Record successful public Gitea SSH authentication 2026-10-02 10:19:23 +02:00
Fabio Scotto di Santolo
dd33a4f55d Move Atlas Gitea Quadlet to admin with internal gitea user 2026-10-02 10:06:51 +02:00
Fabio Scotto di Santolo
0028fe8c4d Cut over Gitea HTTPS to Atlas with managed NPM upstream 2026-10-02 09:36:45 +02:00
Fabio Scotto di Santolo
12037fcc9a Enable restored rootless Gitea on Atlas 2026-10-02 09:35:44 +02:00
Fabio Scotto di Santolo
3f9a626759 Validate Gitea USB backup restore 2026-10-02 09:13:06 +02:00
Fabio Scotto di Santolo
31fedb8d44 Prepare gated Gitea HTTPS and SSH cutover 2026-10-01 22:09:57 +02:00
Fabio Scotto di Santolo
9b5ee77905 Prepare guarded final Gitea restore on Atlas 2026-10-01 22:00:41 +02:00
Fabio Scotto di Santolo
54fb7d46d7 Prepare consistent final Gitea export on Prometheus 2026-10-01 21:57:29 +02:00
Fabio Scotto di Santolo
a609e68f42 Record ZFS and Borg coverage for staged Gitea 2026-10-01 21:38:39 +02:00
Fabio Scotto di Santolo
06d3b175cb Rehearse rootless Gitea restore from verified backup 2026-10-01 21:33:18 +02:00
Fabio Scotto di Santolo
256d758b1a Prepare isolated rootless Gitea Quadlet on Atlas 2026-10-01 21:23:40 +02:00
Fabio Scotto di Santolo
9d0013769c Correct Atlas Gitea migration to rootless Quadlet 2026-10-01 21:15:42 +02:00
Fabio Scotto di Santolo
5b0f415163 Document staged Gitea migration to Atlas 2026-10-01 21:10:32 +02:00
Fabio Scotto di Santolo
3401b6137d Merge pull request #12 from fscotto/feature/atlas-navidrome-music 2026-10-01 20:59:49 +02:00
Fabio Scotto di Santolo
bae6a9f554 Record first scheduled Prometheus backup pull 2026-10-01 19:58:46 +02:00
Fabio Scotto di Santolo
c37483ba38 Track daily Atlas music copy separately in checklist 2026-10-01 10:35:41 +02:00
Fabio Scotto di Santolo
f491362365 Schedule daily Atlas Navidrome music copy 2026-10-01 10:32:48 +02:00
Fabio Scotto di Santolo
5a1047adde Document validated Atlas Navidrome music copy 2026-09-30 22:55:47 +02:00
Fabio Scotto di Santolo
802cb8c7ba Merge pull request #10 from fscotto/feature/atlas-priority2-recovery
Feature/atlas priority2 recovery
2026-09-30 21:25:36 +02:00
Fabio Scotto di Santolo
0144600a4a Add verified Prometheus backup pull to Atlas 2026-09-30 21:21:48 +02:00
Fabio Scotto di Santolo
3d2ef02c98 Keep Atlas SMB and NFS namespaces separate 2026-09-30 21:20:15 +02:00
Fabio Scotto di Santolo
8844d00e24 Define controlled Atlas kernel and OpenZFS updates 2026-09-30 21:19:48 +02:00
Fabio Scotto di Santolo
9798fe3a12 Document and rehearse Atlas disaster recovery 2026-09-30 21:19:28 +02:00
69 changed files with 4296 additions and 218 deletions

250
AGENTS.md
View File

@@ -25,6 +25,9 @@ Ansible-driven personal infrastructure repo for Fedora and Void desktops, Fedora
- Preserve layering `all -> platform -> role -> desktop -> host`.
- Keep `ansible/site.yml` small; orchestration belongs there, implementation belongs in roles.
- Prefer minimal, targeted edits. Preserve idempotency and existing ordering.
- Keep completed one-time cleanup operations out of the playbook. Execute them directly
with explicit authorization; retain only the ongoing desired-state configuration and
historical documentation, not permanent cleanup flags or tasks.
- Use Git Flow branch prefixes: `feature/` for new functionality, `bugfix/` for non-urgent fixes,
`hotfix/` for urgent production fixes, `release/` for release preparation, and `support/` for
maintained release lines. Do not use abbreviated prefixes such as `feat/`.
@@ -54,9 +57,30 @@ Ansible-driven personal infrastructure repo for Fedora and Void desktops, Fedora
- Emacs is disabled by default; temporary Emacs check: `ansible-playbook ansible/site.yml --limit <host> --tags emacs --check --diff -e emacs_enabled=true`
- AI coding agents: `ansible-playbook ansible/site.yml --limit <host> --tags ai_agents --check --diff`
- Mail bootstrap: `sh -n scripts/bootstrap_mail.sh` and `shellcheck scripts/bootstrap_mail.sh`
- Server compose render: `podman-compose -f /opt/docker/server/docker-compose.yml config` and `systemctl status podman-compose-server`
- Server NPM Quadlet: `systemctl status prometheus-npm.service`; the Compose fallback is retired.
- Explicit Prometheus legacy cleanup (destructive only without check mode):
`ansible-playbook ansible/site.yml --limit prometheus --tags server_legacy_cleanup --check --diff -e server_legacy_cleanup=true`
- Atlas media stack:
`ansible-playbook ansible/site.yml --limit atlas --tags storage,sharing,containers --check --diff`
- Atlas rootless Gitea staging (does not start Gitea):
`ansible-playbook ansible/site.yml --limit atlas --tags gitea --check --diff`
- Atlas canonical Gitea domain (restarts only Gitea on a real configuration change):
`ansible-playbook ansible/site.yml --limit atlas --tags gitea_public_domain --check --diff`
- Atlas iCloudPD storage and boot-started Quadlet:
`ansible-playbook ansible/site.yml --limit atlas --tags icloudpd --check --diff`
- Atlas explicit Gitea host-owner migration (live outage; never a normal run):
`ansible-playbook ansible/site.yml --limit atlas --tags gitea_owner_migration -e atlas_gitea_owner_migration=true`
- Atlas explicit isolated Gitea restore rehearsal (not part of normal runs):
`ansible-playbook ansible/site.yml --limit atlas --tags gitea_restore -e atlas_gitea_restore_test=true`
- Atlas final Gitea replacement gate (dry-run only until a stopped-source export is pulled):
`ansible-playbook ansible/site.yml --limit atlas --tags gitea_final_restore --check --diff -e atlas_gitea_final_restore=true`
- Prometheus final Gitea export helper (dry-run installs only; outage action remains opt-in):
`ansible-playbook ansible/site.yml --limit prometheus --tags gitea_final_export --check --diff`
- Gitea cutover network configuration before activation:
`ansible-playbook ansible/site.yml --limit prometheus --tags gitea_cutover,prometheus_backup --check --diff -e server_gitea_on_atlas=true`
and `ansible-playbook ansible/site.yml --limit atlas --tags gitea --check --diff`
- Atlas daily Navidrome music copy:
`ansible-playbook ansible/site.yml --limit atlas --tags music_sync --check --diff`
- Atlas network/share hardening:
`ansible-playbook ansible/site.yml --limit atlas --tags hardening,sharing --check --diff`
- Atlas ZFS snapshot retention and scrub timers:
@@ -73,7 +97,9 @@ Ansible-driven personal infrastructure repo for Fedora and Void desktops, Fedora
`ansible-playbook ansible/site.yml --limit atlas --tags restorecon --check -e '{"atlas_restorecon_paths":["/zpool/archive"]}'`
- Prometheus/Aegis WireGuard gateway:
`ansible-playbook ansible/site.yml --limit prometheus,aegis --tags wireguard --check --diff`
- DuckDNS config only: `ansible-playbook ansible/site.yml --limit prometheus --tags duckdns --check --diff`
- Prometheus NPM Quadlet steady state (does not perform a cutover):
`ansible-playbook ansible/site.yml --limit prometheus --tags npm_quadlet --check --diff`
- DuckDNS config only (skipped on Prometheus): `ansible-playbook ansible/site.yml --limit prometheus --tags duckdns --check --diff`
## Conventions
- Use FQCN Ansible modules.
@@ -117,16 +143,24 @@ The dotfile vars follow the same split: `desktop_common_dotfiles` carries mode-i
- Windows applications are installed manually and are not managed from the WSL profile.
## Rocky Server Notes
- DuckDNS is rendered by `profile_server` from host-local `server_duckdns_domain` and
- Prometheus disables DuckDNS provisioning with `server_duckdns_enabled: false`. Its updater,
log and five-minute cron entry were explicitly retired; the external DuckDNS name and Vault
token remain untouched. The completed one-time cleanup has no remaining playbook tasks.
- When enabled, DuckDNS is rendered by `profile_server` from host-local `server_duckdns_domain` and
`vault_duckdns_token`. Keep the rotated token in encrypted Vault or untracked local vars, never in
dotfiles. The private `~/duckdns/duck.sh` keeps the existing entrypoint; rendering uses `no_log`
and disables diffs. Provisioning does not execute the updater or change its external schedule.
- `rocky_server` is a child of both `platform_rocky` and `server`; `prometheus` is its active target.
- The target must already provide `server_username` with local sudo access before the profile runs.
- The Rocky profile installs Podman and podman-compose, uses firewalld, preserves SELinux enforcement, and renders the
existing Nginx Proxy Manager/Gitea Compose stack with a `podman-compose-server` systemd unit. PostgreSQL and
Navidrome are no longer part of the desired Prometheus configuration. The role does not stop or remove legacy
containers, delete `/opt/postgres/data`, start the Compose stack, update DNS, or cut over traffic.
- The Rocky profile installs Podman and podman-compose. Prometheus explicitly retires the legacy
Compose unit, files and final-export helper with `server_legacy_stack_retired: true`.
Its approved opt-in cleanup removed old application data on 2026-10-03; normal runs do not
delete data or recreate the retired files. On Prometheus, Nginx Proxy Manager is now the rootful
`prometheus-npm.service` Quadlet with a pinned image digest and the existing `/opt/npm/data` and
`/opt/npm/letsencrypt` bind mounts. The rootful `server_web` bridge remains `10.89.0.0/24`.
Gitea runs on Atlas; PostgreSQL and Navidrome are absent from the desired Prometheus stack.
Normal runs do not delete legacy data, update DNS, or perform an implicit cutover;
destructive cleanup requires its explicit tag and opt-in extra-var.
- Firewalld enables SSH, Cockpit (`9090/tcp`), HTTP and HTTPS. Nginx Proxy Manager publishes `80/tcp` and
`443/tcp`; bind its administration interface only to `127.0.0.1:81` and use `npm-tunnel` from Ikaros or Nymph.
Nextcloud remains disabled; do not provision `/srv/nextcloud` directories.
@@ -164,7 +198,9 @@ The dotfile vars follow the same split: `desktop_common_dotfiles` carries mode-i
- `profile_backend_phase1` temporarily runs rootless Navidrome and Syncthing on Atlas until Uranus replaces
them. It binds only to Atlas' LAN IP, never `wg0`; Navidrome and the Syncthing GUI admit only Aegis as
the source-NAT gateway, while native Syncthing ports admit the configured LAN. It initializes fresh
state only and never migrates or deletes source application data.
state only and never migrates or deletes source application data. The enabled rootless
`atlas-music-sync.timer` copies `/zpool/archive/Music` to `/zpool/media/music` daily at 00:45
Europe/Rome without deleting destination files; it requires both datasets to be mounted.
- `wireguard_overlay` manages `wg0` between Prometheus (`10.0.0.1`) and Aegis (`10.0.0.2`). It persists private
keys only on their respective hosts, exchanges only derived public keys through Ansible, and verifies a real peer
handshake. Prometheus opens `51820/udp`; Aegis is the LAN gateway. Its persistent IPv4 forwarding, narrowly scoped
@@ -226,7 +262,7 @@ successfully. The first monthly scrub remains a runtime check.
read-only ZFS snapshot test restored one file to `/var/tmp`, confirmed matching contents, ownership,
mode, mtime and ACL, then removed its temporary copy and on-demand mount. This is a file-level smoke
test, not full dataset recovery. An independent USB file restore passed on 2026-09-25 with matching
content and metadata; full disaster recovery remains a separate Priority 2 task.
content and metadata; the later scaled OS-rebuild rehearsal is documented under Priority 2.
- [x] Add monitoring and alerting for pool health, scrub/resilver, SMART data, temperatures, free space,
snapshot/local-backup growth, Hetzner Storage Box quota, and failed maintenance or backup timers.
The half-hourly Atlas health monitor and systemd final-failure hooks are deployed. A live probe
@@ -237,31 +273,195 @@ successfully. The first monthly scrub remains a runtime check.
on Atlas, but a new real failure notification has not been deliberately triggered.
### Priority 2 - NAS operability and recovery
- [ ] Document and test disaster recovery: rebuild Atlas with Ansible, import the existing pool, restore
from snapshot/USB/Hetzner, preserve Vault and Borg recovery material offline, and define RPO/RTO.
- [ ] Define a controlled Rocky kernel/OpenZFS update and reboot procedure.
- [ ] Add the Atlas-initiated least-privilege Prometheus backup pull: Prometheus exposes only prepared
- [x] Document and test disaster recovery in `docs/atlas-recovery.md`: the operator confirmed Vault
and Borg recovery material is available offline; provisional targets are RPO 24h/RTO 72h. On
2026-09-30 an isolated small Rocky VM was rebuilt with the Atlas Ansible roles, imported its
preserved RAIDZ2 pool without force/rewind, and restored a file from the preserved snapshot;
the second Ansible run was idempotent. Earlier independent production ZFS, USB, and Borg file
restore tests remain separate evidence. A production-size full restore, unclean import, and
measured 24h/72h compliance are not claimed.
- [x] Define a controlled Rocky kernel/OpenZFS update and reboot procedure in
`docs/atlas-updates.md`. The first real change-window execution is not yet
validated; the procedure never reboots automatically or upgrades pool features.
- [x] Add the Atlas-initiated least-privilege Prometheus backup pull: Prometheus exposes only prepared
read-only dumps through a dedicated account and Atlas retains the private SSH key, pinned host key,
atomic pull, verification, retention and systemd service/timer.
- [ ] Decide whether a common SMB/NFS namespace is required. `Archive` (SMB) and `photobook` (NFS) are
intentionally distinct today; only if a shared namespace is selected, finalize its UID/GID, group,
and POSIX ACL model and test the same files through both protocols.
atomic pull, verification, retention and systemd service/timer. The dedicated key/account and unit
files are deployed; live read-only SSH, shell denial, and write denial were verified. On 2026-09-30
a manual export, Atlas pull, checksum verification, and temporary restore passed; both SQLite
databases passed integrity checks and a restored Git repository passed `git fsck`. Both daily
timers are enabled for 02:00/03:00 Europe/Rome. On 2026-10-01 their first scheduled export and
pull succeeded: Atlas verified the payload checksum and published `20261001T000001Z` as `latest`.
- [x] Decide whether a common SMB/NFS namespace is required: no. `Archive` (SMB) and `photobook` (NFS)
remain intentionally distinct; `docs/atlas-sharing-decision.md` records the decision. No ACL or export
change is authorized by this decision.
### Priority 3 - Service expansion
- [ ] After data protection and recovery are validated, populate `/zpool/media/music` and validate Navidrome.
- [x] Populate `/zpool/media/music` and validate Navidrome. On 2026-09-30, 21,158 files
(93,937,810,350 regular-file bytes) were copied from `/zpool/archive/Music` using a temporary
ZFS snapshot; a checksum-based rsync dry run found no differences or extra files. Navidrome saw
all files through its read-only mount, completed a scan, indexed 18,168 tracks, and responded
over HTTP. Some imported playlists still reference obsolete Windows paths. The source was left
intact and the temporary snapshot was removed.
- [x] Schedule a daily, non-deleting copy from `Archive/Music` to the separate Navidrome music
dataset. The rootless `atlas-music-sync.timer` is enabled for 00:45 Europe/Rome; a manual
idempotent service run succeeded on 2026-10-01. The first scheduled run triggered at
00:45 CEST on 2026-10-02 and exited successfully (`Result=success`, status 0); the next
run is scheduled for 2026-10-03 00:45 CEST.
- [x] Design the staged Prometheus-to-Atlas Gitea migration in `docs/atlas-gitea-migration.md`.
The approved topology keeps NPM on Prometheus and moves HTTPS and public SSH (TCP/2222) together;
Gitea runs as an `admin`-owned rootless user Quadlet on Atlas with an internal `gitea` user.
The rootful-to-rootless data-layout
conversion passed an isolated restore rehearsal. The later partial cutover is tracked below.
- [x] Prepare the dedicated Atlas Gitea dataset, non-login UID/GID 1101 with a separate rootless Podman
sub-ID range, and disabled user Quadlet. On 2026-10-01 the targeted Ansible run and a second idempotent
run passed; the generated unit was inactive, with no staging HTTP/SSH listener. POSIX ACLs on only the
service-namespace parents grant this account traversal without access to sibling datasets.
- [x] Perform an isolated rootless restore rehearsal from the verified Prometheus backup. On 2026-10-01
the SHA-256-checked selective extraction and path/SSH conversion succeeded; SQLite `quick_check`
passed, all 33 repositories passed `git fsck`, and source/target public SSH host-key fingerprints
matched. The pinned rootless image answered HTTP and listened on internal SSH/2222 with
`--network none`; the temporary container was removed and the Quadlet stayed inactive. A second
restore run made no changes. This is a rehearsal copy, not the final consistent cutover copy.
- [x] Verify ZFS and Borg coverage of the staged Gitea dataset. On 2026-10-01 the managed recursive
hourly snapshot `atlas-auto-hourly-20261001T193401Z` included it, and the managed incremental
Borg archive `atlas-20261001T193420Z` included its database. A private one-file restore from
each independently matched the staged database and passed SQLite `quick_check`; temporary files
and snapshot mounts were removed, the Borg service ended successfully, and the pool was healthy.
- [x] Include the new Gitea dataset in a UUID-bound offline USB version and test a file restore
before accepting production writes. The operator's 2026-10-01 manual run published version
`20261001T201220Z-254397` successfully on 2026-10-02. Its Gitea database was restored to a
temporary directory from a read-only mount: contents, owner, group, mode, size, mtime and POSIX
ACL matched, and SQLite `quick_check` passed. Temporary files and mounts were removed, LUKS
was closed, and the pool remained healthy. A redundant run was stopped during verification;
its temporary snapshot was cleaned up and the service's resulting failed state was reset.
- [x] Install a separate opt-in final Gitea export helper on Prometheus. Its 2026-10-01 targeted
deployment and `bash -n` passed while Gitea and NPM stayed running. It refuses an active export
timer, stops only Gitea, verifies SQLite, publishes a checksum-verified Gitea-only version for
Atlas' existing pull, and leaves the source stopped on success. It was invoked on 2026-10-02
after the export timer was stopped; version `20261002T071525Z` was pulled and verified on Atlas.
- [x] Prepare the Atlas final-restore gate without replacing the rehearsal: it accepts only a
checksum-verified `gitea-cutover` export, refuses a running target, stages and validates the new
layout before replacing the marked rehearsal, and rolls back a failed swap. Synthetic success
and rollback tests passed on 2026-10-01. On 2026-10-02 the final gate replaced the rehearsal;
SQLite `quick_check`, all 33 repository `git fsck` checks, checksum and SSH host-key comparison passed.
- [x] Start the rootless Atlas Gitea Quadlet and move the primary HTTPS route. On 2026-10-02 Atlas
answered HTTP 200 through the Aegis gateway. NPM stayed on Prometheus; its variable upstream
required a managed Nginx `server_proxy.conf` override because runtime DNS ignores Compose
`extra_hosts`. The primary public HTTPS page and API returned 200, and `git ls-remote` succeeded
for a representative repository after NPM restart; the Navidrome and Syncthing Proxy Hosts also
responded. The source
Gitea container was removed from the desired Compose stack without deleting its data; the
Prometheus backup export timer resumed for NPM only. A post-cutover recursive ZFS snapshot and
encrypted Borg archive `atlas-20261002T073044Z` completed successfully.
- [x] Move the live Gitea Quadlet and dataset from the legacy host `gitea` account to `admin`
after a disposable snapshot-copy test of the pinned derived image. On 2026-10-02 the explicit
outage run stopped only legacy Gitea, made safety snapshot
`zpool/services/data/gitea@gitea-owner-migration-20261002T100104`, changed dataset ownership,
and validated loopback staging (HTTP 200, internal `gitea` UID/GID 1000, SQLite `quick_check`)
before promoting the `admin` Quadlet. Production LAN and public HTTPS returned 200; Navidrome
and Syncthing remained active, the pool was healthy, and the normal Gitea run changed nothing.
The old host account and data on Prometheus remain preserved; the old Atlas Quadlet and its
parent-dataset traverse ACL were removed. A subsequent normal run changed nothing.
- [x] Validate public Gitea SSH/2222 and an authenticated read from Ikaros. After the VPS
firewall was opened on 2026-10-02, TCP/2222 connected, the public ED25519 host-key
fingerprint matched Atlas, Gitea authenticated `fscotto` using the `ikaros` key, and
`git ls-remote` returned HEAD for `fscotto/infra.git` over public SSH.
- [x] Validate authenticated SSH pull and push. On 2026-10-02 the operator reported both
operations working through the public SSH endpoint; the earlier agent-run `git ls-remote`
remains the independent read-only check. The agent did not perform a test push.
- [x] Validate Gitea login and write via HTTPS. On 2026-10-03 the operator confirmed
authenticated web login and Git clone/pull/push through the public HTTPS endpoint. Do not
restart the stale source Gitea after Atlas has accepted writes.
- [ ] Design and deploy Nextcloud as another explicitly temporary Atlas service before Uranus. Give it
separate persistent application, database, and cache storage; keep credentials in Vault; publish it only
through NPM over the Prometheus--Aegis gateway; and define backup, upgrade, and eventual Uranus-migration
procedures before exposing user data. Do not deploy Nextcloud before the data-protection checklist is complete.
- [x] Move Gitea canonical HTTPS and SSH hostname to `git.fscotto.co` on
2026-10-03 through Ansible. Only Gitea restarted; second run changed nothing.
HTTPS and authenticated SSH reads returned the same repository HEAD.
The new NPM hostnames passed TLS/HTTP checks; old DuckDNS Proxy Hosts were
observed disabled. Details are in `docs/domain-fscotto-co.md`.
- [x] Confirm login on the new Gitea hostname and update remaining client remotes/integrations.
The operator confirmed completion on 2026-10-03; the agent did not perform a test push.
- [x] Remove obsolete DuckDNS NPM Proxy Hosts, unused certificates and the old upstream override.
The operator confirmed completion on 2026-10-03; no new agent runtime check was performed.
- [x] Review and remove completed one-time procedures from the playbook.
The operator confirmed completion on 2026-10-03.
- [x] Retire Prometheus' local DuckDNS updater on 2026-10-03 through Ansible:
the five-minute cron entry and private updater/log directory were removed.
Provisioning is disabled; repeat cleanup changed nothing. HTTPS services, private NPM
administration and the export timer stayed healthy. The external name and Vault token
remain untouched for possible future use on a local host.
- [ ] Keep `atlas_manage_media_stack` disabled until the future Immich deployment has validated `/dev/dri`,
container paths, and the required Vault database secret.
### Priority 4 - Optional workflows
- [ ] After data protection is validated, move iCloudPD photo ingestion from Aegis to Atlas as a
temporary service until Uranus is ready. Plan to store photos in `/zpool/archive/Pictures` and
persistent application/MFA state outside `Archive`; validate permissions, SELinux, backups and
recovery before cutover. Keep the current Aegis service and Photobook NFS export unchanged until
the Atlas workflow is tested, then retire them explicitly if no longer needed.
- [x] Deploy the declared Atlas iCloudPD state dataset and inactive rootless `admin` Quadlet.
Photos belong under `/zpool/archive/Pictures/iCloudPD`; private config/MFA state belongs in
`zpool/services/data/icloudpd`. Photobook remains reserved for Immich. Ansible now renders
`icloudpd.conf` with the Apple ID from the existing Vault key, but does not store the password
or manage MFA. Automatic startup was approved on 2026-10-03; the Quadlet now
uses `WantedBy=default.target` and Ansible keeps the service running.
The isolated no-network layout test is documented in
`docs/atlas-icloudpd-migration.md`. On 2026-10-02 Atlas deployment and a second idempotent run
passed; no app config existed at deployment. A manual first start on 2026-10-02 generated
`icloudpd.conf`; an Ansible run then replaced it with a private mode-0600 Vault-backed template
and an idempotent second run. The image later expanded the config, so Ansible now seeds it
only when absent and maintains the declared fields. Its launcher requires `traceroute`; the
rootless Quadlet grants only `NET_RAW`, tested in isolation and after restart. The service
was subsequently initialized interactively; initial ingestion is tracked below.
- [x] Retire Aegis iCloudPD completely. The operator authorized deleting its Quadlet,
`/var/lib/icloudpd` data, and MFA state despite an unaudited container overlay. After two
interactive-sudo runs on 2026-10-02, the unit is `not-found`/`inactive`, the Quadlet and state
directory are absent, and AdGuard remains active. The temporary retirement tasks have since
been removed from the Aegis role; it no longer manages iCloudPD.
- [x] Validate Atlas iCloudPD authentication and initial ingestion. On 2026-10-03 the active
rootless service logged `All photos and videos have been downloaded` at 02:16 and reported
completion for the user. The destination held 11,658 files (86,020,430,015 bytes); the preceding 24h
logs showed download activity without authentication failures or errors. A later read-only check
found the service still active. This confirms the initial download, not the next daily cycle.
- [x] Declare HEIC decoding for Fedora graphical desktops without converting the originals on Atlas.
The Fedora role installs RPM Fusion Free with a pinned signing-key fingerprint and
`libheif-freeworld` on Ikaros and Nymph. The package was confirmed installed on Ikaros on
2026-10-03; Nymph deployment and an actual image-opening test were not observed.
- [ ] Validate Atlas iCloudPD filesystem/SELinux/SMB access, the next daily sync, ZFS/Borg/USB
backup inclusion, and isolated restore of photos and private state. A recursive hourly snapshot
of `zpool/archive` exists after ingestion, but no iCloudPD-specific backup version or restore
has been verified. The first monthly scrub remains a separate open data-protection check.
## Prometheus NPM Quadlet cutover
- [x] Stage a rootful NPM Quadlet using the exact running image and the existing data/certificate
mounts, bridge subnet, public HTTP/HTTPS ports, and loopback-only administration port.
The generated service depends on `server-web-network.service` and is wanted by `multi-user.target`.
- [x] Take and verify the stopped-source export before switching owners. Version
`20261003T091009Z` was pulled to Atlas and its NPM SQLite database checked in isolation.
- [x] Cut over NPM to `prometheus-npm.service` on 2026-10-03. The legacy Compose unit is inactive
and disabled; the Quadlet is active with zero recorded restarts. Public Gitea and Syncthing
HTTPS returned 200 with valid TLS, while public TCP/81 remained unreachable.
- [x] Validate the post-cutover backup path. The export and Atlas pull published
`20261003T091633Z`; checksum, SQLite `quick_check`, ten proxy hosts, six certificate records,
both Quadlet files were present, and the complete Let's Encrypt tree (70 regular files plus
12 symlinks) matched the live data. A targeted normal Ansible run changed nothing. Details and rollback
boundaries are in `docs/prometheus-npm-quadlet.md`.
- [x] Remove only unused Gitea, Navidrome and PostgreSQL images with opt-in
Ansible tasks on 2026-10-03. Second run changed nothing; NPM stayed active
with zero restarts, HTTP/HTTPS passed, backup timer and SSH proxy stayed active.
This image-only step preserved data and fallback; the later approved deletion is tracked below. Validation:
`ansible-playbook ansible/site.yml --limit prometheus --tags server_image_cleanup --check --diff -e server_legacy_image_cleanup=true`
- [x] Complete explicitly approved old-data and Compose fallback removal on 2026-10-03.
Backup paths and mount dependencies were reconciled before deletion; repeat cleanup changed
nothing. The normal Compose/template/helper check did not recreate retired files.
A separately approved manual export/pull published `20261003T112906Z`; checksum and isolated
SQLite restore passed with ten proxy hosts and both Quadlet definitions. NPM, primary HTTPS,
WireGuard, SSH proxy and backup timer remained healthy; existing backup archives were preserved.
- [x] Retire the unused secondary Gitea hostname `git.ov-ad3410.infomaniak.ch`
on 2026-10-03. Its NPM Proxy Host was already soft-deleted and had no
associated certificate. Its Ansible domain and runtime override were removed;
nginx -t and reload passed without restarting NPM. Primary HTTPS returned 200
with valid TLS. At that step only `git.fscotto.duckdns.org` remained declared;
the subsequent domain transition and operator-confirmed cleanup are tracked above.
- [ ] Observe the first scheduled export and Atlas pull after the cutover; the manual end-to-end
cycle passed, but the next unattended cycle has not yet occurred.
## Cerberus Management Node (Deferred)
`cerberus` is postponed until the office in the new house is physically set up. It is not an inventory
@@ -337,5 +537,5 @@ validated exports of older historical data will use a dedicated Atlas NFS datase
`/etc/resolv.conf` linked to `/run/systemd/resolve/resolv.conf`. LAN clients may use AdGuard, but
Aegis must use the independent upstream DNS declared by `aegis_host_dns_servers` so Greenboot does
not depend on the AdGuard container during startup.
- iCloudPD requires post-deployment interactive MFA initialization; its cookie/configuration state is
persisted in `/var/lib/icloudpd/config`.
- Aegis iCloudPD has been retired and is no longer managed by this role. Its service, Quadlet,
data, and MFA state were removed with the operator's explicit authorization.

View File

@@ -182,6 +182,10 @@ Le applicazioni Windows sono installate e gestite manualmente; il profilo WSL no
## Server
La migrazione dei servizi pubblici a `fscotto.co`, la gestione Ansible
degli URL Gitea e i passaggi ancora aperti per ritirare DuckDNS sono in
[`docs/domain-fscotto-co.md`](docs/domain-fscotto-co.md).
Sistema operativo:
- Rocky Linux 9
@@ -201,28 +205,37 @@ Lo stato attuale del profilo server include:
- installazione pacchetti Rocky via DNF, EPEL e CRB
- installazione di Podman e podman-compose
- abilitazione dei servizi systemd dichiarati in inventory/group vars
- copia dei dotfiles server e rendering del `docker-compose.yml` per Nginx Proxy Manager e Gitea,
piu l'unita `podman-compose-server` (attivazione manuale)
- copia dei dotfiles server e rendering del Quadlet rootful `prometheus-npm.service` per Nginx Proxy
Manager; il vecchio fallback Compose è stato rimosso con autorizzazione esplicita
- attivazione di firewalld con SSH, Cockpit (`9090/tcp`), HTTP e HTTPS abilitati
- Syncthing escluso dal profilo server Rocky
Il Compose desiderato su Prometheus non include piu Navidrome ne il database PostgreSQL obsoleto.
Navidrome e Syncthing appartengono ad Atlas; Navidrome ufficiale usa invece SQLite. Il profilo non
arresta o rimuove automaticamente eventuali container legacy e non elimina `/opt/postgres/data`.
Il 2026-10-03 la pulizia opt-in autorizzata ha rimosso dati e immagini precedenti di Gitea,
Navidrome e PostgreSQL, directory obsolete vuote, helper finale Gitea e fallback Compose NPM.
I servizi migrati restano su Atlas. `server_legacy_stack_retired: true` evita che i normali task
ricreino i residui; la cancellazione richiede `--tags server_legacy_cleanup` e
`-e server_legacy_cleanup=true`. NPM attivo e archivi di backup restano intatti.
Export, pull Atlas e restore SQLite isolato post-pulizia sono riusciti; il primo ciclo automatico
resta da osservare. Evidenze e confini del recovery:
[`docs/prometheus-npm-quadlet.md`](docs/prometheus-npm-quadlet.md).
Nginx Proxy Manager pubblica solo `80/tcp` e `443/tcp`; la sua interfaccia di amministrazione e
associata a `127.0.0.1:81` ed e raggiungibile da Ikaros o Nymph con l'alias Bash `npm-tunnel`.
Nextcloud resta disabilitato e il profilo non crea directory `/srv/nextcloud`.
La fase 1 su Atlas non modifica questo deployment NPM ne i suoi dati persistenti. Dopo aver attivato
WireGuard e i servizi Atlas, configurare i proxy host NPM correnti con upstream Navidrome
`http://10.0.0.2:4533` e upstream per la GUI Syncthing `http://10.0.0.2:8384`. Solo la GUI web di
Syncthing usa NPM; il traffico di sincronizzazione resta sulle porte native pubblicate esplicitamente solo
sull'indirizzo WireGuard di Atlas. Configurare l'autenticazione Syncthing e una policy di accesso NPM adeguata prima di pubblicare la GUI.
La fase 1 su Atlas non modifica i dati persistenti NPM. I proxy host NPM usano gli upstream LAN
`http://192.168.178.55:4533` per Navidrome e `http://192.168.178.55:8384` per la GUI Syncthing;
Prometheus li raggiunge attraverso Aegis come gateway WireGuard. Solo la GUI web di Syncthing usa
NPM; il traffico di sincronizzazione resta sulle porte native esposte sulla LAN dichiarata.
Mantenere l'autenticazione Syncthing e una policy di accesso NPM adeguata.
### DuckDNS
`profile_server` genera `~/duckdns/duck.sh` con permessi `0700`, mantenendo il percorso dello
`server_duckdns_enabled: false` disabilita il provisioning su Prometheus, che usa IP statico
e `fscotto.co`. Updater, log e cron ogni cinque minuti sono stati rimossi una sola volta;
non restano task o flag di pulizia. Il nome DuckDNS esterno e il token Vault restano invariati.
Sui server con `server_duckdns_enabled: true`, `profile_server` genera `~/duckdns/duck.sh` con permessi `0700`, mantenendo il percorso dello
script e `duck.log`. Definire `server_duckdns_domain` negli host vars del server e salvare il
**nuovo token rigenerato** in `vault_duckdns_token`, nel Vault cifrato `secrets/vault.yml`
(`ansible-vault edit secrets/vault.yml`) oppure negli override non versionati `secrets/vault.local.yml`.
@@ -311,7 +324,12 @@ l'interfaccia amministrativa resta su `127.0.0.1:81`, raggiungibile via tunnel S
Atlas ospita temporaneamente Navidrome e Syncthing rootless fino alla sostituzione con Uranus. I
servizi sono inizializzati **ex novo**, senza migrare lo stato precedente, rispettivamente sotto
`/zpool/services/data/navidrome` e `/zpool/services/data/syncthing`; la musica in
`/zpool/media/music` viene popolata separatamente. Sono vincolati all'indirizzo LAN di Atlas
`/zpool/media/music` è stata popolata separatamente da `/zpool/archive/Music` il 2026-09-30;
Navidrome ha completato la scansione. Il timer rootless `atlas-music-sync.timer` copia i file nuovi
o modificati ogni giorno alle 00:45 Europe/Rome, senza eliminare quelli presenti solo nella
destinazione; entrambi i dataset ZFS devono essere montati. La prima esecuzione schedulata è
riuscita il 2026-10-02. Alcune playlist originali contengono
ancora vecchi percorsi Windows. I servizi sono vincolati all'indirizzo LAN di Atlas
(`192.168.178.55`), mai a WireGuard. `wireguard_overlay` collega invece Prometheus (`10.0.0.1`)
e Aegis (`10.0.0.2`): le chiavi private restano sui rispettivi host e Ansible scambia solo le pubbliche.
Prometheus apre `51820/udp`; Aegis inoltra soltanto il traffico overlay→LAN dichiarato e applica
@@ -322,6 +340,14 @@ alla LAN. Dopo la verifica dei servizi, configurare manualmente i Proxy Host NPM
negli `AllowedIPs`; aggiungere la VIP Uranus quando esisterà. Dopo il reload di firewalld, Ansible
ricarica le reti Podman rootful di Prometheus per conservare DNS e connettività del proxy.
La migrazione Gitea da Prometheus ad Atlas è descritta in
[`docs/atlas-gitea-migration.md`](docs/atlas-gitea-migration.md). Gitea usa un Quadlet rootless
di `admin` su un dataset dedicato; l'immagine derivata mantiene UID/GID 1000 ma chiama l'utente
interno `gitea`. NPM resta su Prometheus e l'HTTPS pubblico primario serve Atlas. L'SSH pubblico
su TCP/2222 autentica la chiave `ikaros` e un `git ls-remote` è riuscito; l'operatore ha
confermato pull e push SSH. Login e scrittura Git via HTTPS sono stati confermati il 2026-10-03. I dati sorgente restano
conservati su Prometheus senza avviarne il vecchio container.
Validare il gateway con:
```bash
@@ -485,7 +511,7 @@ etichettata di 45Drives Alerts usare
### Timer systemd di Atlas
Tutti i nove timer gestiti sono abilitati. Gli orari sono locali ad Atlas (`Europe/Rome`); Borg e
Tutti i dieci timer gestiti sono abilitati. Gli orari sono locali ad Atlas (`Europe/Rome`); Borg e
monitoraggio aggiungono il ritardo casuale indicato. Tutti hanno `Persistent=true`: un evento perso
viene recuperato quando il timer torna attivo.
@@ -500,10 +526,13 @@ viene recuperato quando il timer torna attivo.
| `atlas-borg-check.timer` | `*-*-15 06:00:00` — giorno 15 alle 06:00, più 0–30 min casuali | Controllo repository Borg |
| `atlas-usb-reminder.timer` | `Sat *-*-01..07 10:00:00 Europe/Rome` — primo sabato alle 10:00 | Solo promemoria 45Drives Alerts |
| `atlas-health-monitor.timer` | `*:0/30` — ogni mezz'ora, più 0–5 min casuali | Controlli di salute in sola lettura |
| `atlas-prometheus-pull.timer` | `*-*-* 03:00:00 Europe/Rome` — ogni giorno alle 03:00 | Pull e verifica del backup preparato su Prometheus |
`atlas-usb-backup.service` **non ha timer** e va avviato manualmente. Il timer del fornitore
`zfs-scrub-weekly@zpool.timer` è disabilitato a favore dello scrub mensile. Il futuro pull del backup
Prometheus non ha ancora un timer, perché non è implementato. Durante un backup Borg attivo,
`zfs-scrub-weekly@zpool.timer` è disabilitato a favore dello scrub mensile. Il timer di preparazione
su Prometheus è attivo alle 02:00 Europe/Rome; il primo ciclo pianificato è riuscito il 2026-10-01.
Un export, pull e ripristino temporaneo post-cutover NPM Quadlet sono riusciti il 2026-10-03;
il primo ciclo pianificato dopo quel cutover resta da osservare. Durante un backup Borg attivo,
`systemctl list-timers` può mostrare `-` per il prossimo evento senza che il timer sia disabilitato.
Per vedere la pianificazione corrente: `systemctl list-timers --all` su Atlas.
@@ -512,14 +541,20 @@ della protezione dei dati: richiede storage applicativo, database e cache separa
pubblicazione solo tramite NPM e Aegis, procedure di backup, aggiornamento e migrazione. Non
distribuirlo prima di completare la checklist di protezione dei dati.
La destinazione futura per l'importazione foto iCloud è Atlas, non Aegis. Dopo la validazione dei
backup, pianificare una migrazione esplicita di iCloudPD con foto sotto `/zpool/archive/Pictures` e
stato applicativo/MFA fuori da `Archive`; testare permessi, SELinux, backup e restore prima del
cutover. L'attuale iCloudPD su Aegis e l'export NFS Photobook restano configurati fino
all'approvazione e alla verifica di questa migrazione separata. Anche il servizio Atlas sarà
temporaneo in attesa di Uranus.
Atlas è la destinazione dichiarata per iCloudPD. Ansible gestisce dataset, Quadlet rootless e
`icloudpd.conf` privato con Apple ID dal Vault: foto in `/zpool/archive/Pictures/iCloudPD`,
stato in `zpool/services/data/icloudpd`. Il primo avvio è stato manuale; password e MFA restano
gestiti interattivamente, senza avvio automatico al boot. L'inizializzazione è stata completata e
il download iniziale di foto e video è terminato il 2026-10-03. Su Aegis
il servizio, il Quadlet e `/var/lib/icloudpd` sono stati rimossi e verificati; il playbook Aegis
non contiene più task iCloudPD. L'accesso SMB e il ripristino dai backup dei nuovi dati restano
da verificare. L'export NFS Photobook resta
invariato. Dettagli in [`docs/atlas-icloudpd-migration.md`](docs/atlas-icloudpd-migration.md).
Il pull dei backup di Prometheus e i test completi di disaster recovery restano da fare. Il backlog
Il primo ciclo pianificato del backup di Prometheus e una prova di disaster recovery a dimensione reale
restano da verificare. Il 2026-09-30 una VM Rocky isolata ha superato ricostruzione OS con Ansible,
import del pool RAIDZ2 fittizio e ripristino da snapshot; RPO 24 ore/RTO 72 ore restano obiettivi
provvisori, non tempi misurati. Dettagli e limiti sono in `docs/atlas-recovery.md`. Il backlog
prioritizzato è in `AGENTS.md`.
---
@@ -617,8 +652,8 @@ Questo significa che, allo stato attuale:
- `deadalus` riceve il profilo Fedora WSL tramite play dev dedicati
- il server Rocky (`prometheus`) e gestito con pacchetti, servizi, dotfiles server e firewalld
- il NAS Rocky (`atlas`) usa un pool ZFS gia esistente, condivisioni NFSv4/SMB limitate alla LAN e Cockpit/45Drives
- lo stack Compose server include soltanto `gitea` e `nginx-proxy-manager`; Navidrome e Syncthing
della fase 1 sono Quadlet rootless su Atlas
- NPM è un Quadlet rootful su Prometheus, mentre Gitea, Navidrome e Syncthing sono Quadlet
rootless su Atlas; il fallback Compose server è stato rimosso
# Dotfiles
@@ -724,9 +759,10 @@ ansible-playbook ansible/site.yml --limit <host> --tags <tag1>,<tag2> --check --
ansible-playbook ansible/site.yml --limit <host> --start-at-task "<task name>" --check --diff
ansible-lint ansible/roles/<role>
yamllint ansible/path/to/file.yml
podman-compose -f /opt/docker/server/docker-compose.yml config
ansible-playbook ansible/site.yml --limit prometheus --tags server_legacy_cleanup --check --diff -e server_legacy_cleanup=true
ansible-playbook ansible/site.yml --limit atlas --tags storage,sharing,containers --check --diff
ansible-playbook ansible/site.yml --limit atlas --tags backend_phase1 --check --diff
ansible-playbook ansible/site.yml --limit atlas --tags music_sync --check --diff
```
## Tag supportati dal playbook

View File

@@ -125,16 +125,26 @@ That gives it Fedora packages through DNF, Docker from the official repository,
## Server
The public service domain transition to `fscotto.co`, Gitea canonical URL
management, and remaining DuckDNS retirement steps are documented in
[`docs/domain-fscotto-co.md`](docs/domain-fscotto-co.md).
`prometheus` is the Rocky Linux 9 server. It has no graphical environment and gets server-specific
dotfiles and templates. The profile provisions configuration only: it does not transfer data, start
the Compose stack, update DNS, or perform a cutover.
dotfiles and templates. The profile does not transfer application data, update DNS, or perform an
implicit service cutover.
The server profile installs platform-specific packages, Podman and podman-compose, declared systemd
services, and firewalld. The manually activated `podman-compose-server` unit contains the existing
Nginx Proxy Manager and Gitea services. The desired Compose file no longer includes Navidrome,
Syncthing, or the obsolete Navidrome PostgreSQL database; their temporary Atlas deployment is managed
by `profile_backend_phase1`. Applying the profile does not stop or remove legacy containers and does
not delete `/opt/postgres/data`.
services, and firewalld. Nginx Proxy Manager runs as the rootful `prometheus-npm.service` Quadlet.
On 2026-10-03 the operator-approved opt-in cleanup removed old Gitea, Navidrome and PostgreSQL
data/images, empty legacy directories, the Gitea final-export helper and the Compose rollback files.
The migrated services stay on Atlas. `server_legacy_stack_retired: true` prevents normal runs from
recreating retired files. Data deletion requires `--tags server_legacy_cleanup` and
`-e server_legacy_cleanup=true`; image-only cleanup has its own `server_image_cleanup` tag and flag.
Active NPM resources and existing backup archives remain preserved.
The post-cleanup export/pull and isolated SQLite restore passed; the first unattended cycle remains
pending. Evidence and recovery boundaries:
[`docs/prometheus-npm-quadlet.md`](docs/prometheus-npm-quadlet.md).
Firewalld enables SSH, Cockpit (`9090/tcp`), HTTP and HTTPS. Nginx Proxy Manager publishes only
`80/tcp` and `443/tcp`; its administration interface is bound to `127.0.0.1:81` and can be reached
@@ -161,7 +171,11 @@ Prometheus authorizes its declared SSH public keys through separate files below
### DuckDNS
`profile_server` renders `~/duckdns/duck.sh` with mode `0700`, keeping the existing updater path
`server_duckdns_enabled: false` disables provisioning on Prometheus, which uses its static IP
and `fscotto.co`. The local updater, log and five-minute cron job were removed once;
no cleanup tasks or flags remain. The external DuckDNS name and Vault token remain untouched.
For servers with `server_duckdns_enabled: true`, `profile_server` renders `~/duckdns/duck.sh` with mode `0700`, keeping the existing updater path
and `duck.log`. Set `server_duckdns_domain` in the server's host vars and store the **rotated**
`vault_duckdns_token` in encrypted `secrets/vault.yml` (using `ansible-vault edit secrets/vault.yml`)
or untracked `secrets/vault.local.yml`. Never commit the rendered script or put the token on a
@@ -210,8 +224,8 @@ ansible/bootstrap/generate-aegis-ign.sh --write IMAGE DEVICE
```
The controller manages it remotely as `pi@aegis`; unlike local desktop profiles, Aegis is
intentionally an SSH inventory target. `profile_aegis` manages rootful Podman Quadlets for AdGuard
Home and iCloudPD, persistent data under `/var/lib`, the Podman auto-update timer, LAN-restricted
intentionally an SSH inventory target. `profile_aegis` manages a rootful Podman Quadlet for AdGuard
Home, its persistent data under `/var/lib`, the Podman auto-update timer, LAN-restricted
firewalld rules, SSH key-only access for `pi`, the `nfs-utils` and `wireguard-tools` rpm-ostree layers,
and `wake-ikaros`. `wireguard_overlay` makes Aegis the internal endpoint and LAN gateway for Prometheus:
it enables persistent IPv4 forwarding, installs a scoped WireGuard-to-LAN firewalld policy, and source-NATs
@@ -224,9 +238,8 @@ opened and closed manually during initial setup. The profile disables the local
stub and points `/etc/resolv.conf` to its full resolver data, freeing port 53 for AdGuard. LAN clients
may use AdGuard on Aegis, while Aegis itself uses the independent upstream DNS declared by
`aegis_host_dns_servers`; this prevents Greenboot from depending on the AdGuard container during
startup. Reboot Aegis after changing its NetworkManager DNS profile. Define
`vault_aegis_icloudpd_apple_id` in Vault before applying it. iCloudPD still requires interactive MFA
initialization after its first deployment.
startup. Reboot Aegis after changing its NetworkManager DNS profile. iCloudPD was retired from Aegis;
the Aegis role no longer contains iCloudPD tasks. Atlas iCloudPD config is Vault-backed; MFA is manual.
New Aegis images create the `admin` account in Butane. Before configuring a newly imaged node, run its
first playbook execution with `-e ansible_user=admin`; the SSH hardening role then permits that same
@@ -296,7 +309,20 @@ Atlas temporarily hosts rootless Navidrome and Syncthing until Uranus replaces t
Atlas' LAN address (`192.168.178.55`); WireGuard remains exclusively between Prometheus (`10.0.0.1`)
and Aegis (`10.0.0.2`). Their state is initialized ex novo in `/zpool/services/data/navidrome` and
`/zpool/services/data/syncthing`; no source application state is migrated. The music library at
`/zpool/media/music` is populated separately.
`/zpool/media/music` was populated separately from `/zpool/archive/Music` on 2026-09-30;
Navidrome completed its library scan. The rootless `atlas-music-sync.timer` copies new and changed
files daily at 00:45 Europe/Rome, without deleting destination-only files. Both ZFS datasets must
be mounted. Its first scheduled run succeeded on 2026-10-02. Some source playlists still contain
obsolete Windows paths.
The Gitea move from Prometheus to Atlas is tracked in
[`docs/atlas-gitea-migration.md`](docs/atlas-gitea-migration.md). The final consistent copy runs in
Atlas' dedicated dataset under `admin`'s rootless user Quadlet. Its pinned derived image uses an
internal Unix user named `gitea` (UID/GID 1000), while clone URLs keep `git@`. NPM remains on Prometheus and the primary
public HTTPS route serves Atlas. Public SSH/2222 authenticates the `ikaros` key and serves
`git ls-remote`; the operator also confirmed SSH pull and push. HTTPS login and Git writes were
confirmed on 2026-10-03. The old Gitea data remains on Prometheus, but its container
is absent from the desired stack.
The separate `wireguard_overlay` role manages `wg0` between Prometheus (`10.0.0.1`) and Aegis
(`10.0.0.2`), generating private keys once on their respective hosts and exchanging only public keys
@@ -500,7 +526,7 @@ monitoring. For a labelled 45Drives Alerts delivery test, use
### Atlas systemd timers
All nine managed timers below are enabled. Times are local to Atlas (`Europe/Rome`); Borg and monitoring
All ten managed timers below are enabled. Times are local to Atlas (`Europe/Rome`); Borg and monitoring
add the indicated randomized delay. Every timer has `Persistent=true`, so a missed calendar run is
scheduled after the timer becomes active again.
@@ -515,10 +541,13 @@ scheduled after the timer becomes active again.
| `atlas-borg-check.timer` | `*-*-15 06:00:00` — 15th of the month at 06:00, plus 0–30 min random delay | Borg repository check |
| `atlas-usb-reminder.timer` | `Sat *-*-01..07 10:00:00 Europe/Rome` — first Saturday at 10:00 | 45Drives Alerts reminder only |
| `atlas-health-monitor.timer` | `*:0/30` — every half-hour, plus 0–5 min random delay | Read-only health checks |
| `atlas-prometheus-pull.timer` | `*-*-* 03:00:00 Europe/Rome` — daily at 03:00 | Pull and verify the prepared Prometheus backup |
`atlas-usb-backup.service` has **no timer**: the encrypted USB backup must be started manually.
The vendor's `zfs-scrub-weekly@zpool.timer` is intentionally disabled in favor of the monthly scrub.
The future Prometheus backup pull has no timer yet because that workflow is not implemented. While a
The Prometheus export timer runs at 02:00 Europe/Rome. Its first scheduled export and Atlas pull
passed on 2026-10-01; a manual post-NPM-Quadlet export, pull, and temporary restore passed on
2026-10-03. The first scheduled cycle after that cutover remains to be observed. While a
Borg backup is still running, `systemctl list-timers` may show `-` for its next trigger; this does not
mean the timer has been disabled. Inspect the current schedule on Atlas with
`systemctl list-timers --all`.
@@ -528,15 +557,30 @@ declared persistent application, database, and cache storage, Vault-backed crede
publishing through Aegis, and defined backup, upgrade, and eventual migration procedures. Do not deploy
it before the data-protection checklist is complete.
The desired future iCloud photo-ingestion host is Atlas, not Aegis. After data-protection validation,
plan an explicit iCloudPD migration with photos under `/zpool/archive/Pictures` and application/MFA
state outside `Archive`, then test permissions, SELinux, backups and recovery before cutting over.
The current Aegis iCloudPD service and Atlas Photobook NFS export remain configured until that
separate migration is approved and validated; the eventual Atlas service is temporary until Uranus.
Atlas is the declared iCloud photo-ingestion host. Ansible manages the rootless Quadlet, a private
Vault-backed `icloudpd.conf`, photos under `/zpool/archive/Pictures/iCloudPD`, and separate state in
`zpool/services/data/icloudpd`. The service was started manually; Ansible does not enable automatic
startup or manage the password and MFA keyring. The operator initialized MFA interactively; on
2026-10-03 the initial photo/video download completed. Aegis iCloudPD, including its service data,
has been removed and verified; the Aegis role no longer manages it. Backup/restore and SMB access
for the new data remain unverified. The Photobook NFS export remains untouched. See
[`docs/atlas-icloudpd-migration.md`](docs/atlas-icloudpd-migration.md).
Prometheus backup pulls and full disaster-recovery tests remain follow-up work. The prioritized
The first scheduled Prometheus backup runs and production-size disaster-recovery tests remain follow-up work. The prioritized
operational backlog is kept in `AGENTS.md`.
Priority 2 procedures and decisions are recorded in
[`docs/atlas-recovery.md`](docs/atlas-recovery.md),
[`docs/atlas-updates.md`](docs/atlas-updates.md), and
[`docs/atlas-sharing-decision.md`](docs/atlas-sharing-decision.md).
The provisional Atlas recovery objectives are RPO 24 hours and RTO 72 hours;
an isolated small-VM OS rebuild, pool import, Ansible reapplication, and
snapshot restore passed, but full-size recovery time is unmeasured. `Archive` (SMB) and
`photobook` (NFS) remain deliberately separate.
The Prometheus pull architecture and manual export/pull/restore evidence are in
[`docs/prometheus-backup.md`](docs/prometheus-backup.md). Both daily timers are
enabled; their first scheduled runs remain to be verified.
## How layering works
A host can intentionally belong to more than one inventory group. The final configuration is the combination of the host and its groups, not a one-host/one-play mapping.
@@ -706,8 +750,9 @@ ansible-playbook ansible/site.yml --limit <host> --tags <tag1>,<tag2> --check --
ansible-playbook ansible/site.yml --limit <host> --start-at-task "<task name>" --check --diff
ansible-lint ansible/roles/<role>
yamllint ansible/path/to/file.yml
podman-compose -f /opt/docker/server/docker-compose.yml config
ansible-playbook ansible/site.yml --limit prometheus --tags server_legacy_cleanup --check --diff -e server_legacy_cleanup=true
ansible-playbook ansible/site.yml --limit atlas --tags storage,sharing,containers --check --diff
ansible-playbook ansible/site.yml --limit atlas --tags music_sync --check --diff
```
## Tags

View File

@@ -6,6 +6,11 @@ effective_username: "{{ server_username }}"
effective_user_group: "{{ server_user_group }}"
effective_user_home: "{{ server_user_home }}"
server_container_stack_dir: /opt/docker/server
server_npm_quadlet_stage: false
server_npm_quadlet_cutover: false
server_legacy_stack_retired: false
server_legacy_cleanup: false
server_duckdns_enabled: true
ai_agents: {}
vim_plugins_enabled: false
@@ -80,5 +85,37 @@ server_sshd_settings:
server_sshd_allow_users:
- "{{ server_username }}"
server_backup_export_enabled: false
server_backup_username: prometheus-backup
server_backup_public_key_name: atlas-pull
server_backup_export_root: /var/lib/prometheus-backup-export
server_backup_rrsync_path: /usr/share/doc/rsync/support/rrsync
server_backup_export_calendar: "*-*-* 02:00:00 Europe/Rome"
server_backup_export_start_timer: false
# Explicit Gitea cutover helper: installed separately from any outage action.
server_gitea_cutover_tools_enabled: false
server_gitea_final_export: false
server_gitea_on_atlas: false
server_gitea_atlas_address: "{{ hostvars['atlas'].ansible_host }}"
server_gitea_npm_domains: []
server_gitea_ssh_public_port: 2222
server_gitea_ssh_target_port: 2222
server_backup_export_source_keep: 3
server_backup_export_paths: >-
{{ ['opt/npm/data', 'opt/npm/letsencrypt']
+ ([] if server_gitea_on_atlas | bool else ['opt/gitea/data', 'home/git/.ssh'])
+ ([] if server_legacy_stack_retired | bool else
['opt/docker/server/docker-compose.yml',
'etc/systemd/system/podman-compose-server.service'])
+ (['etc/containers/systemd/prometheus-npm.container',
'etc/containers/systemd/server-web.network']
if server_npm_quadlet_stage | bool else [])
+ ['etc/ssh/sshd_config', 'etc/ssh/sshd_config.d',
'etc/firewalld', 'etc/wireguard/wg0.conf'] }}
server_backup_export_excludes: >-
{{ ['opt/npm/data/logs']
+ ([] if server_gitea_on_atlas | bool else
['opt/gitea/data/gitea/log', 'opt/gitea/data/gitea/tmp',
'opt/gitea/data/gitea/sessions', 'opt/gitea/data/gitea/indexers']) }}
server_ssh_authorized_keys: []
server_ssh_authorized_key_directory: "{{ server_user_home }}/.ssh/authorized_keys.d"

View File

@@ -42,5 +42,3 @@ aegis_ssh_authorized_keys:
key: "ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIEH/7GJfGt0ZVmKeEzceoFkFkeCXFryKK9vAbaip+HCx nymph"
- name: siren
key: "ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIA95wYlzpfN3rjUhpMeP4KHn8I6ZrjQXoDTgwgRIa++b siren"
aegis_icloudpd_apple_id: "{{ vault_aegis_icloudpd_apple_id | default('') }}"

View File

@@ -49,6 +49,11 @@ atlas_zfs_backup_reservation: 500G
atlas_zfs_dataset_photobook: media/photobook
atlas_mount_root: /zpool
atlas_manage_storage: true
# Rootless Gitea was restored from the stopped-source export before production activation.
atlas_manage_gitea: true
atlas_gitea_production_enabled: true
atlas_gitea_public_domain: git.fscotto.co
atlas_prometheus_pull_start_timer: true
atlas_manage_zfs_snapshots: true
atlas_zfs_snapshot_prefix: atlas-auto
atlas_zfs_snapshot_policies:
@@ -91,6 +96,12 @@ atlas_usb_backup_mapper_name: zpool-backup
atlas_manage_usb_reminder: true
atlas_usb_reminder_calendar: "Sat *-*-01..07 10:00:00 Europe/Rome"
atlas_manage_monitoring: true
atlas_manage_prometheus_backup_pull: true
# Prometheus ED25519 host key read through the controller's strict SSH trust on 2026-09-30.
# Fingerprint: SHA256:rfedk7DHI9mLB3UHk/4F3HHlSIiswtCAFsAXvfh6iXk
atlas_prometheus_ssh_host_key: >-
179.237.102.172 ssh-ed25519
AAAAC3NzaC1lZDI1NTE5AAAAIC4b+QXlPupoEx71W9NKs9tTeYjBqTkVMqbGB97nMNWv
# Physical pool disks and the system NVMe; the disconnected USB disk is intentionally excluded.
atlas_monitor_smart_devices:
- { name: pool-1, path: "{{ atlas_zpool_disks[0] }}", warning_c: 50, critical_c: 55 }
@@ -130,8 +141,8 @@ atlas_monitor_remote_capacity:
atlas_manage_sharing: true
atlas_manage_media_stack: false
# Planned after data-protection validation: move iCloudPD photo ingestion from
# Aegis to Atlas, with photos under /zpool/archive/Pictures and persistent
# application/MFA state outside Archive. Do not deploy or cut over yet.
# Aegis to Atlas, with photos under /zpool/archive/Pictures/iCloudPD and
# application/MFA state in a separate dataset. Do not deploy or cut over yet.
# WireGuard is retired on Atlas. These rootless services are a temporary home
# until Uranus replaces them.
@@ -141,6 +152,7 @@ backend_phase1_bind_address: "{{ ansible_host }}"
backend_phase1_firewalld_zone: "{{ atlas_firewalld_zone }}"
backend_phase1_npm_source_ip: "{{ atlas_aegis_ip }}"
backend_phase1_syncthing_native_subnet: "{{ atlas_lan_subnet }}"
backend_phase1_music_sync_enabled: true
rocky_manage_openzfs_repo: true
rocky_manage_syncthing_binary: false

View File

@@ -6,7 +6,28 @@ ansible_port: 22
ansible_ssh_private_key_file: /home/fscotto/.ssh/id_ed25519
server_username: rocky
server_legacy_stack_retired: true
# Destructive deletion runs only with an explicit extra-var and cleanup tag.
server_legacy_cleanup: false
# Explicit opt-in cleanup; no data, volumes, networks or NPM images are removed.
server_legacy_image_cleanup: false
server_legacy_images:
- docker.gitea.com/gitea:1.25.2
- docker.io/deluan/navidrome:latest
- docker.io/library/postgres:13
server_npm_quadlet_stage: true
server_npm_quadlet_image: docker.io/jc21/nginx-proxy-manager@sha256:52b2c59994f3d36acfcf70a1626f29734df0ed8c71bacc0269f78b6f939858bb
# The stopped-source export and live Quadlet cutover passed on 2026-10-03.
server_npm_quadlet_cutover: true
server_backup_export_enabled: true
server_backup_export_start_timer: true
# Install the final-copy helper only; it is never run by a normal playbook invocation.
server_gitea_cutover_tools_enabled: true
server_gitea_on_atlas: true
server_gitea_npm_domains:
- git.fscotto.duckdns.org
server_duckdns_domain: fscotto
server_duckdns_enabled: false
server_ssh_authorized_keys:
- name: ikaros
key: "ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAINrIxXjA3ffPwziKGR5gzc4gAoBehQPlnEMcXF4Wl0ZS ikaros"

View File

@@ -39,11 +39,41 @@
state: enabled
when: "'workstation_dev_wsl' in group_names"
- name: Install distribution signing keys for Fedora desktop codecs
tags: [packages, heic]
ansible.builtin.dnf:
name: distribution-gpg-keys
state: present
when: "'graphical_desktop' in group_names"
- name: Import RPM Fusion Free signing key for Fedora desktop codecs
tags: [packages, heic]
ansible.builtin.rpm_key:
key: /usr/share/distribution-gpg-keys/rpmfusion/RPM-GPG-KEY-rpmfusion-free-fedora-2020
fingerprint: E9A491A3DE247814E7E067EAE06F8ECDD651FF2E
state: present
when: "'graphical_desktop' in group_names"
- name: Enable RPM Fusion Free for Fedora desktop codecs
tags: [packages, heic]
ansible.builtin.dnf:
name: "https://download1.rpmfusion.org/free/fedora/rpmfusion-free-release-{{ ansible_facts['distribution_major_version'] }}.noarch.rpm"
state: present
when: "'graphical_desktop' in group_names"
- name: Refresh dnf package metadata
tags: [packages]
ansible.builtin.dnf:
update_cache: true
- name: Install HEIC decoder on Fedora desktops
tags: [packages, heic]
ansible.builtin.dnf:
name: libheif-freeworld
state: present
update_cache: true
when: "'graphical_desktop' in group_names"
- name: Install packages on Fedora
tags: [packages]
ansible.builtin.dnf:

View File

@@ -8,10 +8,6 @@ aegis_network_connection_uuid: ""
aegis_host_dns_servers: []
aegis_host_dns_search_domains: []
aegis_adguard_image: docker.io/adguard/adguardhome:latest
aegis_icloudpd_image: docker.io/boredazfcuk/icloudpd:latest
aegis_icloudpd_folder_structure: '{:%Y/%m/%d}'
aegis_icloudpd_synchronisation_interval: 86400
aegis_icloudpd_apple_id: ""
aegis_ikaros_mac_address: aa:bb:cc:dd:ee:ff
aegis_wol_port: 9

View File

@@ -9,13 +9,12 @@
name: sshd.service
state: reloaded
- name: Restart Aegis Quadlet services
- name: Restart Aegis AdGuard Quadlet
ansible.builtin.systemd:
name: "{{ item }}"
state: restarted
daemon_reload: true
loop:
- adguardhome.service
- icloudpd.service
loop_control:
label: "{{ item }}"

View File

@@ -13,14 +13,6 @@
msg: Reboot Aegis to activate the newly layered packages, then rerun the playbook.
when: aegis_layered_packages_result.needs_reboot | default(false)
- name: Require Aegis iCloudPD Apple ID
tags: [aegis, icloudpd]
ansible.builtin.assert:
that:
- aegis_icloudpd_apple_id | length > 0
fail_msg: Define vault_aegis_icloudpd_apple_id before applying the Aegis profile.
no_log: true
- name: Require completed Aegis network placeholders
tags: [aegis, dns, firewall, network, services]
ansible.builtin.assert:
@@ -119,8 +111,6 @@
loop:
- /var/lib/adguard/work
- /var/lib/adguard/conf
- /var/lib/icloudpd/data
- /var/lib/icloudpd/config
- name: Create Quadlet configuration directory
tags: [aegis, containers]
@@ -142,12 +132,9 @@
loop:
- src: adguardhome.container.j2
dest: adguardhome.container
- src: icloudpd.container.j2
dest: icloudpd.container
loop_control:
label: "{{ item.dest }}"
no_log: "{{ item.dest == 'icloudpd.container' }}"
notify: Restart Aegis Quadlet services
notify: Restart Aegis AdGuard Quadlet
- name: Create Aegis systemd-resolved configuration directory
tags: [aegis, adguard, dns, services]
@@ -168,7 +155,7 @@
mode: "0644"
notify:
- Restart Aegis systemd-resolved
- Restart Aegis Quadlet services
- Restart Aegis AdGuard Quadlet
- name: Point Aegis resolver at the full systemd-resolved configuration
tags: [aegis, adguard, dns, services]
@@ -361,7 +348,7 @@
group: root
mode: "0755"
- name: Enable Aegis Quadlet services and automatic updates
- name: Enable Aegis AdGuard Quadlet and automatic updates
tags: [aegis, containers, services]
ansible.builtin.systemd:
name: "{{ item }}"
@@ -370,7 +357,6 @@
daemon_reload: true
loop:
- adguardhome.service
- icloudpd.service
- podman-auto-update.timer
loop_control:
label: "{{ item }}"

View File

@@ -1,20 +0,0 @@
# Managed by Ansible. Do not edit manually.
[Unit]
Description=iCloud Photos Downloader
Wants=network-online.target
After=network-online.target
[Container]
Image={{ aegis_icloudpd_image }}
Environment=apple_id={{ aegis_icloudpd_apple_id }}
Environment=folder_structure={{ aegis_icloudpd_folder_structure }}
Environment=synchronisation_interval={{ aegis_icloudpd_synchronisation_interval }}
Volume=/var/lib/icloudpd/data:/home/root/iCloud:Z
Volume=/var/lib/icloudpd/config:/config:Z
AutoUpdate=registry
[Service]
Restart=always
[Install]
WantedBy=multi-user.target

View File

@@ -115,6 +115,15 @@ atlas_monitor_notifier: "{{ atlas_usb_reminder_notifier }}"
atlas_monitor_smart_devices: []
atlas_monitor_timers: []
atlas_monitor_failure_units: []
atlas_monitor_effective_timers: >-
{{ atlas_monitor_timers
+ ([{'name': 'atlas-prometheus-pull.timer', 'max_age_hours': 26}]
if atlas_manage_prometheus_backup_pull | bool and atlas_prometheus_pull_start_timer | bool
else []) }}
atlas_monitor_effective_failure_units: >-
{{ atlas_monitor_failure_units
+ (['atlas-prometheus-pull.service']
if atlas_manage_prometheus_backup_pull | bool else []) }}
atlas_monitor_remote_capacity: {}
atlas_monitor_pool_warning_percent: 80
atlas_monitor_pool_critical_percent: 90
@@ -141,8 +150,62 @@ atlas_music_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_music }}"
atlas_backup_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_backup }}"
atlas_host_backups_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_host_backups }}"
atlas_backup_prometheus_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_backup_prometheus }}"
atlas_manage_prometheus_backup_pull: false
atlas_prometheus_pull_ssh_dir: /etc/atlas-prometheus-pull
atlas_prometheus_pull_private_key_path: "{{ atlas_prometheus_pull_ssh_dir }}/id_ed25519"
atlas_prometheus_pull_known_hosts_path: "{{ atlas_prometheus_pull_ssh_dir }}/known_hosts"
atlas_prometheus_ssh_host_key: ""
atlas_prometheus_pull_source_user: prometheus-backup
atlas_prometheus_pull_source_port: 22
atlas_prometheus_pull_calendar: "*-*-* 03:00:00 Europe/Rome"
atlas_prometheus_pull_start_timer: false
atlas_prometheus_pull_keep_daily: 30
atlas_prometheus_pull_keep_weekly: 8
atlas_prometheus_pull_keep_monthly: 12
atlas_prometheus_pull_max_age_hours: 24
atlas_photobook_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_photobook }}"
# Rootless Gitea runs in admin's user manager; the image maps internal gitea to UID/GID 1000.
atlas_manage_gitea: false
atlas_gitea_username: "{{ atlas_admin_username }}"
atlas_gitea_group: "{{ atlas_admin_group }}"
atlas_gitea_uid: "{{ atlas_admin_uid }}"
atlas_gitea_gid: "{{ atlas_admin_gid }}"
atlas_gitea_home: "{{ atlas_admin_home }}"
atlas_gitea_container_uid: 1000
atlas_gitea_container_gid: 1000
atlas_gitea_legacy_username: gitea
atlas_gitea_legacy_uid: 1101
atlas_gitea_legacy_home: /var/lib/atlas-gitea
atlas_gitea_owner_migration: false
atlas_gitea_dataset: "{{ atlas_zfs_pool }}/services/data/gitea"
atlas_gitea_mountpoint: "{{ atlas_app_data_mountpoint }}/gitea"
atlas_gitea_quadlet_dir: "{{ atlas_gitea_home }}/.config/containers/systemd"
atlas_gitea_image: localhost/atlas-gitea:1.25.2-user-gitea-v1
atlas_gitea_image_build_dir: "{{ atlas_gitea_home }}/.local/share/atlas-gitea-image"
atlas_gitea_production_enabled: false
atlas_gitea_public_domain: ""
atlas_gitea_bind_address: "{{ ansible_host }}"
atlas_gitea_http_port: 3000
atlas_gitea_ssh_port: 2222
atlas_gitea_staging_bind_address: 127.0.0.1
atlas_gitea_staging_http_port: 3001
atlas_gitea_staging_ssh_port: 2223
atlas_gitea_restore_test: false
atlas_gitea_final_restore: false
atlas_gitea_restore_helper: /usr/local/libexec/atlas-gitea-restore-test
# Declare storage and an inactive Quadlet only. The operator supplies the
# private configuration, handles MFA, and starts the user service manually.
atlas_icloudpd_dataset: "{{ atlas_zfs_pool }}/services/data/icloudpd"
atlas_icloudpd_state_dir: "{{ atlas_app_data_mountpoint }}/icloudpd"
atlas_icloudpd_config_dir: "{{ atlas_icloudpd_state_dir }}/config"
atlas_icloudpd_photos_dir: "{{ atlas_archive_mountpoint }}/Pictures/iCloudPD"
atlas_icloudpd_image: >-
docker.io/boredazfcuk/icloudpd@sha256:9966c31ddf0b5b306ac2410b4edd5d626806d96e80c92b83cbb689972dc9389f
atlas_icloudpd_quadlet_dir: "{{ atlas_admin_home }}/.config/containers/systemd"
atlas_icloudpd_timezone: Europe/Rome
atlas_45drives_repo_url: https://repo.45drives.com/repofiles/rocky/45drives-enterprise.repo
atlas_45drives_repo_file: /etc/yum.repos.d/45drives-enterprise.repo
atlas_45drives_packages:

View File

@@ -0,0 +1,10 @@
FROM docker.gitea.com/gitea@sha256:f1943db2d2f1e447e857b3f0aee4ebb7b184500f86e5b80eae110fd435435906
# Preserve the official image's UID/GID, paths and entrypoint; change only the
# internal Unix identity. The host-side rootless owner is Atlas admin.
USER 0
RUN sed -i 's/^git:x:1000:1000:/gitea:x:1000:1000:/' /etc/passwd \
&& sed -i 's/^git:x:1000:/gitea:x:1000:/' /etc/group \
&& grep -q '^gitea:x:1000:1000:' /etc/passwd \
&& grep -q '^gitea:x:1000:' /etc/group
USER 1000:1000

View File

@@ -0,0 +1,249 @@
#!/usr/bin/python3
"""Rehearse a selective rootful-to-rootless Gitea restore, never a cutover."""
import argparse
import hashlib
import json
import os
from pathlib import Path, PurePosixPath
import re
import shutil
import sqlite3
import tarfile
import tempfile
SOURCE_PREFIX = PurePosixPath("opt/gitea/data")
HOST_KEYS = (
"ssh_host_ed25519_key",
"ssh_host_rsa_key",
"ssh_host_ecdsa_key",
)
SERVER_SETTINGS = {
"START_SSH_SERVER": "true",
"BUILTIN_SSH_SERVER_USER": "git",
"SSH_USER": "git",
"SSH_PORT": "2222",
"SSH_LISTEN_PORT": "2222",
"SSH_SERVER_HOST_KEYS": ", ".join(
f"/var/lib/gitea/ssh/{key}" for key in HOST_KEYS
),
}
def sha256(path):
digest = hashlib.sha256()
with path.open("rb") as stream:
for chunk in iter(lambda: stream.read(1024 * 1024), b""):
digest.update(chunk)
return digest.hexdigest()
def expected_digest(backup):
checksum = (backup / "payload.sha256").read_text().strip().split()
if len(checksum) != 2 or checksum[1] != "payload.tar":
raise ValueError("Unexpected Prometheus backup checksum manifest")
if not re.fullmatch(r"[0-9a-f]{64}", checksum[0]):
raise ValueError("Invalid Prometheus backup SHA-256")
return checksum[0]
def convert_config(config):
original = config.read_text()
output = []
section = ""
server_seen = set()
server_found = False
run_user_seen = False
def append_missing_server_settings():
for key, value in SERVER_SETTINGS.items():
if key not in server_seen:
output.append(f"{key} = {value}\n")
for line in original.splitlines(keepends=True):
match = re.match(r"^\s*\[([^]]+)\]\s*$", line)
if match:
if not run_user_seen:
output.append("RUN_USER = gitea\n")
run_user_seen = True
if section == "server":
append_missing_server_settings()
section = match.group(1).lower()
server_found |= section == "server"
output.append(line)
continue
setting = re.match(r"^(\s*)([A-Z_]+)(\s*=\s*)(.*?)(\r?\n?)$", line)
if setting and section == "" and setting.group(2) == "RUN_USER":
run_user_seen = True
line = f"{setting.group(1)}RUN_USER{setting.group(3)}gitea{setting.group(5)}"
elif setting and section == "server" and setting.group(2) in SERVER_SETTINGS:
key = setting.group(2)
server_seen.add(key)
line = f"{setting.group(1)}{key}{setting.group(3)}{SERVER_SETTINGS[key]}{setting.group(5)}"
else:
line = line.replace("/data/", "/var/lib/gitea/")
output.append(line)
if section == "server":
append_missing_server_settings()
if not server_found:
raise ValueError("Gitea server configuration missing")
config.write_text("".join(output))
config.chmod(0o600)
def extract_gitea(tar_path, staged_data):
count = 0
with tarfile.open(tar_path, mode="r") as archive:
for member in archive:
name = PurePosixPath(member.name)
if name == SOURCE_PREFIX:
continue
if SOURCE_PREFIX not in name.parents:
continue
relative = name.relative_to(SOURCE_PREFIX)
if not relative.parts or any(part in (".", "..") for part in relative.parts):
raise ValueError("Unsafe Gitea backup path")
if not (member.isdir() or member.isfile()):
raise ValueError("Unexpected Gitea backup member type")
destination = staged_data.joinpath(*relative.parts)
if member.isdir():
destination.mkdir(parents=True, exist_ok=True)
destination.chmod(0o700)
continue
destination.parent.mkdir(parents=True, exist_ok=True)
with archive.extractfile(member) as source, destination.open("xb") as target:
shutil.copyfileobj(source, target)
destination.chmod(member.mode & 0o777)
count += 1
if count == 0:
raise ValueError("No Gitea files in backup")
def validate(staged_data, staged_config):
database = staged_data / "gitea/gitea.db"
repositories = staged_data / "git/repositories"
if not database.is_file() or not repositories.is_dir():
raise ValueError("Missing SQLite database or Git repositories")
with sqlite3.connect(f"file:{database}?mode=ro", uri=True) as connection:
if connection.execute("PRAGMA quick_check").fetchone()[0] != "ok":
raise ValueError("Gitea SQLite quick_check failed")
if connection.execute("SELECT count(*) FROM repository").fetchone()[0] < 1:
raise ValueError("Gitea backup contains no repository records")
if not any(repositories.rglob("*.git")):
raise ValueError("Gitea backup contains no Git repository directories")
if not (staged_config / "app.ini").is_file():
raise ValueError("Gitea app.ini missing")
for name in HOST_KEYS:
if not (staged_data / "ssh" / name).is_file():
raise ValueError("Gitea SSH host key missing")
def chown_tree(root, uid, gid):
for directory, dirs, files in os.walk(root):
os.chown(directory, uid, gid)
for name in dirs + files:
os.chown(os.path.join(directory, name), uid, gid)
def replace_rehearsal(target, stage, digest, uid, gid):
previous_data = target / ".previous-rehearsal-data"
previous_config = target / ".previous-rehearsal-config"
if previous_data.exists() or previous_config.exists():
raise ValueError("An interrupted Gitea replacement needs manual recovery")
os.rename(target / "data", previous_data)
try:
os.rename(target / "config", previous_config)
os.rename(stage / "data", target / "data")
os.rename(stage / "config", target / "config")
final_marker = target / ".final-sha256"
final_marker.write_text(digest + "\n")
final_marker.chmod(0o600)
os.chown(final_marker, uid, gid)
(target / ".rehearsal-sha256").unlink()
except Exception:
for name, previous in (("data", previous_data), ("config", previous_config)):
current = target / name
if previous.exists():
if current.exists():
shutil.rmtree(current)
os.rename(previous, current)
(target / ".final-sha256").unlink(missing_ok=True)
raise
shutil.rmtree(previous_data)
shutil.rmtree(previous_config)
def main():
parser = argparse.ArgumentParser()
parser.add_argument("--backup", type=Path, required=True)
parser.add_argument("--target", type=Path, required=True)
parser.add_argument("--uid", type=int, required=True)
parser.add_argument("--gid", type=int, required=True)
parser.add_argument("--replace-rehearsal", action="store_true")
args = parser.parse_args()
backup = args.backup.resolve(strict=True)
target = args.target.resolve(strict=True)
if not str(backup).startswith("/zpool/backup/hosts/prometheus/snapshots/"):
raise ValueError("Refusing backup outside the Atlas Prometheus snapshots")
if str(target) != "/zpool/services/data/gitea":
raise ValueError("Refusing target outside the dedicated Gitea dataset")
if args.uid != 1000 or args.gid != 1000:
raise ValueError("Unexpected admin-owned Gitea account IDs")
expected = expected_digest(backup)
if sha256(backup / "payload.tar") != expected:
raise ValueError("Prometheus backup SHA-256 mismatch")
marker = target / (".final-sha256" if args.replace_rehearsal else ".rehearsal-sha256")
if marker.exists():
if marker.read_text().strip() != expected:
raise ValueError("A different Gitea restore already occupies this dataset")
validate(target / "data", target / "config")
print("unchanged")
return
if args.replace_rehearsal:
metadata = json.loads((backup / "metadata.json").read_text())
if metadata.get("purpose") != "gitea-cutover":
raise ValueError("Final restore requires an explicit Gitea cutover export")
if not (target / ".rehearsal-sha256").is_file():
raise ValueError("Only a marked rehearsal may be replaced")
if not all((target / name).is_dir() for name in ("data", "config")):
raise ValueError("Prepared Gitea volume paths are missing")
else:
if (target / ".final-sha256").exists():
raise ValueError("Refusing a rehearsal restore over final Gitea data")
for name in ("data", "config"):
directory = target / name
if not directory.is_dir() or any(directory.iterdir()):
raise ValueError("Gitea target is not empty; refusing overwrite")
with tempfile.TemporaryDirectory(prefix=".rehearsal-", dir=target) as temporary:
stage = Path(temporary)
staged_data = stage / "data"
staged_config = stage / "config"
staged_data.mkdir()
staged_config.mkdir()
extract_gitea(backup / "payload.tar", staged_data)
source_config = staged_data / "gitea/conf/app.ini"
if not source_config.is_file():
raise ValueError("Source Gitea app.ini missing")
shutil.copy2(source_config, staged_config / "app.ini")
source_config.unlink()
convert_config(staged_config / "app.ini")
validate(staged_data, staged_config)
chown_tree(stage, args.uid, args.gid)
if args.replace_rehearsal:
replace_rehearsal(target, stage, expected, args.uid, args.gid)
else:
for name in ("data", "config"):
(target / name).rmdir()
os.rename(stage / name, target / name)
marker.write_text(expected + "\n")
marker.chmod(0o600)
os.chown(marker, args.uid, args.gid)
print("restored")
if __name__ == "__main__":
main()

View File

@@ -0,0 +1,56 @@
#!/usr/bin/env python3
"""Prune only verified, named Prometheus backup versions after publication."""
import datetime as dt
import pathlib
import re
import shutil
import sys
def main() -> None:
if len(sys.argv) != 5:
raise SystemExit("Usage: atlas-prometheus-prune SNAPSHOTS DAILY WEEKLY MONTHLY")
root = pathlib.Path(sys.argv[1])
counts = [int(value) for value in sys.argv[2:]]
if not root.is_dir() or root.is_symlink() or min(counts) < 1:
raise SystemExit("Invalid backup directory or retention counts")
versions = []
for entry in root.iterdir():
if not entry.is_dir() or entry.is_symlink():
continue
if not re.fullmatch(r"[0-9]{8}T[0-9]{6}Z", entry.name):
continue
try:
when = dt.datetime.strptime(entry.name, "%Y%m%dT%H%M%SZ")
except ValueError:
continue
if not all((entry / name).is_file() for name in ("payload.tar", "payload.sha256", "metadata.json")):
continue
versions.append((when, entry))
versions.sort(reverse=True)
if not versions:
raise SystemExit("No published backup versions found; refusing to prune")
keep = {entry for _, entry in versions[: counts[0]]}
for count, key in (
(counts[1], lambda when: when.isocalendar()[:2]),
(counts[2], lambda when: (when.year, when.month)),
):
periods = set()
for when, entry in versions:
period = key(when)
if period in periods:
continue
periods.add(period)
keep.add(entry)
if len(periods) >= count:
break
for _, entry in versions:
if entry not in keep:
shutil.rmtree(entry)
if __name__ == "__main__":
main()

View File

@@ -1,4 +1,14 @@
---
- name: Reload Atlas admin user manager
become_user: "{{ atlas_admin_username }}"
ansible.builtin.systemd:
scope: user
daemon_reload: true
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus"
when: not ansible_check_mode
- name: Reload SSH service
ansible.builtin.systemd:
name: sshd

View File

@@ -0,0 +1,159 @@
---
- name: Prepare the isolated rootless Atlas Gitea target
tags: [atlas, gitea]
when: atlas_manage_gitea | bool
block:
- name: Require the existing Atlas application-data dataset
ansible.builtin.assert:
that:
- atlas_manage_storage | bool
- atlas_gitea_dataset == atlas_zfs_pool ~ '/services/data/gitea'
- atlas_gitea_mountpoint == atlas_app_data_mountpoint ~ '/gitea'
- atlas_gitea_username == atlas_admin_username
- atlas_gitea_group == atlas_admin_group
- atlas_gitea_uid | int == atlas_admin_uid | int
- atlas_gitea_gid | int == atlas_admin_gid | int
- atlas_gitea_container_uid | int == 1000
- atlas_gitea_container_gid | int == 1000
- atlas_gitea_staging_bind_address == '127.0.0.1'
- not (atlas_gitea_production_enabled | bool) or atlas_manage_firewall | bool
- not (atlas_gitea_production_enabled | bool) or atlas_gitea_bind_address == ansible_host
fail_msg: >-
Rootless Gitea requires Atlas storage, the admin user manager, the
dedicated dataset, and loopback-only staging ports.
- name: Inspect the final-restore marker before production activation
ansible.builtin.stat:
path: "{{ atlas_gitea_mountpoint }}/.final-sha256"
register: atlas_gitea_final_marker
when: atlas_gitea_production_enabled | bool
- name: Refuse production activation without the final consistent restore
ansible.builtin.assert:
that:
- atlas_gitea_final_marker.stat.isreg | default(false)
fail_msg: Restore the final stopped-source Gitea export before enabling production.
when: atlas_gitea_production_enabled | bool
- name: Verify the production Gitea dataset belongs to admin
ansible.builtin.stat:
path: "{{ atlas_gitea_mountpoint }}"
register: atlas_gitea_dataset_owner
when: atlas_gitea_production_enabled | bool
- name: Refuse to overlap the legacy host-account service
ansible.builtin.assert:
that:
- atlas_gitea_dataset_owner.stat.uid | int == atlas_admin_uid | int
- atlas_gitea_dataset_owner.stat.gid | int == atlas_admin_gid | int
fail_msg: >-
Run the explicit Gitea owner migration before enabling the admin
Quadlet; never chown an active legacy service in a normal run.
when: atlas_gitea_production_enabled | bool
- name: Remove the retired account's parent-dataset traverse ACL
ansible.posix.acl:
path: "{{ item }}"
etype: user
entity: "{{ atlas_gitea_legacy_username }}"
state: absent
loop:
- "{{ atlas_services_mountpoint }}"
- "{{ atlas_app_data_mountpoint }}"
when: atlas_gitea_production_enabled | bool
- name: Enable POSIX ACLs only on the service-namespace parents
community.general.zfs:
name: "{{ item }}"
state: present
extra_zfs_properties:
acltype: posix
loop:
- "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_services }}"
- "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_app_data }}"
- name: Create the dedicated Gitea ZFS dataset
community.general.zfs:
name: "{{ atlas_gitea_dataset }}"
state: present
extra_zfs_properties:
compression: zstd
mountpoint: "{{ atlas_gitea_mountpoint }}"
- name: Restrict the Gitea dataset and create rootless volume paths
ansible.builtin.file:
path: "{{ item }}"
state: directory
owner: "{{ atlas_gitea_username }}"
group: "{{ atlas_gitea_group }}"
mode: "0700"
loop:
- "{{ atlas_gitea_mountpoint }}"
- "{{ atlas_gitea_mountpoint }}/data"
- "{{ atlas_gitea_mountpoint }}/config"
- "{{ atlas_gitea_home }}/.config"
- "{{ atlas_gitea_home }}/.config/containers"
- "{{ atlas_gitea_quadlet_dir }}"
- name: Ensure lingering for the admin rootless account
ansible.builtin.command:
argv:
- loginctl
- enable-linger
- "{{ atlas_gitea_username }}"
creates: "/var/lib/systemd/linger/{{ atlas_gitea_username }}"
- name: Start the admin rootless user manager
ansible.builtin.systemd:
name: "user@{{ atlas_gitea_uid }}.service"
state: started
when: not ansible_check_mode
- name: Prepare the admin-owned Gitea image
ansible.builtin.import_tasks: gitea_image.yml
- name: Render the rootless Gitea Quadlet
ansible.builtin.template:
src: atlas-gitea.container.j2
dest: "{{ atlas_gitea_quadlet_dir }}/atlas-gitea.container"
owner: "{{ atlas_gitea_username }}"
group: "{{ atlas_gitea_group }}"
mode: "0644"
- name: Permit only Aegis to reach production Gitea HTTP and SSH
ansible.posix.firewalld:
rich_rule: >-
rule family="ipv4" source address="{{ atlas_aegis_ip }}"
port port="{{ item }}" protocol="tcp" accept
zone: "{{ atlas_firewalld_zone }}"
state: "{{ 'enabled' if atlas_gitea_production_enabled | bool else 'disabled' }}"
permanent: true
immediate: true
loop:
- "{{ atlas_gitea_http_port }}"
- "{{ atlas_gitea_ssh_port }}"
when: atlas_manage_firewall | bool
- name: Reload the rootless Gitea user manager without starting Gitea
become_user: "{{ atlas_gitea_username }}"
ansible.builtin.systemd:
scope: user
daemon_reload: true
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_gitea_uid }}/bus"
when: not ansible_check_mode
- name: Start and enable the rootless Gitea user Quadlet after final restore
become_user: "{{ atlas_gitea_username }}"
ansible.builtin.systemd:
name: atlas-gitea.service
scope: user
state: started
enabled: true
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_gitea_uid }}/bus"
when:
- atlas_gitea_production_enabled | bool
- not ansible_check_mode

View File

@@ -0,0 +1,51 @@
---
- name: Create the admin-owned Gitea image build directory
ansible.builtin.file:
path: "{{ atlas_gitea_image_build_dir }}"
state: directory
owner: "{{ atlas_admin_username }}"
group: "{{ atlas_admin_group }}"
mode: "0700"
- name: Install the pinned rootless Gitea Containerfile
ansible.builtin.copy:
src: Containerfile.gitea-rootless
dest: "{{ atlas_gitea_image_build_dir }}/Containerfile"
owner: "{{ atlas_admin_username }}"
group: "{{ atlas_admin_group }}"
mode: "0644"
- name: Check the admin-owned Gitea image
become_user: "{{ atlas_admin_username }}"
ansible.builtin.command:
argv: [podman, image, exists, "{{ atlas_gitea_image }}"]
args:
chdir: "{{ atlas_gitea_image_build_dir }}"
environment:
HOME: "{{ atlas_admin_home }}"
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
register: atlas_gitea_image_present
changed_when: false
failed_when: false
check_mode: false
- name: Build the pinned Gitea image with the internal gitea identity
become_user: "{{ atlas_admin_username }}"
ansible.builtin.command:
argv:
- podman
- build
- --pull=always
- --tag
- "{{ atlas_gitea_image }}"
- --file
- Containerfile
- .
args:
chdir: "{{ atlas_gitea_image_build_dir }}"
environment:
HOME: "{{ atlas_admin_home }}"
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
when:
- atlas_gitea_image_present.rc != 0
- not ansible_check_mode

View File

@@ -0,0 +1,275 @@
---
# Run only in an approved outage with -e atlas_gitea_owner_migration=true.
- name: Move live Gitea from the legacy host account to admin
tags: [atlas, gitea_owner_migration]
when: atlas_gitea_owner_migration | bool
block:
- name: Refuse a check-mode owner migration
ansible.builtin.assert:
that: not ansible_check_mode
fail_msg: The owner migration requires an explicit live outage.
- name: Inspect the Gitea dataset owner
ansible.builtin.stat:
path: "{{ atlas_gitea_mountpoint }}"
register: atlas_gitea_migration_owner
- name: Require either the legacy owner or an already migrated dataset
ansible.builtin.assert:
that:
- atlas_gitea_migration_owner.stat.isdir | default(false)
- atlas_gitea_migration_owner.stat.uid | int in [atlas_gitea_legacy_uid | int, atlas_admin_uid | int]
fail_msg: Refusing to modify a Gitea dataset with an unexpected owner.
- name: Migrate only a legacy-owned Gitea dataset
when: atlas_gitea_migration_owner.stat.uid | int == atlas_gitea_legacy_uid | int
block:
- name: Require the final cutover marker and configuration
ansible.builtin.stat:
path: "{{ item }}"
loop:
- "{{ atlas_gitea_mountpoint }}/.final-sha256"
- "{{ atlas_gitea_mountpoint }}/config/app.ini"
register: atlas_gitea_migration_files
- name: Refuse migration without both final data and configuration
ansible.builtin.assert:
that: atlas_gitea_migration_files.results | map(attribute='stat.isreg') | min
- name: Check that admin has no existing Gitea Quadlet
ansible.builtin.stat:
path: "{{ atlas_gitea_quadlet_dir }}/atlas-gitea.container"
register: atlas_gitea_admin_quadlet
- name: Refuse to overwrite an existing admin Quadlet
ansible.builtin.assert:
that: not atlas_gitea_admin_quadlet.stat.exists
- name: Check pool health before the outage
ansible.builtin.command:
argv: [zpool, status, -x, "{{ atlas_zfs_pool }}"]
register: atlas_gitea_pool_before
changed_when: false
failed_when: "'is healthy' not in atlas_gitea_pool_before.stdout"
- name: Ensure the admin Gitea image is available before stopping the source
ansible.builtin.import_tasks: gitea_image.yml
- name: Stop, snapshot and test the admin-owned staging service
block:
- name: Stop and disable the legacy Gitea user service
become_user: "{{ atlas_gitea_legacy_username }}"
ansible.builtin.systemd:
name: atlas-gitea.service
scope: user
state: stopped
enabled: false
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_legacy_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_gitea_legacy_uid }}/bus"
- name: Record the migration snapshot name
ansible.builtin.set_fact:
atlas_gitea_migration_snapshot: >-
{{ atlas_gitea_dataset }}@gitea-owner-migration-{{ ansible_facts.date_time.iso8601_basic_short }}
- name: Snapshot the stopped Gitea dataset for manual recovery
ansible.builtin.command:
argv: [zfs, snapshot, "{{ atlas_gitea_migration_snapshot }}"]
- name: Transfer only the Gitea dataset to admin
ansible.builtin.file:
path: "{{ atlas_gitea_mountpoint }}"
state: directory
owner: "{{ atlas_admin_username }}"
group: "{{ atlas_admin_group }}"
recurse: true
- name: Set the actual internal Unix process user
ansible.builtin.lineinfile:
path: "{{ atlas_gitea_mountpoint }}/config/app.ini"
regexp: '^RUN_USER\s*='
line: RUN_USER = gitea
mode: "0600"
no_log: true
diff: false
- name: Preserve public git clone URLs independently of the Unix user
community.general.ini_file:
path: "{{ atlas_gitea_mountpoint }}/config/app.ini"
section: server
option: "{{ item }}"
value: git
mode: "0600"
no_extra_spaces: false
loop: [BUILTIN_SSH_SERVER_USER, SSH_USER]
no_log: true
diff: false
- name: Render admin's loopback-only staging Quadlet
ansible.builtin.template:
src: atlas-gitea.container.j2
dest: "{{ atlas_gitea_quadlet_dir }}/atlas-gitea.container"
owner: "{{ atlas_admin_username }}"
group: "{{ atlas_admin_group }}"
mode: "0644"
vars:
atlas_gitea_production_enabled: false
- name: Reload the admin user manager for staging
become_user: "{{ atlas_admin_username }}"
ansible.builtin.systemd:
scope: user
daemon_reload: true
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus"
- name: Start admin's loopback-only staging service
become_user: "{{ atlas_admin_username }}"
ansible.builtin.systemd:
name: atlas-gitea.service
scope: user
state: started
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus"
- name: Verify staging HTTP before promotion
ansible.builtin.uri:
url: "http://127.0.0.1:{{ atlas_gitea_staging_http_port }}/"
status_code: 200
register: atlas_gitea_staging_http
retries: 30
delay: 2
until: atlas_gitea_staging_http is succeeded
- name: Verify the container really runs as internal gitea
become_user: "{{ atlas_admin_username }}"
ansible.builtin.command:
argv: [podman, exec, atlas-gitea, id, -un]
environment:
HOME: "{{ atlas_admin_home }}"
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
register: atlas_gitea_internal_user
changed_when: false
failed_when: atlas_gitea_internal_user.stdout != 'gitea'
- name: Verify the migrated SQLite database
ansible.builtin.command:
argv:
- sqlite3
- "{{ atlas_gitea_mountpoint }}/data/gitea/gitea.db"
- PRAGMA quick_check;
register: atlas_gitea_migration_sqlite
changed_when: false
failed_when: atlas_gitea_migration_sqlite.stdout != 'ok'
rescue:
- name: Stop admin's failed staging service
become_user: "{{ atlas_admin_username }}"
ansible.builtin.systemd:
name: atlas-gitea.service
scope: user
state: stopped
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus"
failed_when: false
- name: Restore the original Gitea configuration from the safety snapshot
ansible.builtin.command:
argv:
- cp
- -a
- "{{ atlas_gitea_mountpoint }}/.zfs/snapshot/{{ atlas_gitea_migration_snapshot.split('@')[1] }}/config/app.ini"
- "{{ atlas_gitea_mountpoint }}/config/app.ini"
when: atlas_gitea_migration_snapshot is defined
- name: Return the Gitea dataset to the legacy account
ansible.builtin.file:
path: "{{ atlas_gitea_mountpoint }}"
state: directory
owner: "{{ atlas_gitea_legacy_username }}"
group: "{{ atlas_gitea_legacy_username }}"
recurse: true
- name: Restart the legacy Gitea service
become_user: "{{ atlas_gitea_legacy_username }}"
ansible.builtin.systemd:
name: atlas-gitea.service
scope: user
state: started
enabled: true
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_legacy_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_gitea_legacy_uid }}/bus"
- name: Report the failed migration and preserved snapshot
ansible.builtin.fail:
msg: >-
Admin staging failed; legacy Gitea was restarted. Inspect
{{ atlas_gitea_migration_snapshot | default('the host journal') }}.
- name: Stop admin's validated staging service
become_user: "{{ atlas_admin_username }}"
ansible.builtin.systemd:
name: atlas-gitea.service
scope: user
state: stopped
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus"
- name: Render admin's production Gitea Quadlet
ansible.builtin.template:
src: atlas-gitea.container.j2
dest: "{{ atlas_gitea_quadlet_dir }}/atlas-gitea.container"
owner: "{{ atlas_admin_username }}"
group: "{{ atlas_admin_group }}"
mode: "0644"
vars:
atlas_gitea_production_enabled: true
- name: Reload admin's production user manager
become_user: "{{ atlas_admin_username }}"
ansible.builtin.systemd:
scope: user
daemon_reload: true
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus"
- name: Enable and start admin's production Gitea
become_user: "{{ atlas_admin_username }}"
ansible.builtin.systemd:
name: atlas-gitea.service
scope: user
state: started
enabled: true
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus"
- name: Verify production HTTP before retiring the old Quadlet
ansible.builtin.uri:
url: "http://{{ atlas_gitea_bind_address }}:{{ atlas_gitea_http_port }}/"
status_code: 200
register: atlas_gitea_production_http
retries: 30
delay: 2
until: atlas_gitea_production_http is succeeded
- name: Remove only the disabled legacy Quadlet
ansible.builtin.file:
path: "{{ atlas_gitea_legacy_home }}/.config/containers/systemd/atlas-gitea.container"
state: absent
- name: Reload the legacy user manager after Quadlet removal
become_user: "{{ atlas_gitea_legacy_username }}"
ansible.builtin.systemd:
scope: user
daemon_reload: true
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_legacy_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_gitea_legacy_uid }}/bus"

View File

@@ -0,0 +1,58 @@
---
- name: Manage the public domain of the restored production Gitea
tags: [atlas, gitea, gitea_public_domain]
when:
- atlas_manage_gitea | bool
- atlas_gitea_production_enabled | bool
- atlas_gitea_public_domain | length > 0
block:
- name: Require an explicit public Gitea hostname
ansible.builtin.assert:
that:
- atlas_gitea_public_domain is match('^[a-zA-Z0-9][a-zA-Z0-9.-]*\.[a-zA-Z]{2,}$')
- name: Inspect the restored private Gitea configuration
ansible.builtin.stat:
path: "{{ atlas_gitea_mountpoint }}/config/app.ini"
follow: false
register: atlas_gitea_public_config
- name: Refuse to create or replace an unprepared Gitea configuration
ansible.builtin.assert:
that:
- atlas_gitea_public_config.stat.isreg | default(false)
- atlas_gitea_public_config.stat.uid | int == atlas_gitea_uid | int
- atlas_gitea_public_config.stat.mode == '0600'
# app.ini contains secrets: preserve all unrelated settings and suppress diffs.
- name: Set only the declared public Gitea server fields
community.general.ini_file:
path: "{{ atlas_gitea_mountpoint }}/config/app.ini"
section: server
option: "{{ item.option }}"
value: "{{ item.value }}"
create: false
backup: true
owner: "{{ atlas_gitea_username }}"
group: "{{ atlas_gitea_group }}"
mode: "0600"
loop:
- { option: DOMAIN, value: "{{ atlas_gitea_public_domain }}" }
- { option: ROOT_URL, value: "https://{{ atlas_gitea_public_domain }}/" }
- { option: SSH_DOMAIN, value: "{{ atlas_gitea_public_domain }}" }
register: atlas_gitea_public_domain_update
no_log: true
diff: false
- name: Restart only Gitea when its public configuration changes
become_user: "{{ atlas_gitea_username }}"
ansible.builtin.systemd:
name: atlas-gitea.service
scope: user
state: restarted
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_gitea_uid }}/bus"
when:
- atlas_gitea_public_domain_update is changed
- not ansible_check_mode

View File

@@ -0,0 +1,107 @@
---
- name: Restore Gitea from a verified Prometheus backup only on explicit request
tags: [atlas, gitea_restore, gitea_final_restore]
when: atlas_gitea_restore_test | bool or atlas_gitea_final_restore | bool
block:
- name: Require the prepared rootless Gitea target
ansible.builtin.assert:
that:
- atlas_manage_gitea | bool
- not (atlas_gitea_restore_test | bool and atlas_gitea_final_restore | bool)
- atlas_gitea_staging_bind_address == '127.0.0.1'
- atlas_gitea_mountpoint == atlas_app_data_mountpoint ~ '/gitea'
fail_msg: Prepare the isolated, loopback-only rootless Gitea target first.
- name: Confirm the rootless Gitea service is inactive
become_user: "{{ atlas_gitea_username }}"
ansible.builtin.command:
argv:
- systemctl
- --user
- is-active
- atlas-gitea.service
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_gitea_uid }}/bus"
register: atlas_gitea_restore_service_state
changed_when: false
failed_when: false
when: not ansible_check_mode
- name: Refuse to overwrite an active rootless Gitea service
ansible.builtin.assert:
that:
- atlas_gitea_restore_service_state.stdout == 'inactive'
fail_msg: The rootless Gitea user service must be known and inactive before restoring data.
when: not ansible_check_mode
- name: Check for a manually running rootless Gitea container
become_user: "{{ atlas_gitea_username }}"
ansible.builtin.command:
argv:
- podman
- ps
- --quiet
- --filter
- name=atlas-gitea
args:
chdir: "{{ atlas_gitea_home }}"
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_uid }}"
register: atlas_gitea_restore_container_state
changed_when: false
when: not ansible_check_mode
- name: Refuse to overwrite a running rootless Gitea container
ansible.builtin.assert:
that:
- atlas_gitea_restore_container_state.stdout | length == 0
fail_msg: Stop every rootless Atlas Gitea container before restoring data.
when: not ansible_check_mode
- name: Install the selective rootless Gitea restore helper
ansible.builtin.copy:
src: atlas-gitea-restore-test.py
dest: "{{ atlas_gitea_restore_helper }}"
owner: root
group: root
mode: "0700"
- name: Restore only Gitea data into the isolated target
ansible.builtin.command:
argv:
- "{{ atlas_gitea_restore_helper }}"
- --backup
- "{{ atlas_backup_prometheus_mountpoint }}/latest"
- --target
- "{{ atlas_gitea_mountpoint }}"
- --uid
- "{{ atlas_gitea_uid | string }}"
- --gid
- "{{ atlas_gitea_gid | string }}"
register: atlas_gitea_restore_result
changed_when: atlas_gitea_restore_result.stdout == 'restored'
no_log: true
when:
- atlas_gitea_restore_test | bool
- not ansible_check_mode
- name: Replace the marked rehearsal with the final consistent Gitea export
ansible.builtin.command:
argv:
- "{{ atlas_gitea_restore_helper }}"
- --backup
- "{{ atlas_backup_prometheus_mountpoint }}/latest"
- --target
- "{{ atlas_gitea_mountpoint }}"
- --uid
- "{{ atlas_gitea_uid | string }}"
- --gid
- "{{ atlas_gitea_gid | string }}"
- --replace-rehearsal
register: atlas_gitea_final_restore_result
changed_when: atlas_gitea_final_restore_result.stdout == 'restored'
no_log: true
when:
- atlas_gitea_final_restore | bool
- not ansible_check_mode

View File

@@ -0,0 +1,188 @@
---
- name: Require exact Atlas iCloudPD paths and rootless identity
tags: [atlas, icloudpd]
ansible.builtin.assert:
that:
- atlas_manage_storage | bool
- atlas_icloudpd_dataset == atlas_zfs_pool ~ '/services/data/icloudpd'
- atlas_icloudpd_state_dir == atlas_app_data_mountpoint ~ '/icloudpd'
- atlas_icloudpd_config_dir == atlas_icloudpd_state_dir ~ '/config'
- atlas_icloudpd_photos_dir == atlas_archive_mountpoint ~ '/Pictures/iCloudPD'
- atlas_admin_uid | int == 1000
- atlas_admin_gid | int == 1000
- atlas_icloudpd_image is search('@sha256:[0-9a-f]{64}$')
fail_msg: Verify the fixed, separate Atlas iCloudPD photo and state paths.
- name: Declare rootless Atlas iCloudPD storage and boot-started Quadlet
tags: [atlas, icloudpd]
block:
- name: Inspect the existing Archive and application-data datasets
community.general.zfs_facts:
name: "{{ item.dataset }}"
properties: name,mounted,mountpoint
loop:
- dataset: "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_archive }}"
mountpoint: "{{ atlas_archive_mountpoint }}"
- dataset: "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_app_data }}"
mountpoint: "{{ atlas_app_data_mountpoint }}"
loop_control:
label: "{{ item.dataset }}"
register: atlas_icloudpd_parent_datasets
- name: Refuse missing or unmounted iCloudPD parent datasets
ansible.builtin.assert:
that:
- item.ansible_facts.ansible_zfs_datasets | length == 1
- item.ansible_facts.ansible_zfs_datasets[0].mounted == 'yes'
- item.ansible_facts.ansible_zfs_datasets[0].mountpoint == item.item.mountpoint
loop: "{{ atlas_icloudpd_parent_datasets.results }}"
loop_control:
label: "{{ item.item.dataset }}"
- name: Inspect the existing Pictures namespace and proposed target
ansible.builtin.stat:
path: "{{ item }}"
follow: false
loop:
- "{{ atlas_archive_mountpoint }}/Pictures"
- "{{ atlas_icloudpd_photos_dir }}"
- "{{ atlas_icloudpd_photos_dir }}/.atlas-icloudpd-managed"
register: atlas_icloudpd_photo_paths
- name: Refuse to adopt unrelated Pictures data or a symlink
ansible.builtin.assert:
that:
- atlas_icloudpd_photo_paths.results[0].stat.isdir | default(false)
- atlas_icloudpd_photo_paths.results[0].stat.uid | int == atlas_admin_uid | int
- >-
not atlas_icloudpd_photo_paths.results[1].stat.exists or
(atlas_icloudpd_photo_paths.results[1].stat.isdir | default(false) and
atlas_icloudpd_photo_paths.results[2].stat.isreg | default(false))
fail_msg: >-
Pictures must exist and be admin-owned; an existing iCloudPD target
must carry its managed marker. Never adopt or replace unrelated data.
- name: Create a dedicated ZFS dataset for iCloudPD configuration and MFA
community.general.zfs:
name: "{{ atlas_icloudpd_dataset }}"
state: present
extra_zfs_properties:
compression: zstd
mountpoint: "{{ atlas_icloudpd_state_dir }}"
- name: Restrict iCloudPD state and the new photo subtree
ansible.builtin.file:
path: "{{ item.path }}"
state: directory
owner: "{{ atlas_admin_username }}"
group: "{{ atlas_admin_group }}"
mode: "{{ item.mode }}"
loop:
- path: "{{ atlas_icloudpd_state_dir }}"
mode: "0700"
- path: "{{ atlas_icloudpd_config_dir }}"
mode: "0700"
- path: "{{ atlas_icloudpd_photos_dir }}"
mode: "0750"
- path: "{{ atlas_icloudpd_quadlet_dir }}"
mode: "0700"
loop_control:
label: "{{ item.path }}"
- name: Mark only the newly managed iCloudPD photo subtree
ansible.builtin.copy:
content: "Atlas iCloudPD photo subtree; do not remove source photos.\n"
dest: "{{ atlas_icloudpd_photos_dir }}/.atlas-icloudpd-managed"
owner: "{{ atlas_admin_username }}"
group: "{{ atlas_admin_group }}"
mode: "0600"
force: false
- name: Install the image's required mounted-filesystem failsafe
ansible.builtin.copy:
content: ""
dest: "{{ atlas_icloudpd_photos_dir }}/.mounted"
owner: "{{ atlas_admin_username }}"
group: "{{ atlas_admin_group }}"
mode: "0644"
force: false
- name: Require the Vault-backed iCloudPD Apple ID
ansible.builtin.assert:
that:
- vault_atlas_icloudpd_apple_id is defined
- vault_atlas_icloudpd_apple_id | length > 0
- vault_atlas_icloudpd_apple_id != 'REPLACE_ME'
- vault_atlas_icloudpd_apple_id.splitlines() | length == 1
fail_msg: Configure the existing iCloudPD Apple ID in Vault.
no_log: true
- name: Seed private Atlas iCloudPD configuration when absent
ansible.builtin.template:
src: atlas-icloudpd.conf.j2
dest: "{{ atlas_icloudpd_config_dir }}/icloudpd.conf"
owner: "{{ atlas_admin_username }}"
group: "{{ atlas_admin_group }}"
mode: "0600"
force: false
no_log: true
diff: false
- name: Keep declared iCloudPD options in the image-managed configuration
ansible.builtin.lineinfile:
path: "{{ atlas_icloudpd_config_dir }}/icloudpd.conf"
regexp: "^{{ item.key }}="
line: "{{ item.key }}={{ item.value }}"
owner: "{{ atlas_admin_username }}"
group: "{{ atlas_admin_group }}"
mode: "0600"
loop:
- {key: apple_id, value: "{{ vault_atlas_icloudpd_apple_id }}"}
- {key: authentication_type, value: MFA}
- {key: user, value: user}
- {key: user_id, value: "1000"}
- {key: group, value: group}
- {key: group_id, value: "1000"}
- {key: download_path, value: /home/user/iCloud}
- {key: folder_structure, value: "{:%Y/%m/%d}"}
- {key: directory_permissions, value: "750"}
- {key: file_permissions, value: "640"}
- {key: download_interval, value: "86400"}
- {key: auto_delete, value: "false"}
- {key: delete_after_download, value: "false"}
loop_control:
label: "{{ item.key }}"
no_log: true
diff: false
- name: Render the rootless Atlas iCloudPD Quadlet
ansible.builtin.template:
src: atlas-icloudpd.container.j2
dest: "{{ atlas_icloudpd_quadlet_dir }}/atlas-icloudpd.container"
owner: "{{ atlas_admin_username }}"
group: "{{ atlas_admin_group }}"
mode: "0644"
register: atlas_icloudpd_quadlet
- name: Reload the Atlas admin user manager after iCloudPD Quadlet changes
become_user: "{{ atlas_admin_username }}"
ansible.builtin.systemd:
scope: user
daemon_reload: true
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus"
when:
- atlas_icloudpd_quadlet.changed
- not ansible_check_mode
- name: Keep the rootless Atlas iCloudPD service running
become_user: "{{ atlas_admin_username }}"
ansible.builtin.systemd:
name: atlas-icloudpd.service
scope: user
state: started
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus"
when: not ansible_check_mode

View File

@@ -14,6 +14,21 @@
- name: Import Atlas storage tasks
ansible.builtin.import_tasks: storage.yml
- name: Import explicit Atlas Gitea owner migration
ansible.builtin.import_tasks: gitea_owner_migration.yml
- name: Import staged Atlas rootless Gitea tasks
ansible.builtin.import_tasks: gitea.yml
- name: Import the declared Atlas Gitea public domain
ansible.builtin.import_tasks: gitea_public_domain.yml
- name: Import Atlas iCloudPD storage and boot-started Quadlet tasks
ansible.builtin.import_tasks: icloudpd.yml
- name: Import explicit Atlas Gitea restore rehearsal tasks
ansible.builtin.import_tasks: gitea_restore.yml
- name: Import Atlas ZFS maintenance tasks
ansible.builtin.import_tasks: zfs_maintenance.yml
@@ -23,6 +38,12 @@
- name: Import Atlas offline USB backup tasks
ansible.builtin.import_tasks: usb_backup.yml
- name: Import Atlas Prometheus backup pull identity tasks
ansible.builtin.import_tasks: prometheus_pull_identity.yml
- name: Import Atlas Prometheus backup pull job tasks
ansible.builtin.import_tasks: prometheus_pull_job.yml
- name: Import Atlas health monitoring tasks
ansible.builtin.import_tasks: monitoring.yml

View File

@@ -7,8 +7,8 @@
- atlas_zfs_pool != 'CHANGEME_ZFS_POOL'
- atlas_monitor_calendar | length > 0
- atlas_monitor_smart_devices | length > 0
- atlas_monitor_timers | length > 0
- atlas_monitor_failure_units | length > 0
- atlas_monitor_effective_timers | length > 0
- atlas_monitor_effective_failure_units | length > 0
- atlas_monitor_remote_capacity.user == atlas_borg_repository_user
- atlas_monitor_remote_capacity.host == atlas_borg_repository_host
- atlas_monitor_remote_capacity.run_as == atlas_borg_username
@@ -49,7 +49,7 @@
that:
- item.name is match('^[a-zA-Z0-9@_.-]+\\.timer$')
- item.max_age_hours | int >= 0
loop: "{{ atlas_monitor_timers }}"
loop: "{{ atlas_monitor_effective_timers }}"
loop_control:
label: "{{ item.name }}"
when: atlas_manage_monitoring | bool
@@ -59,7 +59,7 @@
ansible.builtin.assert:
that:
- item is match('^[a-zA-Z0-9@_.-]+\\.service$')
loop: "{{ atlas_monitor_failure_units }}"
loop: "{{ atlas_monitor_effective_failure_units }}"
when: atlas_manage_monitoring | bool
- name: Validate Atlas health monitor calendar
@@ -144,7 +144,7 @@
owner: root
group: root
mode: "0755"
loop: "{{ atlas_monitor_failure_units }}"
loop: "{{ atlas_monitor_effective_failure_units }}"
when: atlas_manage_monitoring | bool
- name: Notify 45Drives Alerts when an Atlas job fails
@@ -155,7 +155,7 @@
owner: root
group: root
mode: "0644"
loop: "{{ atlas_monitor_failure_units }}"
loop: "{{ atlas_monitor_effective_failure_units }}"
when: atlas_manage_monitoring | bool
- name: Reload systemd after installing Atlas monitoring

View File

@@ -0,0 +1,66 @@
---
- name: Validate Atlas Prometheus pull identity inputs
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
ansible.builtin.assert:
that:
- atlas_prometheus_pull_ssh_dir.startswith('/etc/')
- atlas_prometheus_pull_private_key_path.startswith(atlas_prometheus_pull_ssh_dir ~ '/')
- atlas_prometheus_pull_known_hosts_path.startswith(atlas_prometheus_pull_ssh_dir ~ '/')
- atlas_prometheus_ssh_host_key.startswith(
(hostvars['prometheus'].ansible_host | string) ~ ' ssh-ed25519 '
)
fail_msg: Pin the verified Prometheus ED25519 SSH host key before enabling the pull.
when: atlas_manage_prometheus_backup_pull | bool
- name: Create private Atlas Prometheus pull SSH directory
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
ansible.builtin.file:
path: "{{ atlas_prometheus_pull_ssh_dir }}"
state: directory
owner: root
group: root
mode: "0700"
when: atlas_manage_prometheus_backup_pull | bool
- name: Generate Atlas-only Prometheus pull SSH identity
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
ansible.builtin.command:
argv:
- ssh-keygen
- -q
- -t
- ed25519
- -N
- ""
- -C
- atlas-prometheus-pull@atlas
- -f
- "{{ atlas_prometheus_pull_private_key_path }}"
creates: "{{ atlas_prometheus_pull_private_key_path }}"
when: atlas_manage_prometheus_backup_pull | bool
- name: Protect Atlas-only Prometheus pull SSH identity
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
ansible.builtin.file:
path: "{{ item.path }}"
owner: root
group: root
mode: "{{ item.mode }}"
loop:
- { path: "{{ atlas_prometheus_pull_private_key_path }}", mode: "0600" }
- { path: "{{ atlas_prometheus_pull_private_key_path }}.pub", mode: "0644" }
loop_control:
label: "{{ item.path }}"
when:
- atlas_manage_prometheus_backup_pull | bool
- not ansible_check_mode
- name: Pin Prometheus SSH host key on Atlas
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
ansible.builtin.copy:
content: "{{ atlas_prometheus_ssh_host_key }}\n"
dest: "{{ atlas_prometheus_pull_known_hosts_path }}"
owner: root
group: root
mode: "0600"
when: atlas_manage_prometheus_backup_pull | bool

View File

@@ -0,0 +1,90 @@
---
- name: Validate Atlas Prometheus backup pull inputs
tags: [atlas, backup, prometheus_backup]
ansible.builtin.assert:
that:
- atlas_manage_storage | bool
- atlas_prometheus_pull_source_user is match('^[a-z_][a-z0-9_-]*$')
- atlas_prometheus_pull_source_port | int > 0
- atlas_prometheus_pull_source_port | int < 65536
- atlas_prometheus_pull_keep_daily | int > 0
- atlas_prometheus_pull_keep_weekly | int > 0
- atlas_prometheus_pull_keep_monthly | int > 0
- atlas_prometheus_pull_max_age_hours | int > 0
- atlas_backup_prometheus_mountpoint.startswith(atlas_mount_root ~ '/')
fail_msg: Define the Atlas backup destination, source account, and retention before enabling the pull.
when: atlas_manage_prometheus_backup_pull | bool
- name: Validate Atlas Prometheus backup pull calendar
tags: [atlas, backup, prometheus_backup]
ansible.builtin.command:
argv: [systemd-analyze, calendar, "{{ atlas_prometheus_pull_calendar }}"]
changed_when: false
check_mode: false
when: atlas_manage_prometheus_backup_pull | bool
- name: Create private Atlas Prometheus backup version directory
tags: [atlas, backup, prometheus_backup]
ansible.builtin.file:
path: "{{ atlas_backup_prometheus_mountpoint }}/snapshots"
state: directory
owner: root
group: root
mode: "0700"
when: atlas_manage_prometheus_backup_pull | bool
- name: Install Atlas Prometheus backup pull helper
tags: [atlas, backup, prometheus_backup]
ansible.builtin.template:
src: atlas-prometheus-pull.sh.j2
dest: /usr/local/sbin/atlas-prometheus-pull
owner: root
group: root
mode: "0750"
when: atlas_manage_prometheus_backup_pull | bool
- name: Install Atlas Prometheus backup retention helper
tags: [atlas, backup, prometheus_backup]
ansible.builtin.copy:
src: atlas-prometheus-prune.py
dest: /usr/local/libexec/atlas-prometheus-prune
owner: root
group: root
mode: "0750"
when: atlas_manage_prometheus_backup_pull | bool
- name: Install Atlas Prometheus backup pull systemd units
tags: [atlas, backup, prometheus_backup]
ansible.builtin.template:
src: "{{ item }}.j2"
dest: "/etc/systemd/system/{{ item }}"
owner: root
group: root
mode: "0644"
loop:
- atlas-prometheus-pull.service
- atlas-prometheus-pull.timer
loop_control:
label: "{{ item }}"
register: atlas_prometheus_pull_units
when: atlas_manage_prometheus_backup_pull | bool
- name: Reload systemd after Atlas Prometheus pull unit changes
tags: [atlas, backup, prometheus_backup]
ansible.builtin.systemd:
daemon_reload: true
when:
- atlas_manage_prometheus_backup_pull | bool
- atlas_prometheus_pull_units is changed
- not ansible_check_mode
- name: Enable Atlas Prometheus pull timer only after explicit activation
tags: [atlas, backup, prometheus_backup]
ansible.builtin.systemd:
name: atlas-prometheus-pull.timer
enabled: true
state: started
when:
- atlas_manage_prometheus_backup_pull | bool
- atlas_prometheus_pull_start_timer | bool
- not ansible_check_mode

View File

@@ -0,0 +1,30 @@
# Managed by Ansible. Staging does not start automatically.
[Unit]
Description=Atlas rootless Gitea
RequiresMountsFor={{ atlas_gitea_mountpoint }}
[Container]
ContainerName=atlas-gitea
Image={{ atlas_gitea_image }}
UserNS=keep-id:uid={{ atlas_gitea_container_uid }},gid={{ atlas_gitea_container_gid }}
{% if atlas_gitea_production_enabled | bool %}
PublishPort={{ atlas_gitea_bind_address }}:{{ atlas_gitea_http_port }}:3000
PublishPort={{ atlas_gitea_bind_address }}:{{ atlas_gitea_ssh_port }}:2222
{% else %}
PublishPort={{ atlas_gitea_staging_bind_address }}:{{ atlas_gitea_staging_http_port }}:3000
PublishPort={{ atlas_gitea_staging_bind_address }}:{{ atlas_gitea_staging_ssh_port }}:2222
{% endif %}
Volume={{ atlas_gitea_mountpoint }}/data:/var/lib/gitea:Z
Volume={{ atlas_gitea_mountpoint }}/config:/etc/gitea:Z
NoNewPrivileges=true
DropCapability=all
[Service]
Restart=on-failure
RestartSec=10
TimeoutStartSec=900
{% if atlas_gitea_production_enabled | bool %}
[Install]
WantedBy=default.target
{% endif %}

View File

@@ -3,8 +3,8 @@
"backup_dataset": {{ (atlas_zfs_pool ~ '/' ~ atlas_zfs_dataset_backup) | to_json }},
"notifier": {{ atlas_monitor_notifier | to_json }},
"smart_devices": {{ atlas_monitor_smart_devices | to_json }},
"timers": {{ atlas_monitor_timers | to_json }},
"failure_units": {{ atlas_monitor_failure_units | to_json }},
"timers": {{ atlas_monitor_effective_timers | to_json }},
"failure_units": {{ atlas_monitor_effective_failure_units | to_json }},
"remote_capacity": {{ atlas_monitor_remote_capacity | to_json }},
"pool_warning_percent": {{ atlas_monitor_pool_warning_percent | int }},
"pool_critical_percent": {{ atlas_monitor_pool_critical_percent | int }},

View File

@@ -0,0 +1,14 @@
# Managed by Ansible. Password, keyring and MFA cookies are stored separately in /config.
apple_id={{ vault_atlas_icloudpd_apple_id }}
authentication_type=MFA
user=user
user_id=1000
group=group
group_id=1000
download_path=/home/user/iCloud
folder_structure={:%Y/%m/%d}
directory_permissions=750
file_permissions=640
download_interval=86400
auto_delete=false
delete_after_download=false

View File

@@ -0,0 +1,25 @@
# Managed by Ansible. Start automatically with the lingering admin user manager.
[Unit]
Description=Atlas rootless iCloud Photos Downloader
RequiresMountsFor={{ atlas_icloudpd_state_dir }} {{ atlas_icloudpd_photos_dir }}
[Container]
ContainerName=atlas-icloudpd
Image={{ atlas_icloudpd_image }}
UserNS=keep-id:uid=1000,gid=1000
# The image initialises its unprivileged UID 1000 account as container root.
User=0
# Upstream launcher requires traceroute for its iCloud reachability check.
AddCapability=NET_RAW
Environment=TZ={{ atlas_icloudpd_timezone }}
Volume={{ atlas_icloudpd_photos_dir }}:/home/user/iCloud:z
Volume={{ atlas_icloudpd_config_dir }}:/config:Z
NoNewPrivileges=true
[Service]
Restart=on-failure
RestartSec=300
TimeoutStartSec=900
[Install]
WantedBy=default.target

View File

@@ -0,0 +1,19 @@
[Unit]
Description=Pull a prepared read-only Prometheus backup to Atlas
RequiresMountsFor={{ atlas_backup_prometheus_mountpoint }}
Wants=network-online.target
After=network-online.target zfs.target
ConditionFileIsExecutable=/usr/local/sbin/atlas-prometheus-pull
ConditionPathExists={{ atlas_prometheus_pull_private_key_path }}
ConditionPathExists={{ atlas_prometheus_pull_known_hosts_path }}
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/atlas-prometheus-pull
User=root
Group=root
UMask=0077
TimeoutStartSec=infinity
Nice=15
IOSchedulingClass=best-effort
IOSchedulingPriority=7

View File

@@ -0,0 +1,75 @@
#!/usr/bin/env bash
set -Eeuo pipefail
umask 077
backup_root={{ atlas_backup_prometheus_mountpoint | quote }}
snapshots="$backup_root/snapshots"
stage=''
exec 9>/run/lock/atlas-prometheus-pull.lock
flock -n 9 || { echo 'A Prometheus pull is already running' >&2; exit 1; }
cleanup() {
local rc=$?
trap - EXIT
if (( rc != 0 )) && [[ -n "$stage" && -d "$stage" ]]; then
rm -rf -- "$stage"
fi
exit "$rc"
}
trap cleanup EXIT
zpool list -H -o name {{ atlas_zfs_pool | quote }} >/dev/null
findmnt -rn --mountpoint "$backup_root" >/dev/null
stage=$(mktemp -d "$backup_root/.staging.XXXXXXXX")
ssh_cmd='/usr/bin/ssh -F /dev/null -o BatchMode=yes -o StrictHostKeyChecking=yes -o UserKnownHostsFile={{ atlas_prometheus_pull_known_hosts_path }} -o IdentitiesOnly=yes -i {{ atlas_prometheus_pull_private_key_path }} -p {{ atlas_prometheus_pull_source_port }}'
rsync -a --partial --delay-updates -e "$ssh_cmd" \
{{ (atlas_prometheus_pull_source_user ~ '@' ~ hostvars['prometheus'].ansible_host ~ ':current/') | quote }} \
"$stage/"
test -s "$stage/payload.tar"
test -s "$stage/payload.sha256"
test -s "$stage/metadata.json"
(cd "$stage" && sha256sum -c payload.sha256)
tar -tf "$stage/payload.tar" >/dev/null
stamp=$(python3 - "$stage/metadata.json" <<'PY'
import json
import datetime as dt
import re
import sys
with open(sys.argv[1], encoding="utf-8") as stream:
metadata = json.load(stream)
stamp = metadata.get("created_utc", "")
if metadata.get("schema") != 1 or metadata.get("host") != "prometheus":
raise SystemExit("Unexpected Prometheus backup metadata")
if not re.fullmatch(r"[0-9]{8}T[0-9]{6}Z", stamp):
raise SystemExit("Invalid Prometheus backup timestamp")
created = dt.datetime.strptime(stamp, "%Y%m%dT%H%M%SZ").replace(tzinfo=dt.timezone.utc)
age = dt.datetime.now(dt.timezone.utc) - created
if age.total_seconds() < -300 or age > dt.timedelta(hours={{ atlas_prometheus_pull_max_age_hours }}):
raise SystemExit("Prometheus backup is outside the configured freshness window")
print(stamp)
PY
)
if [[ -e "$snapshots/$stamp" ]]; then
cmp "$stage/payload.sha256" "$snapshots/$stamp/payload.sha256"
cmp "$stage/metadata.json" "$snapshots/$stamp/metadata.json"
(cd "$snapshots/$stamp" && sha256sum -c payload.sha256)
rm -rf -- "${stage:?}"
stage=''
else
chown -R root:root "$stage"
chmod 0700 "$stage"
chmod 0600 "$stage/payload.tar" "$stage/payload.sha256" "$stage/metadata.json"
mv -- "$stage" "$snapshots/$stamp"
stage=''
fi
latest_link=$(readlink "$backup_root/latest" 2>/dev/null || true)
latest_stamp=${latest_link##*/}
if [[ -z "$latest_stamp" || "$stamp" > "$latest_stamp" ]]; then
ln -s "snapshots/$stamp" "$backup_root/.latest.new"
mv -Tf -- "$backup_root/.latest.new" "$backup_root/latest"
fi
python3 /usr/local/libexec/atlas-prometheus-prune "$snapshots" \
{{ atlas_prometheus_pull_keep_daily }} {{ atlas_prometheus_pull_keep_weekly }} {{ atlas_prometheus_pull_keep_monthly }}
echo "Verified and published Prometheus backup $stamp"

View File

@@ -0,0 +1,10 @@
[Unit]
Description=Schedule Atlas pull of prepared Prometheus backups
[Timer]
OnCalendar={{ atlas_prometheus_pull_calendar }}
Persistent=true
Unit=atlas-prometheus-pull.service
[Install]
WantedBy=timers.target

View File

@@ -30,3 +30,7 @@ backend_phase1_timezone: Europe/Rome
backend_phase1_services:
- atlas-navidrome.service
- atlas-syncthing.service
backend_phase1_music_sync_enabled: false
backend_phase1_music_source_dir: "{{ backend_phase1_archive_dir }}/Music"
backend_phase1_music_sync_calendar: "*-*-* 00:45:00 Europe/Rome"
backend_phase1_user_systemd_dir: "{{ backend_phase1_user_home }}/.config/systemd/user"

View File

@@ -18,21 +18,28 @@
- backend_phase1_app_data_root.startswith('/')
- backend_phase1_navidrome_data_dir.startswith(backend_phase1_app_data_root + '/')
- backend_phase1_syncthing_root.startswith(backend_phase1_app_data_root + '/')
- >-
not (backend_phase1_music_sync_enabled | bool) or
(backend_phase1_music_source_dir.startswith(backend_phase1_archive_dir + '/')
and backend_phase1_music_sync_calendar | length > 0)
fail_msg: >-
Disable the rootful media-stack gate and provide the Atlas LAN bind
address, firewall sources, and absolute ZFS-backed paths before
enabling phase one. This role does not manage Prometheus or migrate
application data.
tags: [music_sync]
- name: Read the rootless service account
ansible.builtin.getent:
database: passwd
key: "{{ backend_phase1_username }}"
tags: [music_sync]
- name: Record rootless service account IDs
ansible.builtin.set_fact:
backend_phase1_uid: "{{ ansible_facts['getent_passwd'][backend_phase1_username][1] }}"
backend_phase1_gid: "{{ ansible_facts['getent_passwd'][backend_phase1_username][2] }}"
tags: [music_sync]
- name: Read system service state before starting rootless Syncthing
ansible.builtin.service_facts:
@@ -65,6 +72,7 @@
loop_control:
label: "{{ item.dataset }}"
register: backend_phase1_zfs_facts
tags: [music_sync]
- name: Require mounted datasets at the declared paths
ansible.builtin.assert:
@@ -79,6 +87,23 @@
loop: "{{ backend_phase1_zfs_facts.results }}"
loop_control:
label: "{{ item.item.dataset }}"
tags: [music_sync]
- name: Inspect the music copy source
ansible.builtin.stat:
path: "{{ backend_phase1_music_source_dir }}"
register: backend_phase1_music_source_stat
when: backend_phase1_music_sync_enabled | bool
tags: [music_sync]
- name: Require an existing music source directory
ansible.builtin.assert:
that:
- backend_phase1_music_source_stat.stat.isdir | default(false)
fail_msg: >-
{{ backend_phase1_music_source_dir }} must exist before enabling the daily music copy.
when: backend_phase1_music_sync_enabled | bool
tags: [music_sync]
- name: Enable lingering for the rootless service account
ansible.builtin.command:
@@ -87,12 +112,14 @@
- enable-linger
- "{{ backend_phase1_username }}"
creates: "/var/lib/systemd/linger/{{ backend_phase1_username }}"
tags: [music_sync]
- name: Start the rootless user systemd manager
ansible.builtin.systemd:
name: "user@{{ backend_phase1_uid }}.service"
state: started
when: not ansible_check_mode
tags: [music_sync]
- name: Create rootless Quadlet and application directories
ansible.builtin.file:
@@ -113,6 +140,43 @@
loop_control:
label: "{{ item.path }}"
- name: Install rsync for the daily music copy
ansible.builtin.dnf:
name: rsync
state: present
when: backend_phase1_music_sync_enabled | bool
tags: [music_sync]
- name: Create the rootless user systemd directory
ansible.builtin.file:
path: "{{ backend_phase1_user_systemd_dir }}"
state: directory
owner: "{{ backend_phase1_username }}"
group: "{{ backend_phase1_user_group }}"
mode: "0700"
when: backend_phase1_music_sync_enabled | bool
tags: [music_sync]
- name: Install the daily music copy service
ansible.builtin.template:
src: atlas-music-sync.service.j2
dest: "{{ backend_phase1_user_systemd_dir }}/atlas-music-sync.service"
owner: "{{ backend_phase1_username }}"
group: "{{ backend_phase1_user_group }}"
mode: "0644"
when: backend_phase1_music_sync_enabled | bool
tags: [music_sync]
- name: Install the daily music copy timer
ansible.builtin.template:
src: atlas-music-sync.timer.j2
dest: "{{ backend_phase1_user_systemd_dir }}/atlas-music-sync.timer"
owner: "{{ backend_phase1_username }}"
group: "{{ backend_phase1_user_group }}"
mode: "0644"
when: backend_phase1_music_sync_enabled | bool
tags: [music_sync]
- name: Render the rootless Navidrome Quadlet
ansible.builtin.template:
src: atlas-navidrome.container.j2
@@ -140,6 +204,7 @@
XDG_RUNTIME_DIR: "/run/user/{{ backend_phase1_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ backend_phase1_uid }}/bus"
when: not ansible_check_mode
tags: [music_sync]
- name: Permit NPM access to phase-one web interfaces through Aegis
ansible.posix.firewalld:
@@ -186,3 +251,19 @@
when:
- backend_phase1_start_services | bool
- not ansible_check_mode
- name: Enable the daily music copy timer
become_user: "{{ backend_phase1_username }}"
ansible.builtin.systemd:
name: atlas-music-sync.timer
scope: user
state: started
enabled: true
daemon_reload: true
environment:
XDG_RUNTIME_DIR: "/run/user/{{ backend_phase1_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ backend_phase1_uid }}/bus"
when:
- backend_phase1_music_sync_enabled | bool
- not ansible_check_mode
tags: [music_sync]

View File

@@ -0,0 +1,10 @@
# Managed by Ansible. Do not edit manually.
[Unit]
Description=Copy Atlas Archive music to the Navidrome library
[Service]
Type=oneshot
ExecStartPre=/usr/bin/mountpoint -q {{ backend_phase1_archive_dir }}
ExecStartPre=/usr/bin/mountpoint -q {{ backend_phase1_music_dir }}
ExecStartPre=/usr/bin/test -d {{ backend_phase1_music_source_dir }}
ExecStart=/usr/bin/rsync -aH --no-perms --no-owner --no-group --delay-updates --stats -- {{ backend_phase1_music_source_dir }}/ {{ backend_phase1_music_dir }}/

View File

@@ -0,0 +1,11 @@
# Managed by Ansible. Do not edit manually.
[Unit]
Description=Schedule the daily Atlas Navidrome music copy
[Timer]
OnCalendar={{ backend_phase1_music_sync_calendar }}
Persistent=true
Unit=atlas-music-sync.service
[Install]
WantedBy=timers.target

View File

@@ -0,0 +1,103 @@
---
- name: Validate Prometheus backup export identity inputs
tags: [services, backup, prometheus_backup]
ansible.builtin.assert:
that:
- inventory_hostname == 'prometheus'
- server_backup_username is match('^[a-z_][a-z0-9_-]*$')
- server_backup_username not in ['root', server_username]
- server_backup_export_root.startswith('/var/lib/')
- server_backup_public_key_name is match('^[a-z0-9_-]+$')
- hostvars['atlas'].atlas_manage_prometheus_backup_pull | default(false) | bool
fail_msg: Enable Atlas and Prometheus backup roles together with dedicated identity settings.
when: server_backup_export_enabled | bool
- name: Create dedicated Prometheus backup export group
tags: [services, backup, prometheus_backup]
ansible.builtin.group:
name: "{{ server_backup_username }}"
system: true
state: present
when: server_backup_export_enabled | bool
- name: Create locked Prometheus backup export account
tags: [services, backup, prometheus_backup]
ansible.builtin.user:
name: "{{ server_backup_username }}"
group: "{{ server_backup_username }}"
groups: []
append: false
comment: Read-only prepared backup export for Atlas
home: "{{ server_backup_export_root }}"
create_home: false
shell: /bin/bash
password_lock: true
system: true
state: present
when: server_backup_export_enabled | bool
- name: Require restricted rrsync helper on Prometheus
tags: [services, backup, prometheus_backup]
ansible.builtin.stat:
path: "{{ server_backup_rrsync_path }}"
register: server_backup_rrsync_file
when: server_backup_export_enabled | bool
- name: Validate restricted rrsync helper
tags: [services, backup, prometheus_backup]
ansible.builtin.assert:
that:
- server_backup_rrsync_file.stat.exists
- server_backup_rrsync_file.stat.isreg
- server_backup_rrsync_file.stat.pw_name == 'root'
fail_msg: Rocky rsync must provide the root-owned rrsync support script.
when: server_backup_export_enabled | bool
- name: Create prepared backup export root
tags: [services, backup, prometheus_backup]
ansible.builtin.file:
path: "{{ server_backup_export_root }}"
state: directory
owner: root
group: "{{ server_backup_username }}"
mode: "0750"
when: server_backup_export_enabled | bool
- name: Create restricted Prometheus backup SSH directories
tags: [services, backup, prometheus_backup]
ansible.builtin.file:
path: "{{ item }}"
state: directory
owner: root
group: "{{ server_backup_username }}"
mode: "0750"
loop:
- "{{ server_backup_export_root }}/.ssh"
- "{{ server_backup_export_root }}/.ssh/authorized_keys.d"
when: server_backup_export_enabled | bool
- name: Read Atlas public key for Prometheus backup pull
tags: [services, backup, prometheus_backup]
ansible.builtin.slurp:
src: "{{ hostvars['atlas'].atlas_prometheus_pull_private_key_path | default('/etc/atlas-prometheus-pull/id_ed25519') }}.pub"
delegate_to: atlas
become: true
register: server_backup_atlas_public_key
when:
- server_backup_export_enabled | bool
- not ansible_check_mode
- name: Authorize only restricted read-only backup access from Atlas
tags: [services, backup, prometheus_backup]
ansible.builtin.copy:
content: >-
{{ 'command="/usr/bin/python3 ' ~ server_backup_rrsync_path ~ ' -ro '
~ server_backup_export_root ~ '/versions",restrict '
~ (server_backup_atlas_public_key.content | b64decode | trim) ~ '\n' }}
dest: "{{ server_backup_export_root }}/.ssh/authorized_keys.d/{{ server_backup_public_key_name }}"
owner: root
group: "{{ server_backup_username }}"
mode: "0640"
when:
- server_backup_export_enabled | bool
- not ansible_check_mode

View File

@@ -0,0 +1,91 @@
---
- name: Validate Prometheus backup export job inputs
tags: [services, backup, prometheus_backup]
ansible.builtin.assert:
that:
- server_backup_export_source_keep | int >= 2
- server_backup_export_paths | length > 0
- server_backup_export_paths | unique | length == server_backup_export_paths | length
- >-
server_backup_export_paths
| select('match', '^[a-zA-Z0-9][a-zA-Z0-9._/-]*$') | list | length
== server_backup_export_paths | length
- >-
server_backup_export_paths
| reject('search', '(^|/)\.\.(/|$)') | list | length
== server_backup_export_paths | length
- >-
server_backup_export_excludes
| select('match', '^[a-zA-Z0-9][a-zA-Z0-9._/-]*$') | list | length
== server_backup_export_excludes | length
- >-
server_backup_export_excludes
| reject('search', '(^|/)\.\.(/|$)') | list | length
== server_backup_export_excludes | length
fail_msg: Define safe relative paths and at least two prepared export versions.
when: server_backup_export_enabled | bool
- name: Validate Prometheus backup export calendar
tags: [services, backup, prometheus_backup]
ansible.builtin.command:
argv: [systemd-analyze, calendar, "{{ server_backup_export_calendar }}"]
changed_when: false
check_mode: false
when: server_backup_export_enabled | bool
- name: Ensure prepared Prometheus backup versions directory exists
tags: [services, backup, prometheus_backup]
ansible.builtin.file:
path: "{{ server_backup_export_root }}/versions"
state: directory
owner: root
group: "{{ server_backup_username }}"
mode: "0750"
when: server_backup_export_enabled | bool
- name: Install Prometheus backup export helper
tags: [services, backup, prometheus_backup, gitea_cutover, npm_quadlet_backup]
ansible.builtin.template:
src: prometheus-backup-export.sh.j2
dest: /usr/local/sbin/prometheus-backup-export
owner: root
group: root
mode: "0750"
validate: "bash -n %s"
when: server_backup_export_enabled | bool
- name: Install Prometheus backup export systemd units
tags: [services, backup, prometheus_backup]
ansible.builtin.template:
src: "{{ item }}.j2"
dest: "/etc/systemd/system/{{ item }}"
owner: root
group: root
mode: "0644"
loop:
- prometheus-backup-export.service
- prometheus-backup-export.timer
loop_control:
label: "{{ item }}"
register: server_backup_export_units
when: server_backup_export_enabled | bool
- name: Reload systemd after Prometheus backup export unit changes
tags: [services, backup, prometheus_backup]
ansible.builtin.systemd:
daemon_reload: true
when:
- server_backup_export_enabled | bool
- server_backup_export_units is changed
- not ansible_check_mode
- name: Enable Prometheus backup export timer only after explicit activation
tags: [services, backup, prometheus_backup]
ansible.builtin.systemd:
name: prometheus-backup-export.timer
enabled: true
state: started
when:
- server_backup_export_enabled | bool
- server_backup_export_start_timer | bool
- not ansible_check_mode

View File

@@ -0,0 +1,38 @@
---
- name: Install the explicit Gitea final-export helper
tags: [services, gitea_final_export]
ansible.builtin.template:
src: prometheus-gitea-final-export.sh.j2
dest: /usr/local/sbin/prometheus-gitea-final-export
owner: root
group: root
mode: "0750"
when:
- server_gitea_cutover_tools_enabled | bool
- not server_legacy_stack_retired | bool
- name: Require the prepared source and explicit final-export approval
tags: [services, gitea_final_export]
ansible.builtin.assert:
that:
- server_gitea_cutover_tools_enabled | bool
- server_backup_export_enabled | bool
- not server_legacy_stack_retired | bool
- not ansible_check_mode
fail_msg: >-
Install the cutover helper and perform an explicit non-check-mode run
only after the Gitea outage gate has been approved.
when: server_gitea_final_export | bool
- name: Stop source Gitea and publish the final consistent export
tags: [services, gitea_final_export]
ansible.builtin.command:
argv:
- /usr/local/sbin/prometheus-gitea-final-export
register: server_gitea_final_export_result
changed_when: server_gitea_final_export_result.rc == 0
no_log: true
when:
- server_gitea_final_export | bool
- not server_legacy_stack_retired | bool
- not ansible_check_mode

View File

@@ -0,0 +1,59 @@
---
- name: Validate the NPM Gitea cutover override
tags: [services, gitea_cutover]
ansible.builtin.assert:
that:
- server_gitea_cutover_tools_enabled | bool
- server_gitea_npm_domains | length > 0
- server_gitea_npm_domains | select('match', '^[a-zA-Z0-9.-]+$') | list | length == server_gitea_npm_domains | length
fail_msg: Declare the exact NPM Gitea hostnames before enabling the Atlas upstream.
when: server_gitea_on_atlas | bool
- name: Ensure the NPM custom configuration directory exists
tags: [services, gitea_cutover]
ansible.builtin.file:
path: /opt/npm/data/nginx/custom
state: directory
owner: root
group: root
mode: "0755"
when: server_gitea_cutover_tools_enabled | bool
- name: Render the Gitea-only NPM runtime upstream override
tags: [services, gitea_cutover]
ansible.builtin.template:
src: prometheus-gitea-npm-proxy.conf.j2
dest: /opt/npm/data/nginx/custom/server_proxy.conf
owner: root
group: root
mode: "0644"
register: server_gitea_npm_override
when: server_gitea_on_atlas | bool
- name: Remove the Gitea NPM override when source routing is selected
tags: [services, gitea_cutover]
ansible.builtin.file:
path: /opt/npm/data/nginx/custom/server_proxy.conf
state: absent
when:
- server_gitea_cutover_tools_enabled | bool
- not server_gitea_on_atlas | bool
- name: Validate NPM configuration after a Gitea upstream change
tags: [services, gitea_cutover]
ansible.builtin.command:
argv: [podman, exec, nginx-proxy-manager, nginx, -t]
changed_when: false
when:
- server_gitea_on_atlas | bool
- server_gitea_npm_override is changed
- not ansible_check_mode
- name: Reload NPM after validating the Gitea upstream change
tags: [services, gitea_cutover]
ansible.builtin.command:
argv: [podman, exec, nginx-proxy-manager, nginx, -s, reload]
when:
- server_gitea_on_atlas | bool
- server_gitea_npm_override is changed
- not ansible_check_mode

View File

@@ -0,0 +1,61 @@
---
- name: Validate the Prometheus Gitea SSH cutover inputs
tags: [services, gitea_cutover]
ansible.builtin.assert:
that:
- server_gitea_cutover_tools_enabled | bool
- server_gitea_atlas_address is match('^[0-9]{1,3}(\.[0-9]{1,3}){3}$')
- server_gitea_ssh_public_port | int > 1024
- server_gitea_ssh_public_port | int < 65536
- server_gitea_ssh_target_port | int > 1024
- server_gitea_ssh_target_port | int < 65536
- server_gitea_ssh_public_port | int != 22
fail_msg: Keep administrative SSH on 22 and provide the Atlas rootless Gitea SSH endpoint.
when: server_gitea_on_atlas | bool
- name: Install the Gitea SSH socket proxy units without activating them
tags: [services, gitea_cutover]
ansible.builtin.template:
src: "{{ item }}.j2"
dest: "/etc/systemd/system/{{ item }}"
owner: root
group: root
mode: "0644"
loop:
- prometheus-gitea-ssh-proxy.socket
- prometheus-gitea-ssh-proxy.service
loop_control:
label: "{{ item }}"
register: server_gitea_ssh_proxy_units
when: server_gitea_cutover_tools_enabled | bool
- name: Reload systemd after Gitea SSH proxy unit changes
tags: [services, gitea_cutover]
ansible.builtin.systemd:
daemon_reload: true
when:
- server_gitea_cutover_tools_enabled | bool
- server_gitea_ssh_proxy_units is changed
- not ansible_check_mode
- name: Manage the public Gitea SSH socket separately from administrative SSH
tags: [services, gitea_cutover]
ansible.builtin.systemd:
name: prometheus-gitea-ssh-proxy.socket
state: "{{ 'started' if server_gitea_on_atlas | bool else 'stopped' }}"
enabled: "{{ server_gitea_on_atlas | bool }}"
when:
- server_gitea_cutover_tools_enabled | bool
- not ansible_check_mode
- name: Open only the public Gitea SSH port after cutover
tags: [services, gitea_cutover]
ansible.posix.firewalld:
port: "{{ server_gitea_ssh_public_port }}/tcp"
zone: "{{ server_firewalld_zone }}"
state: "{{ 'enabled' if server_gitea_on_atlas | bool else 'disabled' }}"
permanent: true
immediate: true
when:
- server_gitea_cutover_tools_enabled | bool
- server_firewall_backend == 'firewalld'

View File

@@ -0,0 +1,112 @@
---
- name: Require explicit retirement of the migrated Prometheus source
ansible.builtin.assert:
that:
- inventory_hostname == 'prometheus'
- server_legacy_stack_retired | bool
- server_gitea_on_atlas | bool
- server_npm_quadlet_cutover | bool
- server_backup_export_enabled | bool
- name: Verify legacy paths have no mounts or container users
ansible.builtin.command:
argv:
- python3
- -c
- |
import json, os, pathlib, subprocess
def run(*args):
return subprocess.check_output(args, text=True).strip()
paths = ['/opt/gitea', '/home/git/.ssh', '/opt/navidrome',
'/opt/postgres', '/opt/music', '/opt/containerd', '/opt/docker']
mounts = json.loads(run('findmnt', '--json', '--list', '-o', 'TARGET'))['filesystems']
for path in paths:
assert os.path.realpath(path) == path, 'Symlink in cleanup path: ' + path
for mount in mounts:
target = mount['target']
assert target != path and not target.startswith(path + '/'), 'Mounted cleanup path: ' + path
ids = run('podman', 'ps', '-aq').split()
containers = json.loads(run('podman', 'inspect', *ids)) if ids else []
for container in containers:
assert container['Name'].lstrip('/') == 'nginx-proxy-manager', 'Unexpected container; review before cleanup'
for mount in container.get('Mounts', []):
source = os.path.realpath(mount['Source'])
for path in paths:
assert source != path and not source.startswith(path + '/'), 'Container uses cleanup path: ' + path
for path in ['/opt/music', '/opt/containerd']:
if os.path.isdir(path):
for entry in pathlib.Path(path).rglob('*'):
assert entry.is_dir() and not entry.is_symlink(), 'Unexpected file in empty legacy path: ' + str(entry)
if os.path.isdir('/opt/docker'):
allowed = {'/opt/docker/server', '/opt/docker/server/docker-compose.yml'}
for entry in pathlib.Path('/opt/docker').rglob('*'):
assert str(entry) in allowed and not entry.is_symlink(), 'Unexpected legacy Docker content: ' + str(entry)
assert run('systemctl', 'is-active', 'prometheus-npm.service') == 'active'
assert subprocess.run(['systemctl', 'is-active', '--quiet', 'podman-compose-server.service']).returncode != 0
assert subprocess.run(['systemctl', 'is-active', '--quiet', 'prometheus-backup-export.service']).returncode != 0
print('Legacy cleanup preflight passed')
changed_when: false
check_mode: false
- name: Require the updated backup configuration before deleting fallback files
ansible.builtin.command:
argv:
- python3
- -c
- |
import pathlib, subprocess
unit = subprocess.check_output(['systemctl', 'show', 'prometheus-backup-export.service',
'-p', 'RequiresMountsFor', '--value'], text=True)
assert '/opt/gitea' not in unit, 'Backup unit still depends on legacy Gitea'
helper = pathlib.Path('/usr/local/sbin/prometheus-backup-export').read_text()
assert 'podman-compose-server' not in helper and 'opt/docker/server' not in helper
subprocess.run(['bash', '-n', '/usr/local/sbin/prometheus-backup-export'], check=True)
changed_when: false
when: not ansible_check_mode
- name: Delete only the explicitly approved legacy data and fallback files
ansible.builtin.file:
path: "{{ item }}"
state: absent
loop:
- /opt/gitea
- /home/git/.ssh
- /opt/navidrome
- /opt/postgres
- /opt/music
- /opt/containerd
- /opt/docker
- /usr/local/sbin/prometheus-gitea-final-export
- /etc/systemd/system/podman-compose-server.service
register: server_legacy_deleted
diff: false
- name: Reload systemd after removing the inactive legacy unit
ansible.builtin.systemd:
daemon_reload: true
when:
- server_legacy_deleted is changed
- not ansible_check_mode
- name: Inspect the obsolete Git home without following symlinks
ansible.builtin.stat:
path: /home/git
follow: false
register: server_legacy_git_home
- name: Require the obsolete Git account to be absent before removing its empty home
ansible.builtin.command:
argv: [getent, passwd, git]
register: server_legacy_git_account
changed_when: false
failed_when: server_legacy_git_account.rc != 2
check_mode: false
when: server_legacy_git_home.stat.exists
# rmdir refuses any nonempty directory; never recursively delete this parent.
- name: Remove only the empty obsolete Git home
ansible.builtin.command:
argv: [rmdir, /home/git]
register: server_legacy_git_home_removed
changed_when: server_legacy_git_home_removed.rc == 0
when: server_legacy_git_home.stat.exists

View File

@@ -0,0 +1,33 @@
---
- name: Require the migrated Prometheus topology for image cleanup
ansible.builtin.assert:
that:
- inventory_hostname == 'prometheus'
- server_gitea_on_atlas | bool
- server_npm_quadlet_cutover | bool
- server_legacy_images | default([]) | length > 0
- >-
server_legacy_images | difference([
'docker.gitea.com/gitea:1.25.2',
'docker.io/deluan/navidrome:latest',
'docker.io/library/postgres:13']) | length == 0
- name: Check whether the explicitly selected legacy images exist
ansible.builtin.command:
argv: [podman, image, exists, "{{ item }}"]
loop: "{{ server_legacy_images }}"
register: server_legacy_image_presence
changed_when: false
failed_when: server_legacy_image_presence.rc not in [0, 1]
check_mode: false
# No --force: Podman must refuse images referenced by any existing container.
- name: Remove only unused explicitly selected legacy images
ansible.builtin.command:
argv: [podman, image, rm, "{{ item.item }}"]
loop: "{{ server_legacy_image_presence.results }}"
loop_control:
label: "{{ item.item }}"
when: item.rc == 0
register: server_legacy_image_removal
changed_when: server_legacy_image_removal.rc == 0

View File

@@ -11,6 +11,7 @@
- name: Configure DuckDNS updater
tags: [dotfiles, dotfiles:server, duckdns]
ansible.builtin.import_tasks: duckdns.yml
when: server_duckdns_enabled | bool
- name: Ensure server directories exist
tags: [dotfiles, services]
@@ -23,6 +24,9 @@
loop: "{{ server_directories | default([]) }}"
loop_control:
label: "{{ item.path }}"
when:
- item.path != '/opt/gitea/data' or not server_gitea_on_atlas | bool
- item.path != server_container_stack_dir or not server_legacy_stack_retired | bool
- name: Copy server dotfiles
tags: [dotfiles, dotfiles:server]
@@ -37,7 +41,7 @@
label: "{{ item.dest }}"
- name: Render server templates
tags: [dotfiles, dotfiles:server]
tags: [dotfiles, dotfiles:server, gitea_cutover]
ansible.builtin.template:
src: "{{ item.src }}"
dest: "{{ item.dest if item.dest.startswith('/') else server_user_home ~ '/' ~ item.dest }}"
@@ -48,10 +52,41 @@
loop_control:
label: "{{ item.dest }}"
no_log: "{{ item.no_log | default(false) }}"
when: item.src != 'server/docker-compose.yml.j2' or not server_legacy_stack_retired | bool
- name: Manage Podman Compose stack
tags: [services, podman]
ansible.builtin.include_tasks: podman-compose.yml
when: not server_legacy_stack_retired | bool
- name: Import staged NPM Quadlet tasks
ansible.builtin.import_tasks: npm_quadlet.yml
- name: Import explicit legacy server image cleanup
ansible.builtin.import_tasks: legacy_image_cleanup.yml
tags: [never, server_image_cleanup]
when: server_legacy_image_cleanup | default(false) | bool
- name: Import Prometheus backup export identity tasks
ansible.builtin.import_tasks: backup_export_identity.yml
- name: Import Prometheus backup export job tasks
ansible.builtin.import_tasks: backup_export_job.yml
tags: [server_legacy_cleanup]
- name: Import explicitly approved legacy server data cleanup
ansible.builtin.import_tasks: legacy_cleanup.yml
tags: [never, server_legacy_cleanup]
when: server_legacy_cleanup | bool
- name: Import explicit Prometheus Gitea final-export tasks
ansible.builtin.import_tasks: gitea_final_export.yml
- name: Import Prometheus Gitea SSH proxy tasks
ansible.builtin.import_tasks: gitea_ssh_proxy.yml
- name: Import Prometheus Gitea NPM proxy override tasks
ansible.builtin.import_tasks: gitea_npm_proxy.yml
- name: Ensure server SSH authorized key fragments directory exists
tags: [services, ssh]
@@ -77,13 +112,17 @@
when: server_ssh_authorized_keys | length > 0
- name: Configure server SSH authorized key fragments
tags: [services, ssh]
tags: [services, ssh, prometheus_backup]
ansible.builtin.lineinfile:
path: /etc/ssh/sshd_config
regexp: '^\s*AuthorizedKeysFile\s+'
line: >-
AuthorizedKeysFile {{ server_ssh_authorized_keys | map(attribute='name')
| map('regex_replace', '^', '%h/.ssh/authorized_keys.d/') | join(' ') }}
AuthorizedKeysFile {{
((server_ssh_authorized_keys | map(attribute='name')
| map('regex_replace', '^', '%h/.ssh/authorized_keys.d/') | list)
+ (['%h/.ssh/authorized_keys.d/' ~ server_backup_public_key_name]
if server_backup_export_enabled | bool else [])) | join(' ')
}}
state: present
validate: "sshd -t -f %s"
notify: Reload SSH service
@@ -100,11 +139,13 @@
notify: Reload SSH service
- name: Restrict SSH login to allowed users on server
tags: [services]
tags: [services, prometheus_backup]
ansible.builtin.lineinfile:
path: /etc/ssh/sshd_config
regexp: '^\s*AllowUsers\s+'
line: "AllowUsers {{ server_sshd_allow_users | join(' ') }}"
line: >-
AllowUsers {{ (server_sshd_allow_users
+ ([server_backup_username] if server_backup_export_enabled | bool else [])) | join(' ') }}
state: present
validate: "sshd -t -f %s"
notify: Reload SSH service

View File

@@ -0,0 +1,71 @@
---
- name: Require staged NPM Quadlet for an active cutover
tags: [services, npm_quadlet]
ansible.builtin.assert:
that:
- not server_npm_quadlet_cutover | bool or server_npm_quadlet_stage | bool
fail_msg: The NPM Quadlet cutover requires the staged container and network.
- name: Validate staged NPM Quadlet inputs
tags: [services, npm_quadlet]
ansible.builtin.assert:
that:
- server_npm_quadlet_image is defined
- server_npm_quadlet_image is match('^docker\.io/jc21/nginx-proxy-manager@sha256:[a-f0-9]{64}$')
- server_gitea_on_atlas | bool
fail_msg: Stage the exact running NPM image only after Gitea has left Compose.
when: server_npm_quadlet_stage | bool
- name: Ensure rootful Quadlet directory exists for NPM
tags: [services, npm_quadlet]
ansible.builtin.file:
path: /etc/containers/systemd
state: directory
owner: root
group: root
mode: "0755"
when: server_npm_quadlet_stage | bool
- name: Render staged NPM container and network Quadlets
tags: [services, npm_quadlet]
ansible.builtin.template:
src: "{{ item }}.j2"
dest: "/etc/containers/systemd/{{ item }}"
owner: root
group: root
mode: "0644"
loop:
- prometheus-npm.container
- server-web.network
loop_control:
label: "{{ item }}"
register: server_npm_quadlet_units
when: server_npm_quadlet_stage | bool
- name: Reload systemd after staging NPM Quadlets
tags: [services, npm_quadlet]
ansible.builtin.systemd:
daemon_reload: true
when:
- server_npm_quadlet_stage | bool
- server_npm_quadlet_units is changed
- not ansible_check_mode
- name: Verify the staged NPM Quadlet was generated
tags: [services, npm_quadlet]
ansible.builtin.command:
argv: [systemctl, show, prometheus-npm.service, --property=LoadState, --value]
register: server_npm_quadlet_load_state
changed_when: false
when:
- server_npm_quadlet_stage | bool
- not ansible_check_mode
- name: Reject an invalid staged NPM Quadlet
tags: [services, npm_quadlet]
ansible.builtin.assert:
that: server_npm_quadlet_load_state.stdout == 'loaded'
fail_msg: Quadlet generator did not produce prometheus-npm.service.
when:
- server_npm_quadlet_stage | bool
- not ansible_check_mode

View File

@@ -0,0 +1,15 @@
[Unit]
Description=Prepare a read-only Prometheus application backup for Atlas
RequiresMountsFor=/opt/npm {% if not server_gitea_on_atlas | bool %}/opt/gitea {% endif %}{{ server_backup_export_root }}
ConditionFileIsExecutable=/usr/local/sbin/prometheus-backup-export
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/prometheus-backup-export
User=root
Group=root
UMask=0077
TimeoutStartSec=infinity
Nice=10
IOSchedulingClass=best-effort
IOSchedulingPriority=7

View File

@@ -0,0 +1,118 @@
#!/usr/bin/env bash
set -Eeuo pipefail
umask 077
export_root={{ server_backup_export_root | quote }}
versions="$export_root/versions"
{% if server_legacy_stack_retired | bool %}
stack_unit=prometheus-npm.service
{% else %}
stack_unit=''
compose_active=false
quadlet_active=false
systemctl is-active --quiet podman-compose-server.service && compose_active=true
systemctl is-active --quiet prometheus-npm.service && quadlet_active=true
if [[ "$compose_active" == "$quadlet_active" ]]; then
echo 'Expected exactly one active NPM service (Compose or Quadlet)' >&2
exit 1
fi
if "$quadlet_active"; then
stack_unit=prometheus-npm.service
else
stack_unit=podman-compose-server.service
fi
{% endif %}
stamp=$(date -u +%Y%m%dT%H%M%SZ)
stage=''
stack_stopped=false
exec 9>/run/lock/prometheus-backup-export.lock
flock -n 9 || { echo 'A backup export is already running' >&2; exit 1; }
cleanup() {
local rc=$?
trap - EXIT
if "$stack_stopped"; then
if systemctl is-active --quiet "$stack_unit"; then
systemctl restart "$stack_unit" || rc=1
else
systemctl start "$stack_unit" || rc=1
fi
fi
if (( rc != 0 )) && [[ -n "$stage" && -d "$stage" ]]; then
rm -rf -- "$stage"
fi
exit "$rc"
}
trap cleanup EXIT
trap 'exit 129' HUP
trap 'exit 130' INT
trap 'exit 143' TERM
systemctl is-active --quiet "$stack_unit" || {
echo "The managed NPM unit $stack_unit must be active before preparing a backup" >&2
exit 1
}
paths=(
{% for path in server_backup_export_paths %}
{{ path | quote }}
{% endfor %}
)
excludes=(
{% for path in server_backup_export_excludes %}
--exclude={{ path | quote }}
{% endfor %}
)
for path in "${paths[@]}"; do
[[ -e "/$path" ]] || { echo "Required backup path missing: /$path" >&2; exit 1; }
done
[[ ! -e "$versions/$stamp" ]] || { echo "Export version already exists: $stamp" >&2; exit 1; }
stage=$(mktemp -d "$export_root/.staging.XXXXXXXX")
# SQLite databases and their accompanying files are copied while both
# managed containers are stopped. The EXIT trap restarts the stack on error.
stack_stopped=true
systemctl stop "$stack_unit"
tar --acls --xattrs --selinux "${excludes[@]}" -C / -cf "$stage/payload.tar" "${paths[@]}"
systemctl start "$stack_unit"
for container in nginx-proxy-manager{% if not server_gitea_on_atlas | bool %} gitea{% endif %}; do
running=false
for _ in {1..30}; do
if [[ $(podman inspect --format '{{ '{{.State.Running}}' }}' "$container" 2>/dev/null) == true ]]; then
running=true
break
fi
sleep 2
done
"$running" || { echo "Container did not restart: $container" >&2; exit 1; }
done
ready=false
for _ in {1..60}; do
if curl -fsS --connect-timeout 2 --max-time 3 -o /dev/null http://127.0.0.1:81/; then
ready=true
break
fi
sleep 2
done
"$ready" || { echo 'NPM administration did not become ready after backup' >&2; exit 1; }
stack_stopped=false
tar -tf "$stage/payload.tar" >/dev/null
(cd "$stage" && sha256sum payload.tar >payload.sha256)
printf '{"schema":1,"host":"prometheus","created_utc":"%s"}\n' "$stamp" >"$stage/metadata.json"
chown root:{{ server_backup_username }} "$stage" "$stage/payload.tar" "$stage/payload.sha256" "$stage/metadata.json"
chmod 0750 "$stage"
chmod 0640 "$stage/payload.tar" "$stage/payload.sha256" "$stage/metadata.json"
mv -- "$stage" "$versions/$stamp"
stage=''
ln -s "$stamp" "$versions/.current.new"
mv -Tf -- "$versions/.current.new" "$versions/current"
# Keep a small source-side safety window; Atlas owns long-term retention.
mapfile -t old_versions < <(find "$versions" -mindepth 1 -maxdepth 1 -type d \
-printf '%f\n' | grep -E '^[0-9]{8}T[0-9]{6}Z$' | sort -r | tail -n +{{ server_backup_export_source_keep + 1 }})
for old in "${old_versions[@]}"; do
rm -rf -- "${versions:?}/$old"
done
echo "Prepared Prometheus backup export $stamp"

View File

@@ -0,0 +1,10 @@
[Unit]
Description=Prepare daily Prometheus application backup for Atlas
[Timer]
OnCalendar={{ server_backup_export_calendar }}
Persistent=false
Unit=prometheus-backup-export.service
[Install]
WantedBy=timers.target

View File

@@ -0,0 +1,80 @@
#!/usr/bin/env bash
set -Eeuo pipefail
umask 077
export_root={{ server_backup_export_root | quote }}
versions="$export_root/versions"
stamp=$(date -u +%Y%m%dT%H%M%SZ)
stage=''
gitea_stopped=false
exec 9>/run/lock/prometheus-backup-export.lock
flock -n 9 || { echo 'A Prometheus backup export is already running' >&2; exit 1; }
cleanup() {
local rc=$?
trap - EXIT
if (( rc != 0 )) && "$gitea_stopped"; then
podman start gitea >/dev/null || rc=1
fi
if (( rc != 0 )) && [[ -n "$stage" && -d "$stage" ]]; then
rm -rf -- "$stage"
fi
exit "$rc"
}
trap cleanup EXIT
trap 'exit 129' HUP
trap 'exit 130' INT
trap 'exit 143' TERM
systemctl is-active --quiet podman-compose-server.service || {
echo 'Prometheus Compose stack is not active' >&2; exit 1;
}
if systemctl is-active --quiet prometheus-backup-export.timer; then
echo 'Stop the scheduled export timer for the cutover first' >&2
exit 1
fi
[[ $(podman inspect --format '{{ '{{.State.Running}}' }}' gitea) == true ]] || {
echo 'Source Gitea must be running before the final export' >&2; exit 1;
}
[[ -d /opt/gitea/data && -d /home/git/.ssh ]] || {
echo 'Required source Gitea paths are missing' >&2; exit 1;
}
[[ ! -e "$versions/$stamp" ]] || {
echo 'Final export timestamp already exists' >&2; exit 1;
}
gitea_stopped=true
podman stop --time 30 gitea >/dev/null
[[ $(podman inspect --format '{{ '{{.State.Running}}' }}' gitea) == false ]] || {
echo 'Source Gitea did not stop' >&2; exit 1;
}
python3 - <<'PY'
import sqlite3
path = '/opt/gitea/data/gitea/gitea.db'
with sqlite3.connect(f'file:{path}?mode=ro', uri=True) as database:
if database.execute('PRAGMA quick_check').fetchone()[0] != 'ok':
raise SystemExit('Source Gitea SQLite quick_check failed')
PY
stage=$(mktemp -d "$export_root/.staging.XXXXXXXX")
tar --acls --xattrs --selinux -C / -cf "$stage/payload.tar" \
opt/gitea/data home/git/.ssh
tar -tf "$stage/payload.tar" >/dev/null
(cd "$stage" && sha256sum payload.tar >payload.sha256)
printf '{"schema":1,"host":"prometheus","purpose":"gitea-cutover","created_utc":"%s"}\n' \
"$stamp" >"$stage/metadata.json"
[[ $(podman inspect --format '{{ '{{.State.Running}}' }}' gitea) == false ]] || {
echo 'Source Gitea restarted during final export' >&2; exit 1;
}
chown root:{{ server_backup_username }} "$stage" "$stage/payload.tar" \
"$stage/payload.sha256" "$stage/metadata.json"
chmod 0750 "$stage"
chmod 0640 "$stage/payload.tar" "$stage/payload.sha256" "$stage/metadata.json"
mv -- "$stage" "$versions/$stamp"
stage=''
ln -s "$stamp" "$versions/.current.new"
mv -Tf -- "$versions/.current.new" "$versions/current"
echo "Prepared final Gitea export $stamp; source Gitea remains stopped"

View File

@@ -0,0 +1,6 @@
# Managed by Ansible. NPM's variable proxy upstream uses Nginx DNS, not /etc/hosts.
{% for domain in server_gitea_npm_domains %}
if ($host = {{ domain }}) {
set $server {{ server_gitea_atlas_address }};
}
{% endfor %}

View File

@@ -0,0 +1,12 @@
[Unit]
Description=Forward public Gitea SSH to Atlas through Aegis
Requires=prometheus-gitea-ssh-proxy.socket
After=network-online.target wg-quick@wg0.service
[Service]
ExecStart=/usr/lib/systemd/systemd-socket-proxyd {{ server_gitea_atlas_address }}:{{ server_gitea_ssh_target_port }}
DynamicUser=true
NoNewPrivileges=true
ProtectSystem=strict
ProtectHome=true
PrivateTmp=true

View File

@@ -0,0 +1,9 @@
[Unit]
Description=Public Gitea SSH socket on Prometheus
[Socket]
ListenStream=0.0.0.0:{{ server_gitea_ssh_public_port }}
NoDelay=true
[Install]
WantedBy=sockets.target

View File

@@ -0,0 +1,26 @@
[Unit]
Description=Nginx Proxy Manager on Prometheus
RequiresMountsFor=/opt/npm/data /opt/npm/letsencrypt
[Container]
Image={{ server_npm_quadlet_image }}
ContainerName=nginx-proxy-manager
Network=server-web.network
NetworkAlias=nginx-proxy-manager
AddHost=host.containers.internal:host-gateway
PublishPort=80:80
PublishPort=443:443
PublishPort=127.0.0.1:81:81
Volume=/opt/npm/data:/data
Volume=/opt/npm/letsencrypt:/etc/letsencrypt
Pull=missing
[Service]
Restart=always
TimeoutStartSec=180
TimeoutStopSec=120
{% if server_npm_quadlet_cutover | bool %}
[Install]
WantedBy=multi-user.target
{% endif %}

View File

@@ -0,0 +1,5 @@
[Network]
NetworkName=server_web
Driver=bridge
Subnet=10.89.0.0/24
Gateway=10.89.0.1

View File

@@ -4,7 +4,7 @@ name: server
services:
nginx-proxy-manager:
image: docker.io/jc21/nginx-proxy-manager:latest
image: {{ server_npm_quadlet_image if server_npm_quadlet_stage | bool else 'docker.io/jc21/nginx-proxy-manager:latest' }}
container_name: nginx-proxy-manager
restart: unless-stopped
ports:
@@ -38,6 +38,7 @@ services:
# networks:
# - web
{% if not server_gitea_on_atlas | bool %}
gitea:
image: docker.gitea.com/gitea:1.25.2
container_name: gitea
@@ -55,6 +56,7 @@ services:
ports:
- "3000:3000"
- "127.0.0.1:222:22"
{% endif %}
networks:

74
docs/atlas-dr-lab.md Normal file
View File

@@ -0,0 +1,74 @@
# Isolated Atlas DR lab
This is a **scaled rehearsal**, not a substitute for a full-data restore. The
`atlas-dr-lab` libvirt VM on Ikaros was left **shut off** on 2026-09-30. Its
persistent volumes are in the default libvirt pool: the current 30 GiB OS
volume `atlas-dr-lab-os-rebuild2.qcow2`, the pre-rebuild OS volume
`atlas-dr-lab-os.qcow2`, and four independent 4 GiB
`atlas-dr-lab-data{1,2,3,4}.qcow2` volumes. The VM uses libvirt's `default`
NAT network (last DHCP address `192.168.122.168`), 2 vCPU, and 4 GiB RAM.
The data disks have `virtio-atlasdrdata{1,2,3,4}` serials. No physical disk or
production Atlas storage is attached. The VM has no autostart.
## Rebuild inputs and isolation
- Use Rocky's **9.8 GenericCloud Base x86_64** image
`Rocky-9-GenericCloud-Base-9.8-20260525.0.x86_64.qcow2` from
`https://download.rockylinux.org/pub/rocky/9.8/images/x86_64/`.
Verify its `.CHECKSUM` file; the observed SHA-256 was
`92c206cc6f790c61583247eefe87890f8828420662c17cacf247cec78ab4eec8`.
- Use a dedicated lab-only inventory merged **after** the repository
inventory, and always `--limit atlas_dr_lab`. The temporary 2026-09-30
inventory/playbook and logs are in `/tmp/atlas-dr-lab-image/`; copy a
sanitized inventory to durable private storage before `/tmp` is cleared if
the lab will be repeated. Never reuse `host_vars/atlas.yml`, production
Vault secrets, or production disk by-id paths for the lab.
- The lab host belongs to `platform_rocky` and `atlas`. It uses `dradmin`
(UID/GID 1000) with the operator's **public** SSH key and a random,
unknown password hash, the libvirt DHCP address, pool `zpool`, mount root
`/zpool`, the four `virtio-atlasdrdata*` by-id paths, a 1 GiB backup
reservation, and `rocky_manage_openzfs_repo: true` with only `zfs` in
`host_packages`. The following gates remain false: sharing, firewall,
media stack, ZFS timers, Borg, USB, monitoring, and Prometheus pull.
`atlas_manage_storage` is true. Set `atlas_create_pool: true` **only for the
first disposable pool creation**, then set it false before any later run.
- A minimal lab playbook selects `atlas_dr_lab`, `become: true`, and the
existing `packages_rocky` and `profile_atlas` roles. Use a separate
`ANSIBLE_CONFIG` without the production Vault password script, and keep
host-key checking on with a lab-specific known-hosts file. The 2026-09-30
runs used `-i ansible/inventory/hosts.yml -i <lab-inventory.yml>` and
`--limit atlas_dr_lab` throughout.
## Rehearsal and narrow checks
1. Before any pool operation, compare `virsh -c qemu:///system domblklist
atlas-dr-lab` with the four intended qcow2 paths, and in the guest compare
`/dev/disk/by-id/virtio-atlasdrdata*` with `lsblk`. Do not proceed if a
physical disk or production identity appears.
2. For a first-time disposable build only, run the lab playbook with
`--tags pool` and `atlas_create_pool: true`, then immediately set the gate
false. Run the full lab playbook and check `zpool status -P zpool`,
`zfs list -r zpool`, SELinux, and failed systemd units.
3. Write a non-sensitive canary under the lab `/zpool/archive` and snapshot
it. Record the pool GUID and canary SHA-256. Export the lab pool cleanly,
shut down the VM, and replace **only the OS volume** with a fresh verified
Rocky image. Preserve all four data volumes. Reconfigure cloud-init for a
new instance; the seed CD-ROM must use **SATA**. The SCSI seed attachment
tried during this rehearsal was not detected by cloud-init and was
replaced with a SATA attachment before proceeding.
4. On the new OS, apply `packages_rocky` to reinstall OpenZFS. First run
`zpool import -d /dev/disk/by-id` **without importing**, compare GUID and
vdev membership, then use ordinary `zpool import -d /dev/disk/by-id zpool`.
Do not use `-f`, `-F`, `-X`, rollback, or pool creation.
5. Reapply `profile_atlas` with the lab gates and `atlas_create_pool: false`.
Verify the canary, restored snapshot file in an empty temporary directory,
dataset hierarchy, SELinux, and pool health. A second full playbook run
should report `changed=0`. Remove temporary restored files and shut down
the VM after testing.
The observed 2026-09-30 pool GUID was `8880368391795119587`; the canary
SHA-256 was `949701c7a95fadae1fddc21abe846c4312212dbfeb7477948f3188fc3ec34a78`.
The post-rebuild Ansible run succeeded, a repeat run reported `changed=0`,
12 datasets and the original snapshot were present, and the pool was healthy.
The snapshot-restored file matched content and basic metadata. See
[`atlas-recovery.md`](atlas-recovery.md) for the production runbook and limits.

View File

@@ -0,0 +1,241 @@
# Gitea migration from Prometheus to Atlas
The later 2026-10-03 canonical-domain change to `git.fscotto.co` is recorded
in `docs/domain-fscotto-co.md`. Public SSH remains on TCP/2222; the old
DuckDNS Proxy Host was observed disabled. Earlier domain references below
describe migration evidence, not the current canonical URL.
This records the staged migration and its observed partial cutover. Gitea is
temporary on Atlas until Uranus; NPM remains on Prometheus. On 2026-10-03
the operator explicitly approved removal of the old Prometheus Gitea data,
SSH fragment and final-export helper. NPM now uses a rootful Quadlet with no
installed Compose fallback. The source-retention and rollback steps below
are historical migration gates, not current recovery instructions.
Existing backup archives were preserved; use current Atlas data and verified
backups for recovery. Do not recreate or restart stale source Gitea.
## Observed source before cutover and chosen topology (2026-10-01)
- Prometheus runs the rootful `docker.gitea.com/gitea:1.25.2` image in its
managed Compose stack. `/opt/gitea/data` is about 280 MiB, uses SQLite,
and contains 33 repositories. A live read-only SQLite `quick_check` passed.
`/home/git/.ssh` is a separate small bind mount; `/opt/gitea/data/ssh`
contains the existing SSH host keys. Neither tree may be discarded.
- Gitea answers HTTP 200 on Prometheus port 3000. NPM currently forwards
`git.fscotto.duckdns.org` and `git.ov-ad3410.infomaniak.ch` to the Compose
hostname `gitea:3000`. Public DNS resolves to Prometheus. The container's
SSH port is bound only to `127.0.0.1:222`; this is not a public Gitea SSH
listener. Prometheus' public port 22 remains administrative SSH.
- Atlas has a healthy pool and a verified, private Prometheus backup under
`/zpool/backup/hosts/prometheus/latest`. The 2026-10-01 scheduled export
and pull succeeded. The intended target is a separate
`/zpool/services/data/gitea` dataset, not `Archive` or the backup dataset.
- The approved cutover keeps NPM on Prometheus, changes the two HTTP Proxy
Hosts' effective upstream to Atlas over the Prometheus--Aegis gateway, and offers public Gitea
SSH on port 2222 via the same gateway. Prometheus port 22 is unchanged.
HTTPS and SSH must be validated together before declaring cutover.
- The initial staging ran as a **rootless user Quadlet** under a dedicated,
non-login Atlas account, using the pinned `1.25.2-rootless` image. This was an explicit
rootful-to-rootless **data-layout conversion**, not a drop-in image swap:
the target mounts `/var/lib/gitea` and `/etc/gitea`, and uses Gitea's
built-in SSH server instead of the source image's OpenSSH daemon. Keep the
application version unchanged until the conversion has passed an isolated
restore test. The host's rootful Quadlet directory must not be used.
## Phase 1: prepare without traffic changes
Preparation completed on 2026-10-01: Ansible created
`zpool/services/data/gitea`, a dedicated non-login `gitea` account (UID/GID
1101), separate subordinate IDs, parent-dataset traverse ACLs, and an inactive
user Quadlet under `/var/lib/atlas-gitea/.config/containers/systemd/`. The
Quadlet has no `[Install]` section and, until the final cutover, binds only
loopback staging ports 3001/2223 if started manually. A second targeted
Ansible run changed nothing; the generated service was inactive and neither
staging port listened.
The explicit rehearsal is managed by:
```bash
ansible-playbook ansible/site.yml --limit atlas --tags gitea_restore \
-e atlas_gitea_restore_test=true
```
On 2026-10-01 this selected the latest verified Prometheus backup, checked its
SHA-256, extracted only `opt/gitea/data`, moved `app.ini` into the rootless
config mount, rewrote `/data/` paths, enabled built-in SSH on internal port
2222, and retained the three source SSH host-key pairs. SQLite `quick_check`
passed, all 33 restored repositories passed `git fsck`, and each source/target
public host-key fingerprint matched. A temporary `1.25.2-rootless` container
with `--network none` answered HTTP internally and listened on internal
SSH/2222. The container was removed; the user Quadlet remains inactive, with
no staging listener. The second restore run changed nothing. This copy is
deliberately stale once new source writes occur and **must not** be used as the
final cutover copy.
Target backup checks on 2026-10-01: the managed recursive hourly ZFS snapshot
`atlas-auto-hourly-20261001T193401Z` contains the new dataset. The managed
Borg service completed archive `atlas-20261001T193420Z`, whose contents list
includes the staged Gitea database. A separate one-file restore from each
source into private `/var/tmp` directories matched the live staged database
and passed SQLite `quick_check`. Temporary files and the on-demand snapshot
mount were removed; the Borg temporary snapshot was cleaned up and the pool
remained healthy. This is file-level proof, **not** a full Gitea recovery.
The operator's UUID-bound offline USB run published version
`20261001T201220Z-254397` on 2026-10-02. A separate read-only mount and
temporary restore of `services/data/gitea/data/gitea/gitea.db` matched
contents, owner, group, mode, size, mtime and POSIX ACL; SQLite
`quick_check` returned `ok`. The temporary mount and copy were removed,
LUKS was closed, and the pool was healthy. This is a file-level restore test,
not a complete Gitea recovery rehearsal from USB.
1. Provision a dedicated target dataset and non-login service identity via
Ansible, keeping UID/GID distinct from Atlas' reserved Immich `1100`.
Install the user Quadlet in that identity's
`~/.config/containers/systemd/`, **without** an `[Install]` section;
do not enable, start, or expose it yet.
2. Verify the selected Atlas backup SHA-256 and metadata, then extract **only**
`opt/gitea/data` to private staging. Keep `home/git/.ssh` in the source
backup for rollback; the rootless image does not consume its OpenSSH mount.
Never unpack NPM,
WireGuard, or other host configuration from this sensitive tarball into a
live namespace. Convert the rootful `/data` tree on a disposable copy:
place application data under `/var/lib/gitea`, move `app.ini` to
`/etc/gitea`, and rewrite every absolute `/data/...` path for the new
layout. Enable `START_SSH_SERVER`, use internal SSH port 2222, and retain
the source host-key pairs for the built-in server only after verifying
their fingerprints and compatibility. Do not rely on the old
`/home/git/.ssh` OpenSSH mount in the rootless image. Set only the target
copy's ownership and path-scoped SELinux labels.
3. Validate SQLite integrity, repository count and representative `git fsck`,
LFS/attachment presence, permissions, and an isolated rootless test
container with no production ingress or outbound network. Because the
source stays active, this is a rehearsal copy, not the final cutover copy.
Regenerate Git hooks if the changed installation path requires it.
4. ZFS, Borg and UUID-bound offline USB inclusion and one-file restores have
passed. These do not replace the final consistent source copy.
## Phase 2: explicit final cutover
The opt-in `/usr/local/sbin/prometheus-gitea-final-export` helper was installed
on 2026-10-01 and passed `bash -n`. It refuses to
run while the scheduled Prometheus export timer is active. When explicitly
triggered, it stops only the source Gitea container, checks SQLite, publishes
a checksum-verified Gitea-only version for Atlas' existing pull, and leaves
the source stopped on success. NPM remains running. A failure before
completion restarts source Gitea. Its Ansible gate is
`--tags gitea_final_export -e server_gitea_final_export=true`.
After Atlas pulls that version, its separate
`--tags gitea_final_restore -e atlas_gitea_final_restore=true` gate accepts
only metadata marked `gitea-cutover`, validates a private staged replacement,
and swaps it for the marked rehearsal. The swap and its rollback path passed
synthetic tests on 2026-10-01; the live gate succeeded on 2026-10-02.
On 2026-10-02 the operator approved the outage. The final stopped-source
export `20261002T071525Z` passed the Atlas pull checksum; the guarded restore
replaced the rehearsal. SQLite `quick_check`, all 33 repository `git fsck`
checks, and the source/target SSH host-key comparison passed. The rootless
Atlas Quadlet serves LAN HTTP/3000 and SSH/2222, reachable from Prometheus
through Aegis; its firewall admits only Aegis. The final marker gates startup.
Prometheus now runs the NPM-only Compose stack. Both NPM database records still
say `gitea:3000`, but Nginx evaluates this variable upstream through its
runtime DNS resolver, which **does not** use a Compose `extra_hosts` alias.
The initial alias attempt returned 502. A managed `server_proxy.conf` override
sets `$server` to Atlas' IP for only the two declared Gitea domains; it passed
`nginx -t` and primary HTTPS/API returned 200 after a clean NPM restart
without the alias; a representative public `git ls-remote` also succeeded.
Navidrome and Syncthing Proxy Hosts still responded. No NPM SQLite records
or credentials were changed. The
secondary hostname `git.ov-ad3410.infomaniak.ch` did not resolve from Ikaros
and had no generated NPM config file at the time of inspection.
On 2026-10-03 the operator retired this unused secondary hostname. Its NPM
Proxy Host was already soft-deleted; Ansible now declares only
`git.fscotto.duckdns.org` and removes the secondary runtime override.
Prometheus' public TCP/2222 socket proxies to Atlas without changing admin
SSH/22. The local socket presents the preserved Gitea ED25519 host key, but
an external TCP/2222 connection from Ikaros initially timed out. During that
test no SYN reached Prometheus `eth0`; its socket and firewalld port were active.
After the VPS firewall was opened later on 2026-10-02, the public port connected,
its ED25519 host-key fingerprint matched Atlas, Gitea authenticated the `ikaros`
key as `fscotto`, and a public SSH `git ls-remote` for `fscotto/infra.git`
returned HEAD. The operator subsequently reported successful authenticated
SSH pull and push; the agent did not perform a write test. HTTPS write/login
remain untested. Do not
restart the stale source after public HTTPS has accepted target writes.
The Prometheus export timer resumed with NPM-only paths. A recursive ZFS
snapshot at `20261002T073032Z` and encrypted Borg archive
`atlas-20261002T073044Z` captured the Atlas target after cutover; Borg exited
successfully, cleaned its temporary snapshot, and the pool was healthy.
## Corrected Atlas service owner (2026-10-02)
The operator required the host Quadlet to belong to `admin`, while the Unix
user **inside** the container must be named `gitea`. The pinned derived
`Containerfile.gitea-rootless` changes only the base image's UID/GID 1000
passwd/group names from `git` to `gitea`; it retains the rootless image's
paths and entrypoint. Gitea's `RUN_USER` is `gitea`, while its built-in SSH
user and advertised clone user remain `git`, preserving `git@` URLs. The
selective restore helper now generates the same three settings for any future
explicit restore, instead of recreating a `RUN_USER = git` target.
A disposable, loopback-only container using a copy of a Gitea ZFS snapshot
passed HTTP, SQLite, internal-user and SSH host-key checks without touching
live data. After explicit outage approval, the opt-in
`--tags gitea_owner_migration -e atlas_gitea_owner_migration=true` run stopped
the old user service, took safety snapshot
`zpool/services/data/gitea@gitea-owner-migration-20261002T100104`, transferred
only the Gitea dataset to `admin`, tested an `admin` staging Quadlet on
loopback, then promoted it to the production LAN ports. The old Atlas Quadlet
was removed. The old host `gitea` account and its sub-ID range are retained
for a deliberate rollback; they must not restart stale Gitea. The parent
traverse ACL is removed by the normal Gitea role once the new owner is live.
The new service returned HTTP 200 locally and through public primary HTTPS;
Navidrome and Syncthing remained active under `admin`, the pool was healthy,
and a second normal Gitea Ansible run was idempotent. This does **not** close
the separate external TCP/2222 or authenticated clone/push validation gap.
1. Agree on an outage and record source/target versions, pool health, the
latest backups, SSH host-key fingerprints, and both current NPM routes.
Stop the Prometheus export timer for the change window so it cannot
restart the old Compose stack unexpectedly.
2. Quiesce source writes with the final-export helper: it stops Gitea before
the consistent export and leaves it stopped after success. Pull that export
to Atlas and verify checksum and timestamp. Keep
`/opt/gitea/data` and `/home/git/.ssh` intact for rollback. Do not allow
source Gitea to restart after accepting writes on Atlas.
3. Restore the final Gitea-only payload to the target and repeat integrity
checks. Verify its advertised SSH port is 2222, its existing HTTPS
`ROOT_URL`, repositories, LFS/attachments, and SSH host-key identity. Enable
the production Atlas Quadlet only after the final-restore marker exists;
its firewall permits only Aegis to reach HTTP and SSH. Validate local HTTP
and the target service before switching NPM.
4. Enable the public TCP/2222 socket proxy on Prometheus to Atlas over Aegis
without changing administrative TCP/22. Switch Prometheus to the desired
NPM-only Compose stack and use the managed Gitea-only NPM runtime upstream
override. Do not use Compose `extra_hosts`: Nginx bypasses it for the
variable upstream. The old Gitea data stays intact. Do not change public DNS.
5. Test HTTPS login, representative clone/push, LFS, and public SSH clone/push
on port 2222 from outside the Atlas LAN. Record the last source write and
first healthy target service times; do not claim RPO/RTO without measuring.
6. Resume the Prometheus NPM-only backup export timer after the desired stack
is active and verify its next result. Verify the next Atlas snapshot/Borg
run covers Gitea and test a restored target copy. Do not delete old source
data.
## Rollback gate
Before Atlas accepts writes, restore the old Compose definition and remove the
NPM override, disable the public 2222 proxy, and restart the unchanged source
Gitea if target validation fails. **After Atlas accepts writes, do not blindly restart the source:** its
SQLite database and repositories are stale. Quiesce Atlas, capture its new
data, and decide a reverse migration or an extended outage explicitly.
Upstream references: [rootful container layout](https://docs.gitea.com/1.25/installation/install-with-docker/),
[rootless image layout and incompatibility](https://docs.gitea.com/installation/install-with-docker-rootless/),
[rootless Podman Quadlet](https://docs.gitea.com/installation/install-with-podman-quadlet/),
[standard-image conversion](https://docs.gitea.com/1.24/installation/install-with-docker-rootless/),
and [restore and hook regeneration](https://docs.gitea.com/1.26/administration/backup-and-restore/).

View File

@@ -0,0 +1,173 @@
# iCloudPD: Aegis to Atlas
Atlas is the temporary ingestion host until Uranus. Aegis iCloudPD and its
state were retired. Ansible declares Atlas storage, the rootless Quadlet,
and a private `icloudpd.conf` with the Apple ID from the existing Vault key.
The password, keyring and MFA cookies remain application-managed; initialization
is interactive.
Do not place cookies, keyring files, passwords, or the Apple ID in this document,
unencrypted repository content, or a terminal transcript.
## Historical source and current destination (2026-10-02)
- Before retirement, Aegis' rootful `icloudpd.service` was active (no reported restarts, running
since 2026-07-25), but its declared data bind `/var/lib/icloudpd/data`
has **zero top-level entries** and is 4 KiB as observed on 2026-10-02.
Its persistent config has two top-level entries. `pi` cannot run passwordless
sudo, so the container's internal filesystem and root-only state have **not**
been audited. Do not conclude there are no photos to preserve: they could be
inside the container overlay because the declared bind targets the wrong
home. The current
Quadlet mounts that data directory at `/home/root/iCloud`; the image's
documented default is `/home/user/iCloud` with its default `user=user`.
- The non-secret `folder_structure` value in the persisted Aegis config is a
systemd generator path, **not** `{:%Y/%m/%d}`. The Quadlet passes percent
characters in `Environment=` without systemd escaping; that is the likely
cause. A running unit therefore does not prove that Aegis ingests photos.
Do not copy this config or assume that its MFA state is usable on Atlas.
- Atlas' `zpool` is healthy. `/zpool/archive/Pictures` already contains about
25 GiB of unrelated data; iCloudPD gets only a new managed
`/zpool/archive/Pictures/iCloudPD` subtree. Both that subtree and
`zpool/services/data/icloudpd` were created on 2026-10-02. Never rsync with `--delete` into
Pictures or adopt its existing contents. `/zpool/media/photobook` is reserved
for Immich and remains untouched, including its Aegis-only NFS export.
The upstream image documents `/config/icloudpd.conf` as its primary
configuration (environment configuration is deprecated), an exact
`/home/${user}/iCloud/.mounted` failsafe, and an interactive `--Initialise`
step for keyring and MFA cookies. The configuration must use the same download
path, user/UID, and folder format as the bind mounts. References:
[image configuration](https://github.com/boredazfcuk/docker-icloudpd/blob/master/CONFIGURATION.md),
[Podman user namespaces](https://docs.podman.io/en/latest/markdown/podman-pod.unit.5.html).
## Declared Atlas target
| Item | Location or policy |
| --- | --- |
| Downloaded photos | `/zpool/archive/Pictures/iCloudPD`, a new managed subtree of the SMB `Archive` dataset |
| Config, keyring, MFA cookies | `zpool/services/data/icloudpd` at `/zpool/services/data/icloudpd/config`, outside Archive |
| Host service owner | `admin` rootless user manager; no rootful Quadlet or published port |
| Container identity | Entry process root in its user namespace; downloader UID/GID 1000 maps to host `admin` |
| Image | Digest-pinned `docker.io/boredazfcuk/icloudpd`, with no registry auto-update |
| SELinux | Private `:Z` config bind; shared `:z` photo bind because Archive is also exposed through SMB and used by Syncthing. The label and SMB behavior require runtime testing. |
| Access | The new subtree is `admin:admin` mode 0750. No Photobook ownership, ACL, or export changes. |
| Sync policy | Daily interval; explicit directory/file modes 750/640; no iCloud deletion and no deletion of destination-only files |
The photo subtree receives a managed marker and the image's `.mounted` file.
An existing unmarked path is refused rather than taken over. The existing
Pictures tree is not chowned or emptied. The Quadlet now has `[Install]` with
`WantedBy=default.target`, so the lingering admin user manager starts it at boot.
Ansible keeps the service running. Ansible renders a mode-0600
`icloudpd.conf` with `no_log` and no diff, but does not pull the image,
initialize MFA, or run a cutover task. Boot startup was approved on 2026-10-03
after a reboot left the previously manual-started service inactive.
The previous gated check-mode tests and isolated Quadlet-generator test proved
only the proposed layout; they predate the simplified declarative role. They
were not a production deployment or an authentication test.
## Evidence already gathered without production writes
The digest-pinned image was pulled into **admin's** Atlas Podman store. An
isolated `/var/tmp` test ran with no network, a fake Apple ID, private temporary
config/photo mounts, `keep-id:uid=1000,gid=1000`, and no new privileges. Both
container root and UID 1000 wrote to the mounts; UID
1000's files mapped to host `admin`. A short-lived container remained running,
retained the intended `/home/user/iCloud` and literal `{:%Y/%m/%d}` config,
and saw an admin-owned `.mounted` marker. The container and temporary files
were removed. A second isolated test showed that dropping **all** container
capabilities prevents its root entrypoint from reading an admin-owned 0600
config; with the default rootless user-namespace capabilities it could read
and write that file. The Quadlet retains `NoNewPrivileges=true` but does not
drop every capability. This proves only the container layout and namespace mapping,
**not** Apple authentication, a real download, SMB visibility, scheduled
operation, backup coverage, or recovery.
The earlier disposable Photobook ACL test is superseded by the operator's
clarification that Photobook belongs to Immich. It is not evidence for the
current Archive destination, and the proposed Photobook ACL change was never
deployed.
Backup path review on 2026-10-02: the managed Borg and USB scripts snapshot
the pool recursively and bind every mounted child dataset, so both
`archive` and the proposed `services/data/icloudpd` fall within their
declared source scope. Borg's runner switches to the dedicated `borg` account
with only `CAP_DAC_READ_SEARCH`; a read-only check using those exact `setpriv`
capability flags could traverse/read Archive, whereas plain
`sudo -u borg` could not. USB copies as root and preserves POSIX ACLs, but not
generic xattrs/SELinux labels. **This was scope and permission evidence, not a
completed backup or restore of iCloudPD data**, which did not exist at the time.
## Validation status and remaining checks
- Aegis retirement is complete: `icloudpd.service` is `not-found`/`inactive`,
the rootful Quadlet and `/var/lib/icloudpd` are absent, and AdGuard is active.
The temporary retirement tasks are no longer in the Aegis role. The Podman
image cache may remain; it is not service data.
- Atlas storage and the `admin` Quadlet are deployed. The second Ansible
run changed nothing and did not start the service; a later manual start
generated the config. `/zpool/media/photobook` was unchanged.
- The image generated `/zpool/services/data/icloudpd/config/icloudpd.conf`
on first start. Ansible replaced that default file with a private template
using the Apple ID already in Vault. The operator initialized password
and MFA interactively; never put credentials or codes in the repository,
chat, or Ansible extra-vars. Automatic boot startup was separately approved
on 2026-10-03; this does not change the interactive MFA procedure.
- Initial ingestion completed on 2026-10-03. Still check folder structure,
ownership, SELinux and SMB access, no unintended deletions, the next daily
cycle, completed Borg and USB versions, and isolated restore of photos and
private state. A recursive hourly `zpool/archive` snapshot exists after
ingestion, but no iCloudPD-specific backup restore has passed. The first
real scrub and measured recovery targets are separate open items.
On 2026-10-02 Atlas storage and the inactive Quadlet were deployed; a second
Ansible run made zero changes. The generated service was inactive, and no
`icloudpd.conf` existed. Two interactive-sudo Aegis runs removed its service,
Quadlet and `/var/lib/icloudpd`, then cleared the failed-unit record left by a
SIGKILL during shutdown. Read-only verification found `LoadState=not-found`,
`ActiveState=inactive`, both paths absent, and AdGuard active.
On 2026-10-02 the operator requested the first manual start. The rootless
service stayed active, and the image generated `icloudpd.conf` under the
private config dataset. Its mode was tightened from 0644 to 0600. The generated
`apple_id` field is empty; no MFA or download is verified. The service has no
boot-time install target, so it is not configured for automatic startup.
The 2026-10-02 Atlas `icloudpd` run rendered the Vault-backed template without
printing its contents; the second run made zero changes. File owner is
`admin:admin`, mode 0600, and the Apple ID field is nonempty. The rootless
service remained active with zero restarts. At that point keyring initialization,
cookie creation and a real download were unverified. The template now reads
`vault_atlas_icloudpd_apple_id`, which is already present in the encrypted
Vault; no password or MFA code was added to the template.
The attempted interactive initialization then lost its container. Diagnosis
found that the image launcher requires `traceroute` to pass its iCloud
reachability check. Rootless Podman without `NET_RAW` returned `Operation not
permitted` despite working Atlas/container DNS and host HTTPS. An isolated
container with only `CAP_NET_RAW` passed the same check. The Quadlet now grants
that single capability while keeping `NoNewPrivileges=true`; a manual restart
passed `traceroute`, and the app stayed running. Logs then showed only the missing
keyring and a wait for `--Initialise` again. The app expanded the generated config
on startup, so Ansible now seeds it only when absent and idempotently maintains
only its declared options. A second live Ansible run made zero changes. At
that point MFA, actual ingestion, and backup/restore were unverified.
On 2026-10-03, after interactive initialization, the rootless service was
active and the previous 24h of logs showed download activity with no
authentication failures or errors. At 02:16 the application reported `All
photos and videos have been downloaded` and `Download complete for user`.
The destination contained 11,658 files totaling 86,020,430,015 bytes; this
is a filesystem file count, not a count of distinct iCloud assets. A later
read-only check found the service still active. This closes initial
authentication and ingestion only: a subsequent daily cycle and end-to-end
recovery of the new photos and private state remain untested.
On 2026-10-03 Atlas rebooted at 10:17 CEST; iCloudPD stayed inactive because
its Quadlet had no install target. A manual start restored the running service
and the application began listing iCloud files. The operator then approved
persistent boot startup. The managed Quadlet now declares
`WantedBy=default.target`; the live generator created
`default.target.wants/atlas-icloudpd.service`, admin has `Linger=yes`, and the
service remained active with zero restarts. No NAS reboot was performed to
test this change; actual post-reboot startup remains untested.

141
docs/atlas-recovery.md Normal file
View File

@@ -0,0 +1,141 @@
# Atlas recovery runbook
This runbook is for a **replacement Rocky Linux 9 installation**, not a normal
playbook run. A scaled whole-OS rebuild with a disposable pool passed in an
isolated VM on 2026-09-30, but no production-size whole-host recovery has been
tested. The existing production pool must be imported, never created or
rewritten. The provisional targets are **RPO 24 hours**
and **RTO 72 hours**, for Archive and Atlas services alike. They are planning
objectives, not demonstrated recovery times. The manual USB cadence may leave
an older copy; a recent Borg archive is needed to meet the RPO after total
pool loss.
## Before an incident
- Keep an offline copy of the encrypted Ansible Vault, its unlock material,
the exported Borg repository key, and the Borg passphrase. Do not store
unlock material in this repository or in a recovery command line.
On 2026-09-30 the operator confirmed these are available independently of
Atlas and the Ansible controller; their usability has not been tested here.
- Keep the Atlas installation media and a reproducible checkout of this
repository available independently of Atlas. Record the exact Git revision
used for a successful deployment.
- Record the pool's current disk identities with `zpool status -P zpool` and
`lsblk -o NAME,SIZE,MODEL,SERIAL,FSTYPE,UUID`. Compare these with
`atlas_zpool_disks` before touching a replacement host. The `host_vars`
values are historical identifiers, not evidence that a newly attached disk
is the same device.
- Verify that the latest hourly/daily snapshots, Borg archive, and offline USB
version exist and note their timestamps. A timer being enabled is not proof
that a backup completed.
## Incident gate
1. Identify whether the fault is the OS disk, one or more pool disks, accidental
deletion, or an unavailable host. Preserve failed media when possible.
2. Stop writes to affected services and capture the last known good backup
timestamps. Do not run `zpool create`, `zpool destroy`, `zfs rollback`,
`zpool import -F`, `zpool import -X`, `zpool import -f`, or disk formatting
as a diagnostic shortcut.
3. Choose one recovery source below. Do not merge several sources into the
production namespace without comparing their timestamps and content.
## Rebuild the OS and import the existing pool
1. Install Rocky Linux 9 on a **separate system disk**. Configure basic network,
SSH, a temporary sudo administrator, SELinux enforcing, and the current
OpenZFS kmod repository. Keep the pool drives untouched.
2. Run read-only identification: `lsblk -f`, `zpool import`, and
`zpool import -d /dev/disk/by-id`. Check the pool GUID, vdev layout, and
stable drive identities against the incident record. If any differ, stop.
3. Import only after matching the expected pool and host ownership. A pool
cleanly exported from the old host can be imported with
`zpool import -d /dev/disk/by-id zpool`. If it reports that the pool is
active elsewhere or needs a rewind/force, stop and investigate rather than
adding flags. Verify with `zpool status -v zpool`, `zfs list -r zpool`,
`zfs get -r mountpoint,canmount zpool`, and `findmnt -R /zpool`.
4. Leave `atlas_create_pool: false`. Ensure `host_vars/atlas.yml` reflects the
replacement host's actual SSH address and disk identities before running
Ansible. Apply `ansible/site.yml --limit atlas` with the bootstrap admin
connection override as documented in the Atlas setup section of README.
This may start shares/services, so keep clients disconnected or services
gated until data and permissions are verified.
5. Check `getenforce`, `zpool status -v zpool`, `systemctl --failed`, SSH,
firewalld, Cockpit, NFS, SMB, and the backup/monitoring timers. Do not
report recovery complete on the basis of Ansible success alone.
## Choose the data source
- **Local snapshot, pool intact:** inspect `zfs list -t snapshot -r zpool`.
Mount/access the chosen snapshot read-only and copy selected files to an
empty staging directory; compare content, owner, mode, mtime, and POSIX ACL.
Move into the live namespace only after an operator-approved scope review.
Do not use an automatic rollback: it can discard newer changes in the
dataset and descendants.
- **Offline USB:** verify the configured LUKS and ext4 UUIDs from
`host_vars/atlas.yml` before unlocking. Mount ext4 read-only with `ro,noload`,
use only a published `atlas/latest` version, and restore to an empty staging
directory. Compare checksums and metadata. The USB copy intentionally omits
generic xattrs and SELinux labels; relabel only the restored destination.
Never run the backup service to perform a restore.
- **Hetzner Borg:** use the dedicated pinned host key, repository path,
offline exported recovery key, and Vault-backed passphrase. List archives
and extract a selected archive into an empty staging directory, never the
live `/zpool` tree. A repository check and sample restore were previously
performed; that does not prove this incident's archive is complete. Compare
content and metadata before publication. Avoid `borg break-lock` while any
backup/check job may still be active.
After publishing restored files, run the explicit Ansible `restorecon` tag only
for the paths actually restored, for example:
```bash
ansible-playbook ansible/site.yml --limit atlas --tags restorecon \
-e '{"atlas_restorecon_paths":["/zpool/archive"]}'
```
Then check ownership/ACLs, application-specific integrity, SMB/NFS client
access, backup service health, and `zpool status -v zpool`. Reconnect clients
only after these checks pass. Record the last recoverable timestamp (actual
RPO) and elapsed service outage (actual RTO) in the incident log.
## Scaled isolated rehearsal (2026-09-30)
The lab setup, repeatable checks, and preserved VM state are recorded in
[`atlas-dr-lab.md`](atlas-dr-lab.md).
On Ikaros, a local libvirt `atlas-dr-lab` VM used a 30 GiB Rocky 9.8 system
disk and four separate, disposable 4 GiB virtio data disks with stable
`/dev/disk/by-id` identities. The official Rocky cloud image matched its
published SHA-256. The lab inventory was separate from production, used a
fresh lab-only password hash and the operator's public SSH key, and disabled
sharing, Borg, USB backup, monitoring, media services, and the Prometheus pull.
No production disk, Vault secret, or production data was attached or copied.
1. The existing `packages_rocky` and `profile_atlas` roles installed OpenZFS,
created a RAIDZ2 `zpool` through the explicit one-time pool gate, and built
all 12 declared datasets with a lab-sized 1 GiB backup reservation. The
pool creation gate was set false immediately afterward.
2. A 4 MiB canary file was written under the lab `archive` dataset and a ZFS
snapshot created. The pool was cleanly exported and the VM shut down.
3. Only the system-disk volume was replaced by a fresh Rocky cloud image;
the four virtio data volumes were retained. Ansible reinstalled OpenZFS.
Read-only `zpool import -d /dev/disk/by-id` showed the expected RAIDZ2
topology and pool GUID `8880368391795119587` before an ordinary import
without `-f`, rewind, or rollback.
4. The imported pool was healthy. The canary SHA-256 matched its pre-rebuild
value. `profile_atlas` completed against the imported pool and a second
run reported `changed=0`. A file restored from the preserved snapshot into
`/var/tmp` matched SHA-256, owner, group, mode, size, and mtime; the temporary
copy was removed. Final checks found SELinux Enforcing, 12 datasets, the
snapshot, no failed units, and a healthy pool. The VM was shut down while
retaining its disposable volumes for a future rehearsal.
This proves the **sequence** for a cleanly exported, small pool and the tested
Ansible subset, not recovery duration or capacity at 2 TB. The earlier
2026-09-25 independent production ZFS/USB file restores and the earlier Borg
temporary-directory restore remain separate evidence. The VM did not restore
production USB/Borg archives, exercise services with production data, test an
unclean import, or prove the provisional RPO/RTO. Before relying on 24h/72h,
measure a representative full restore and service cutover in a suitably sized
future change window. Never use the production Atlas pool for a rehearsal.

View File

@@ -0,0 +1,20 @@
# Atlas SMB/NFS namespace decision
Decision date: 2026-09-30. Keep the current namespaces **separate**.
- `/zpool/archive` is the SMB3 `Archive` share for authorized Samba accounts.
- `/zpool/media/photobook` is the Aegis-only NFSv4 export, `all_squash`-mapped
to UID/GID `1100`.
- No new dual-protocol namespace, broad export, group, or ACL model is needed.
Existing permissions and client access remain unchanged.
The two paths serve different ownership and exposure needs. A common namespace
would expand the permissions design and require same-file SMB/NFS interoperability
testing without a present requirement. Revisit only when a specific workflow
needs both protocols on the same files; then decide UID/GID, group, POSIX ACL,
SELinux policy and client behavior before changing exports or permissions.
Read-only Atlas verification on 2026-09-30 confirmed that Samba `Archive` points
to `/zpool/archive`, NFS exports `/zpool/media/photobook` only to
`192.168.178.54` with `all_squash` and anonymous UID/GID `1100`, both datasets
are distinct, and `zpool` is healthy. No sharing configuration was changed.

63
docs/atlas-updates.md Normal file
View File

@@ -0,0 +1,63 @@
# Atlas Rocky/OpenZFS update and reboot procedure (draft)
This is an operator-controlled maintenance procedure. The playbook does not
reboot Atlas, replace a pool device, or perform a pool feature upgrade.
## Preflight
1. Schedule an outage and confirm no Borg, USB, snapshot, scrub, or resilver
job is active. A service in `activating` is still active; do not interrupt it.
2. Check `zpool status -v zpool` (including scrub status), `zfs list -r zpool`,
`systemctl --failed`, and `systemctl list-timers --all`. Resolve pool errors
first. Record current `uname -r`, `modinfo zfs | grep '^version:'`,
`rpm -q kernel-core kmod-zfs zfs`, and the current boot entry.
3. Confirm a recent successful Borg archive and a usable snapshot. Confirm
the latest published offline USB version and its physical availability;
do not start a USB backup merely to satisfy a checklist without capacity,
UUID, and operator checks. Record timestamps, not just timer state.
4. Ensure console/KVM or another independent recovery route is available.
Check free space in `/boot` and the root filesystem. Review proposed DNF
transactions before consenting to package changes.
## Change window
1. Stop client writes and quiesce stateful applications deliberately. Record
which services were stopped; do not assume `ansible-playbook --check` does
this. Avoid updating during a running scrub or backup.
2. Use `dnf upgrade --assumeno` first to review the kernel, `kmod-zfs`, `zfs`,
and dependencies. Confirm a matching kmod will be available for the target
kernel. If compatibility is uncertain, defer the update.
3. Apply the approved DNF transaction. Do not run `zpool upgrade` or enable
new pool feature flags as part of ordinary OS maintenance; that can remove
downgrade options. Preserve at least one known-good boot entry.
4. Reboot **manually** during the agreed outage. Ansible must not trigger it.
## Post-boot gate
1. Verify `uname -r`, `modinfo zfs`, `rpm -q kernel-core kmod-zfs zfs`,
`zpool status -v zpool`, `zfs list -r zpool`, and `findmnt -R /zpool`.
2. Verify SELinux remains enforcing; inspect `systemctl --failed` and the
journal for ZFS, mount, SSH, NFS, SMB, Cockpit, Podman, and backup errors.
3. Validate a read-only file listing through SMB and an NFS client access
check before reopening writes. Check the rootless temporary services and
all backup/monitoring timers. Run the Atlas health monitor in `--dry-run`
mode, then a real check after inspection.
4. Re-enable clients and record versions, downtime, anomalies, and next
successful snapshot/Borg run. A green boot alone is not a completed update.
## Failure response
If the new kernel cannot load ZFS, boot the previous known-good kernel from
the console and inspect package/kmod matching before trying another reboot.
Do not force-import, rewind, clear errors, or upgrade pool features to make a
failed OS update appear successful. Preserve logs and stop for a recovery
decision if the pool does not import cleanly.
The procedure-definition item is complete, but the procedure is **not yet
rehearsed** on a replacement host or during a real Atlas update. Record the
first controlled execution and its post-boot evidence separately.
Read-only preflight on 2026-09-30 observed kernel
`5.14.0-687.52.1.el9_8.x86_64`, ZFS module/package `2.2.11-1`, a healthy
`zpool`, enforcing SELinux, and no failed systemd units. This did not review
an upgrade transaction, stop services, or reboot the host.

82
docs/domain-fscotto-co.md Normal file
View File

@@ -0,0 +1,82 @@
# fscotto.co domain transition
## Observed state (2026-10-03)
Namecheap remains the DNS provider. The operator moved GitHub Pages to
`blog.fscotto.co` in `fscotto/fscotto.github.io`, aligned Hugo and Pages
settings, and changed the apex A record to `179.237.102.172`. The blog
remains a CNAME to `fscotto.github.io`; mail records were left unchanged.
A new Hugo deployment and cache clearing resolved the initial stale DNS
and generated URLs. Blog HTTPS returned 200 with valid TLS.
The `git`, `music` and `syncthing` subdomains are CNAMEs to `fscotto.co`.
The operator added NPM Proxy Hosts with certificates, WebSocket support
and Force SSL:
| Hostname | HTTP upstream |
| --- | --- |
| git.fscotto.co | 192.168.178.55:3000 |
| music.fscotto.co | 192.168.178.55:4533 |
| syncthing.fscotto.co | 192.168.178.55:8384 |
All three redirected HTTP to HTTPS and returned final HTTPS 200 with valid
TLS. Only the Syncthing GUI uses NPM; native synchronization is unchanged.
NPM administration remains loopback-only on port 81 via SSH tunnel.
## Gitea canonical hostname
Atlas declares `atlas_gitea_public_domain: git.fscotto.co`. Ansible manages
only `[server] DOMAIN`, `ROOT_URL` and `SSH_DOMAIN` in the existing private
app.ini, preserving unrelated settings and mode 0600. Private configuration
backups are created; diffs and secret-bearing results are suppressed.
Only Gitea restarts when these fields change; a repeat run changed nothing.
HTTPS uses `https://git.fscotto.co/`; public SSH remains TCP/2222.
Agent read-only checks returned the same HEAD from `fscotto/infra.git`
over HTTPS and authenticated SSH. SSH host identity was checked against
the already-trusted old endpoint key. No test push or user-authenticated
web login was performed by the agent.
```bash
ansible-playbook ansible/site.yml --limit atlas --tags gitea_public_domain --check --diff
```
Client remotes do not update automatically. Update them deliberately after
checking repository paths; integrations and webhooks are separate operations.
For the verified infrastructure repository only:
```bash
git remote set-url origin ssh://git@git.fscotto.co:2222/fscotto/infra.git
```
Do not copy this path into unrelated clones. Verify Gitea's known SSH key
before accepting the new hostname's identity.
## Local DuckDNS retirement
Prometheus declares `server_duckdns_enabled: false`. On 2026-10-03 the explicit
Ansible cleanup removed the five-minute rocky cron entry and the private
`~/duckdns` directory containing only `duck.sh` and `duck.log`. The temporary
cleanup tasks and flag were subsequently removed from the playbook at the
operator's request. Only the disabled provisioning state remains; ordinary
provisioning cannot recreate the updater.
The external DuckDNS name, Vault token, disabled NPM hosts and certificates
remain untouched for a separate future decision.
The repeat cleanup changed nothing; ordinary DuckDNS provisioning was skipped.
The cron table had no remaining entries, NPM and the export timer were active,
and NPM administration still listened only on `127.0.0.1:81`.
## Operator-confirmed transition completion
On 2026-10-03 the operator confirmed completion of:
- Web login on the new Gitea hostname.
- Updates to remaining Git remotes, webhooks and integrations.
- Removal of obsolete DuckDNS NPM Proxy Hosts, unused certificates and the old upstream override.
- Review and removal of completed one-time procedures from the playbook.
These are operator confirmations, not new agent runtime checks or a test push.
At the earlier inspection the three old DuckDNS Proxy Hosts were disabled,
not deleted; that observation predates the confirmed cleanup. Existing backup
archives remain preserved. DNS/Pages/NPM changes were operator actions;
the Gitea application configuration change was deployed through Ansible.

131
docs/prometheus-backup.md Normal file
View File

@@ -0,0 +1,131 @@
# Prometheus to Atlas backup pull
The playbook and both hosts have the dedicated identity, restricted SSH
access, helpers, and systemd units. A manual export, pull, and temporary
restore passed on 2026-09-30. The first scheduled export and pull passed on
2026-10-01. After NPM moved to its Quadlet, another manual export, pull, and
isolated restore passed on 2026-10-03. The first scheduled cycle after that
cutover is still pending. See `docs/prometheus-npm-quadlet.md`.
## Declared design
- Prometheus prepares a tar archive of Nginx Proxy Manager data and certificates,
its active Quadlet and network definitions,
and SSH/firewalld/WireGuard configuration. Gitea now runs on Atlas and is no
longer included in new Prometheus exports. NPM access logs are excluded.
The archive contains credentials, certificates, and the WireGuard private
key: protect both copies accordingly.
- The approved consistency mode stops the NPM Quadlet for local tar creation
at 02:00 Europe/Rome, then restarts it even if archiving fails. After the
approved legacy cleanup, the helper requires the Quadlet active and has
no Compose dependency. A manual test outside that window requires separate approval.
- Prometheus publishes the archive with its checksum as a versioned, read-only
source under `/var/lib/prometheus-backup-export`. A locked service account
has no sudo or supplementary groups. Its only authorized SSH key is forced
through Rocky's `rrsync -ro`; root owns the key file and export directories,
so the account cannot add an unrestricted key or change prepared data.
- Atlas generates and retains the private Ed25519 identity under
`/etc/atlas-prometheus-pull`. Its pinned Prometheus host key came through
the controller's already strict SSH trust; the observed fingerprint was
`SHA256:rfedk7DHI9mLB3UHk/4F3HHlSIiswtCAFsAXvfh6iXk` on 2026-09-30.
Atlas pulls only the prepared `current/` version, verifies SHA-256, tar
readability, metadata, and source freshness, then publishes atomically
below `/zpool/backup/hosts/prometheus/snapshots`. Long-term retention runs
only after publication. A local `rrsync` fixture verified the in-tree
`current` symlink. A live Atlas-to-Prometheus SSH test verified that the
account could list only the prepared versions directory,
cannot obtain a shell, and cannot write to the export. The key is restricted
to `/var/lib/prometheus-backup-export/versions`, not the account's `.ssh`.
- Approved source preparation is 02:00 Europe/Rome, pull 03:00, three source
versions, and 30 daily/8 weekly/12 monthly Atlas versions. The source
timer is non-persistent to avoid an unexpected outage after a missed run.
Atlas rejects a prepared source older than 24 hours.
- The Atlas pull joins the existing health monitor's timer/failure checks
only when enabled. Its failure hook uses 45Drives Alerts; email delivery
is not claimed. A failed source preparation should produce a stale-source
pull failure, not a silently successful reuse of an old archive.
## Activation and verification
1. The user confirmed downtime/consistency mode, schedule, retention, and
targeted configuration scope. Review the tar path list and exclusions
against the actual containers.
2. The identity and units are deployed. Re-run the targeted
check, confirm the Atlas public key remains only the restricted Prometheus
account's key, and verify `sshd -T -C user=prometheus-backup,...` plus
read-only SSH denial tests after any SSH configuration change.
3. During an agreed window, start the Prometheus export service manually.
Confirm the active NPM service is healthy afterward, inspect the archive
without exposing file contents, and verify the checksum/metadata.
4. Start the Atlas pull service manually. Confirm the SSH host pin, source
freshness, checksum, tar listing, published `latest`, retention behavior,
clean temporary directories, and healthy pool.
5. Independently restore the selected archive to an empty staging directory
(never `/`) and compare NPM SQLite, data, active Quadlet files, certificates,
permissions, and representative files. Historical pre-Gitea-cutover
versions also include Gitea repositories; current versions do not. Test
application startup only in an isolated environment or an approved restore
window.
6. Both timers are enabled. Verify their calendars and the next actual run
after any service-ownership change. A successful manual test is not proof
of a later scheduled cycle.
Narrow static validation:
```bash
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
ansible-playbook ansible/site.yml --syntax-check
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
ansible-playbook ansible/site.yml --limit prometheus,atlas \
--tags prometheus_backup --check --diff
```
Do not run the export service as part of a routine playbook deployment. The
service restart and any restore/cutover require separate operator decisions.
On 2026-09-30 the initial targeted `--check --diff` run ended `changed=0`
with gates false. After enabling **implementation only**, a targeted real run
installed the identities and units; both timers were confirmed `disabled` and
`inactive`, the Compose stack stayed active, and the new account was locked
with no supplementary groups. No application was stopped.
The rendered shell helpers passed `bash -n` and ShellCheck; the retention
helper passed an isolated 400-version fixture. These static/isolated checks
were followed by live SSH, export, pull, and temporary restore checks.
Read-only preflight on 2026-09-30 found the Compose service active, all
declared source paths present, both timers inactive, and no prepared versions.
The source filesystem had about 6.0 GB free. Of the 2.1 GB NPM data tree,
2.1 GB was excluded access logs, so the expected archive is much smaller than
the raw tree size; capacity still needs verification after actual exports.
The manual export produced a 285,777,920-byte tar (273 MiB allocated at the
source), and Prometheus retained about 5.8 GB free. NPM and Gitea restarted;
both containers were running and their local HTTP endpoints returned 200.
Atlas pulled the same version, verified SHA-256, published `latest`, and kept
the pool healthy. A full extract to `/var/tmp` yielded 4,747 files; both
SQLite databases passed `PRAGMA integrity_check`, and one restored Gitea Git
repository passed `git fsck`. The temporary restore directory was removed.
This did not test application startup on an isolated host.
After these checks, Ansible enabled the Prometheus 02:00 Europe/Rome export
timer and Atlas 03:00 Europe/Rome pull timer. Their first scheduled run passed
on 2026-10-01; Atlas verified and published `20261001T000001Z` as `latest`.
Atlas' health monitor includes the pull timer.
On 2026-10-03 the stopped-source version `20261003T091009Z` was verified and
pulled before the NPM cutover. The post-cutover version `20261003T091633Z`
was exported by the Quadlet-aware helper, checksum-verified, pulled to Atlas,
and restored to an isolated temporary directory. NPM SQLite `quick_check`
passed with ten proxy hosts and six certificate records. The archive contains
both Quadlet definitions. A manifest of all 70 regular Let's Encrypt files
and 12 symlinks, including content hashes and link targets, matched the live
Prometheus tree. No private key or secret content was printed. The next
scheduled export/pull is still pending observation.
## Post-cleanup validation (2026-10-03)
The operator-approved removal of legacy data and Compose fallback also
removed those backup input paths and the obsolete Gitea mount dependency.
A separately approved export and Atlas pull published `20261003T112906Z`.
Both SHA-256 checks passed; an isolated SQLite restore passed `quick_check`
and contained ten proxy hosts. Both active Quadlet definitions were present;
retired paths were absent. Existing backup archives were not deleted by cleanup.
The first scheduled cycle after these changes remains unverified.

View File

@@ -0,0 +1,145 @@
# Prometheus NPM Quadlet cutover
## Current state (2026-10-03)
Nginx Proxy Manager runs as the **rootful** generated
`prometheus-npm.service` on Prometheus. The Quadlet files are
`/etc/containers/systemd/prometheus-npm.container` and
`/etc/containers/systemd/server-web.network`; the image is pinned by digest
in `ansible/inventory/host_vars/prometheus.yml`. The generated service is
wanted by `multi-user.target` and requires the generated network service.
The old Compose unit, Compose file and Gitea final-export helper were
removed by the operator-approved cleanup on 2026-10-03. The retired
application data and empty legacy directories were also removed.
Prometheus host vars set `server_legacy_stack_retired: true` so normal runs
do not recreate those files. Destructive deletion still requires a separate
cleanup tag and explicit extra-var.
There was **no data copy** in this cutover. The Quadlet reuses the existing
`/opt/npm/data:/data` and `/opt/npm/letsencrypt:/etc/letsencrypt` bind mounts
with the same container name and `server_web` bridge (`10.89.0.0/24`). Ports
80 and 443 remain public; administration port 81 remains bound to
`127.0.0.1`. Gitea stays on Atlas, and NPM remains on Prometheus. The
Compose fallback is no longer installed. The Quadlet uses `Pull=missing`,
not an automatic floating-tag update.
## Cutover and recovery boundaries
The separate `scripts/cutover_prometheus_npm_quadlet.sh` was run **once** in
the approved outage window, after source backup version
`20261003T091009Z` was checksum-verified and pulled to Atlas. Its preflight
required exactly the Compose owner, an inactive generated Quadlet, the
expected image, and the current backup version. The execution held the
backup-export lock, stopped the export timer, stopped and disabled Compose,
started the Quadlet, checked the exact image ID, SQLite database counts,
certificate content, Nginx configuration, and local Gitea/Syncthing HTTPS,
then restarted the timer. Its failure trap would have restarted Compose.
**Do not rerun that forward-cutover script after success**: its preconditions
intentionally reject an active Quadlet.
Recovery is now a Quadlet rebuild and restoration from a verified Atlas
backup, with an explicit outage decision before replacing live NPM state.
The old Compose owner is no longer installed; reintroducing it would require
a separately reviewed configuration and outage plan. The historical
in-window rollback trap is not a supported post-cleanup rollback procedure.
Do not restore an old database over a live instance or remove NPM bind mounts.
## Verified evidence
- Immediately after cutover, `prometheus-npm.service` was active with zero
recorded restarts; Compose was inactive/disabled. The generated
`multi-user.target.wants` link and network dependency were present. An
actual reboot has not been performed solely for this test.
- The running image ID matched the prior Compose image. Podman showed the
original two bind mounts, `server_web`, public 80/443, and loopback-only 81.
External HTTPS to Gitea and Syncthing returned 200 with TLS verification
result 0. External access to TCP/81 timed out.
- The first **manual post-cutover** export `20261003T091633Z` succeeded with
the Quadlet as its active owner. The Atlas pull published that version;
its SHA-256 payload check passed. An isolated restore passed NPM SQLite
`quick_check` with ten proxy hosts and six certificate records. Both
Quadlet definitions were present in the tar archive.
- A path/content manifest of all 70 regular Let's Encrypt files and the
path/target manifest of all 12 symlinks in the Atlas archive exactly
matched the live Prometheus tree (aggregate SHA-256
`ce0965fbd3ff44bb8502ed9f314e0131edd86d822039de115b39f6a2273c2da8`).
The earlier apparent 70-vs-82 count was only a regular-file-versus-symlink
counting difference, not missing certificate data. No certificate key
contents were exposed during comparison.
- The targeted `--tags npm_quadlet` normal Ansible run completed with
`changed=0`, and the backup export timer remained active/enabled.
The first unattended 02:00 Europe/Rome export and 03:00 Atlas pull **after**
this cutover have not yet occurred. Check their service results and the
published version after the next cycle; the successful manual cycle proves
the new path works but not its next scheduled execution.
```bash
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
ansible-playbook ansible/site.yml --limit prometheus --tags npm_quadlet --check --diff
sudo systemctl status prometheus-npm.service prometheus-backup-export.timer
sudo systemctl show podman-compose-server.service -p LoadState # expected: not-found
```
The backup archive includes credentials, certificates, and WireGuard
configuration. Do not publish it or print its contents in diagnostics; see
`docs/prometheus-backup.md` for the restricted pull and restore procedure.
## Selective legacy image cleanup
On 2026-10-03 opt-in Ansible tasks removed only the unused Gitea 1.25.2,
Navidrome latest and PostgreSQL 13 rootful images, without force or global
prune. Podman refuses images referenced by existing containers. The second
run changed nothing. NPM remained active with zero restarts; local admin
and public Gitea HTTPS returned 200. Backup timer and SSH proxy stayed active.
Validation:
```bash
ansible-playbook ansible/site.yml --limit prometheus --tags server_image_cleanup --check --diff -e server_legacy_image_cleanup=true
```
The image cleanup defaults to disabled and carries the `never` tag.
Check mode probes image presence but skips removal; it does not prove
Podman would accept deletion. It never removes NPM resources.
## Approved legacy data and fallback cleanup
The operator explicitly approved deletion on 2026-10-03. The separate
`server_legacy_cleanup` tasks removed `/opt/gitea`, `/home/git/.ssh`,
`/opt/navidrome`, `/opt/postgres`, `/opt/music`, `/opt/containerd`,
`/opt/docker`, the old Compose unit and the final Gitea export helper.
The empty `/home/git` parent is removed only with `rmdir`, after confirming
the Git account is absent. Guards reject symlinked paths, nested mounts,
unexpected containers, container users of these paths, unexpected content
in the empty legacy trees, and an active Compose or export service.
The second cleanup run changed nothing.
Before deletion, Ansible removed obsolete backup input paths and the
Gitea mount dependency. Normal Compose/template/final-export task checks
changed nothing and did not recreate the retired files. Deletion is opt-in:
```bash
ansible-playbook ansible/site.yml --limit prometheus --tags server_legacy_cleanup --check --diff -e server_legacy_cleanup=true
```
Remove check mode only for approved deletion. No active NPM data, certificate,
image, network, volume, SSH proxy, WireGuard configuration or backup archive
is removed. No services were restarted by the cleanup.
After separate approval for the brief managed NPM pause, the new export
`20261003T112906Z` completed successfully and was pulled to Atlas. SHA-256
passed on both hosts; an isolated SQLite restore passed `quick_check` and
contained ten proxy hosts. Both Quadlet definitions were present, and
retired paths were absent. Temporary restore files were removed.
NPM was active with zero automatic restarts; primary public Gitea HTTPS
returned 200 with valid TLS. Backup timer, SSH proxy and WireGuard stayed active.
The first scheduled post-cleanup cycle remains unverified.
After separate operator approval on 2026-10-03, the unused secondary hostname
`git.ov-ad3410.infomaniak.ch` was removed from the declared domains and
the managed NPM runtime override. Its Proxy Host (id 10) was already
soft-deleted, with no generated config or associated certificate. Historical
deleted records and backup archives are preserved; no DNS changes were made.
Only `git.fscotto.duckdns.org` remains declared for the Gitea override.
Nginx validation and reload passed without restarting NPM; the primary
public HTTPS endpoint returned 200 with valid TLS.

View File

@@ -0,0 +1,106 @@
#!/usr/bin/env bash
# Run on Prometheus as root with the exact verified source-export version.
set -Eeuo pipefail
expected_export=${1:?Pass the verified Prometheus backup export version}
mode=${2:---preflight}
[[ $expected_export =~ ^[0-9]{8}T[0-9]{6}Z$ ]] || exit 2
[[ $mode == --preflight || $mode == --execute ]] || exit 2
[[ $EUID -eq 0 ]] || { echo 'Run as root on Prometheus' >&2; exit 2; }
compose_unit=podman-compose-server.service
quadlet_unit=prometheus-npm.service
backup_timer=prometheus-backup-export.timer
versions=/var/lib/prometheus-backup-export/versions
quadlet_file=/etc/containers/systemd/prometheus-npm.container
exec 9>/run/lock/prometheus-backup-export.lock
flock -n 9 || { echo 'Backup/export lock is busy' >&2; exit 1; }
systemctl is-active --quiet "$compose_unit"
if systemctl is-active --quiet "$quadlet_unit"; then
echo 'NPM Quadlet is already active; refusing overlapping cutover' >&2
exit 1
fi
[[ $(systemctl show "$quadlet_unit" -p LoadState --value) == loaded ]]
[[ $(systemctl is-enabled "$compose_unit") == enabled ]]
[[ $(readlink "$versions/current") == "$expected_export" ]]
image=$(sed -n 's/^Image=//p' "$quadlet_file")
[[ $image =~ ^docker\.io/jc21/nginx-proxy-manager@sha256:[a-f0-9]{64}$ ]]
podman image exists "$image"
(cd "$versions/current" && sha256sum -c payload.sha256 && tar -tf payload.tar >/dev/null)
curl -fsS --connect-timeout 2 --max-time 5 -o /dev/null http://127.0.0.1:81/
old_image=$(podman inspect nginx-proxy-manager --format '{{.Image}}')
data_signature() {
python3 - <<'PY'
import hashlib, os, sqlite3
db = sqlite3.connect('file:/opt/npm/data/database.sqlite?mode=ro', uri=True)
assert db.execute('pragma quick_check').fetchone()[0] == 'ok'
counts = [db.execute('select count(*) from ' + table).fetchone()[0]
for table in ('proxy_host', 'certificate', 'user')]
db.close()
digest = hashlib.sha256()
for root, dirs, files in os.walk('/opt/npm/letsencrypt'):
dirs.sort()
for name in sorted(files):
path = os.path.join(root, name)
with open(path, 'rb') as stream:
digest.update(path.encode() + b'\0' + stream.read())
print(*counts, digest.hexdigest())
PY
}
before=$(data_signature)
if [[ $mode == --preflight ]]; then
echo 'NPM Quadlet cutover preflight passed; no service was changed'
exit 0
fi
stopped_old=false
rollback() {
rc=$?
trap - EXIT
if (( rc != 0 )) && "$stopped_old"; then
echo 'NPM Quadlet cutover failed; restoring Compose' >&2
systemctl stop "$quadlet_unit" || true
systemctl enable "$compose_unit" || true
systemctl start "$compose_unit" || true
systemctl start "$backup_timer" || true
curl -fsS --connect-timeout 2 --max-time 10 -o /dev/null http://127.0.0.1:81/ || true
fi
exit "$rc"
}
trap rollback EXIT
stopped_old=true
systemctl stop "$backup_timer"
systemctl stop "$compose_unit"
if podman container exists nginx-proxy-manager; then
echo 'Compose left the NPM container behind; refusing duplicate ownership' >&2
exit 1
fi
systemctl disable "$compose_unit"
systemctl start "$quadlet_unit"
ready=false
for _ in {1..60}; do
if curl -fsS --connect-timeout 2 --max-time 3 -o /dev/null http://127.0.0.1:81/; then
ready=true
break
fi
sleep 2
done
"$ready"
systemctl is-active --quiet "$quadlet_unit"
[[ $(podman inspect nginx-proxy-manager --format '{{.Image}}') == "$old_image" ]]
podman exec nginx-proxy-manager nginx -t
[[ $(data_signature) == "$before" ]]
for hostname in git.fscotto.duckdns.org syncthing.fscotto.duckdns.org; do
status=$(curl -ksS --connect-timeout 3 --max-time 10 \
--resolve "$hostname:443:127.0.0.1" -o /dev/null -w '%{http_code}' \
"https://$hostname/")
[[ $status == 200 ]]
done
systemctl start "$backup_timer"
stopped_old=false
echo 'NPM Quadlet cutover passed local application and data checks'

View File

@@ -1,83 +1,71 @@
$ANSIBLE_VAULT;1.1;AES256
61353065386233646137323235306631353635663530363237636231316265643562353465323430
6165646466623962313835313537633137633766373930380a316335323962616265643136346666
63336133336131346336383534356637623831363138323165633262386333363535393365383233
6234393835653439370a313963313365373633323464343263383661383336363662633133643232
34366634383862363635653034313531623330396639616462343630326162316535643465653532
36326534333637376462353561343964633636366331363833313263353133383636623537303663
35393032316439336666343161653439643638376134363535656262343963393365623432336433
35383934313762313037326430316666363731666231336534326661353034333063643364343230
65333739303566366263333565333465613136646237623937393733623438613832393634663463
39376131313234333039633735613233373931613232653036663665316636303961653834366339
36353730316132316233303964303839363161346564396163336137663134353062363733656430
37643339326661653031376265646132623162373562393437373437313732396537383939333666
62353036316633306666313461663033303830393765396131643035353730383931646239663935
32626461316364386135303761383837613063336466363162323332663764616464373565383231
61346463336566346533326535376439643133613762383633396131323632356533636139336365
62393838316634623932643034376631333539343965383436613364643962363834346337353334
32656439366439313734353963343133333533653839613632323338336131373566613835393536
31663433616334373432376531346435336530303936356461303163646463613661643161313661
66663866343565616631616338353737356164353562366164383736346131666662623132333466
39383865653631373232393433663430643961646265386166333137643966303834363262373636
62396434373363353636376133666133663162653265313139313732353639336232333862643036
64386231336561396537326139346566306434633934343038663165396665363032383466633662
62336163633964363435386630343966333162333730336138333239646631633132663931376462
33663139356261313065376636613930353735396131306538306664646135636336643032623131
38346264333331353633326535326431626563323036313665643337353563333339646430386564
31613435383036313430316366323636663735326336393338353835323861333564363832656462
35336435623261326363633033316130393062616339353263643062633331646137376135656365
35636139336564346164616235616431326531333433646330386134323932373339646536356464
66343533326534326165323564663533653666633035343163633832393361336462343937623165
62383931326630363036396333313931393836366439653433623165666166356338653364336534
35333936653833386163633738326164386166613561333530633937343230363366333662666539
39333361633933663735303438663239303536363433313962643137386533633539326365383765
37636538386339333935386132353265353031643662616330316463623661663738353433313830
36373963633166333464653338343830373063323536383364393033393235326639613662343737
38663362636331343061646465313237313431373433353361353265333766633463353632646536
31323231306138323031396630656538363930373439336234343963616334363632653738316465
63653938373830336362313238656266613362636634616537653863336132343931616262396130
66393239303866656232653832343132366537333537343635666563343639323433383163613335
39613533376634316133633430303535306266656333626264343733666335393661666561396633
39346265316137326465326635396362333565393133623637633132616232326263663662343137
33363733306135363361643031306265363733656362386666306334333035393839636533343363
35396638616636633639343930373136376339346162393061393765363837646365383866636131
33653465666239393133616232636231333332396138376332393664343364643835306530393238
34663237303530303837663535646263393931373531393039356336316561653130356262636562
38336362326639653237626634376334666565653036353236313634376364626338646538386536
38626636386466373566646166393963643164343536373236396138303532393161363335386638
32633032393737626363613463323366366637616361313537356136626661626633613739323338
35383963666431343566356562333234663936376562616638636261303466633539376334303331
39303834663234663063356233313962326664383839393832303462643636393034383434303465
64333635376135326333356435373734643430623736373234643335343130383066326436356664
63346663326364343634303930343338336139313864316165366232643537366635653764353763
31363863633261643263303433373330366161323166366462336332313135366338393334653764
66353733653137663835663731373364613030373334663061313433373861613665363236633130
65613965366636343465336533613438373466383737373366653965633437323562643966396431
39303033643438633762633263326132663466643438656366363431616237633031333936313831
30323930383233313032323638356333626230333764363662313662646536643839353032353462
30326166653937353130623133303533343934633565393831623033303234316330353432313266
30636536633933376365623665616262663236383731633633346232613366333137396139306363
35633336643266326335303261666666653536666630613639376336373237646134306462616537
33343561373162666332613634643837343566646161373065366637653135613632353334636363
63363232303963646530333366663862323264326536643337323266396566316233613630303637
66646366376466373931613734363931316230323063373666653062373364396433633762633762
38613933323733653238383935623230383562646563363833653838636165626365646537383639
33666535656363393562316336633439636138373365623431393965653765306138646234663938
65653133663663393731646337386535333261643932336132396237323930306136643534353930
65636438396432623034626561613137336138623265393064383034623863303166356138393564
37373164626634653662326234333539663735323464613334616130643937373730363263633366
31393937326432386165343338313031376565313866363731643534313233303064373935303538
31343832336230393636653432653162336361383963633766343461653466316337353931333363
63313137303564336630343937356564643763383764613362366634373362666465626334336539
64366533376165306532343461613265366266383862323032333465336161663161376630316465
30306562666163646235656664653635366461366435663961623635383437663564356563346462
31636234663765623838333237393239373564366262613637363938653463396530613963643837
38636634376637366332623035313465393762653865623130336263343663303066366135616639
63333964356466613038303263366462346261353030646532366361393965306435613131316463
65366266376637323764643239323730366565633335666638666334663635373961303637383861
35313431646434656562333937663837393038386361616630626532636339306432353434656165
33663261383166386432383465666136376237346565303164363461666663346130346162316338
62373061353034316234303835663439396434343738303764376665336239626238386436386234
61306166383637366266393730323732386163366261393630336431633862353761343763363665
61323039396234393835303633363339373633653334343766653032313230343464326664356566
3462623830666664626633373966363866333337383730313066
37646664613266633436346262613633613830623366383138613432366365373765353230333134
3332333764313337396637323133623937343738373133370a333930356365653034323235643230
36633864343161653833356636373931383761663864663334336236373733326266386639366335
3131313661313637320a313937633361646333333962303335333233346166343831373039663964
33366532386135663463643965363766643063616436316463666232666138323236346231303537
34303535333866376430363063623934623761373865656231656661383935393866353566346430
64333434613432376436343438343561383235366631623730653533633535326237666265653439
34366264653665643063663361313339663034323932326233366636326336323432303434373765
36316532316265343434383438623239666232373633626330333464303361643630303635643834
39313136623830303762313462343637633763626333393033346637663931663238653734626131
38393963646563333732353531653239643330326539643538323164343934356166343034316565
33346431333735636537613930383331393265313962626234363237373562313231393061326439
64363765323935316661353531366165343139633963336139313737306332613364643031666161
30386362643930316265616564306336633133303166363665333462316265313364393939306162
31303639313933356337386134623934663461643161306666633261653538633232343036653833
66316466636233343136393765636333353230353738313833333265663238303730313936326664
38373239353162363438323964333030666563346161643437326335666162356264396135393532
63363862373136346532653734336335616132386237303031363433663132343861633937386130
30633938616364303462303030303966303939633066393264303462393730363233373937356439
36663533376232663737613734653532313136343939663539373866333638396266666163383864
63613532393334373539346338616163383637633237666234613437663966653733616361353830
61656666376133363330633863346637376266343134633037313132313361366638616261363839
39393062396237666333303937363536346561343763663133323236393037383532396465336138
35613463356532376534386433626337613030343266353332306462306463336336343830666138
33656138363837633337393865643633623261613335366263643162663637623636666162653632
30626238616266323332616234393838343330663662393433366630393566316336636530303165
38343665623437356636643236393734396264356632326133623264633862633333626330336663
61376263656665653731636133316161653635323138303866623862303065366232633736623336
39656666386435343062656138313061616661313966326432663236626631316162623961616636
35343939613262303066626537396164616666316265643065373638663436643961336138313862
39666163646538356338356631346534633139643636393866646462646533363265663234633761
65363661336138353239656165393836386134666331663036653132306433343764643666306333
33623661626565633333306337303263633335386632386330353730316436313931326164363862
37616265653161633632353865346639653961653836353962303762336535666266386535363165
39653138646663376634323131613463333035326639313266613830616431316131383464353533
62656634346637636164626461613137303461633761336232373133653532323566303136663030
30633337346534636566343934306662356238396365306563336666623435353731613136333036
36343436373932323265306639363761353364383635333136366231373166613861633032343233
61376338616630343639333964356162613332323835333730333135356665383431626138643534
66393966666465303763316230386538393863303063386564303165303962346139373338303436
39373032663538323532323766353864643338326561313564373562616430326264386362666532
36613132306462336631363035343732636465343562643430343035373961366566383130656165
64613938393265343037633161653937323933646637653036306532366237313838346361333932
34663565653264626137323239336532643262356166633665313761336162303635346666383863
61383930383033626337383366353766393536653135383062656639323361353539356232613736
63646235663363333333623463313961326533653236363938383765663439613832653039386436
38393734633536313731323437336332353564363564333736663037386530333639326338656561
66336637353238383231613666313261383234336531666132396230373931623363323832633064
39613036346166393936613939363865616135653830366435643538336365353333613831353962
36623537363434373137633063373934383439333462646361613737303239643834303535366138
34656465316431656461373737643537303936636539383934373831616438343965373765373535
63653164363731653030303466646539636361383664343763646163663238383435653035653666
63383165626365653261303834333234626534396333353231303261396361616233363334383336
33316462636133336132656364613439396131613565646565396365316238323962353462653736
64386139313266663963643962363133386133393166306163626632646463333363323830306164
35353936653137383761326132373739306163613764386531613032313235373331303530383633
64646533313434653734366233633535323564386431306538633666383661303038613330653832
34643463396137643034353439653334653836333161396130363637326339383363303037306330
61353635633334343432646461396439393439383639336139316161373737333961653731393333
61636164343838346365373736356161386430356533303331333838333732363233613931613863
66633662383466306332366563373865323861323833353238356563363635313463366333653432
65303839653963376566383737346231343663363363313332383365646363373737323839613564
34613362303335316363363661653639386538326337386537333765643161613961316531613563
38386564636637643762643830666138383361396233303339643665343261356462393830376662
32656334346536636536343263336565333234353831616565366538393661353561376538346334
61396135623230366433303932396130636331333263316333643861626564343330386636613063
32383061616435643736653264313839363232346332343565336464353138396339623533393237
38353632646565323735643462626239663736643033643231613464663866663262366632353434
37363866343239363131633464316133396462353336613962306332343563333962333934616330
3536356634376131633039373834376533633065303533653333

View File

@@ -8,7 +8,7 @@ vault_git_work_email: "REPLACE_ME"
vault_git_work_gpg: "REPLACE_ME"
vault_ikaros_authorized_ssh_keys:
- "ssh-ed25519 REPLACE_ME"
vault_aegis_icloudpd_apple_id: "REPLACE_ME"
vault_atlas_icloudpd_apple_id: "REPLACE_ME"
vault_atlas_admin_password_hash: "REPLACE_WITH_A_SHADOW_COMPATIBLE_HASH"
vault_atlas_samba_password: "REPLACE_ME"
vault_atlas_immich_db_password: "REPLACE_ME"