mirror of
https://github.com/fscotto/infra.git
synced 2026-10-03 13:29:58 +00:00
Compare commits
33 Commits
702283b430
...
main
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
18eb2d2eb2 | ||
|
|
269fb13665 | ||
|
|
2dfe766b7b | ||
|
|
755f24bc72 | ||
|
|
7bc7f0e645 | ||
|
|
1577eec19d | ||
|
|
e30683c3d1 | ||
|
|
4bd6aafb53 | ||
|
|
9e76309833 | ||
|
|
ed3fee06e8 | ||
|
|
309d64b4ed | ||
|
|
dd33a4f55d | ||
|
|
0028fe8c4d | ||
|
|
12037fcc9a | ||
|
|
3f9a626759 | ||
|
|
31fedb8d44 | ||
|
|
9b5ee77905 | ||
|
|
54fb7d46d7 | ||
|
|
a609e68f42 | ||
|
|
06d3b175cb | ||
|
|
256d758b1a | ||
|
|
9d0013769c | ||
|
|
5b0f415163 | ||
|
|
3401b6137d | ||
|
|
bae6a9f554 | ||
|
|
c37483ba38 | ||
|
|
f491362365 | ||
|
|
5a1047adde | ||
|
|
802cb8c7ba | ||
|
|
0144600a4a | ||
|
|
3d2ef02c98 | ||
|
|
8844d00e24 | ||
|
|
9798fe3a12 |
250
AGENTS.md
250
AGENTS.md
@@ -25,6 +25,9 @@ Ansible-driven personal infrastructure repo for Fedora and Void desktops, Fedora
|
||||
- Preserve layering `all -> platform -> role -> desktop -> host`.
|
||||
- Keep `ansible/site.yml` small; orchestration belongs there, implementation belongs in roles.
|
||||
- Prefer minimal, targeted edits. Preserve idempotency and existing ordering.
|
||||
- Keep completed one-time cleanup operations out of the playbook. Execute them directly
|
||||
with explicit authorization; retain only the ongoing desired-state configuration and
|
||||
historical documentation, not permanent cleanup flags or tasks.
|
||||
- Use Git Flow branch prefixes: `feature/` for new functionality, `bugfix/` for non-urgent fixes,
|
||||
`hotfix/` for urgent production fixes, `release/` for release preparation, and `support/` for
|
||||
maintained release lines. Do not use abbreviated prefixes such as `feat/`.
|
||||
@@ -54,9 +57,30 @@ Ansible-driven personal infrastructure repo for Fedora and Void desktops, Fedora
|
||||
- Emacs is disabled by default; temporary Emacs check: `ansible-playbook ansible/site.yml --limit <host> --tags emacs --check --diff -e emacs_enabled=true`
|
||||
- AI coding agents: `ansible-playbook ansible/site.yml --limit <host> --tags ai_agents --check --diff`
|
||||
- Mail bootstrap: `sh -n scripts/bootstrap_mail.sh` and `shellcheck scripts/bootstrap_mail.sh`
|
||||
- Server compose render: `podman-compose -f /opt/docker/server/docker-compose.yml config` and `systemctl status podman-compose-server`
|
||||
- Server NPM Quadlet: `systemctl status prometheus-npm.service`; the Compose fallback is retired.
|
||||
- Explicit Prometheus legacy cleanup (destructive only without check mode):
|
||||
`ansible-playbook ansible/site.yml --limit prometheus --tags server_legacy_cleanup --check --diff -e server_legacy_cleanup=true`
|
||||
- Atlas media stack:
|
||||
`ansible-playbook ansible/site.yml --limit atlas --tags storage,sharing,containers --check --diff`
|
||||
- Atlas rootless Gitea staging (does not start Gitea):
|
||||
`ansible-playbook ansible/site.yml --limit atlas --tags gitea --check --diff`
|
||||
- Atlas canonical Gitea domain (restarts only Gitea on a real configuration change):
|
||||
`ansible-playbook ansible/site.yml --limit atlas --tags gitea_public_domain --check --diff`
|
||||
- Atlas iCloudPD storage and boot-started Quadlet:
|
||||
`ansible-playbook ansible/site.yml --limit atlas --tags icloudpd --check --diff`
|
||||
- Atlas explicit Gitea host-owner migration (live outage; never a normal run):
|
||||
`ansible-playbook ansible/site.yml --limit atlas --tags gitea_owner_migration -e atlas_gitea_owner_migration=true`
|
||||
- Atlas explicit isolated Gitea restore rehearsal (not part of normal runs):
|
||||
`ansible-playbook ansible/site.yml --limit atlas --tags gitea_restore -e atlas_gitea_restore_test=true`
|
||||
- Atlas final Gitea replacement gate (dry-run only until a stopped-source export is pulled):
|
||||
`ansible-playbook ansible/site.yml --limit atlas --tags gitea_final_restore --check --diff -e atlas_gitea_final_restore=true`
|
||||
- Prometheus final Gitea export helper (dry-run installs only; outage action remains opt-in):
|
||||
`ansible-playbook ansible/site.yml --limit prometheus --tags gitea_final_export --check --diff`
|
||||
- Gitea cutover network configuration before activation:
|
||||
`ansible-playbook ansible/site.yml --limit prometheus --tags gitea_cutover,prometheus_backup --check --diff -e server_gitea_on_atlas=true`
|
||||
and `ansible-playbook ansible/site.yml --limit atlas --tags gitea --check --diff`
|
||||
- Atlas daily Navidrome music copy:
|
||||
`ansible-playbook ansible/site.yml --limit atlas --tags music_sync --check --diff`
|
||||
- Atlas network/share hardening:
|
||||
`ansible-playbook ansible/site.yml --limit atlas --tags hardening,sharing --check --diff`
|
||||
- Atlas ZFS snapshot retention and scrub timers:
|
||||
@@ -73,7 +97,9 @@ Ansible-driven personal infrastructure repo for Fedora and Void desktops, Fedora
|
||||
`ansible-playbook ansible/site.yml --limit atlas --tags restorecon --check -e '{"atlas_restorecon_paths":["/zpool/archive"]}'`
|
||||
- Prometheus/Aegis WireGuard gateway:
|
||||
`ansible-playbook ansible/site.yml --limit prometheus,aegis --tags wireguard --check --diff`
|
||||
- DuckDNS config only: `ansible-playbook ansible/site.yml --limit prometheus --tags duckdns --check --diff`
|
||||
- Prometheus NPM Quadlet steady state (does not perform a cutover):
|
||||
`ansible-playbook ansible/site.yml --limit prometheus --tags npm_quadlet --check --diff`
|
||||
- DuckDNS config only (skipped on Prometheus): `ansible-playbook ansible/site.yml --limit prometheus --tags duckdns --check --diff`
|
||||
|
||||
## Conventions
|
||||
- Use FQCN Ansible modules.
|
||||
@@ -117,16 +143,24 @@ The dotfile vars follow the same split: `desktop_common_dotfiles` carries mode-i
|
||||
- Windows applications are installed manually and are not managed from the WSL profile.
|
||||
|
||||
## Rocky Server Notes
|
||||
- DuckDNS is rendered by `profile_server` from host-local `server_duckdns_domain` and
|
||||
- Prometheus disables DuckDNS provisioning with `server_duckdns_enabled: false`. Its updater,
|
||||
log and five-minute cron entry were explicitly retired; the external DuckDNS name and Vault
|
||||
token remain untouched. The completed one-time cleanup has no remaining playbook tasks.
|
||||
- When enabled, DuckDNS is rendered by `profile_server` from host-local `server_duckdns_domain` and
|
||||
`vault_duckdns_token`. Keep the rotated token in encrypted Vault or untracked local vars, never in
|
||||
dotfiles. The private `~/duckdns/duck.sh` keeps the existing entrypoint; rendering uses `no_log`
|
||||
and disables diffs. Provisioning does not execute the updater or change its external schedule.
|
||||
- `rocky_server` is a child of both `platform_rocky` and `server`; `prometheus` is its active target.
|
||||
- The target must already provide `server_username` with local sudo access before the profile runs.
|
||||
- The Rocky profile installs Podman and podman-compose, uses firewalld, preserves SELinux enforcement, and renders the
|
||||
existing Nginx Proxy Manager/Gitea Compose stack with a `podman-compose-server` systemd unit. PostgreSQL and
|
||||
Navidrome are no longer part of the desired Prometheus configuration. The role does not stop or remove legacy
|
||||
containers, delete `/opt/postgres/data`, start the Compose stack, update DNS, or cut over traffic.
|
||||
- The Rocky profile installs Podman and podman-compose. Prometheus explicitly retires the legacy
|
||||
Compose unit, files and final-export helper with `server_legacy_stack_retired: true`.
|
||||
Its approved opt-in cleanup removed old application data on 2026-10-03; normal runs do not
|
||||
delete data or recreate the retired files. On Prometheus, Nginx Proxy Manager is now the rootful
|
||||
`prometheus-npm.service` Quadlet with a pinned image digest and the existing `/opt/npm/data` and
|
||||
`/opt/npm/letsencrypt` bind mounts. The rootful `server_web` bridge remains `10.89.0.0/24`.
|
||||
Gitea runs on Atlas; PostgreSQL and Navidrome are absent from the desired Prometheus stack.
|
||||
Normal runs do not delete legacy data, update DNS, or perform an implicit cutover;
|
||||
destructive cleanup requires its explicit tag and opt-in extra-var.
|
||||
- Firewalld enables SSH, Cockpit (`9090/tcp`), HTTP and HTTPS. Nginx Proxy Manager publishes `80/tcp` and
|
||||
`443/tcp`; bind its administration interface only to `127.0.0.1:81` and use `npm-tunnel` from Ikaros or Nymph.
|
||||
Nextcloud remains disabled; do not provision `/srv/nextcloud` directories.
|
||||
@@ -164,7 +198,9 @@ The dotfile vars follow the same split: `desktop_common_dotfiles` carries mode-i
|
||||
- `profile_backend_phase1` temporarily runs rootless Navidrome and Syncthing on Atlas until Uranus replaces
|
||||
them. It binds only to Atlas' LAN IP, never `wg0`; Navidrome and the Syncthing GUI admit only Aegis as
|
||||
the source-NAT gateway, while native Syncthing ports admit the configured LAN. It initializes fresh
|
||||
state only and never migrates or deletes source application data.
|
||||
state only and never migrates or deletes source application data. The enabled rootless
|
||||
`atlas-music-sync.timer` copies `/zpool/archive/Music` to `/zpool/media/music` daily at 00:45
|
||||
Europe/Rome without deleting destination files; it requires both datasets to be mounted.
|
||||
- `wireguard_overlay` manages `wg0` between Prometheus (`10.0.0.1`) and Aegis (`10.0.0.2`). It persists private
|
||||
keys only on their respective hosts, exchanges only derived public keys through Ansible, and verifies a real peer
|
||||
handshake. Prometheus opens `51820/udp`; Aegis is the LAN gateway. Its persistent IPv4 forwarding, narrowly scoped
|
||||
@@ -226,7 +262,7 @@ successfully. The first monthly scrub remains a runtime check.
|
||||
read-only ZFS snapshot test restored one file to `/var/tmp`, confirmed matching contents, ownership,
|
||||
mode, mtime and ACL, then removed its temporary copy and on-demand mount. This is a file-level smoke
|
||||
test, not full dataset recovery. An independent USB file restore passed on 2026-09-25 with matching
|
||||
content and metadata; full disaster recovery remains a separate Priority 2 task.
|
||||
content and metadata; the later scaled OS-rebuild rehearsal is documented under Priority 2.
|
||||
- [x] Add monitoring and alerting for pool health, scrub/resilver, SMART data, temperatures, free space,
|
||||
snapshot/local-backup growth, Hetzner Storage Box quota, and failed maintenance or backup timers.
|
||||
The half-hourly Atlas health monitor and systemd final-failure hooks are deployed. A live probe
|
||||
@@ -237,31 +273,195 @@ successfully. The first monthly scrub remains a runtime check.
|
||||
on Atlas, but a new real failure notification has not been deliberately triggered.
|
||||
|
||||
### Priority 2 - NAS operability and recovery
|
||||
- [ ] Document and test disaster recovery: rebuild Atlas with Ansible, import the existing pool, restore
|
||||
from snapshot/USB/Hetzner, preserve Vault and Borg recovery material offline, and define RPO/RTO.
|
||||
- [ ] Define a controlled Rocky kernel/OpenZFS update and reboot procedure.
|
||||
- [ ] Add the Atlas-initiated least-privilege Prometheus backup pull: Prometheus exposes only prepared
|
||||
- [x] Document and test disaster recovery in `docs/atlas-recovery.md`: the operator confirmed Vault
|
||||
and Borg recovery material is available offline; provisional targets are RPO 24h/RTO 72h. On
|
||||
2026-09-30 an isolated small Rocky VM was rebuilt with the Atlas Ansible roles, imported its
|
||||
preserved RAIDZ2 pool without force/rewind, and restored a file from the preserved snapshot;
|
||||
the second Ansible run was idempotent. Earlier independent production ZFS, USB, and Borg file
|
||||
restore tests remain separate evidence. A production-size full restore, unclean import, and
|
||||
measured 24h/72h compliance are not claimed.
|
||||
- [x] Define a controlled Rocky kernel/OpenZFS update and reboot procedure in
|
||||
`docs/atlas-updates.md`. The first real change-window execution is not yet
|
||||
validated; the procedure never reboots automatically or upgrades pool features.
|
||||
- [x] Add the Atlas-initiated least-privilege Prometheus backup pull: Prometheus exposes only prepared
|
||||
read-only dumps through a dedicated account and Atlas retains the private SSH key, pinned host key,
|
||||
atomic pull, verification, retention and systemd service/timer.
|
||||
- [ ] Decide whether a common SMB/NFS namespace is required. `Archive` (SMB) and `photobook` (NFS) are
|
||||
intentionally distinct today; only if a shared namespace is selected, finalize its UID/GID, group,
|
||||
and POSIX ACL model and test the same files through both protocols.
|
||||
atomic pull, verification, retention and systemd service/timer. The dedicated key/account and unit
|
||||
files are deployed; live read-only SSH, shell denial, and write denial were verified. On 2026-09-30
|
||||
a manual export, Atlas pull, checksum verification, and temporary restore passed; both SQLite
|
||||
databases passed integrity checks and a restored Git repository passed `git fsck`. Both daily
|
||||
timers are enabled for 02:00/03:00 Europe/Rome. On 2026-10-01 their first scheduled export and
|
||||
pull succeeded: Atlas verified the payload checksum and published `20261001T000001Z` as `latest`.
|
||||
- [x] Decide whether a common SMB/NFS namespace is required: no. `Archive` (SMB) and `photobook` (NFS)
|
||||
remain intentionally distinct; `docs/atlas-sharing-decision.md` records the decision. No ACL or export
|
||||
change is authorized by this decision.
|
||||
|
||||
### Priority 3 - Service expansion
|
||||
- [ ] After data protection and recovery are validated, populate `/zpool/media/music` and validate Navidrome.
|
||||
- [x] Populate `/zpool/media/music` and validate Navidrome. On 2026-09-30, 21,158 files
|
||||
(93,937,810,350 regular-file bytes) were copied from `/zpool/archive/Music` using a temporary
|
||||
ZFS snapshot; a checksum-based rsync dry run found no differences or extra files. Navidrome saw
|
||||
all files through its read-only mount, completed a scan, indexed 18,168 tracks, and responded
|
||||
over HTTP. Some imported playlists still reference obsolete Windows paths. The source was left
|
||||
intact and the temporary snapshot was removed.
|
||||
- [x] Schedule a daily, non-deleting copy from `Archive/Music` to the separate Navidrome music
|
||||
dataset. The rootless `atlas-music-sync.timer` is enabled for 00:45 Europe/Rome; a manual
|
||||
idempotent service run succeeded on 2026-10-01. The first scheduled run triggered at
|
||||
00:45 CEST on 2026-10-02 and exited successfully (`Result=success`, status 0); the next
|
||||
run is scheduled for 2026-10-03 00:45 CEST.
|
||||
- [x] Design the staged Prometheus-to-Atlas Gitea migration in `docs/atlas-gitea-migration.md`.
|
||||
The approved topology keeps NPM on Prometheus and moves HTTPS and public SSH (TCP/2222) together;
|
||||
Gitea runs as an `admin`-owned rootless user Quadlet on Atlas with an internal `gitea` user.
|
||||
The rootful-to-rootless data-layout
|
||||
conversion passed an isolated restore rehearsal. The later partial cutover is tracked below.
|
||||
- [x] Prepare the dedicated Atlas Gitea dataset, non-login UID/GID 1101 with a separate rootless Podman
|
||||
sub-ID range, and disabled user Quadlet. On 2026-10-01 the targeted Ansible run and a second idempotent
|
||||
run passed; the generated unit was inactive, with no staging HTTP/SSH listener. POSIX ACLs on only the
|
||||
service-namespace parents grant this account traversal without access to sibling datasets.
|
||||
- [x] Perform an isolated rootless restore rehearsal from the verified Prometheus backup. On 2026-10-01
|
||||
the SHA-256-checked selective extraction and path/SSH conversion succeeded; SQLite `quick_check`
|
||||
passed, all 33 repositories passed `git fsck`, and source/target public SSH host-key fingerprints
|
||||
matched. The pinned rootless image answered HTTP and listened on internal SSH/2222 with
|
||||
`--network none`; the temporary container was removed and the Quadlet stayed inactive. A second
|
||||
restore run made no changes. This is a rehearsal copy, not the final consistent cutover copy.
|
||||
- [x] Verify ZFS and Borg coverage of the staged Gitea dataset. On 2026-10-01 the managed recursive
|
||||
hourly snapshot `atlas-auto-hourly-20261001T193401Z` included it, and the managed incremental
|
||||
Borg archive `atlas-20261001T193420Z` included its database. A private one-file restore from
|
||||
each independently matched the staged database and passed SQLite `quick_check`; temporary files
|
||||
and snapshot mounts were removed, the Borg service ended successfully, and the pool was healthy.
|
||||
- [x] Include the new Gitea dataset in a UUID-bound offline USB version and test a file restore
|
||||
before accepting production writes. The operator's 2026-10-01 manual run published version
|
||||
`20261001T201220Z-254397` successfully on 2026-10-02. Its Gitea database was restored to a
|
||||
temporary directory from a read-only mount: contents, owner, group, mode, size, mtime and POSIX
|
||||
ACL matched, and SQLite `quick_check` passed. Temporary files and mounts were removed, LUKS
|
||||
was closed, and the pool remained healthy. A redundant run was stopped during verification;
|
||||
its temporary snapshot was cleaned up and the service's resulting failed state was reset.
|
||||
- [x] Install a separate opt-in final Gitea export helper on Prometheus. Its 2026-10-01 targeted
|
||||
deployment and `bash -n` passed while Gitea and NPM stayed running. It refuses an active export
|
||||
timer, stops only Gitea, verifies SQLite, publishes a checksum-verified Gitea-only version for
|
||||
Atlas' existing pull, and leaves the source stopped on success. It was invoked on 2026-10-02
|
||||
after the export timer was stopped; version `20261002T071525Z` was pulled and verified on Atlas.
|
||||
- [x] Prepare the Atlas final-restore gate without replacing the rehearsal: it accepts only a
|
||||
checksum-verified `gitea-cutover` export, refuses a running target, stages and validates the new
|
||||
layout before replacing the marked rehearsal, and rolls back a failed swap. Synthetic success
|
||||
and rollback tests passed on 2026-10-01. On 2026-10-02 the final gate replaced the rehearsal;
|
||||
SQLite `quick_check`, all 33 repository `git fsck` checks, checksum and SSH host-key comparison passed.
|
||||
- [x] Start the rootless Atlas Gitea Quadlet and move the primary HTTPS route. On 2026-10-02 Atlas
|
||||
answered HTTP 200 through the Aegis gateway. NPM stayed on Prometheus; its variable upstream
|
||||
required a managed Nginx `server_proxy.conf` override because runtime DNS ignores Compose
|
||||
`extra_hosts`. The primary public HTTPS page and API returned 200, and `git ls-remote` succeeded
|
||||
for a representative repository after NPM restart; the Navidrome and Syncthing Proxy Hosts also
|
||||
responded. The source
|
||||
Gitea container was removed from the desired Compose stack without deleting its data; the
|
||||
Prometheus backup export timer resumed for NPM only. A post-cutover recursive ZFS snapshot and
|
||||
encrypted Borg archive `atlas-20261002T073044Z` completed successfully.
|
||||
- [x] Move the live Gitea Quadlet and dataset from the legacy host `gitea` account to `admin`
|
||||
after a disposable snapshot-copy test of the pinned derived image. On 2026-10-02 the explicit
|
||||
outage run stopped only legacy Gitea, made safety snapshot
|
||||
`zpool/services/data/gitea@gitea-owner-migration-20261002T100104`, changed dataset ownership,
|
||||
and validated loopback staging (HTTP 200, internal `gitea` UID/GID 1000, SQLite `quick_check`)
|
||||
before promoting the `admin` Quadlet. Production LAN and public HTTPS returned 200; Navidrome
|
||||
and Syncthing remained active, the pool was healthy, and the normal Gitea run changed nothing.
|
||||
The old host account and data on Prometheus remain preserved; the old Atlas Quadlet and its
|
||||
parent-dataset traverse ACL were removed. A subsequent normal run changed nothing.
|
||||
- [x] Validate public Gitea SSH/2222 and an authenticated read from Ikaros. After the VPS
|
||||
firewall was opened on 2026-10-02, TCP/2222 connected, the public ED25519 host-key
|
||||
fingerprint matched Atlas, Gitea authenticated `fscotto` using the `ikaros` key, and
|
||||
`git ls-remote` returned HEAD for `fscotto/infra.git` over public SSH.
|
||||
- [x] Validate authenticated SSH pull and push. On 2026-10-02 the operator reported both
|
||||
operations working through the public SSH endpoint; the earlier agent-run `git ls-remote`
|
||||
remains the independent read-only check. The agent did not perform a test push.
|
||||
- [x] Validate Gitea login and write via HTTPS. On 2026-10-03 the operator confirmed
|
||||
authenticated web login and Git clone/pull/push through the public HTTPS endpoint. Do not
|
||||
restart the stale source Gitea after Atlas has accepted writes.
|
||||
- [ ] Design and deploy Nextcloud as another explicitly temporary Atlas service before Uranus. Give it
|
||||
separate persistent application, database, and cache storage; keep credentials in Vault; publish it only
|
||||
through NPM over the Prometheus--Aegis gateway; and define backup, upgrade, and eventual Uranus-migration
|
||||
procedures before exposing user data. Do not deploy Nextcloud before the data-protection checklist is complete.
|
||||
- [x] Move Gitea canonical HTTPS and SSH hostname to `git.fscotto.co` on
|
||||
2026-10-03 through Ansible. Only Gitea restarted; second run changed nothing.
|
||||
HTTPS and authenticated SSH reads returned the same repository HEAD.
|
||||
The new NPM hostnames passed TLS/HTTP checks; old DuckDNS Proxy Hosts were
|
||||
observed disabled. Details are in `docs/domain-fscotto-co.md`.
|
||||
- [x] Confirm login on the new Gitea hostname and update remaining client remotes/integrations.
|
||||
The operator confirmed completion on 2026-10-03; the agent did not perform a test push.
|
||||
- [x] Remove obsolete DuckDNS NPM Proxy Hosts, unused certificates and the old upstream override.
|
||||
The operator confirmed completion on 2026-10-03; no new agent runtime check was performed.
|
||||
- [x] Review and remove completed one-time procedures from the playbook.
|
||||
The operator confirmed completion on 2026-10-03.
|
||||
- [x] Retire Prometheus' local DuckDNS updater on 2026-10-03 through Ansible:
|
||||
the five-minute cron entry and private updater/log directory were removed.
|
||||
Provisioning is disabled; repeat cleanup changed nothing. HTTPS services, private NPM
|
||||
administration and the export timer stayed healthy. The external name and Vault token
|
||||
remain untouched for possible future use on a local host.
|
||||
- [ ] Keep `atlas_manage_media_stack` disabled until the future Immich deployment has validated `/dev/dri`,
|
||||
container paths, and the required Vault database secret.
|
||||
|
||||
### Priority 4 - Optional workflows
|
||||
- [ ] After data protection is validated, move iCloudPD photo ingestion from Aegis to Atlas as a
|
||||
temporary service until Uranus is ready. Plan to store photos in `/zpool/archive/Pictures` and
|
||||
persistent application/MFA state outside `Archive`; validate permissions, SELinux, backups and
|
||||
recovery before cutover. Keep the current Aegis service and Photobook NFS export unchanged until
|
||||
the Atlas workflow is tested, then retire them explicitly if no longer needed.
|
||||
- [x] Deploy the declared Atlas iCloudPD state dataset and inactive rootless `admin` Quadlet.
|
||||
Photos belong under `/zpool/archive/Pictures/iCloudPD`; private config/MFA state belongs in
|
||||
`zpool/services/data/icloudpd`. Photobook remains reserved for Immich. Ansible now renders
|
||||
`icloudpd.conf` with the Apple ID from the existing Vault key, but does not store the password
|
||||
or manage MFA. Automatic startup was approved on 2026-10-03; the Quadlet now
|
||||
uses `WantedBy=default.target` and Ansible keeps the service running.
|
||||
The isolated no-network layout test is documented in
|
||||
`docs/atlas-icloudpd-migration.md`. On 2026-10-02 Atlas deployment and a second idempotent run
|
||||
passed; no app config existed at deployment. A manual first start on 2026-10-02 generated
|
||||
`icloudpd.conf`; an Ansible run then replaced it with a private mode-0600 Vault-backed template
|
||||
and an idempotent second run. The image later expanded the config, so Ansible now seeds it
|
||||
only when absent and maintains the declared fields. Its launcher requires `traceroute`; the
|
||||
rootless Quadlet grants only `NET_RAW`, tested in isolation and after restart. The service
|
||||
was subsequently initialized interactively; initial ingestion is tracked below.
|
||||
- [x] Retire Aegis iCloudPD completely. The operator authorized deleting its Quadlet,
|
||||
`/var/lib/icloudpd` data, and MFA state despite an unaudited container overlay. After two
|
||||
interactive-sudo runs on 2026-10-02, the unit is `not-found`/`inactive`, the Quadlet and state
|
||||
directory are absent, and AdGuard remains active. The temporary retirement tasks have since
|
||||
been removed from the Aegis role; it no longer manages iCloudPD.
|
||||
- [x] Validate Atlas iCloudPD authentication and initial ingestion. On 2026-10-03 the active
|
||||
rootless service logged `All photos and videos have been downloaded` at 02:16 and reported
|
||||
completion for the user. The destination held 11,658 files (86,020,430,015 bytes); the preceding 24h
|
||||
logs showed download activity without authentication failures or errors. A later read-only check
|
||||
found the service still active. This confirms the initial download, not the next daily cycle.
|
||||
- [x] Declare HEIC decoding for Fedora graphical desktops without converting the originals on Atlas.
|
||||
The Fedora role installs RPM Fusion Free with a pinned signing-key fingerprint and
|
||||
`libheif-freeworld` on Ikaros and Nymph. The package was confirmed installed on Ikaros on
|
||||
2026-10-03; Nymph deployment and an actual image-opening test were not observed.
|
||||
- [ ] Validate Atlas iCloudPD filesystem/SELinux/SMB access, the next daily sync, ZFS/Borg/USB
|
||||
backup inclusion, and isolated restore of photos and private state. A recursive hourly snapshot
|
||||
of `zpool/archive` exists after ingestion, but no iCloudPD-specific backup version or restore
|
||||
has been verified. The first monthly scrub remains a separate open data-protection check.
|
||||
|
||||
## Prometheus NPM Quadlet cutover
|
||||
- [x] Stage a rootful NPM Quadlet using the exact running image and the existing data/certificate
|
||||
mounts, bridge subnet, public HTTP/HTTPS ports, and loopback-only administration port.
|
||||
The generated service depends on `server-web-network.service` and is wanted by `multi-user.target`.
|
||||
- [x] Take and verify the stopped-source export before switching owners. Version
|
||||
`20261003T091009Z` was pulled to Atlas and its NPM SQLite database checked in isolation.
|
||||
- [x] Cut over NPM to `prometheus-npm.service` on 2026-10-03. The legacy Compose unit is inactive
|
||||
and disabled; the Quadlet is active with zero recorded restarts. Public Gitea and Syncthing
|
||||
HTTPS returned 200 with valid TLS, while public TCP/81 remained unreachable.
|
||||
- [x] Validate the post-cutover backup path. The export and Atlas pull published
|
||||
`20261003T091633Z`; checksum, SQLite `quick_check`, ten proxy hosts, six certificate records,
|
||||
both Quadlet files were present, and the complete Let's Encrypt tree (70 regular files plus
|
||||
12 symlinks) matched the live data. A targeted normal Ansible run changed nothing. Details and rollback
|
||||
boundaries are in `docs/prometheus-npm-quadlet.md`.
|
||||
- [x] Remove only unused Gitea, Navidrome and PostgreSQL images with opt-in
|
||||
Ansible tasks on 2026-10-03. Second run changed nothing; NPM stayed active
|
||||
with zero restarts, HTTP/HTTPS passed, backup timer and SSH proxy stayed active.
|
||||
This image-only step preserved data and fallback; the later approved deletion is tracked below. Validation:
|
||||
`ansible-playbook ansible/site.yml --limit prometheus --tags server_image_cleanup --check --diff -e server_legacy_image_cleanup=true`
|
||||
- [x] Complete explicitly approved old-data and Compose fallback removal on 2026-10-03.
|
||||
Backup paths and mount dependencies were reconciled before deletion; repeat cleanup changed
|
||||
nothing. The normal Compose/template/helper check did not recreate retired files.
|
||||
A separately approved manual export/pull published `20261003T112906Z`; checksum and isolated
|
||||
SQLite restore passed with ten proxy hosts and both Quadlet definitions. NPM, primary HTTPS,
|
||||
WireGuard, SSH proxy and backup timer remained healthy; existing backup archives were preserved.
|
||||
- [x] Retire the unused secondary Gitea hostname `git.ov-ad3410.infomaniak.ch`
|
||||
on 2026-10-03. Its NPM Proxy Host was already soft-deleted and had no
|
||||
associated certificate. Its Ansible domain and runtime override were removed;
|
||||
nginx -t and reload passed without restarting NPM. Primary HTTPS returned 200
|
||||
with valid TLS. At that step only `git.fscotto.duckdns.org` remained declared;
|
||||
the subsequent domain transition and operator-confirmed cleanup are tracked above.
|
||||
- [ ] Observe the first scheduled export and Atlas pull after the cutover; the manual end-to-end
|
||||
cycle passed, but the next unattended cycle has not yet occurred.
|
||||
|
||||
## Cerberus Management Node (Deferred)
|
||||
`cerberus` is postponed until the office in the new house is physically set up. It is not an inventory
|
||||
@@ -337,5 +537,5 @@ validated exports of older historical data will use a dedicated Atlas NFS datase
|
||||
`/etc/resolv.conf` linked to `/run/systemd/resolve/resolv.conf`. LAN clients may use AdGuard, but
|
||||
Aegis must use the independent upstream DNS declared by `aegis_host_dns_servers` so Greenboot does
|
||||
not depend on the AdGuard container during startup.
|
||||
- iCloudPD requires post-deployment interactive MFA initialization; its cookie/configuration state is
|
||||
persisted in `/var/lib/icloudpd/config`.
|
||||
- Aegis iCloudPD has been retired and is no longer managed by this role. Its service, Quadlet,
|
||||
data, and MFA state were removed with the operator's explicit authorization.
|
||||
|
||||
86
README.it.md
86
README.it.md
@@ -182,6 +182,10 @@ Le applicazioni Windows sono installate e gestite manualmente; il profilo WSL no
|
||||
|
||||
## Server
|
||||
|
||||
La migrazione dei servizi pubblici a `fscotto.co`, la gestione Ansible
|
||||
degli URL Gitea e i passaggi ancora aperti per ritirare DuckDNS sono in
|
||||
[`docs/domain-fscotto-co.md`](docs/domain-fscotto-co.md).
|
||||
|
||||
Sistema operativo:
|
||||
|
||||
- Rocky Linux 9
|
||||
@@ -201,28 +205,37 @@ Lo stato attuale del profilo server include:
|
||||
- installazione pacchetti Rocky via DNF, EPEL e CRB
|
||||
- installazione di Podman e podman-compose
|
||||
- abilitazione dei servizi systemd dichiarati in inventory/group vars
|
||||
- copia dei dotfiles server e rendering del `docker-compose.yml` per Nginx Proxy Manager e Gitea,
|
||||
piu l'unita `podman-compose-server` (attivazione manuale)
|
||||
- copia dei dotfiles server e rendering del Quadlet rootful `prometheus-npm.service` per Nginx Proxy
|
||||
Manager; il vecchio fallback Compose è stato rimosso con autorizzazione esplicita
|
||||
- attivazione di firewalld con SSH, Cockpit (`9090/tcp`), HTTP e HTTPS abilitati
|
||||
- Syncthing escluso dal profilo server Rocky
|
||||
|
||||
Il Compose desiderato su Prometheus non include piu Navidrome ne il database PostgreSQL obsoleto.
|
||||
Navidrome e Syncthing appartengono ad Atlas; Navidrome ufficiale usa invece SQLite. Il profilo non
|
||||
arresta o rimuove automaticamente eventuali container legacy e non elimina `/opt/postgres/data`.
|
||||
Il 2026-10-03 la pulizia opt-in autorizzata ha rimosso dati e immagini precedenti di Gitea,
|
||||
Navidrome e PostgreSQL, directory obsolete vuote, helper finale Gitea e fallback Compose NPM.
|
||||
I servizi migrati restano su Atlas. `server_legacy_stack_retired: true` evita che i normali task
|
||||
ricreino i residui; la cancellazione richiede `--tags server_legacy_cleanup` e
|
||||
`-e server_legacy_cleanup=true`. NPM attivo e archivi di backup restano intatti.
|
||||
Export, pull Atlas e restore SQLite isolato post-pulizia sono riusciti; il primo ciclo automatico
|
||||
resta da osservare. Evidenze e confini del recovery:
|
||||
[`docs/prometheus-npm-quadlet.md`](docs/prometheus-npm-quadlet.md).
|
||||
|
||||
Nginx Proxy Manager pubblica solo `80/tcp` e `443/tcp`; la sua interfaccia di amministrazione e
|
||||
associata a `127.0.0.1:81` ed e raggiungibile da Ikaros o Nymph con l'alias Bash `npm-tunnel`.
|
||||
Nextcloud resta disabilitato e il profilo non crea directory `/srv/nextcloud`.
|
||||
|
||||
La fase 1 su Atlas non modifica questo deployment NPM ne i suoi dati persistenti. Dopo aver attivato
|
||||
WireGuard e i servizi Atlas, configurare i proxy host NPM correnti con upstream Navidrome
|
||||
`http://10.0.0.2:4533` e upstream per la GUI Syncthing `http://10.0.0.2:8384`. Solo la GUI web di
|
||||
Syncthing usa NPM; il traffico di sincronizzazione resta sulle porte native pubblicate esplicitamente solo
|
||||
sull'indirizzo WireGuard di Atlas. Configurare l'autenticazione Syncthing e una policy di accesso NPM adeguata prima di pubblicare la GUI.
|
||||
La fase 1 su Atlas non modifica i dati persistenti NPM. I proxy host NPM usano gli upstream LAN
|
||||
`http://192.168.178.55:4533` per Navidrome e `http://192.168.178.55:8384` per la GUI Syncthing;
|
||||
Prometheus li raggiunge attraverso Aegis come gateway WireGuard. Solo la GUI web di Syncthing usa
|
||||
NPM; il traffico di sincronizzazione resta sulle porte native esposte sulla LAN dichiarata.
|
||||
Mantenere l'autenticazione Syncthing e una policy di accesso NPM adeguata.
|
||||
|
||||
### DuckDNS
|
||||
|
||||
`profile_server` genera `~/duckdns/duck.sh` con permessi `0700`, mantenendo il percorso dello
|
||||
`server_duckdns_enabled: false` disabilita il provisioning su Prometheus, che usa IP statico
|
||||
e `fscotto.co`. Updater, log e cron ogni cinque minuti sono stati rimossi una sola volta;
|
||||
non restano task o flag di pulizia. Il nome DuckDNS esterno e il token Vault restano invariati.
|
||||
|
||||
Sui server con `server_duckdns_enabled: true`, `profile_server` genera `~/duckdns/duck.sh` con permessi `0700`, mantenendo il percorso dello
|
||||
script e `duck.log`. Definire `server_duckdns_domain` negli host vars del server e salvare il
|
||||
**nuovo token rigenerato** in `vault_duckdns_token`, nel Vault cifrato `secrets/vault.yml`
|
||||
(`ansible-vault edit secrets/vault.yml`) oppure negli override non versionati `secrets/vault.local.yml`.
|
||||
@@ -311,7 +324,12 @@ l'interfaccia amministrativa resta su `127.0.0.1:81`, raggiungibile via tunnel S
|
||||
Atlas ospita temporaneamente Navidrome e Syncthing rootless fino alla sostituzione con Uranus. I
|
||||
servizi sono inizializzati **ex novo**, senza migrare lo stato precedente, rispettivamente sotto
|
||||
`/zpool/services/data/navidrome` e `/zpool/services/data/syncthing`; la musica in
|
||||
`/zpool/media/music` viene popolata separatamente. Sono vincolati all'indirizzo LAN di Atlas
|
||||
`/zpool/media/music` è stata popolata separatamente da `/zpool/archive/Music` il 2026-09-30;
|
||||
Navidrome ha completato la scansione. Il timer rootless `atlas-music-sync.timer` copia i file nuovi
|
||||
o modificati ogni giorno alle 00:45 Europe/Rome, senza eliminare quelli presenti solo nella
|
||||
destinazione; entrambi i dataset ZFS devono essere montati. La prima esecuzione schedulata è
|
||||
riuscita il 2026-10-02. Alcune playlist originali contengono
|
||||
ancora vecchi percorsi Windows. I servizi sono vincolati all'indirizzo LAN di Atlas
|
||||
(`192.168.178.55`), mai a WireGuard. `wireguard_overlay` collega invece Prometheus (`10.0.0.1`)
|
||||
e Aegis (`10.0.0.2`): le chiavi private restano sui rispettivi host e Ansible scambia solo le pubbliche.
|
||||
Prometheus apre `51820/udp`; Aegis inoltra soltanto il traffico overlay→LAN dichiarato e applica
|
||||
@@ -322,6 +340,14 @@ alla LAN. Dopo la verifica dei servizi, configurare manualmente i Proxy Host NPM
|
||||
negli `AllowedIPs`; aggiungere la VIP Uranus quando esisterà. Dopo il reload di firewalld, Ansible
|
||||
ricarica le reti Podman rootful di Prometheus per conservare DNS e connettività del proxy.
|
||||
|
||||
La migrazione Gitea da Prometheus ad Atlas è descritta in
|
||||
[`docs/atlas-gitea-migration.md`](docs/atlas-gitea-migration.md). Gitea usa un Quadlet rootless
|
||||
di `admin` su un dataset dedicato; l'immagine derivata mantiene UID/GID 1000 ma chiama l'utente
|
||||
interno `gitea`. NPM resta su Prometheus e l'HTTPS pubblico primario serve Atlas. L'SSH pubblico
|
||||
su TCP/2222 autentica la chiave `ikaros` e un `git ls-remote` è riuscito; l'operatore ha
|
||||
confermato pull e push SSH. Login e scrittura Git via HTTPS sono stati confermati il 2026-10-03. I dati sorgente restano
|
||||
conservati su Prometheus senza avviarne il vecchio container.
|
||||
|
||||
Validare il gateway con:
|
||||
|
||||
```bash
|
||||
@@ -485,7 +511,7 @@ etichettata di 45Drives Alerts usare
|
||||
|
||||
### Timer systemd di Atlas
|
||||
|
||||
Tutti i nove timer gestiti sono abilitati. Gli orari sono locali ad Atlas (`Europe/Rome`); Borg e
|
||||
Tutti i dieci timer gestiti sono abilitati. Gli orari sono locali ad Atlas (`Europe/Rome`); Borg e
|
||||
monitoraggio aggiungono il ritardo casuale indicato. Tutti hanno `Persistent=true`: un evento perso
|
||||
viene recuperato quando il timer torna attivo.
|
||||
|
||||
@@ -500,10 +526,13 @@ viene recuperato quando il timer torna attivo.
|
||||
| `atlas-borg-check.timer` | `*-*-15 06:00:00` — giorno 15 alle 06:00, più 0–30 min casuali | Controllo repository Borg |
|
||||
| `atlas-usb-reminder.timer` | `Sat *-*-01..07 10:00:00 Europe/Rome` — primo sabato alle 10:00 | Solo promemoria 45Drives Alerts |
|
||||
| `atlas-health-monitor.timer` | `*:0/30` — ogni mezz'ora, più 0–5 min casuali | Controlli di salute in sola lettura |
|
||||
| `atlas-prometheus-pull.timer` | `*-*-* 03:00:00 Europe/Rome` — ogni giorno alle 03:00 | Pull e verifica del backup preparato su Prometheus |
|
||||
|
||||
`atlas-usb-backup.service` **non ha timer** e va avviato manualmente. Il timer del fornitore
|
||||
`zfs-scrub-weekly@zpool.timer` è disabilitato a favore dello scrub mensile. Il futuro pull del backup
|
||||
Prometheus non ha ancora un timer, perché non è implementato. Durante un backup Borg attivo,
|
||||
`zfs-scrub-weekly@zpool.timer` è disabilitato a favore dello scrub mensile. Il timer di preparazione
|
||||
su Prometheus è attivo alle 02:00 Europe/Rome; il primo ciclo pianificato è riuscito il 2026-10-01.
|
||||
Un export, pull e ripristino temporaneo post-cutover NPM Quadlet sono riusciti il 2026-10-03;
|
||||
il primo ciclo pianificato dopo quel cutover resta da osservare. Durante un backup Borg attivo,
|
||||
`systemctl list-timers` può mostrare `-` per il prossimo evento senza che il timer sia disabilitato.
|
||||
Per vedere la pianificazione corrente: `systemctl list-timers --all` su Atlas.
|
||||
|
||||
@@ -512,14 +541,20 @@ della protezione dei dati: richiede storage applicativo, database e cache separa
|
||||
pubblicazione solo tramite NPM e Aegis, procedure di backup, aggiornamento e migrazione. Non
|
||||
distribuirlo prima di completare la checklist di protezione dei dati.
|
||||
|
||||
La destinazione futura per l'importazione foto iCloud è Atlas, non Aegis. Dopo la validazione dei
|
||||
backup, pianificare una migrazione esplicita di iCloudPD con foto sotto `/zpool/archive/Pictures` e
|
||||
stato applicativo/MFA fuori da `Archive`; testare permessi, SELinux, backup e restore prima del
|
||||
cutover. L'attuale iCloudPD su Aegis e l'export NFS Photobook restano configurati fino
|
||||
all'approvazione e alla verifica di questa migrazione separata. Anche il servizio Atlas sarà
|
||||
temporaneo in attesa di Uranus.
|
||||
Atlas è la destinazione dichiarata per iCloudPD. Ansible gestisce dataset, Quadlet rootless e
|
||||
`icloudpd.conf` privato con Apple ID dal Vault: foto in `/zpool/archive/Pictures/iCloudPD`,
|
||||
stato in `zpool/services/data/icloudpd`. Il primo avvio è stato manuale; password e MFA restano
|
||||
gestiti interattivamente, senza avvio automatico al boot. L'inizializzazione è stata completata e
|
||||
il download iniziale di foto e video è terminato il 2026-10-03. Su Aegis
|
||||
il servizio, il Quadlet e `/var/lib/icloudpd` sono stati rimossi e verificati; il playbook Aegis
|
||||
non contiene più task iCloudPD. L'accesso SMB e il ripristino dai backup dei nuovi dati restano
|
||||
da verificare. L'export NFS Photobook resta
|
||||
invariato. Dettagli in [`docs/atlas-icloudpd-migration.md`](docs/atlas-icloudpd-migration.md).
|
||||
|
||||
Il pull dei backup di Prometheus e i test completi di disaster recovery restano da fare. Il backlog
|
||||
Il primo ciclo pianificato del backup di Prometheus e una prova di disaster recovery a dimensione reale
|
||||
restano da verificare. Il 2026-09-30 una VM Rocky isolata ha superato ricostruzione OS con Ansible,
|
||||
import del pool RAIDZ2 fittizio e ripristino da snapshot; RPO 24 ore/RTO 72 ore restano obiettivi
|
||||
provvisori, non tempi misurati. Dettagli e limiti sono in `docs/atlas-recovery.md`. Il backlog
|
||||
prioritizzato è in `AGENTS.md`.
|
||||
|
||||
---
|
||||
@@ -617,8 +652,8 @@ Questo significa che, allo stato attuale:
|
||||
- `deadalus` riceve il profilo Fedora WSL tramite play dev dedicati
|
||||
- il server Rocky (`prometheus`) e gestito con pacchetti, servizi, dotfiles server e firewalld
|
||||
- il NAS Rocky (`atlas`) usa un pool ZFS gia esistente, condivisioni NFSv4/SMB limitate alla LAN e Cockpit/45Drives
|
||||
- lo stack Compose server include soltanto `gitea` e `nginx-proxy-manager`; Navidrome e Syncthing
|
||||
della fase 1 sono Quadlet rootless su Atlas
|
||||
- NPM è un Quadlet rootful su Prometheus, mentre Gitea, Navidrome e Syncthing sono Quadlet
|
||||
rootless su Atlas; il fallback Compose server è stato rimosso
|
||||
|
||||
# Dotfiles
|
||||
|
||||
@@ -724,9 +759,10 @@ ansible-playbook ansible/site.yml --limit <host> --tags <tag1>,<tag2> --check --
|
||||
ansible-playbook ansible/site.yml --limit <host> --start-at-task "<task name>" --check --diff
|
||||
ansible-lint ansible/roles/<role>
|
||||
yamllint ansible/path/to/file.yml
|
||||
podman-compose -f /opt/docker/server/docker-compose.yml config
|
||||
ansible-playbook ansible/site.yml --limit prometheus --tags server_legacy_cleanup --check --diff -e server_legacy_cleanup=true
|
||||
ansible-playbook ansible/site.yml --limit atlas --tags storage,sharing,containers --check --diff
|
||||
ansible-playbook ansible/site.yml --limit atlas --tags backend_phase1 --check --diff
|
||||
ansible-playbook ansible/site.yml --limit atlas --tags music_sync --check --diff
|
||||
```
|
||||
|
||||
## Tag supportati dal playbook
|
||||
|
||||
91
README.md
91
README.md
@@ -125,16 +125,26 @@ That gives it Fedora packages through DNF, Docker from the official repository,
|
||||
|
||||
## Server
|
||||
|
||||
The public service domain transition to `fscotto.co`, Gitea canonical URL
|
||||
management, and remaining DuckDNS retirement steps are documented in
|
||||
[`docs/domain-fscotto-co.md`](docs/domain-fscotto-co.md).
|
||||
|
||||
`prometheus` is the Rocky Linux 9 server. It has no graphical environment and gets server-specific
|
||||
dotfiles and templates. The profile provisions configuration only: it does not transfer data, start
|
||||
the Compose stack, update DNS, or perform a cutover.
|
||||
dotfiles and templates. The profile does not transfer application data, update DNS, or perform an
|
||||
implicit service cutover.
|
||||
|
||||
The server profile installs platform-specific packages, Podman and podman-compose, declared systemd
|
||||
services, and firewalld. The manually activated `podman-compose-server` unit contains the existing
|
||||
Nginx Proxy Manager and Gitea services. The desired Compose file no longer includes Navidrome,
|
||||
Syncthing, or the obsolete Navidrome PostgreSQL database; their temporary Atlas deployment is managed
|
||||
by `profile_backend_phase1`. Applying the profile does not stop or remove legacy containers and does
|
||||
not delete `/opt/postgres/data`.
|
||||
services, and firewalld. Nginx Proxy Manager runs as the rootful `prometheus-npm.service` Quadlet.
|
||||
On 2026-10-03 the operator-approved opt-in cleanup removed old Gitea, Navidrome and PostgreSQL
|
||||
data/images, empty legacy directories, the Gitea final-export helper and the Compose rollback files.
|
||||
The migrated services stay on Atlas. `server_legacy_stack_retired: true` prevents normal runs from
|
||||
recreating retired files. Data deletion requires `--tags server_legacy_cleanup` and
|
||||
`-e server_legacy_cleanup=true`; image-only cleanup has its own `server_image_cleanup` tag and flag.
|
||||
Active NPM resources and existing backup archives remain preserved.
|
||||
The post-cleanup export/pull and isolated SQLite restore passed; the first unattended cycle remains
|
||||
pending. Evidence and recovery boundaries:
|
||||
[`docs/prometheus-npm-quadlet.md`](docs/prometheus-npm-quadlet.md).
|
||||
|
||||
|
||||
Firewalld enables SSH, Cockpit (`9090/tcp`), HTTP and HTTPS. Nginx Proxy Manager publishes only
|
||||
`80/tcp` and `443/tcp`; its administration interface is bound to `127.0.0.1:81` and can be reached
|
||||
@@ -161,7 +171,11 @@ Prometheus authorizes its declared SSH public keys through separate files below
|
||||
|
||||
### DuckDNS
|
||||
|
||||
`profile_server` renders `~/duckdns/duck.sh` with mode `0700`, keeping the existing updater path
|
||||
`server_duckdns_enabled: false` disables provisioning on Prometheus, which uses its static IP
|
||||
and `fscotto.co`. The local updater, log and five-minute cron job were removed once;
|
||||
no cleanup tasks or flags remain. The external DuckDNS name and Vault token remain untouched.
|
||||
|
||||
For servers with `server_duckdns_enabled: true`, `profile_server` renders `~/duckdns/duck.sh` with mode `0700`, keeping the existing updater path
|
||||
and `duck.log`. Set `server_duckdns_domain` in the server's host vars and store the **rotated**
|
||||
`vault_duckdns_token` in encrypted `secrets/vault.yml` (using `ansible-vault edit secrets/vault.yml`)
|
||||
or untracked `secrets/vault.local.yml`. Never commit the rendered script or put the token on a
|
||||
@@ -210,8 +224,8 @@ ansible/bootstrap/generate-aegis-ign.sh --write IMAGE DEVICE
|
||||
```
|
||||
|
||||
The controller manages it remotely as `pi@aegis`; unlike local desktop profiles, Aegis is
|
||||
intentionally an SSH inventory target. `profile_aegis` manages rootful Podman Quadlets for AdGuard
|
||||
Home and iCloudPD, persistent data under `/var/lib`, the Podman auto-update timer, LAN-restricted
|
||||
intentionally an SSH inventory target. `profile_aegis` manages a rootful Podman Quadlet for AdGuard
|
||||
Home, its persistent data under `/var/lib`, the Podman auto-update timer, LAN-restricted
|
||||
firewalld rules, SSH key-only access for `pi`, the `nfs-utils` and `wireguard-tools` rpm-ostree layers,
|
||||
and `wake-ikaros`. `wireguard_overlay` makes Aegis the internal endpoint and LAN gateway for Prometheus:
|
||||
it enables persistent IPv4 forwarding, installs a scoped WireGuard-to-LAN firewalld policy, and source-NATs
|
||||
@@ -224,9 +238,8 @@ opened and closed manually during initial setup. The profile disables the local
|
||||
stub and points `/etc/resolv.conf` to its full resolver data, freeing port 53 for AdGuard. LAN clients
|
||||
may use AdGuard on Aegis, while Aegis itself uses the independent upstream DNS declared by
|
||||
`aegis_host_dns_servers`; this prevents Greenboot from depending on the AdGuard container during
|
||||
startup. Reboot Aegis after changing its NetworkManager DNS profile. Define
|
||||
`vault_aegis_icloudpd_apple_id` in Vault before applying it. iCloudPD still requires interactive MFA
|
||||
initialization after its first deployment.
|
||||
startup. Reboot Aegis after changing its NetworkManager DNS profile. iCloudPD was retired from Aegis;
|
||||
the Aegis role no longer contains iCloudPD tasks. Atlas iCloudPD config is Vault-backed; MFA is manual.
|
||||
|
||||
New Aegis images create the `admin` account in Butane. Before configuring a newly imaged node, run its
|
||||
first playbook execution with `-e ansible_user=admin`; the SSH hardening role then permits that same
|
||||
@@ -296,7 +309,20 @@ Atlas temporarily hosts rootless Navidrome and Syncthing until Uranus replaces t
|
||||
Atlas' LAN address (`192.168.178.55`); WireGuard remains exclusively between Prometheus (`10.0.0.1`)
|
||||
and Aegis (`10.0.0.2`). Their state is initialized ex novo in `/zpool/services/data/navidrome` and
|
||||
`/zpool/services/data/syncthing`; no source application state is migrated. The music library at
|
||||
`/zpool/media/music` is populated separately.
|
||||
`/zpool/media/music` was populated separately from `/zpool/archive/Music` on 2026-09-30;
|
||||
Navidrome completed its library scan. The rootless `atlas-music-sync.timer` copies new and changed
|
||||
files daily at 00:45 Europe/Rome, without deleting destination-only files. Both ZFS datasets must
|
||||
be mounted. Its first scheduled run succeeded on 2026-10-02. Some source playlists still contain
|
||||
obsolete Windows paths.
|
||||
|
||||
The Gitea move from Prometheus to Atlas is tracked in
|
||||
[`docs/atlas-gitea-migration.md`](docs/atlas-gitea-migration.md). The final consistent copy runs in
|
||||
Atlas' dedicated dataset under `admin`'s rootless user Quadlet. Its pinned derived image uses an
|
||||
internal Unix user named `gitea` (UID/GID 1000), while clone URLs keep `git@`. NPM remains on Prometheus and the primary
|
||||
public HTTPS route serves Atlas. Public SSH/2222 authenticates the `ikaros` key and serves
|
||||
`git ls-remote`; the operator also confirmed SSH pull and push. HTTPS login and Git writes were
|
||||
confirmed on 2026-10-03. The old Gitea data remains on Prometheus, but its container
|
||||
is absent from the desired stack.
|
||||
|
||||
The separate `wireguard_overlay` role manages `wg0` between Prometheus (`10.0.0.1`) and Aegis
|
||||
(`10.0.0.2`), generating private keys once on their respective hosts and exchanging only public keys
|
||||
@@ -500,7 +526,7 @@ monitoring. For a labelled 45Drives Alerts delivery test, use
|
||||
|
||||
### Atlas systemd timers
|
||||
|
||||
All nine managed timers below are enabled. Times are local to Atlas (`Europe/Rome`); Borg and monitoring
|
||||
All ten managed timers below are enabled. Times are local to Atlas (`Europe/Rome`); Borg and monitoring
|
||||
add the indicated randomized delay. Every timer has `Persistent=true`, so a missed calendar run is
|
||||
scheduled after the timer becomes active again.
|
||||
|
||||
@@ -515,10 +541,13 @@ scheduled after the timer becomes active again.
|
||||
| `atlas-borg-check.timer` | `*-*-15 06:00:00` — 15th of the month at 06:00, plus 0–30 min random delay | Borg repository check |
|
||||
| `atlas-usb-reminder.timer` | `Sat *-*-01..07 10:00:00 Europe/Rome` — first Saturday at 10:00 | 45Drives Alerts reminder only |
|
||||
| `atlas-health-monitor.timer` | `*:0/30` — every half-hour, plus 0–5 min random delay | Read-only health checks |
|
||||
| `atlas-prometheus-pull.timer` | `*-*-* 03:00:00 Europe/Rome` — daily at 03:00 | Pull and verify the prepared Prometheus backup |
|
||||
|
||||
`atlas-usb-backup.service` has **no timer**: the encrypted USB backup must be started manually.
|
||||
The vendor's `zfs-scrub-weekly@zpool.timer` is intentionally disabled in favor of the monthly scrub.
|
||||
The future Prometheus backup pull has no timer yet because that workflow is not implemented. While a
|
||||
The Prometheus export timer runs at 02:00 Europe/Rome. Its first scheduled export and Atlas pull
|
||||
passed on 2026-10-01; a manual post-NPM-Quadlet export, pull, and temporary restore passed on
|
||||
2026-10-03. The first scheduled cycle after that cutover remains to be observed. While a
|
||||
Borg backup is still running, `systemctl list-timers` may show `-` for its next trigger; this does not
|
||||
mean the timer has been disabled. Inspect the current schedule on Atlas with
|
||||
`systemctl list-timers --all`.
|
||||
@@ -528,15 +557,30 @@ declared persistent application, database, and cache storage, Vault-backed crede
|
||||
publishing through Aegis, and defined backup, upgrade, and eventual migration procedures. Do not deploy
|
||||
it before the data-protection checklist is complete.
|
||||
|
||||
The desired future iCloud photo-ingestion host is Atlas, not Aegis. After data-protection validation,
|
||||
plan an explicit iCloudPD migration with photos under `/zpool/archive/Pictures` and application/MFA
|
||||
state outside `Archive`, then test permissions, SELinux, backups and recovery before cutting over.
|
||||
The current Aegis iCloudPD service and Atlas Photobook NFS export remain configured until that
|
||||
separate migration is approved and validated; the eventual Atlas service is temporary until Uranus.
|
||||
Atlas is the declared iCloud photo-ingestion host. Ansible manages the rootless Quadlet, a private
|
||||
Vault-backed `icloudpd.conf`, photos under `/zpool/archive/Pictures/iCloudPD`, and separate state in
|
||||
`zpool/services/data/icloudpd`. The service was started manually; Ansible does not enable automatic
|
||||
startup or manage the password and MFA keyring. The operator initialized MFA interactively; on
|
||||
2026-10-03 the initial photo/video download completed. Aegis iCloudPD, including its service data,
|
||||
has been removed and verified; the Aegis role no longer manages it. Backup/restore and SMB access
|
||||
for the new data remain unverified. The Photobook NFS export remains untouched. See
|
||||
[`docs/atlas-icloudpd-migration.md`](docs/atlas-icloudpd-migration.md).
|
||||
|
||||
Prometheus backup pulls and full disaster-recovery tests remain follow-up work. The prioritized
|
||||
The first scheduled Prometheus backup runs and production-size disaster-recovery tests remain follow-up work. The prioritized
|
||||
operational backlog is kept in `AGENTS.md`.
|
||||
|
||||
Priority 2 procedures and decisions are recorded in
|
||||
[`docs/atlas-recovery.md`](docs/atlas-recovery.md),
|
||||
[`docs/atlas-updates.md`](docs/atlas-updates.md), and
|
||||
[`docs/atlas-sharing-decision.md`](docs/atlas-sharing-decision.md).
|
||||
The provisional Atlas recovery objectives are RPO 24 hours and RTO 72 hours;
|
||||
an isolated small-VM OS rebuild, pool import, Ansible reapplication, and
|
||||
snapshot restore passed, but full-size recovery time is unmeasured. `Archive` (SMB) and
|
||||
`photobook` (NFS) remain deliberately separate.
|
||||
The Prometheus pull architecture and manual export/pull/restore evidence are in
|
||||
[`docs/prometheus-backup.md`](docs/prometheus-backup.md). Both daily timers are
|
||||
enabled; their first scheduled runs remain to be verified.
|
||||
|
||||
## How layering works
|
||||
|
||||
A host can intentionally belong to more than one inventory group. The final configuration is the combination of the host and its groups, not a one-host/one-play mapping.
|
||||
@@ -706,8 +750,9 @@ ansible-playbook ansible/site.yml --limit <host> --tags <tag1>,<tag2> --check --
|
||||
ansible-playbook ansible/site.yml --limit <host> --start-at-task "<task name>" --check --diff
|
||||
ansible-lint ansible/roles/<role>
|
||||
yamllint ansible/path/to/file.yml
|
||||
podman-compose -f /opt/docker/server/docker-compose.yml config
|
||||
ansible-playbook ansible/site.yml --limit prometheus --tags server_legacy_cleanup --check --diff -e server_legacy_cleanup=true
|
||||
ansible-playbook ansible/site.yml --limit atlas --tags storage,sharing,containers --check --diff
|
||||
ansible-playbook ansible/site.yml --limit atlas --tags music_sync --check --diff
|
||||
```
|
||||
|
||||
## Tags
|
||||
|
||||
@@ -6,6 +6,11 @@ effective_username: "{{ server_username }}"
|
||||
effective_user_group: "{{ server_user_group }}"
|
||||
effective_user_home: "{{ server_user_home }}"
|
||||
server_container_stack_dir: /opt/docker/server
|
||||
server_npm_quadlet_stage: false
|
||||
server_npm_quadlet_cutover: false
|
||||
server_legacy_stack_retired: false
|
||||
server_legacy_cleanup: false
|
||||
server_duckdns_enabled: true
|
||||
ai_agents: {}
|
||||
vim_plugins_enabled: false
|
||||
|
||||
@@ -80,5 +85,37 @@ server_sshd_settings:
|
||||
|
||||
server_sshd_allow_users:
|
||||
- "{{ server_username }}"
|
||||
server_backup_export_enabled: false
|
||||
server_backup_username: prometheus-backup
|
||||
server_backup_public_key_name: atlas-pull
|
||||
server_backup_export_root: /var/lib/prometheus-backup-export
|
||||
server_backup_rrsync_path: /usr/share/doc/rsync/support/rrsync
|
||||
server_backup_export_calendar: "*-*-* 02:00:00 Europe/Rome"
|
||||
server_backup_export_start_timer: false
|
||||
# Explicit Gitea cutover helper: installed separately from any outage action.
|
||||
server_gitea_cutover_tools_enabled: false
|
||||
server_gitea_final_export: false
|
||||
server_gitea_on_atlas: false
|
||||
server_gitea_atlas_address: "{{ hostvars['atlas'].ansible_host }}"
|
||||
server_gitea_npm_domains: []
|
||||
server_gitea_ssh_public_port: 2222
|
||||
server_gitea_ssh_target_port: 2222
|
||||
server_backup_export_source_keep: 3
|
||||
server_backup_export_paths: >-
|
||||
{{ ['opt/npm/data', 'opt/npm/letsencrypt']
|
||||
+ ([] if server_gitea_on_atlas | bool else ['opt/gitea/data', 'home/git/.ssh'])
|
||||
+ ([] if server_legacy_stack_retired | bool else
|
||||
['opt/docker/server/docker-compose.yml',
|
||||
'etc/systemd/system/podman-compose-server.service'])
|
||||
+ (['etc/containers/systemd/prometheus-npm.container',
|
||||
'etc/containers/systemd/server-web.network']
|
||||
if server_npm_quadlet_stage | bool else [])
|
||||
+ ['etc/ssh/sshd_config', 'etc/ssh/sshd_config.d',
|
||||
'etc/firewalld', 'etc/wireguard/wg0.conf'] }}
|
||||
server_backup_export_excludes: >-
|
||||
{{ ['opt/npm/data/logs']
|
||||
+ ([] if server_gitea_on_atlas | bool else
|
||||
['opt/gitea/data/gitea/log', 'opt/gitea/data/gitea/tmp',
|
||||
'opt/gitea/data/gitea/sessions', 'opt/gitea/data/gitea/indexers']) }}
|
||||
server_ssh_authorized_keys: []
|
||||
server_ssh_authorized_key_directory: "{{ server_user_home }}/.ssh/authorized_keys.d"
|
||||
|
||||
@@ -42,5 +42,3 @@ aegis_ssh_authorized_keys:
|
||||
key: "ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIEH/7GJfGt0ZVmKeEzceoFkFkeCXFryKK9vAbaip+HCx nymph"
|
||||
- name: siren
|
||||
key: "ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIA95wYlzpfN3rjUhpMeP4KHn8I6ZrjQXoDTgwgRIa++b siren"
|
||||
|
||||
aegis_icloudpd_apple_id: "{{ vault_aegis_icloudpd_apple_id | default('') }}"
|
||||
|
||||
@@ -49,6 +49,11 @@ atlas_zfs_backup_reservation: 500G
|
||||
atlas_zfs_dataset_photobook: media/photobook
|
||||
atlas_mount_root: /zpool
|
||||
atlas_manage_storage: true
|
||||
# Rootless Gitea was restored from the stopped-source export before production activation.
|
||||
atlas_manage_gitea: true
|
||||
atlas_gitea_production_enabled: true
|
||||
atlas_gitea_public_domain: git.fscotto.co
|
||||
atlas_prometheus_pull_start_timer: true
|
||||
atlas_manage_zfs_snapshots: true
|
||||
atlas_zfs_snapshot_prefix: atlas-auto
|
||||
atlas_zfs_snapshot_policies:
|
||||
@@ -91,6 +96,12 @@ atlas_usb_backup_mapper_name: zpool-backup
|
||||
atlas_manage_usb_reminder: true
|
||||
atlas_usb_reminder_calendar: "Sat *-*-01..07 10:00:00 Europe/Rome"
|
||||
atlas_manage_monitoring: true
|
||||
atlas_manage_prometheus_backup_pull: true
|
||||
# Prometheus ED25519 host key read through the controller's strict SSH trust on 2026-09-30.
|
||||
# Fingerprint: SHA256:rfedk7DHI9mLB3UHk/4F3HHlSIiswtCAFsAXvfh6iXk
|
||||
atlas_prometheus_ssh_host_key: >-
|
||||
179.237.102.172 ssh-ed25519
|
||||
AAAAC3NzaC1lZDI1NTE5AAAAIC4b+QXlPupoEx71W9NKs9tTeYjBqTkVMqbGB97nMNWv
|
||||
# Physical pool disks and the system NVMe; the disconnected USB disk is intentionally excluded.
|
||||
atlas_monitor_smart_devices:
|
||||
- { name: pool-1, path: "{{ atlas_zpool_disks[0] }}", warning_c: 50, critical_c: 55 }
|
||||
@@ -130,8 +141,8 @@ atlas_monitor_remote_capacity:
|
||||
atlas_manage_sharing: true
|
||||
atlas_manage_media_stack: false
|
||||
# Planned after data-protection validation: move iCloudPD photo ingestion from
|
||||
# Aegis to Atlas, with photos under /zpool/archive/Pictures and persistent
|
||||
# application/MFA state outside Archive. Do not deploy or cut over yet.
|
||||
# Aegis to Atlas, with photos under /zpool/archive/Pictures/iCloudPD and
|
||||
# application/MFA state in a separate dataset. Do not deploy or cut over yet.
|
||||
|
||||
# WireGuard is retired on Atlas. These rootless services are a temporary home
|
||||
# until Uranus replaces them.
|
||||
@@ -141,6 +152,7 @@ backend_phase1_bind_address: "{{ ansible_host }}"
|
||||
backend_phase1_firewalld_zone: "{{ atlas_firewalld_zone }}"
|
||||
backend_phase1_npm_source_ip: "{{ atlas_aegis_ip }}"
|
||||
backend_phase1_syncthing_native_subnet: "{{ atlas_lan_subnet }}"
|
||||
backend_phase1_music_sync_enabled: true
|
||||
|
||||
rocky_manage_openzfs_repo: true
|
||||
rocky_manage_syncthing_binary: false
|
||||
|
||||
@@ -6,7 +6,28 @@ ansible_port: 22
|
||||
ansible_ssh_private_key_file: /home/fscotto/.ssh/id_ed25519
|
||||
|
||||
server_username: rocky
|
||||
server_legacy_stack_retired: true
|
||||
# Destructive deletion runs only with an explicit extra-var and cleanup tag.
|
||||
server_legacy_cleanup: false
|
||||
# Explicit opt-in cleanup; no data, volumes, networks or NPM images are removed.
|
||||
server_legacy_image_cleanup: false
|
||||
server_legacy_images:
|
||||
- docker.gitea.com/gitea:1.25.2
|
||||
- docker.io/deluan/navidrome:latest
|
||||
- docker.io/library/postgres:13
|
||||
server_npm_quadlet_stage: true
|
||||
server_npm_quadlet_image: docker.io/jc21/nginx-proxy-manager@sha256:52b2c59994f3d36acfcf70a1626f29734df0ed8c71bacc0269f78b6f939858bb
|
||||
# The stopped-source export and live Quadlet cutover passed on 2026-10-03.
|
||||
server_npm_quadlet_cutover: true
|
||||
server_backup_export_enabled: true
|
||||
server_backup_export_start_timer: true
|
||||
# Install the final-copy helper only; it is never run by a normal playbook invocation.
|
||||
server_gitea_cutover_tools_enabled: true
|
||||
server_gitea_on_atlas: true
|
||||
server_gitea_npm_domains:
|
||||
- git.fscotto.duckdns.org
|
||||
server_duckdns_domain: fscotto
|
||||
server_duckdns_enabled: false
|
||||
server_ssh_authorized_keys:
|
||||
- name: ikaros
|
||||
key: "ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAINrIxXjA3ffPwziKGR5gzc4gAoBehQPlnEMcXF4Wl0ZS ikaros"
|
||||
|
||||
@@ -39,11 +39,41 @@
|
||||
state: enabled
|
||||
when: "'workstation_dev_wsl' in group_names"
|
||||
|
||||
- name: Install distribution signing keys for Fedora desktop codecs
|
||||
tags: [packages, heic]
|
||||
ansible.builtin.dnf:
|
||||
name: distribution-gpg-keys
|
||||
state: present
|
||||
when: "'graphical_desktop' in group_names"
|
||||
|
||||
- name: Import RPM Fusion Free signing key for Fedora desktop codecs
|
||||
tags: [packages, heic]
|
||||
ansible.builtin.rpm_key:
|
||||
key: /usr/share/distribution-gpg-keys/rpmfusion/RPM-GPG-KEY-rpmfusion-free-fedora-2020
|
||||
fingerprint: E9A491A3DE247814E7E067EAE06F8ECDD651FF2E
|
||||
state: present
|
||||
when: "'graphical_desktop' in group_names"
|
||||
|
||||
- name: Enable RPM Fusion Free for Fedora desktop codecs
|
||||
tags: [packages, heic]
|
||||
ansible.builtin.dnf:
|
||||
name: "https://download1.rpmfusion.org/free/fedora/rpmfusion-free-release-{{ ansible_facts['distribution_major_version'] }}.noarch.rpm"
|
||||
state: present
|
||||
when: "'graphical_desktop' in group_names"
|
||||
|
||||
- name: Refresh dnf package metadata
|
||||
tags: [packages]
|
||||
ansible.builtin.dnf:
|
||||
update_cache: true
|
||||
|
||||
- name: Install HEIC decoder on Fedora desktops
|
||||
tags: [packages, heic]
|
||||
ansible.builtin.dnf:
|
||||
name: libheif-freeworld
|
||||
state: present
|
||||
update_cache: true
|
||||
when: "'graphical_desktop' in group_names"
|
||||
|
||||
- name: Install packages on Fedora
|
||||
tags: [packages]
|
||||
ansible.builtin.dnf:
|
||||
|
||||
@@ -8,10 +8,6 @@ aegis_network_connection_uuid: ""
|
||||
aegis_host_dns_servers: []
|
||||
aegis_host_dns_search_domains: []
|
||||
aegis_adguard_image: docker.io/adguard/adguardhome:latest
|
||||
aegis_icloudpd_image: docker.io/boredazfcuk/icloudpd:latest
|
||||
aegis_icloudpd_folder_structure: '{:%Y/%m/%d}'
|
||||
aegis_icloudpd_synchronisation_interval: 86400
|
||||
aegis_icloudpd_apple_id: ""
|
||||
aegis_ikaros_mac_address: aa:bb:cc:dd:ee:ff
|
||||
aegis_wol_port: 9
|
||||
|
||||
|
||||
@@ -9,13 +9,12 @@
|
||||
name: sshd.service
|
||||
state: reloaded
|
||||
|
||||
- name: Restart Aegis Quadlet services
|
||||
- name: Restart Aegis AdGuard Quadlet
|
||||
ansible.builtin.systemd:
|
||||
name: "{{ item }}"
|
||||
state: restarted
|
||||
daemon_reload: true
|
||||
loop:
|
||||
- adguardhome.service
|
||||
- icloudpd.service
|
||||
loop_control:
|
||||
label: "{{ item }}"
|
||||
|
||||
@@ -13,14 +13,6 @@
|
||||
msg: Reboot Aegis to activate the newly layered packages, then rerun the playbook.
|
||||
when: aegis_layered_packages_result.needs_reboot | default(false)
|
||||
|
||||
- name: Require Aegis iCloudPD Apple ID
|
||||
tags: [aegis, icloudpd]
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- aegis_icloudpd_apple_id | length > 0
|
||||
fail_msg: Define vault_aegis_icloudpd_apple_id before applying the Aegis profile.
|
||||
no_log: true
|
||||
|
||||
- name: Require completed Aegis network placeholders
|
||||
tags: [aegis, dns, firewall, network, services]
|
||||
ansible.builtin.assert:
|
||||
@@ -119,8 +111,6 @@
|
||||
loop:
|
||||
- /var/lib/adguard/work
|
||||
- /var/lib/adguard/conf
|
||||
- /var/lib/icloudpd/data
|
||||
- /var/lib/icloudpd/config
|
||||
|
||||
- name: Create Quadlet configuration directory
|
||||
tags: [aegis, containers]
|
||||
@@ -142,12 +132,9 @@
|
||||
loop:
|
||||
- src: adguardhome.container.j2
|
||||
dest: adguardhome.container
|
||||
- src: icloudpd.container.j2
|
||||
dest: icloudpd.container
|
||||
loop_control:
|
||||
label: "{{ item.dest }}"
|
||||
no_log: "{{ item.dest == 'icloudpd.container' }}"
|
||||
notify: Restart Aegis Quadlet services
|
||||
notify: Restart Aegis AdGuard Quadlet
|
||||
|
||||
- name: Create Aegis systemd-resolved configuration directory
|
||||
tags: [aegis, adguard, dns, services]
|
||||
@@ -168,7 +155,7 @@
|
||||
mode: "0644"
|
||||
notify:
|
||||
- Restart Aegis systemd-resolved
|
||||
- Restart Aegis Quadlet services
|
||||
- Restart Aegis AdGuard Quadlet
|
||||
|
||||
- name: Point Aegis resolver at the full systemd-resolved configuration
|
||||
tags: [aegis, adguard, dns, services]
|
||||
@@ -361,7 +348,7 @@
|
||||
group: root
|
||||
mode: "0755"
|
||||
|
||||
- name: Enable Aegis Quadlet services and automatic updates
|
||||
- name: Enable Aegis AdGuard Quadlet and automatic updates
|
||||
tags: [aegis, containers, services]
|
||||
ansible.builtin.systemd:
|
||||
name: "{{ item }}"
|
||||
@@ -370,7 +357,6 @@
|
||||
daemon_reload: true
|
||||
loop:
|
||||
- adguardhome.service
|
||||
- icloudpd.service
|
||||
- podman-auto-update.timer
|
||||
loop_control:
|
||||
label: "{{ item }}"
|
||||
|
||||
@@ -1,20 +0,0 @@
|
||||
# Managed by Ansible. Do not edit manually.
|
||||
[Unit]
|
||||
Description=iCloud Photos Downloader
|
||||
Wants=network-online.target
|
||||
After=network-online.target
|
||||
|
||||
[Container]
|
||||
Image={{ aegis_icloudpd_image }}
|
||||
Environment=apple_id={{ aegis_icloudpd_apple_id }}
|
||||
Environment=folder_structure={{ aegis_icloudpd_folder_structure }}
|
||||
Environment=synchronisation_interval={{ aegis_icloudpd_synchronisation_interval }}
|
||||
Volume=/var/lib/icloudpd/data:/home/root/iCloud:Z
|
||||
Volume=/var/lib/icloudpd/config:/config:Z
|
||||
AutoUpdate=registry
|
||||
|
||||
[Service]
|
||||
Restart=always
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
@@ -115,6 +115,15 @@ atlas_monitor_notifier: "{{ atlas_usb_reminder_notifier }}"
|
||||
atlas_monitor_smart_devices: []
|
||||
atlas_monitor_timers: []
|
||||
atlas_monitor_failure_units: []
|
||||
atlas_monitor_effective_timers: >-
|
||||
{{ atlas_monitor_timers
|
||||
+ ([{'name': 'atlas-prometheus-pull.timer', 'max_age_hours': 26}]
|
||||
if atlas_manage_prometheus_backup_pull | bool and atlas_prometheus_pull_start_timer | bool
|
||||
else []) }}
|
||||
atlas_monitor_effective_failure_units: >-
|
||||
{{ atlas_monitor_failure_units
|
||||
+ (['atlas-prometheus-pull.service']
|
||||
if atlas_manage_prometheus_backup_pull | bool else []) }}
|
||||
atlas_monitor_remote_capacity: {}
|
||||
atlas_monitor_pool_warning_percent: 80
|
||||
atlas_monitor_pool_critical_percent: 90
|
||||
@@ -141,8 +150,62 @@ atlas_music_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_music }}"
|
||||
atlas_backup_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_backup }}"
|
||||
atlas_host_backups_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_host_backups }}"
|
||||
atlas_backup_prometheus_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_backup_prometheus }}"
|
||||
atlas_manage_prometheus_backup_pull: false
|
||||
atlas_prometheus_pull_ssh_dir: /etc/atlas-prometheus-pull
|
||||
atlas_prometheus_pull_private_key_path: "{{ atlas_prometheus_pull_ssh_dir }}/id_ed25519"
|
||||
atlas_prometheus_pull_known_hosts_path: "{{ atlas_prometheus_pull_ssh_dir }}/known_hosts"
|
||||
atlas_prometheus_ssh_host_key: ""
|
||||
atlas_prometheus_pull_source_user: prometheus-backup
|
||||
atlas_prometheus_pull_source_port: 22
|
||||
atlas_prometheus_pull_calendar: "*-*-* 03:00:00 Europe/Rome"
|
||||
atlas_prometheus_pull_start_timer: false
|
||||
atlas_prometheus_pull_keep_daily: 30
|
||||
atlas_prometheus_pull_keep_weekly: 8
|
||||
atlas_prometheus_pull_keep_monthly: 12
|
||||
atlas_prometheus_pull_max_age_hours: 24
|
||||
atlas_photobook_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_photobook }}"
|
||||
|
||||
# Rootless Gitea runs in admin's user manager; the image maps internal gitea to UID/GID 1000.
|
||||
atlas_manage_gitea: false
|
||||
atlas_gitea_username: "{{ atlas_admin_username }}"
|
||||
atlas_gitea_group: "{{ atlas_admin_group }}"
|
||||
atlas_gitea_uid: "{{ atlas_admin_uid }}"
|
||||
atlas_gitea_gid: "{{ atlas_admin_gid }}"
|
||||
atlas_gitea_home: "{{ atlas_admin_home }}"
|
||||
atlas_gitea_container_uid: 1000
|
||||
atlas_gitea_container_gid: 1000
|
||||
atlas_gitea_legacy_username: gitea
|
||||
atlas_gitea_legacy_uid: 1101
|
||||
atlas_gitea_legacy_home: /var/lib/atlas-gitea
|
||||
atlas_gitea_owner_migration: false
|
||||
atlas_gitea_dataset: "{{ atlas_zfs_pool }}/services/data/gitea"
|
||||
atlas_gitea_mountpoint: "{{ atlas_app_data_mountpoint }}/gitea"
|
||||
atlas_gitea_quadlet_dir: "{{ atlas_gitea_home }}/.config/containers/systemd"
|
||||
atlas_gitea_image: localhost/atlas-gitea:1.25.2-user-gitea-v1
|
||||
atlas_gitea_image_build_dir: "{{ atlas_gitea_home }}/.local/share/atlas-gitea-image"
|
||||
atlas_gitea_production_enabled: false
|
||||
atlas_gitea_public_domain: ""
|
||||
atlas_gitea_bind_address: "{{ ansible_host }}"
|
||||
atlas_gitea_http_port: 3000
|
||||
atlas_gitea_ssh_port: 2222
|
||||
atlas_gitea_staging_bind_address: 127.0.0.1
|
||||
atlas_gitea_staging_http_port: 3001
|
||||
atlas_gitea_staging_ssh_port: 2223
|
||||
atlas_gitea_restore_test: false
|
||||
atlas_gitea_final_restore: false
|
||||
atlas_gitea_restore_helper: /usr/local/libexec/atlas-gitea-restore-test
|
||||
|
||||
# Declare storage and an inactive Quadlet only. The operator supplies the
|
||||
# private configuration, handles MFA, and starts the user service manually.
|
||||
atlas_icloudpd_dataset: "{{ atlas_zfs_pool }}/services/data/icloudpd"
|
||||
atlas_icloudpd_state_dir: "{{ atlas_app_data_mountpoint }}/icloudpd"
|
||||
atlas_icloudpd_config_dir: "{{ atlas_icloudpd_state_dir }}/config"
|
||||
atlas_icloudpd_photos_dir: "{{ atlas_archive_mountpoint }}/Pictures/iCloudPD"
|
||||
atlas_icloudpd_image: >-
|
||||
docker.io/boredazfcuk/icloudpd@sha256:9966c31ddf0b5b306ac2410b4edd5d626806d96e80c92b83cbb689972dc9389f
|
||||
atlas_icloudpd_quadlet_dir: "{{ atlas_admin_home }}/.config/containers/systemd"
|
||||
atlas_icloudpd_timezone: Europe/Rome
|
||||
|
||||
atlas_45drives_repo_url: https://repo.45drives.com/repofiles/rocky/45drives-enterprise.repo
|
||||
atlas_45drives_repo_file: /etc/yum.repos.d/45drives-enterprise.repo
|
||||
atlas_45drives_packages:
|
||||
|
||||
@@ -0,0 +1,10 @@
|
||||
FROM docker.gitea.com/gitea@sha256:f1943db2d2f1e447e857b3f0aee4ebb7b184500f86e5b80eae110fd435435906
|
||||
|
||||
# Preserve the official image's UID/GID, paths and entrypoint; change only the
|
||||
# internal Unix identity. The host-side rootless owner is Atlas admin.
|
||||
USER 0
|
||||
RUN sed -i 's/^git:x:1000:1000:/gitea:x:1000:1000:/' /etc/passwd \
|
||||
&& sed -i 's/^git:x:1000:/gitea:x:1000:/' /etc/group \
|
||||
&& grep -q '^gitea:x:1000:1000:' /etc/passwd \
|
||||
&& grep -q '^gitea:x:1000:' /etc/group
|
||||
USER 1000:1000
|
||||
249
ansible/roles/profile_atlas/files/atlas-gitea-restore-test.py
Normal file
249
ansible/roles/profile_atlas/files/atlas-gitea-restore-test.py
Normal file
@@ -0,0 +1,249 @@
|
||||
#!/usr/bin/python3
|
||||
"""Rehearse a selective rootful-to-rootless Gitea restore, never a cutover."""
|
||||
|
||||
import argparse
|
||||
import hashlib
|
||||
import json
|
||||
import os
|
||||
from pathlib import Path, PurePosixPath
|
||||
import re
|
||||
import shutil
|
||||
import sqlite3
|
||||
import tarfile
|
||||
import tempfile
|
||||
|
||||
|
||||
SOURCE_PREFIX = PurePosixPath("opt/gitea/data")
|
||||
HOST_KEYS = (
|
||||
"ssh_host_ed25519_key",
|
||||
"ssh_host_rsa_key",
|
||||
"ssh_host_ecdsa_key",
|
||||
)
|
||||
SERVER_SETTINGS = {
|
||||
"START_SSH_SERVER": "true",
|
||||
"BUILTIN_SSH_SERVER_USER": "git",
|
||||
"SSH_USER": "git",
|
||||
"SSH_PORT": "2222",
|
||||
"SSH_LISTEN_PORT": "2222",
|
||||
"SSH_SERVER_HOST_KEYS": ", ".join(
|
||||
f"/var/lib/gitea/ssh/{key}" for key in HOST_KEYS
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
def sha256(path):
|
||||
digest = hashlib.sha256()
|
||||
with path.open("rb") as stream:
|
||||
for chunk in iter(lambda: stream.read(1024 * 1024), b""):
|
||||
digest.update(chunk)
|
||||
return digest.hexdigest()
|
||||
|
||||
|
||||
def expected_digest(backup):
|
||||
checksum = (backup / "payload.sha256").read_text().strip().split()
|
||||
if len(checksum) != 2 or checksum[1] != "payload.tar":
|
||||
raise ValueError("Unexpected Prometheus backup checksum manifest")
|
||||
if not re.fullmatch(r"[0-9a-f]{64}", checksum[0]):
|
||||
raise ValueError("Invalid Prometheus backup SHA-256")
|
||||
return checksum[0]
|
||||
|
||||
|
||||
def convert_config(config):
|
||||
original = config.read_text()
|
||||
output = []
|
||||
section = ""
|
||||
server_seen = set()
|
||||
server_found = False
|
||||
run_user_seen = False
|
||||
|
||||
def append_missing_server_settings():
|
||||
for key, value in SERVER_SETTINGS.items():
|
||||
if key not in server_seen:
|
||||
output.append(f"{key} = {value}\n")
|
||||
|
||||
for line in original.splitlines(keepends=True):
|
||||
match = re.match(r"^\s*\[([^]]+)\]\s*$", line)
|
||||
if match:
|
||||
if not run_user_seen:
|
||||
output.append("RUN_USER = gitea\n")
|
||||
run_user_seen = True
|
||||
if section == "server":
|
||||
append_missing_server_settings()
|
||||
section = match.group(1).lower()
|
||||
server_found |= section == "server"
|
||||
output.append(line)
|
||||
continue
|
||||
setting = re.match(r"^(\s*)([A-Z_]+)(\s*=\s*)(.*?)(\r?\n?)$", line)
|
||||
if setting and section == "" and setting.group(2) == "RUN_USER":
|
||||
run_user_seen = True
|
||||
line = f"{setting.group(1)}RUN_USER{setting.group(3)}gitea{setting.group(5)}"
|
||||
elif setting and section == "server" and setting.group(2) in SERVER_SETTINGS:
|
||||
key = setting.group(2)
|
||||
server_seen.add(key)
|
||||
line = f"{setting.group(1)}{key}{setting.group(3)}{SERVER_SETTINGS[key]}{setting.group(5)}"
|
||||
else:
|
||||
line = line.replace("/data/", "/var/lib/gitea/")
|
||||
output.append(line)
|
||||
if section == "server":
|
||||
append_missing_server_settings()
|
||||
if not server_found:
|
||||
raise ValueError("Gitea server configuration missing")
|
||||
config.write_text("".join(output))
|
||||
config.chmod(0o600)
|
||||
|
||||
|
||||
def extract_gitea(tar_path, staged_data):
|
||||
count = 0
|
||||
with tarfile.open(tar_path, mode="r") as archive:
|
||||
for member in archive:
|
||||
name = PurePosixPath(member.name)
|
||||
if name == SOURCE_PREFIX:
|
||||
continue
|
||||
if SOURCE_PREFIX not in name.parents:
|
||||
continue
|
||||
relative = name.relative_to(SOURCE_PREFIX)
|
||||
if not relative.parts or any(part in (".", "..") for part in relative.parts):
|
||||
raise ValueError("Unsafe Gitea backup path")
|
||||
if not (member.isdir() or member.isfile()):
|
||||
raise ValueError("Unexpected Gitea backup member type")
|
||||
destination = staged_data.joinpath(*relative.parts)
|
||||
if member.isdir():
|
||||
destination.mkdir(parents=True, exist_ok=True)
|
||||
destination.chmod(0o700)
|
||||
continue
|
||||
destination.parent.mkdir(parents=True, exist_ok=True)
|
||||
with archive.extractfile(member) as source, destination.open("xb") as target:
|
||||
shutil.copyfileobj(source, target)
|
||||
destination.chmod(member.mode & 0o777)
|
||||
count += 1
|
||||
if count == 0:
|
||||
raise ValueError("No Gitea files in backup")
|
||||
|
||||
|
||||
def validate(staged_data, staged_config):
|
||||
database = staged_data / "gitea/gitea.db"
|
||||
repositories = staged_data / "git/repositories"
|
||||
if not database.is_file() or not repositories.is_dir():
|
||||
raise ValueError("Missing SQLite database or Git repositories")
|
||||
with sqlite3.connect(f"file:{database}?mode=ro", uri=True) as connection:
|
||||
if connection.execute("PRAGMA quick_check").fetchone()[0] != "ok":
|
||||
raise ValueError("Gitea SQLite quick_check failed")
|
||||
if connection.execute("SELECT count(*) FROM repository").fetchone()[0] < 1:
|
||||
raise ValueError("Gitea backup contains no repository records")
|
||||
if not any(repositories.rglob("*.git")):
|
||||
raise ValueError("Gitea backup contains no Git repository directories")
|
||||
if not (staged_config / "app.ini").is_file():
|
||||
raise ValueError("Gitea app.ini missing")
|
||||
for name in HOST_KEYS:
|
||||
if not (staged_data / "ssh" / name).is_file():
|
||||
raise ValueError("Gitea SSH host key missing")
|
||||
|
||||
|
||||
def chown_tree(root, uid, gid):
|
||||
for directory, dirs, files in os.walk(root):
|
||||
os.chown(directory, uid, gid)
|
||||
for name in dirs + files:
|
||||
os.chown(os.path.join(directory, name), uid, gid)
|
||||
|
||||
|
||||
def replace_rehearsal(target, stage, digest, uid, gid):
|
||||
previous_data = target / ".previous-rehearsal-data"
|
||||
previous_config = target / ".previous-rehearsal-config"
|
||||
if previous_data.exists() or previous_config.exists():
|
||||
raise ValueError("An interrupted Gitea replacement needs manual recovery")
|
||||
os.rename(target / "data", previous_data)
|
||||
try:
|
||||
os.rename(target / "config", previous_config)
|
||||
os.rename(stage / "data", target / "data")
|
||||
os.rename(stage / "config", target / "config")
|
||||
final_marker = target / ".final-sha256"
|
||||
final_marker.write_text(digest + "\n")
|
||||
final_marker.chmod(0o600)
|
||||
os.chown(final_marker, uid, gid)
|
||||
(target / ".rehearsal-sha256").unlink()
|
||||
except Exception:
|
||||
for name, previous in (("data", previous_data), ("config", previous_config)):
|
||||
current = target / name
|
||||
if previous.exists():
|
||||
if current.exists():
|
||||
shutil.rmtree(current)
|
||||
os.rename(previous, current)
|
||||
(target / ".final-sha256").unlink(missing_ok=True)
|
||||
raise
|
||||
shutil.rmtree(previous_data)
|
||||
shutil.rmtree(previous_config)
|
||||
|
||||
|
||||
def main():
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("--backup", type=Path, required=True)
|
||||
parser.add_argument("--target", type=Path, required=True)
|
||||
parser.add_argument("--uid", type=int, required=True)
|
||||
parser.add_argument("--gid", type=int, required=True)
|
||||
parser.add_argument("--replace-rehearsal", action="store_true")
|
||||
args = parser.parse_args()
|
||||
|
||||
backup = args.backup.resolve(strict=True)
|
||||
target = args.target.resolve(strict=True)
|
||||
if not str(backup).startswith("/zpool/backup/hosts/prometheus/snapshots/"):
|
||||
raise ValueError("Refusing backup outside the Atlas Prometheus snapshots")
|
||||
if str(target) != "/zpool/services/data/gitea":
|
||||
raise ValueError("Refusing target outside the dedicated Gitea dataset")
|
||||
if args.uid != 1000 or args.gid != 1000:
|
||||
raise ValueError("Unexpected admin-owned Gitea account IDs")
|
||||
expected = expected_digest(backup)
|
||||
if sha256(backup / "payload.tar") != expected:
|
||||
raise ValueError("Prometheus backup SHA-256 mismatch")
|
||||
|
||||
marker = target / (".final-sha256" if args.replace_rehearsal else ".rehearsal-sha256")
|
||||
if marker.exists():
|
||||
if marker.read_text().strip() != expected:
|
||||
raise ValueError("A different Gitea restore already occupies this dataset")
|
||||
validate(target / "data", target / "config")
|
||||
print("unchanged")
|
||||
return
|
||||
if args.replace_rehearsal:
|
||||
metadata = json.loads((backup / "metadata.json").read_text())
|
||||
if metadata.get("purpose") != "gitea-cutover":
|
||||
raise ValueError("Final restore requires an explicit Gitea cutover export")
|
||||
if not (target / ".rehearsal-sha256").is_file():
|
||||
raise ValueError("Only a marked rehearsal may be replaced")
|
||||
if not all((target / name).is_dir() for name in ("data", "config")):
|
||||
raise ValueError("Prepared Gitea volume paths are missing")
|
||||
else:
|
||||
if (target / ".final-sha256").exists():
|
||||
raise ValueError("Refusing a rehearsal restore over final Gitea data")
|
||||
for name in ("data", "config"):
|
||||
directory = target / name
|
||||
if not directory.is_dir() or any(directory.iterdir()):
|
||||
raise ValueError("Gitea target is not empty; refusing overwrite")
|
||||
|
||||
with tempfile.TemporaryDirectory(prefix=".rehearsal-", dir=target) as temporary:
|
||||
stage = Path(temporary)
|
||||
staged_data = stage / "data"
|
||||
staged_config = stage / "config"
|
||||
staged_data.mkdir()
|
||||
staged_config.mkdir()
|
||||
extract_gitea(backup / "payload.tar", staged_data)
|
||||
source_config = staged_data / "gitea/conf/app.ini"
|
||||
if not source_config.is_file():
|
||||
raise ValueError("Source Gitea app.ini missing")
|
||||
shutil.copy2(source_config, staged_config / "app.ini")
|
||||
source_config.unlink()
|
||||
convert_config(staged_config / "app.ini")
|
||||
validate(staged_data, staged_config)
|
||||
chown_tree(stage, args.uid, args.gid)
|
||||
if args.replace_rehearsal:
|
||||
replace_rehearsal(target, stage, expected, args.uid, args.gid)
|
||||
else:
|
||||
for name in ("data", "config"):
|
||||
(target / name).rmdir()
|
||||
os.rename(stage / name, target / name)
|
||||
marker.write_text(expected + "\n")
|
||||
marker.chmod(0o600)
|
||||
os.chown(marker, args.uid, args.gid)
|
||||
print("restored")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
56
ansible/roles/profile_atlas/files/atlas-prometheus-prune.py
Normal file
56
ansible/roles/profile_atlas/files/atlas-prometheus-prune.py
Normal file
@@ -0,0 +1,56 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Prune only verified, named Prometheus backup versions after publication."""
|
||||
|
||||
import datetime as dt
|
||||
import pathlib
|
||||
import re
|
||||
import shutil
|
||||
import sys
|
||||
|
||||
|
||||
def main() -> None:
|
||||
if len(sys.argv) != 5:
|
||||
raise SystemExit("Usage: atlas-prometheus-prune SNAPSHOTS DAILY WEEKLY MONTHLY")
|
||||
root = pathlib.Path(sys.argv[1])
|
||||
counts = [int(value) for value in sys.argv[2:]]
|
||||
if not root.is_dir() or root.is_symlink() or min(counts) < 1:
|
||||
raise SystemExit("Invalid backup directory or retention counts")
|
||||
versions = []
|
||||
for entry in root.iterdir():
|
||||
if not entry.is_dir() or entry.is_symlink():
|
||||
continue
|
||||
if not re.fullmatch(r"[0-9]{8}T[0-9]{6}Z", entry.name):
|
||||
continue
|
||||
try:
|
||||
when = dt.datetime.strptime(entry.name, "%Y%m%dT%H%M%SZ")
|
||||
except ValueError:
|
||||
continue
|
||||
if not all((entry / name).is_file() for name in ("payload.tar", "payload.sha256", "metadata.json")):
|
||||
continue
|
||||
versions.append((when, entry))
|
||||
versions.sort(reverse=True)
|
||||
if not versions:
|
||||
raise SystemExit("No published backup versions found; refusing to prune")
|
||||
|
||||
keep = {entry for _, entry in versions[: counts[0]]}
|
||||
for count, key in (
|
||||
(counts[1], lambda when: when.isocalendar()[:2]),
|
||||
(counts[2], lambda when: (when.year, when.month)),
|
||||
):
|
||||
periods = set()
|
||||
for when, entry in versions:
|
||||
period = key(when)
|
||||
if period in periods:
|
||||
continue
|
||||
periods.add(period)
|
||||
keep.add(entry)
|
||||
if len(periods) >= count:
|
||||
break
|
||||
|
||||
for _, entry in versions:
|
||||
if entry not in keep:
|
||||
shutil.rmtree(entry)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -1,4 +1,14 @@
|
||||
---
|
||||
- name: Reload Atlas admin user manager
|
||||
become_user: "{{ atlas_admin_username }}"
|
||||
ansible.builtin.systemd:
|
||||
scope: user
|
||||
daemon_reload: true
|
||||
environment:
|
||||
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
|
||||
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus"
|
||||
when: not ansible_check_mode
|
||||
|
||||
- name: Reload SSH service
|
||||
ansible.builtin.systemd:
|
||||
name: sshd
|
||||
|
||||
159
ansible/roles/profile_atlas/tasks/gitea.yml
Normal file
159
ansible/roles/profile_atlas/tasks/gitea.yml
Normal file
@@ -0,0 +1,159 @@
|
||||
---
|
||||
- name: Prepare the isolated rootless Atlas Gitea target
|
||||
tags: [atlas, gitea]
|
||||
when: atlas_manage_gitea | bool
|
||||
block:
|
||||
- name: Require the existing Atlas application-data dataset
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- atlas_manage_storage | bool
|
||||
- atlas_gitea_dataset == atlas_zfs_pool ~ '/services/data/gitea'
|
||||
- atlas_gitea_mountpoint == atlas_app_data_mountpoint ~ '/gitea'
|
||||
- atlas_gitea_username == atlas_admin_username
|
||||
- atlas_gitea_group == atlas_admin_group
|
||||
- atlas_gitea_uid | int == atlas_admin_uid | int
|
||||
- atlas_gitea_gid | int == atlas_admin_gid | int
|
||||
- atlas_gitea_container_uid | int == 1000
|
||||
- atlas_gitea_container_gid | int == 1000
|
||||
- atlas_gitea_staging_bind_address == '127.0.0.1'
|
||||
- not (atlas_gitea_production_enabled | bool) or atlas_manage_firewall | bool
|
||||
- not (atlas_gitea_production_enabled | bool) or atlas_gitea_bind_address == ansible_host
|
||||
fail_msg: >-
|
||||
Rootless Gitea requires Atlas storage, the admin user manager, the
|
||||
dedicated dataset, and loopback-only staging ports.
|
||||
|
||||
- name: Inspect the final-restore marker before production activation
|
||||
ansible.builtin.stat:
|
||||
path: "{{ atlas_gitea_mountpoint }}/.final-sha256"
|
||||
register: atlas_gitea_final_marker
|
||||
when: atlas_gitea_production_enabled | bool
|
||||
|
||||
- name: Refuse production activation without the final consistent restore
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- atlas_gitea_final_marker.stat.isreg | default(false)
|
||||
fail_msg: Restore the final stopped-source Gitea export before enabling production.
|
||||
when: atlas_gitea_production_enabled | bool
|
||||
|
||||
- name: Verify the production Gitea dataset belongs to admin
|
||||
ansible.builtin.stat:
|
||||
path: "{{ atlas_gitea_mountpoint }}"
|
||||
register: atlas_gitea_dataset_owner
|
||||
when: atlas_gitea_production_enabled | bool
|
||||
|
||||
- name: Refuse to overlap the legacy host-account service
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- atlas_gitea_dataset_owner.stat.uid | int == atlas_admin_uid | int
|
||||
- atlas_gitea_dataset_owner.stat.gid | int == atlas_admin_gid | int
|
||||
fail_msg: >-
|
||||
Run the explicit Gitea owner migration before enabling the admin
|
||||
Quadlet; never chown an active legacy service in a normal run.
|
||||
when: atlas_gitea_production_enabled | bool
|
||||
|
||||
- name: Remove the retired account's parent-dataset traverse ACL
|
||||
ansible.posix.acl:
|
||||
path: "{{ item }}"
|
||||
etype: user
|
||||
entity: "{{ atlas_gitea_legacy_username }}"
|
||||
state: absent
|
||||
loop:
|
||||
- "{{ atlas_services_mountpoint }}"
|
||||
- "{{ atlas_app_data_mountpoint }}"
|
||||
when: atlas_gitea_production_enabled | bool
|
||||
|
||||
- name: Enable POSIX ACLs only on the service-namespace parents
|
||||
community.general.zfs:
|
||||
name: "{{ item }}"
|
||||
state: present
|
||||
extra_zfs_properties:
|
||||
acltype: posix
|
||||
loop:
|
||||
- "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_services }}"
|
||||
- "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_app_data }}"
|
||||
|
||||
- name: Create the dedicated Gitea ZFS dataset
|
||||
community.general.zfs:
|
||||
name: "{{ atlas_gitea_dataset }}"
|
||||
state: present
|
||||
extra_zfs_properties:
|
||||
compression: zstd
|
||||
mountpoint: "{{ atlas_gitea_mountpoint }}"
|
||||
|
||||
- name: Restrict the Gitea dataset and create rootless volume paths
|
||||
ansible.builtin.file:
|
||||
path: "{{ item }}"
|
||||
state: directory
|
||||
owner: "{{ atlas_gitea_username }}"
|
||||
group: "{{ atlas_gitea_group }}"
|
||||
mode: "0700"
|
||||
loop:
|
||||
- "{{ atlas_gitea_mountpoint }}"
|
||||
- "{{ atlas_gitea_mountpoint }}/data"
|
||||
- "{{ atlas_gitea_mountpoint }}/config"
|
||||
- "{{ atlas_gitea_home }}/.config"
|
||||
- "{{ atlas_gitea_home }}/.config/containers"
|
||||
- "{{ atlas_gitea_quadlet_dir }}"
|
||||
|
||||
- name: Ensure lingering for the admin rootless account
|
||||
ansible.builtin.command:
|
||||
argv:
|
||||
- loginctl
|
||||
- enable-linger
|
||||
- "{{ atlas_gitea_username }}"
|
||||
creates: "/var/lib/systemd/linger/{{ atlas_gitea_username }}"
|
||||
|
||||
- name: Start the admin rootless user manager
|
||||
ansible.builtin.systemd:
|
||||
name: "user@{{ atlas_gitea_uid }}.service"
|
||||
state: started
|
||||
when: not ansible_check_mode
|
||||
|
||||
- name: Prepare the admin-owned Gitea image
|
||||
ansible.builtin.import_tasks: gitea_image.yml
|
||||
|
||||
- name: Render the rootless Gitea Quadlet
|
||||
ansible.builtin.template:
|
||||
src: atlas-gitea.container.j2
|
||||
dest: "{{ atlas_gitea_quadlet_dir }}/atlas-gitea.container"
|
||||
owner: "{{ atlas_gitea_username }}"
|
||||
group: "{{ atlas_gitea_group }}"
|
||||
mode: "0644"
|
||||
|
||||
- name: Permit only Aegis to reach production Gitea HTTP and SSH
|
||||
ansible.posix.firewalld:
|
||||
rich_rule: >-
|
||||
rule family="ipv4" source address="{{ atlas_aegis_ip }}"
|
||||
port port="{{ item }}" protocol="tcp" accept
|
||||
zone: "{{ atlas_firewalld_zone }}"
|
||||
state: "{{ 'enabled' if atlas_gitea_production_enabled | bool else 'disabled' }}"
|
||||
permanent: true
|
||||
immediate: true
|
||||
loop:
|
||||
- "{{ atlas_gitea_http_port }}"
|
||||
- "{{ atlas_gitea_ssh_port }}"
|
||||
when: atlas_manage_firewall | bool
|
||||
|
||||
- name: Reload the rootless Gitea user manager without starting Gitea
|
||||
become_user: "{{ atlas_gitea_username }}"
|
||||
ansible.builtin.systemd:
|
||||
scope: user
|
||||
daemon_reload: true
|
||||
environment:
|
||||
XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_uid }}"
|
||||
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_gitea_uid }}/bus"
|
||||
when: not ansible_check_mode
|
||||
|
||||
- name: Start and enable the rootless Gitea user Quadlet after final restore
|
||||
become_user: "{{ atlas_gitea_username }}"
|
||||
ansible.builtin.systemd:
|
||||
name: atlas-gitea.service
|
||||
scope: user
|
||||
state: started
|
||||
enabled: true
|
||||
environment:
|
||||
XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_uid }}"
|
||||
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_gitea_uid }}/bus"
|
||||
when:
|
||||
- atlas_gitea_production_enabled | bool
|
||||
- not ansible_check_mode
|
||||
51
ansible/roles/profile_atlas/tasks/gitea_image.yml
Normal file
51
ansible/roles/profile_atlas/tasks/gitea_image.yml
Normal file
@@ -0,0 +1,51 @@
|
||||
---
|
||||
- name: Create the admin-owned Gitea image build directory
|
||||
ansible.builtin.file:
|
||||
path: "{{ atlas_gitea_image_build_dir }}"
|
||||
state: directory
|
||||
owner: "{{ atlas_admin_username }}"
|
||||
group: "{{ atlas_admin_group }}"
|
||||
mode: "0700"
|
||||
|
||||
- name: Install the pinned rootless Gitea Containerfile
|
||||
ansible.builtin.copy:
|
||||
src: Containerfile.gitea-rootless
|
||||
dest: "{{ atlas_gitea_image_build_dir }}/Containerfile"
|
||||
owner: "{{ atlas_admin_username }}"
|
||||
group: "{{ atlas_admin_group }}"
|
||||
mode: "0644"
|
||||
|
||||
- name: Check the admin-owned Gitea image
|
||||
become_user: "{{ atlas_admin_username }}"
|
||||
ansible.builtin.command:
|
||||
argv: [podman, image, exists, "{{ atlas_gitea_image }}"]
|
||||
args:
|
||||
chdir: "{{ atlas_gitea_image_build_dir }}"
|
||||
environment:
|
||||
HOME: "{{ atlas_admin_home }}"
|
||||
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
|
||||
register: atlas_gitea_image_present
|
||||
changed_when: false
|
||||
failed_when: false
|
||||
check_mode: false
|
||||
|
||||
- name: Build the pinned Gitea image with the internal gitea identity
|
||||
become_user: "{{ atlas_admin_username }}"
|
||||
ansible.builtin.command:
|
||||
argv:
|
||||
- podman
|
||||
- build
|
||||
- --pull=always
|
||||
- --tag
|
||||
- "{{ atlas_gitea_image }}"
|
||||
- --file
|
||||
- Containerfile
|
||||
- .
|
||||
args:
|
||||
chdir: "{{ atlas_gitea_image_build_dir }}"
|
||||
environment:
|
||||
HOME: "{{ atlas_admin_home }}"
|
||||
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
|
||||
when:
|
||||
- atlas_gitea_image_present.rc != 0
|
||||
- not ansible_check_mode
|
||||
275
ansible/roles/profile_atlas/tasks/gitea_owner_migration.yml
Normal file
275
ansible/roles/profile_atlas/tasks/gitea_owner_migration.yml
Normal file
@@ -0,0 +1,275 @@
|
||||
---
|
||||
# Run only in an approved outage with -e atlas_gitea_owner_migration=true.
|
||||
- name: Move live Gitea from the legacy host account to admin
|
||||
tags: [atlas, gitea_owner_migration]
|
||||
when: atlas_gitea_owner_migration | bool
|
||||
block:
|
||||
- name: Refuse a check-mode owner migration
|
||||
ansible.builtin.assert:
|
||||
that: not ansible_check_mode
|
||||
fail_msg: The owner migration requires an explicit live outage.
|
||||
|
||||
- name: Inspect the Gitea dataset owner
|
||||
ansible.builtin.stat:
|
||||
path: "{{ atlas_gitea_mountpoint }}"
|
||||
register: atlas_gitea_migration_owner
|
||||
|
||||
- name: Require either the legacy owner or an already migrated dataset
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- atlas_gitea_migration_owner.stat.isdir | default(false)
|
||||
- atlas_gitea_migration_owner.stat.uid | int in [atlas_gitea_legacy_uid | int, atlas_admin_uid | int]
|
||||
fail_msg: Refusing to modify a Gitea dataset with an unexpected owner.
|
||||
|
||||
- name: Migrate only a legacy-owned Gitea dataset
|
||||
when: atlas_gitea_migration_owner.stat.uid | int == atlas_gitea_legacy_uid | int
|
||||
block:
|
||||
- name: Require the final cutover marker and configuration
|
||||
ansible.builtin.stat:
|
||||
path: "{{ item }}"
|
||||
loop:
|
||||
- "{{ atlas_gitea_mountpoint }}/.final-sha256"
|
||||
- "{{ atlas_gitea_mountpoint }}/config/app.ini"
|
||||
register: atlas_gitea_migration_files
|
||||
|
||||
- name: Refuse migration without both final data and configuration
|
||||
ansible.builtin.assert:
|
||||
that: atlas_gitea_migration_files.results | map(attribute='stat.isreg') | min
|
||||
|
||||
- name: Check that admin has no existing Gitea Quadlet
|
||||
ansible.builtin.stat:
|
||||
path: "{{ atlas_gitea_quadlet_dir }}/atlas-gitea.container"
|
||||
register: atlas_gitea_admin_quadlet
|
||||
|
||||
- name: Refuse to overwrite an existing admin Quadlet
|
||||
ansible.builtin.assert:
|
||||
that: not atlas_gitea_admin_quadlet.stat.exists
|
||||
|
||||
- name: Check pool health before the outage
|
||||
ansible.builtin.command:
|
||||
argv: [zpool, status, -x, "{{ atlas_zfs_pool }}"]
|
||||
register: atlas_gitea_pool_before
|
||||
changed_when: false
|
||||
failed_when: "'is healthy' not in atlas_gitea_pool_before.stdout"
|
||||
|
||||
- name: Ensure the admin Gitea image is available before stopping the source
|
||||
ansible.builtin.import_tasks: gitea_image.yml
|
||||
|
||||
- name: Stop, snapshot and test the admin-owned staging service
|
||||
block:
|
||||
- name: Stop and disable the legacy Gitea user service
|
||||
become_user: "{{ atlas_gitea_legacy_username }}"
|
||||
ansible.builtin.systemd:
|
||||
name: atlas-gitea.service
|
||||
scope: user
|
||||
state: stopped
|
||||
enabled: false
|
||||
environment:
|
||||
XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_legacy_uid }}"
|
||||
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_gitea_legacy_uid }}/bus"
|
||||
|
||||
- name: Record the migration snapshot name
|
||||
ansible.builtin.set_fact:
|
||||
atlas_gitea_migration_snapshot: >-
|
||||
{{ atlas_gitea_dataset }}@gitea-owner-migration-{{ ansible_facts.date_time.iso8601_basic_short }}
|
||||
|
||||
- name: Snapshot the stopped Gitea dataset for manual recovery
|
||||
ansible.builtin.command:
|
||||
argv: [zfs, snapshot, "{{ atlas_gitea_migration_snapshot }}"]
|
||||
|
||||
- name: Transfer only the Gitea dataset to admin
|
||||
ansible.builtin.file:
|
||||
path: "{{ atlas_gitea_mountpoint }}"
|
||||
state: directory
|
||||
owner: "{{ atlas_admin_username }}"
|
||||
group: "{{ atlas_admin_group }}"
|
||||
recurse: true
|
||||
|
||||
- name: Set the actual internal Unix process user
|
||||
ansible.builtin.lineinfile:
|
||||
path: "{{ atlas_gitea_mountpoint }}/config/app.ini"
|
||||
regexp: '^RUN_USER\s*='
|
||||
line: RUN_USER = gitea
|
||||
mode: "0600"
|
||||
no_log: true
|
||||
diff: false
|
||||
|
||||
- name: Preserve public git clone URLs independently of the Unix user
|
||||
community.general.ini_file:
|
||||
path: "{{ atlas_gitea_mountpoint }}/config/app.ini"
|
||||
section: server
|
||||
option: "{{ item }}"
|
||||
value: git
|
||||
mode: "0600"
|
||||
no_extra_spaces: false
|
||||
loop: [BUILTIN_SSH_SERVER_USER, SSH_USER]
|
||||
no_log: true
|
||||
diff: false
|
||||
|
||||
- name: Render admin's loopback-only staging Quadlet
|
||||
ansible.builtin.template:
|
||||
src: atlas-gitea.container.j2
|
||||
dest: "{{ atlas_gitea_quadlet_dir }}/atlas-gitea.container"
|
||||
owner: "{{ atlas_admin_username }}"
|
||||
group: "{{ atlas_admin_group }}"
|
||||
mode: "0644"
|
||||
vars:
|
||||
atlas_gitea_production_enabled: false
|
||||
|
||||
- name: Reload the admin user manager for staging
|
||||
become_user: "{{ atlas_admin_username }}"
|
||||
ansible.builtin.systemd:
|
||||
scope: user
|
||||
daemon_reload: true
|
||||
environment:
|
||||
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
|
||||
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus"
|
||||
|
||||
- name: Start admin's loopback-only staging service
|
||||
become_user: "{{ atlas_admin_username }}"
|
||||
ansible.builtin.systemd:
|
||||
name: atlas-gitea.service
|
||||
scope: user
|
||||
state: started
|
||||
environment:
|
||||
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
|
||||
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus"
|
||||
|
||||
- name: Verify staging HTTP before promotion
|
||||
ansible.builtin.uri:
|
||||
url: "http://127.0.0.1:{{ atlas_gitea_staging_http_port }}/"
|
||||
status_code: 200
|
||||
register: atlas_gitea_staging_http
|
||||
retries: 30
|
||||
delay: 2
|
||||
until: atlas_gitea_staging_http is succeeded
|
||||
|
||||
- name: Verify the container really runs as internal gitea
|
||||
become_user: "{{ atlas_admin_username }}"
|
||||
ansible.builtin.command:
|
||||
argv: [podman, exec, atlas-gitea, id, -un]
|
||||
environment:
|
||||
HOME: "{{ atlas_admin_home }}"
|
||||
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
|
||||
register: atlas_gitea_internal_user
|
||||
changed_when: false
|
||||
failed_when: atlas_gitea_internal_user.stdout != 'gitea'
|
||||
|
||||
- name: Verify the migrated SQLite database
|
||||
ansible.builtin.command:
|
||||
argv:
|
||||
- sqlite3
|
||||
- "{{ atlas_gitea_mountpoint }}/data/gitea/gitea.db"
|
||||
- PRAGMA quick_check;
|
||||
register: atlas_gitea_migration_sqlite
|
||||
changed_when: false
|
||||
failed_when: atlas_gitea_migration_sqlite.stdout != 'ok'
|
||||
|
||||
rescue:
|
||||
- name: Stop admin's failed staging service
|
||||
become_user: "{{ atlas_admin_username }}"
|
||||
ansible.builtin.systemd:
|
||||
name: atlas-gitea.service
|
||||
scope: user
|
||||
state: stopped
|
||||
environment:
|
||||
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
|
||||
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus"
|
||||
failed_when: false
|
||||
|
||||
- name: Restore the original Gitea configuration from the safety snapshot
|
||||
ansible.builtin.command:
|
||||
argv:
|
||||
- cp
|
||||
- -a
|
||||
- "{{ atlas_gitea_mountpoint }}/.zfs/snapshot/{{ atlas_gitea_migration_snapshot.split('@')[1] }}/config/app.ini"
|
||||
- "{{ atlas_gitea_mountpoint }}/config/app.ini"
|
||||
when: atlas_gitea_migration_snapshot is defined
|
||||
|
||||
- name: Return the Gitea dataset to the legacy account
|
||||
ansible.builtin.file:
|
||||
path: "{{ atlas_gitea_mountpoint }}"
|
||||
state: directory
|
||||
owner: "{{ atlas_gitea_legacy_username }}"
|
||||
group: "{{ atlas_gitea_legacy_username }}"
|
||||
recurse: true
|
||||
|
||||
- name: Restart the legacy Gitea service
|
||||
become_user: "{{ atlas_gitea_legacy_username }}"
|
||||
ansible.builtin.systemd:
|
||||
name: atlas-gitea.service
|
||||
scope: user
|
||||
state: started
|
||||
enabled: true
|
||||
environment:
|
||||
XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_legacy_uid }}"
|
||||
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_gitea_legacy_uid }}/bus"
|
||||
|
||||
- name: Report the failed migration and preserved snapshot
|
||||
ansible.builtin.fail:
|
||||
msg: >-
|
||||
Admin staging failed; legacy Gitea was restarted. Inspect
|
||||
{{ atlas_gitea_migration_snapshot | default('the host journal') }}.
|
||||
|
||||
- name: Stop admin's validated staging service
|
||||
become_user: "{{ atlas_admin_username }}"
|
||||
ansible.builtin.systemd:
|
||||
name: atlas-gitea.service
|
||||
scope: user
|
||||
state: stopped
|
||||
environment:
|
||||
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
|
||||
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus"
|
||||
|
||||
- name: Render admin's production Gitea Quadlet
|
||||
ansible.builtin.template:
|
||||
src: atlas-gitea.container.j2
|
||||
dest: "{{ atlas_gitea_quadlet_dir }}/atlas-gitea.container"
|
||||
owner: "{{ atlas_admin_username }}"
|
||||
group: "{{ atlas_admin_group }}"
|
||||
mode: "0644"
|
||||
vars:
|
||||
atlas_gitea_production_enabled: true
|
||||
|
||||
- name: Reload admin's production user manager
|
||||
become_user: "{{ atlas_admin_username }}"
|
||||
ansible.builtin.systemd:
|
||||
scope: user
|
||||
daemon_reload: true
|
||||
environment:
|
||||
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
|
||||
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus"
|
||||
|
||||
- name: Enable and start admin's production Gitea
|
||||
become_user: "{{ atlas_admin_username }}"
|
||||
ansible.builtin.systemd:
|
||||
name: atlas-gitea.service
|
||||
scope: user
|
||||
state: started
|
||||
enabled: true
|
||||
environment:
|
||||
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
|
||||
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus"
|
||||
|
||||
- name: Verify production HTTP before retiring the old Quadlet
|
||||
ansible.builtin.uri:
|
||||
url: "http://{{ atlas_gitea_bind_address }}:{{ atlas_gitea_http_port }}/"
|
||||
status_code: 200
|
||||
register: atlas_gitea_production_http
|
||||
retries: 30
|
||||
delay: 2
|
||||
until: atlas_gitea_production_http is succeeded
|
||||
|
||||
- name: Remove only the disabled legacy Quadlet
|
||||
ansible.builtin.file:
|
||||
path: "{{ atlas_gitea_legacy_home }}/.config/containers/systemd/atlas-gitea.container"
|
||||
state: absent
|
||||
|
||||
- name: Reload the legacy user manager after Quadlet removal
|
||||
become_user: "{{ atlas_gitea_legacy_username }}"
|
||||
ansible.builtin.systemd:
|
||||
scope: user
|
||||
daemon_reload: true
|
||||
environment:
|
||||
XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_legacy_uid }}"
|
||||
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_gitea_legacy_uid }}/bus"
|
||||
58
ansible/roles/profile_atlas/tasks/gitea_public_domain.yml
Normal file
58
ansible/roles/profile_atlas/tasks/gitea_public_domain.yml
Normal file
@@ -0,0 +1,58 @@
|
||||
---
|
||||
- name: Manage the public domain of the restored production Gitea
|
||||
tags: [atlas, gitea, gitea_public_domain]
|
||||
when:
|
||||
- atlas_manage_gitea | bool
|
||||
- atlas_gitea_production_enabled | bool
|
||||
- atlas_gitea_public_domain | length > 0
|
||||
block:
|
||||
- name: Require an explicit public Gitea hostname
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- atlas_gitea_public_domain is match('^[a-zA-Z0-9][a-zA-Z0-9.-]*\.[a-zA-Z]{2,}$')
|
||||
|
||||
- name: Inspect the restored private Gitea configuration
|
||||
ansible.builtin.stat:
|
||||
path: "{{ atlas_gitea_mountpoint }}/config/app.ini"
|
||||
follow: false
|
||||
register: atlas_gitea_public_config
|
||||
|
||||
- name: Refuse to create or replace an unprepared Gitea configuration
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- atlas_gitea_public_config.stat.isreg | default(false)
|
||||
- atlas_gitea_public_config.stat.uid | int == atlas_gitea_uid | int
|
||||
- atlas_gitea_public_config.stat.mode == '0600'
|
||||
|
||||
# app.ini contains secrets: preserve all unrelated settings and suppress diffs.
|
||||
- name: Set only the declared public Gitea server fields
|
||||
community.general.ini_file:
|
||||
path: "{{ atlas_gitea_mountpoint }}/config/app.ini"
|
||||
section: server
|
||||
option: "{{ item.option }}"
|
||||
value: "{{ item.value }}"
|
||||
create: false
|
||||
backup: true
|
||||
owner: "{{ atlas_gitea_username }}"
|
||||
group: "{{ atlas_gitea_group }}"
|
||||
mode: "0600"
|
||||
loop:
|
||||
- { option: DOMAIN, value: "{{ atlas_gitea_public_domain }}" }
|
||||
- { option: ROOT_URL, value: "https://{{ atlas_gitea_public_domain }}/" }
|
||||
- { option: SSH_DOMAIN, value: "{{ atlas_gitea_public_domain }}" }
|
||||
register: atlas_gitea_public_domain_update
|
||||
no_log: true
|
||||
diff: false
|
||||
|
||||
- name: Restart only Gitea when its public configuration changes
|
||||
become_user: "{{ atlas_gitea_username }}"
|
||||
ansible.builtin.systemd:
|
||||
name: atlas-gitea.service
|
||||
scope: user
|
||||
state: restarted
|
||||
environment:
|
||||
XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_uid }}"
|
||||
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_gitea_uid }}/bus"
|
||||
when:
|
||||
- atlas_gitea_public_domain_update is changed
|
||||
- not ansible_check_mode
|
||||
107
ansible/roles/profile_atlas/tasks/gitea_restore.yml
Normal file
107
ansible/roles/profile_atlas/tasks/gitea_restore.yml
Normal file
@@ -0,0 +1,107 @@
|
||||
---
|
||||
- name: Restore Gitea from a verified Prometheus backup only on explicit request
|
||||
tags: [atlas, gitea_restore, gitea_final_restore]
|
||||
when: atlas_gitea_restore_test | bool or atlas_gitea_final_restore | bool
|
||||
block:
|
||||
- name: Require the prepared rootless Gitea target
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- atlas_manage_gitea | bool
|
||||
- not (atlas_gitea_restore_test | bool and atlas_gitea_final_restore | bool)
|
||||
- atlas_gitea_staging_bind_address == '127.0.0.1'
|
||||
- atlas_gitea_mountpoint == atlas_app_data_mountpoint ~ '/gitea'
|
||||
fail_msg: Prepare the isolated, loopback-only rootless Gitea target first.
|
||||
|
||||
- name: Confirm the rootless Gitea service is inactive
|
||||
become_user: "{{ atlas_gitea_username }}"
|
||||
ansible.builtin.command:
|
||||
argv:
|
||||
- systemctl
|
||||
- --user
|
||||
- is-active
|
||||
- atlas-gitea.service
|
||||
environment:
|
||||
XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_uid }}"
|
||||
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_gitea_uid }}/bus"
|
||||
register: atlas_gitea_restore_service_state
|
||||
changed_when: false
|
||||
failed_when: false
|
||||
when: not ansible_check_mode
|
||||
|
||||
- name: Refuse to overwrite an active rootless Gitea service
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- atlas_gitea_restore_service_state.stdout == 'inactive'
|
||||
fail_msg: The rootless Gitea user service must be known and inactive before restoring data.
|
||||
when: not ansible_check_mode
|
||||
|
||||
- name: Check for a manually running rootless Gitea container
|
||||
become_user: "{{ atlas_gitea_username }}"
|
||||
ansible.builtin.command:
|
||||
argv:
|
||||
- podman
|
||||
- ps
|
||||
- --quiet
|
||||
- --filter
|
||||
- name=atlas-gitea
|
||||
args:
|
||||
chdir: "{{ atlas_gitea_home }}"
|
||||
environment:
|
||||
XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_uid }}"
|
||||
register: atlas_gitea_restore_container_state
|
||||
changed_when: false
|
||||
when: not ansible_check_mode
|
||||
|
||||
- name: Refuse to overwrite a running rootless Gitea container
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- atlas_gitea_restore_container_state.stdout | length == 0
|
||||
fail_msg: Stop every rootless Atlas Gitea container before restoring data.
|
||||
when: not ansible_check_mode
|
||||
|
||||
- name: Install the selective rootless Gitea restore helper
|
||||
ansible.builtin.copy:
|
||||
src: atlas-gitea-restore-test.py
|
||||
dest: "{{ atlas_gitea_restore_helper }}"
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0700"
|
||||
|
||||
- name: Restore only Gitea data into the isolated target
|
||||
ansible.builtin.command:
|
||||
argv:
|
||||
- "{{ atlas_gitea_restore_helper }}"
|
||||
- --backup
|
||||
- "{{ atlas_backup_prometheus_mountpoint }}/latest"
|
||||
- --target
|
||||
- "{{ atlas_gitea_mountpoint }}"
|
||||
- --uid
|
||||
- "{{ atlas_gitea_uid | string }}"
|
||||
- --gid
|
||||
- "{{ atlas_gitea_gid | string }}"
|
||||
register: atlas_gitea_restore_result
|
||||
changed_when: atlas_gitea_restore_result.stdout == 'restored'
|
||||
no_log: true
|
||||
when:
|
||||
- atlas_gitea_restore_test | bool
|
||||
- not ansible_check_mode
|
||||
|
||||
- name: Replace the marked rehearsal with the final consistent Gitea export
|
||||
ansible.builtin.command:
|
||||
argv:
|
||||
- "{{ atlas_gitea_restore_helper }}"
|
||||
- --backup
|
||||
- "{{ atlas_backup_prometheus_mountpoint }}/latest"
|
||||
- --target
|
||||
- "{{ atlas_gitea_mountpoint }}"
|
||||
- --uid
|
||||
- "{{ atlas_gitea_uid | string }}"
|
||||
- --gid
|
||||
- "{{ atlas_gitea_gid | string }}"
|
||||
- --replace-rehearsal
|
||||
register: atlas_gitea_final_restore_result
|
||||
changed_when: atlas_gitea_final_restore_result.stdout == 'restored'
|
||||
no_log: true
|
||||
when:
|
||||
- atlas_gitea_final_restore | bool
|
||||
- not ansible_check_mode
|
||||
188
ansible/roles/profile_atlas/tasks/icloudpd.yml
Normal file
188
ansible/roles/profile_atlas/tasks/icloudpd.yml
Normal file
@@ -0,0 +1,188 @@
|
||||
---
|
||||
- name: Require exact Atlas iCloudPD paths and rootless identity
|
||||
tags: [atlas, icloudpd]
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- atlas_manage_storage | bool
|
||||
- atlas_icloudpd_dataset == atlas_zfs_pool ~ '/services/data/icloudpd'
|
||||
- atlas_icloudpd_state_dir == atlas_app_data_mountpoint ~ '/icloudpd'
|
||||
- atlas_icloudpd_config_dir == atlas_icloudpd_state_dir ~ '/config'
|
||||
- atlas_icloudpd_photos_dir == atlas_archive_mountpoint ~ '/Pictures/iCloudPD'
|
||||
- atlas_admin_uid | int == 1000
|
||||
- atlas_admin_gid | int == 1000
|
||||
- atlas_icloudpd_image is search('@sha256:[0-9a-f]{64}$')
|
||||
fail_msg: Verify the fixed, separate Atlas iCloudPD photo and state paths.
|
||||
|
||||
- name: Declare rootless Atlas iCloudPD storage and boot-started Quadlet
|
||||
tags: [atlas, icloudpd]
|
||||
block:
|
||||
- name: Inspect the existing Archive and application-data datasets
|
||||
community.general.zfs_facts:
|
||||
name: "{{ item.dataset }}"
|
||||
properties: name,mounted,mountpoint
|
||||
loop:
|
||||
- dataset: "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_archive }}"
|
||||
mountpoint: "{{ atlas_archive_mountpoint }}"
|
||||
- dataset: "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_app_data }}"
|
||||
mountpoint: "{{ atlas_app_data_mountpoint }}"
|
||||
loop_control:
|
||||
label: "{{ item.dataset }}"
|
||||
register: atlas_icloudpd_parent_datasets
|
||||
|
||||
- name: Refuse missing or unmounted iCloudPD parent datasets
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- item.ansible_facts.ansible_zfs_datasets | length == 1
|
||||
- item.ansible_facts.ansible_zfs_datasets[0].mounted == 'yes'
|
||||
- item.ansible_facts.ansible_zfs_datasets[0].mountpoint == item.item.mountpoint
|
||||
loop: "{{ atlas_icloudpd_parent_datasets.results }}"
|
||||
loop_control:
|
||||
label: "{{ item.item.dataset }}"
|
||||
|
||||
- name: Inspect the existing Pictures namespace and proposed target
|
||||
ansible.builtin.stat:
|
||||
path: "{{ item }}"
|
||||
follow: false
|
||||
loop:
|
||||
- "{{ atlas_archive_mountpoint }}/Pictures"
|
||||
- "{{ atlas_icloudpd_photos_dir }}"
|
||||
- "{{ atlas_icloudpd_photos_dir }}/.atlas-icloudpd-managed"
|
||||
register: atlas_icloudpd_photo_paths
|
||||
|
||||
- name: Refuse to adopt unrelated Pictures data or a symlink
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- atlas_icloudpd_photo_paths.results[0].stat.isdir | default(false)
|
||||
- atlas_icloudpd_photo_paths.results[0].stat.uid | int == atlas_admin_uid | int
|
||||
- >-
|
||||
not atlas_icloudpd_photo_paths.results[1].stat.exists or
|
||||
(atlas_icloudpd_photo_paths.results[1].stat.isdir | default(false) and
|
||||
atlas_icloudpd_photo_paths.results[2].stat.isreg | default(false))
|
||||
fail_msg: >-
|
||||
Pictures must exist and be admin-owned; an existing iCloudPD target
|
||||
must carry its managed marker. Never adopt or replace unrelated data.
|
||||
|
||||
- name: Create a dedicated ZFS dataset for iCloudPD configuration and MFA
|
||||
community.general.zfs:
|
||||
name: "{{ atlas_icloudpd_dataset }}"
|
||||
state: present
|
||||
extra_zfs_properties:
|
||||
compression: zstd
|
||||
mountpoint: "{{ atlas_icloudpd_state_dir }}"
|
||||
|
||||
- name: Restrict iCloudPD state and the new photo subtree
|
||||
ansible.builtin.file:
|
||||
path: "{{ item.path }}"
|
||||
state: directory
|
||||
owner: "{{ atlas_admin_username }}"
|
||||
group: "{{ atlas_admin_group }}"
|
||||
mode: "{{ item.mode }}"
|
||||
loop:
|
||||
- path: "{{ atlas_icloudpd_state_dir }}"
|
||||
mode: "0700"
|
||||
- path: "{{ atlas_icloudpd_config_dir }}"
|
||||
mode: "0700"
|
||||
- path: "{{ atlas_icloudpd_photos_dir }}"
|
||||
mode: "0750"
|
||||
- path: "{{ atlas_icloudpd_quadlet_dir }}"
|
||||
mode: "0700"
|
||||
loop_control:
|
||||
label: "{{ item.path }}"
|
||||
|
||||
- name: Mark only the newly managed iCloudPD photo subtree
|
||||
ansible.builtin.copy:
|
||||
content: "Atlas iCloudPD photo subtree; do not remove source photos.\n"
|
||||
dest: "{{ atlas_icloudpd_photos_dir }}/.atlas-icloudpd-managed"
|
||||
owner: "{{ atlas_admin_username }}"
|
||||
group: "{{ atlas_admin_group }}"
|
||||
mode: "0600"
|
||||
force: false
|
||||
|
||||
- name: Install the image's required mounted-filesystem failsafe
|
||||
ansible.builtin.copy:
|
||||
content: ""
|
||||
dest: "{{ atlas_icloudpd_photos_dir }}/.mounted"
|
||||
owner: "{{ atlas_admin_username }}"
|
||||
group: "{{ atlas_admin_group }}"
|
||||
mode: "0644"
|
||||
force: false
|
||||
|
||||
- name: Require the Vault-backed iCloudPD Apple ID
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- vault_atlas_icloudpd_apple_id is defined
|
||||
- vault_atlas_icloudpd_apple_id | length > 0
|
||||
- vault_atlas_icloudpd_apple_id != 'REPLACE_ME'
|
||||
- vault_atlas_icloudpd_apple_id.splitlines() | length == 1
|
||||
fail_msg: Configure the existing iCloudPD Apple ID in Vault.
|
||||
no_log: true
|
||||
|
||||
- name: Seed private Atlas iCloudPD configuration when absent
|
||||
ansible.builtin.template:
|
||||
src: atlas-icloudpd.conf.j2
|
||||
dest: "{{ atlas_icloudpd_config_dir }}/icloudpd.conf"
|
||||
owner: "{{ atlas_admin_username }}"
|
||||
group: "{{ atlas_admin_group }}"
|
||||
mode: "0600"
|
||||
force: false
|
||||
no_log: true
|
||||
diff: false
|
||||
|
||||
- name: Keep declared iCloudPD options in the image-managed configuration
|
||||
ansible.builtin.lineinfile:
|
||||
path: "{{ atlas_icloudpd_config_dir }}/icloudpd.conf"
|
||||
regexp: "^{{ item.key }}="
|
||||
line: "{{ item.key }}={{ item.value }}"
|
||||
owner: "{{ atlas_admin_username }}"
|
||||
group: "{{ atlas_admin_group }}"
|
||||
mode: "0600"
|
||||
loop:
|
||||
- {key: apple_id, value: "{{ vault_atlas_icloudpd_apple_id }}"}
|
||||
- {key: authentication_type, value: MFA}
|
||||
- {key: user, value: user}
|
||||
- {key: user_id, value: "1000"}
|
||||
- {key: group, value: group}
|
||||
- {key: group_id, value: "1000"}
|
||||
- {key: download_path, value: /home/user/iCloud}
|
||||
- {key: folder_structure, value: "{:%Y/%m/%d}"}
|
||||
- {key: directory_permissions, value: "750"}
|
||||
- {key: file_permissions, value: "640"}
|
||||
- {key: download_interval, value: "86400"}
|
||||
- {key: auto_delete, value: "false"}
|
||||
- {key: delete_after_download, value: "false"}
|
||||
loop_control:
|
||||
label: "{{ item.key }}"
|
||||
no_log: true
|
||||
diff: false
|
||||
|
||||
- name: Render the rootless Atlas iCloudPD Quadlet
|
||||
ansible.builtin.template:
|
||||
src: atlas-icloudpd.container.j2
|
||||
dest: "{{ atlas_icloudpd_quadlet_dir }}/atlas-icloudpd.container"
|
||||
owner: "{{ atlas_admin_username }}"
|
||||
group: "{{ atlas_admin_group }}"
|
||||
mode: "0644"
|
||||
register: atlas_icloudpd_quadlet
|
||||
|
||||
- name: Reload the Atlas admin user manager after iCloudPD Quadlet changes
|
||||
become_user: "{{ atlas_admin_username }}"
|
||||
ansible.builtin.systemd:
|
||||
scope: user
|
||||
daemon_reload: true
|
||||
environment:
|
||||
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
|
||||
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus"
|
||||
when:
|
||||
- atlas_icloudpd_quadlet.changed
|
||||
- not ansible_check_mode
|
||||
|
||||
- name: Keep the rootless Atlas iCloudPD service running
|
||||
become_user: "{{ atlas_admin_username }}"
|
||||
ansible.builtin.systemd:
|
||||
name: atlas-icloudpd.service
|
||||
scope: user
|
||||
state: started
|
||||
environment:
|
||||
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
|
||||
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus"
|
||||
when: not ansible_check_mode
|
||||
@@ -14,6 +14,21 @@
|
||||
- name: Import Atlas storage tasks
|
||||
ansible.builtin.import_tasks: storage.yml
|
||||
|
||||
- name: Import explicit Atlas Gitea owner migration
|
||||
ansible.builtin.import_tasks: gitea_owner_migration.yml
|
||||
|
||||
- name: Import staged Atlas rootless Gitea tasks
|
||||
ansible.builtin.import_tasks: gitea.yml
|
||||
|
||||
- name: Import the declared Atlas Gitea public domain
|
||||
ansible.builtin.import_tasks: gitea_public_domain.yml
|
||||
|
||||
- name: Import Atlas iCloudPD storage and boot-started Quadlet tasks
|
||||
ansible.builtin.import_tasks: icloudpd.yml
|
||||
|
||||
- name: Import explicit Atlas Gitea restore rehearsal tasks
|
||||
ansible.builtin.import_tasks: gitea_restore.yml
|
||||
|
||||
- name: Import Atlas ZFS maintenance tasks
|
||||
ansible.builtin.import_tasks: zfs_maintenance.yml
|
||||
|
||||
@@ -23,6 +38,12 @@
|
||||
- name: Import Atlas offline USB backup tasks
|
||||
ansible.builtin.import_tasks: usb_backup.yml
|
||||
|
||||
- name: Import Atlas Prometheus backup pull identity tasks
|
||||
ansible.builtin.import_tasks: prometheus_pull_identity.yml
|
||||
|
||||
- name: Import Atlas Prometheus backup pull job tasks
|
||||
ansible.builtin.import_tasks: prometheus_pull_job.yml
|
||||
|
||||
- name: Import Atlas health monitoring tasks
|
||||
ansible.builtin.import_tasks: monitoring.yml
|
||||
|
||||
|
||||
@@ -7,8 +7,8 @@
|
||||
- atlas_zfs_pool != 'CHANGEME_ZFS_POOL'
|
||||
- atlas_monitor_calendar | length > 0
|
||||
- atlas_monitor_smart_devices | length > 0
|
||||
- atlas_monitor_timers | length > 0
|
||||
- atlas_monitor_failure_units | length > 0
|
||||
- atlas_monitor_effective_timers | length > 0
|
||||
- atlas_monitor_effective_failure_units | length > 0
|
||||
- atlas_monitor_remote_capacity.user == atlas_borg_repository_user
|
||||
- atlas_monitor_remote_capacity.host == atlas_borg_repository_host
|
||||
- atlas_monitor_remote_capacity.run_as == atlas_borg_username
|
||||
@@ -49,7 +49,7 @@
|
||||
that:
|
||||
- item.name is match('^[a-zA-Z0-9@_.-]+\\.timer$')
|
||||
- item.max_age_hours | int >= 0
|
||||
loop: "{{ atlas_monitor_timers }}"
|
||||
loop: "{{ atlas_monitor_effective_timers }}"
|
||||
loop_control:
|
||||
label: "{{ item.name }}"
|
||||
when: atlas_manage_monitoring | bool
|
||||
@@ -59,7 +59,7 @@
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- item is match('^[a-zA-Z0-9@_.-]+\\.service$')
|
||||
loop: "{{ atlas_monitor_failure_units }}"
|
||||
loop: "{{ atlas_monitor_effective_failure_units }}"
|
||||
when: atlas_manage_monitoring | bool
|
||||
|
||||
- name: Validate Atlas health monitor calendar
|
||||
@@ -144,7 +144,7 @@
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0755"
|
||||
loop: "{{ atlas_monitor_failure_units }}"
|
||||
loop: "{{ atlas_monitor_effective_failure_units }}"
|
||||
when: atlas_manage_monitoring | bool
|
||||
|
||||
- name: Notify 45Drives Alerts when an Atlas job fails
|
||||
@@ -155,7 +155,7 @@
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0644"
|
||||
loop: "{{ atlas_monitor_failure_units }}"
|
||||
loop: "{{ atlas_monitor_effective_failure_units }}"
|
||||
when: atlas_manage_monitoring | bool
|
||||
|
||||
- name: Reload systemd after installing Atlas monitoring
|
||||
|
||||
@@ -0,0 +1,66 @@
|
||||
---
|
||||
- name: Validate Atlas Prometheus pull identity inputs
|
||||
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- atlas_prometheus_pull_ssh_dir.startswith('/etc/')
|
||||
- atlas_prometheus_pull_private_key_path.startswith(atlas_prometheus_pull_ssh_dir ~ '/')
|
||||
- atlas_prometheus_pull_known_hosts_path.startswith(atlas_prometheus_pull_ssh_dir ~ '/')
|
||||
- atlas_prometheus_ssh_host_key.startswith(
|
||||
(hostvars['prometheus'].ansible_host | string) ~ ' ssh-ed25519 '
|
||||
)
|
||||
fail_msg: Pin the verified Prometheus ED25519 SSH host key before enabling the pull.
|
||||
when: atlas_manage_prometheus_backup_pull | bool
|
||||
|
||||
- name: Create private Atlas Prometheus pull SSH directory
|
||||
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
|
||||
ansible.builtin.file:
|
||||
path: "{{ atlas_prometheus_pull_ssh_dir }}"
|
||||
state: directory
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0700"
|
||||
when: atlas_manage_prometheus_backup_pull | bool
|
||||
|
||||
- name: Generate Atlas-only Prometheus pull SSH identity
|
||||
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
|
||||
ansible.builtin.command:
|
||||
argv:
|
||||
- ssh-keygen
|
||||
- -q
|
||||
- -t
|
||||
- ed25519
|
||||
- -N
|
||||
- ""
|
||||
- -C
|
||||
- atlas-prometheus-pull@atlas
|
||||
- -f
|
||||
- "{{ atlas_prometheus_pull_private_key_path }}"
|
||||
creates: "{{ atlas_prometheus_pull_private_key_path }}"
|
||||
when: atlas_manage_prometheus_backup_pull | bool
|
||||
|
||||
- name: Protect Atlas-only Prometheus pull SSH identity
|
||||
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
|
||||
ansible.builtin.file:
|
||||
path: "{{ item.path }}"
|
||||
owner: root
|
||||
group: root
|
||||
mode: "{{ item.mode }}"
|
||||
loop:
|
||||
- { path: "{{ atlas_prometheus_pull_private_key_path }}", mode: "0600" }
|
||||
- { path: "{{ atlas_prometheus_pull_private_key_path }}.pub", mode: "0644" }
|
||||
loop_control:
|
||||
label: "{{ item.path }}"
|
||||
when:
|
||||
- atlas_manage_prometheus_backup_pull | bool
|
||||
- not ansible_check_mode
|
||||
|
||||
- name: Pin Prometheus SSH host key on Atlas
|
||||
tags: [atlas, backup, prometheus_backup, prometheus_backup_key]
|
||||
ansible.builtin.copy:
|
||||
content: "{{ atlas_prometheus_ssh_host_key }}\n"
|
||||
dest: "{{ atlas_prometheus_pull_known_hosts_path }}"
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0600"
|
||||
when: atlas_manage_prometheus_backup_pull | bool
|
||||
90
ansible/roles/profile_atlas/tasks/prometheus_pull_job.yml
Normal file
90
ansible/roles/profile_atlas/tasks/prometheus_pull_job.yml
Normal file
@@ -0,0 +1,90 @@
|
||||
---
|
||||
- name: Validate Atlas Prometheus backup pull inputs
|
||||
tags: [atlas, backup, prometheus_backup]
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- atlas_manage_storage | bool
|
||||
- atlas_prometheus_pull_source_user is match('^[a-z_][a-z0-9_-]*$')
|
||||
- atlas_prometheus_pull_source_port | int > 0
|
||||
- atlas_prometheus_pull_source_port | int < 65536
|
||||
- atlas_prometheus_pull_keep_daily | int > 0
|
||||
- atlas_prometheus_pull_keep_weekly | int > 0
|
||||
- atlas_prometheus_pull_keep_monthly | int > 0
|
||||
- atlas_prometheus_pull_max_age_hours | int > 0
|
||||
- atlas_backup_prometheus_mountpoint.startswith(atlas_mount_root ~ '/')
|
||||
fail_msg: Define the Atlas backup destination, source account, and retention before enabling the pull.
|
||||
when: atlas_manage_prometheus_backup_pull | bool
|
||||
|
||||
- name: Validate Atlas Prometheus backup pull calendar
|
||||
tags: [atlas, backup, prometheus_backup]
|
||||
ansible.builtin.command:
|
||||
argv: [systemd-analyze, calendar, "{{ atlas_prometheus_pull_calendar }}"]
|
||||
changed_when: false
|
||||
check_mode: false
|
||||
when: atlas_manage_prometheus_backup_pull | bool
|
||||
|
||||
- name: Create private Atlas Prometheus backup version directory
|
||||
tags: [atlas, backup, prometheus_backup]
|
||||
ansible.builtin.file:
|
||||
path: "{{ atlas_backup_prometheus_mountpoint }}/snapshots"
|
||||
state: directory
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0700"
|
||||
when: atlas_manage_prometheus_backup_pull | bool
|
||||
|
||||
- name: Install Atlas Prometheus backup pull helper
|
||||
tags: [atlas, backup, prometheus_backup]
|
||||
ansible.builtin.template:
|
||||
src: atlas-prometheus-pull.sh.j2
|
||||
dest: /usr/local/sbin/atlas-prometheus-pull
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0750"
|
||||
when: atlas_manage_prometheus_backup_pull | bool
|
||||
|
||||
- name: Install Atlas Prometheus backup retention helper
|
||||
tags: [atlas, backup, prometheus_backup]
|
||||
ansible.builtin.copy:
|
||||
src: atlas-prometheus-prune.py
|
||||
dest: /usr/local/libexec/atlas-prometheus-prune
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0750"
|
||||
when: atlas_manage_prometheus_backup_pull | bool
|
||||
|
||||
- name: Install Atlas Prometheus backup pull systemd units
|
||||
tags: [atlas, backup, prometheus_backup]
|
||||
ansible.builtin.template:
|
||||
src: "{{ item }}.j2"
|
||||
dest: "/etc/systemd/system/{{ item }}"
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0644"
|
||||
loop:
|
||||
- atlas-prometheus-pull.service
|
||||
- atlas-prometheus-pull.timer
|
||||
loop_control:
|
||||
label: "{{ item }}"
|
||||
register: atlas_prometheus_pull_units
|
||||
when: atlas_manage_prometheus_backup_pull | bool
|
||||
|
||||
- name: Reload systemd after Atlas Prometheus pull unit changes
|
||||
tags: [atlas, backup, prometheus_backup]
|
||||
ansible.builtin.systemd:
|
||||
daemon_reload: true
|
||||
when:
|
||||
- atlas_manage_prometheus_backup_pull | bool
|
||||
- atlas_prometheus_pull_units is changed
|
||||
- not ansible_check_mode
|
||||
|
||||
- name: Enable Atlas Prometheus pull timer only after explicit activation
|
||||
tags: [atlas, backup, prometheus_backup]
|
||||
ansible.builtin.systemd:
|
||||
name: atlas-prometheus-pull.timer
|
||||
enabled: true
|
||||
state: started
|
||||
when:
|
||||
- atlas_manage_prometheus_backup_pull | bool
|
||||
- atlas_prometheus_pull_start_timer | bool
|
||||
- not ansible_check_mode
|
||||
@@ -0,0 +1,30 @@
|
||||
# Managed by Ansible. Staging does not start automatically.
|
||||
[Unit]
|
||||
Description=Atlas rootless Gitea
|
||||
RequiresMountsFor={{ atlas_gitea_mountpoint }}
|
||||
|
||||
[Container]
|
||||
ContainerName=atlas-gitea
|
||||
Image={{ atlas_gitea_image }}
|
||||
UserNS=keep-id:uid={{ atlas_gitea_container_uid }},gid={{ atlas_gitea_container_gid }}
|
||||
{% if atlas_gitea_production_enabled | bool %}
|
||||
PublishPort={{ atlas_gitea_bind_address }}:{{ atlas_gitea_http_port }}:3000
|
||||
PublishPort={{ atlas_gitea_bind_address }}:{{ atlas_gitea_ssh_port }}:2222
|
||||
{% else %}
|
||||
PublishPort={{ atlas_gitea_staging_bind_address }}:{{ atlas_gitea_staging_http_port }}:3000
|
||||
PublishPort={{ atlas_gitea_staging_bind_address }}:{{ atlas_gitea_staging_ssh_port }}:2222
|
||||
{% endif %}
|
||||
Volume={{ atlas_gitea_mountpoint }}/data:/var/lib/gitea:Z
|
||||
Volume={{ atlas_gitea_mountpoint }}/config:/etc/gitea:Z
|
||||
NoNewPrivileges=true
|
||||
DropCapability=all
|
||||
|
||||
[Service]
|
||||
Restart=on-failure
|
||||
RestartSec=10
|
||||
TimeoutStartSec=900
|
||||
{% if atlas_gitea_production_enabled | bool %}
|
||||
|
||||
[Install]
|
||||
WantedBy=default.target
|
||||
{% endif %}
|
||||
@@ -3,8 +3,8 @@
|
||||
"backup_dataset": {{ (atlas_zfs_pool ~ '/' ~ atlas_zfs_dataset_backup) | to_json }},
|
||||
"notifier": {{ atlas_monitor_notifier | to_json }},
|
||||
"smart_devices": {{ atlas_monitor_smart_devices | to_json }},
|
||||
"timers": {{ atlas_monitor_timers | to_json }},
|
||||
"failure_units": {{ atlas_monitor_failure_units | to_json }},
|
||||
"timers": {{ atlas_monitor_effective_timers | to_json }},
|
||||
"failure_units": {{ atlas_monitor_effective_failure_units | to_json }},
|
||||
"remote_capacity": {{ atlas_monitor_remote_capacity | to_json }},
|
||||
"pool_warning_percent": {{ atlas_monitor_pool_warning_percent | int }},
|
||||
"pool_critical_percent": {{ atlas_monitor_pool_critical_percent | int }},
|
||||
|
||||
14
ansible/roles/profile_atlas/templates/atlas-icloudpd.conf.j2
Normal file
14
ansible/roles/profile_atlas/templates/atlas-icloudpd.conf.j2
Normal file
@@ -0,0 +1,14 @@
|
||||
# Managed by Ansible. Password, keyring and MFA cookies are stored separately in /config.
|
||||
apple_id={{ vault_atlas_icloudpd_apple_id }}
|
||||
authentication_type=MFA
|
||||
user=user
|
||||
user_id=1000
|
||||
group=group
|
||||
group_id=1000
|
||||
download_path=/home/user/iCloud
|
||||
folder_structure={:%Y/%m/%d}
|
||||
directory_permissions=750
|
||||
file_permissions=640
|
||||
download_interval=86400
|
||||
auto_delete=false
|
||||
delete_after_download=false
|
||||
@@ -0,0 +1,25 @@
|
||||
# Managed by Ansible. Start automatically with the lingering admin user manager.
|
||||
[Unit]
|
||||
Description=Atlas rootless iCloud Photos Downloader
|
||||
RequiresMountsFor={{ atlas_icloudpd_state_dir }} {{ atlas_icloudpd_photos_dir }}
|
||||
|
||||
[Container]
|
||||
ContainerName=atlas-icloudpd
|
||||
Image={{ atlas_icloudpd_image }}
|
||||
UserNS=keep-id:uid=1000,gid=1000
|
||||
# The image initialises its unprivileged UID 1000 account as container root.
|
||||
User=0
|
||||
# Upstream launcher requires traceroute for its iCloud reachability check.
|
||||
AddCapability=NET_RAW
|
||||
Environment=TZ={{ atlas_icloudpd_timezone }}
|
||||
Volume={{ atlas_icloudpd_photos_dir }}:/home/user/iCloud:z
|
||||
Volume={{ atlas_icloudpd_config_dir }}:/config:Z
|
||||
NoNewPrivileges=true
|
||||
|
||||
[Service]
|
||||
Restart=on-failure
|
||||
RestartSec=300
|
||||
TimeoutStartSec=900
|
||||
|
||||
[Install]
|
||||
WantedBy=default.target
|
||||
@@ -0,0 +1,19 @@
|
||||
[Unit]
|
||||
Description=Pull a prepared read-only Prometheus backup to Atlas
|
||||
RequiresMountsFor={{ atlas_backup_prometheus_mountpoint }}
|
||||
Wants=network-online.target
|
||||
After=network-online.target zfs.target
|
||||
ConditionFileIsExecutable=/usr/local/sbin/atlas-prometheus-pull
|
||||
ConditionPathExists={{ atlas_prometheus_pull_private_key_path }}
|
||||
ConditionPathExists={{ atlas_prometheus_pull_known_hosts_path }}
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
ExecStart=/usr/local/sbin/atlas-prometheus-pull
|
||||
User=root
|
||||
Group=root
|
||||
UMask=0077
|
||||
TimeoutStartSec=infinity
|
||||
Nice=15
|
||||
IOSchedulingClass=best-effort
|
||||
IOSchedulingPriority=7
|
||||
@@ -0,0 +1,75 @@
|
||||
#!/usr/bin/env bash
|
||||
set -Eeuo pipefail
|
||||
umask 077
|
||||
|
||||
backup_root={{ atlas_backup_prometheus_mountpoint | quote }}
|
||||
snapshots="$backup_root/snapshots"
|
||||
stage=''
|
||||
exec 9>/run/lock/atlas-prometheus-pull.lock
|
||||
flock -n 9 || { echo 'A Prometheus pull is already running' >&2; exit 1; }
|
||||
|
||||
cleanup() {
|
||||
local rc=$?
|
||||
trap - EXIT
|
||||
if (( rc != 0 )) && [[ -n "$stage" && -d "$stage" ]]; then
|
||||
rm -rf -- "$stage"
|
||||
fi
|
||||
exit "$rc"
|
||||
}
|
||||
trap cleanup EXIT
|
||||
|
||||
zpool list -H -o name {{ atlas_zfs_pool | quote }} >/dev/null
|
||||
findmnt -rn --mountpoint "$backup_root" >/dev/null
|
||||
stage=$(mktemp -d "$backup_root/.staging.XXXXXXXX")
|
||||
ssh_cmd='/usr/bin/ssh -F /dev/null -o BatchMode=yes -o StrictHostKeyChecking=yes -o UserKnownHostsFile={{ atlas_prometheus_pull_known_hosts_path }} -o IdentitiesOnly=yes -i {{ atlas_prometheus_pull_private_key_path }} -p {{ atlas_prometheus_pull_source_port }}'
|
||||
rsync -a --partial --delay-updates -e "$ssh_cmd" \
|
||||
{{ (atlas_prometheus_pull_source_user ~ '@' ~ hostvars['prometheus'].ansible_host ~ ':current/') | quote }} \
|
||||
"$stage/"
|
||||
|
||||
test -s "$stage/payload.tar"
|
||||
test -s "$stage/payload.sha256"
|
||||
test -s "$stage/metadata.json"
|
||||
(cd "$stage" && sha256sum -c payload.sha256)
|
||||
tar -tf "$stage/payload.tar" >/dev/null
|
||||
stamp=$(python3 - "$stage/metadata.json" <<'PY'
|
||||
import json
|
||||
import datetime as dt
|
||||
import re
|
||||
import sys
|
||||
|
||||
with open(sys.argv[1], encoding="utf-8") as stream:
|
||||
metadata = json.load(stream)
|
||||
stamp = metadata.get("created_utc", "")
|
||||
if metadata.get("schema") != 1 or metadata.get("host") != "prometheus":
|
||||
raise SystemExit("Unexpected Prometheus backup metadata")
|
||||
if not re.fullmatch(r"[0-9]{8}T[0-9]{6}Z", stamp):
|
||||
raise SystemExit("Invalid Prometheus backup timestamp")
|
||||
created = dt.datetime.strptime(stamp, "%Y%m%dT%H%M%SZ").replace(tzinfo=dt.timezone.utc)
|
||||
age = dt.datetime.now(dt.timezone.utc) - created
|
||||
if age.total_seconds() < -300 or age > dt.timedelta(hours={{ atlas_prometheus_pull_max_age_hours }}):
|
||||
raise SystemExit("Prometheus backup is outside the configured freshness window")
|
||||
print(stamp)
|
||||
PY
|
||||
)
|
||||
if [[ -e "$snapshots/$stamp" ]]; then
|
||||
cmp "$stage/payload.sha256" "$snapshots/$stamp/payload.sha256"
|
||||
cmp "$stage/metadata.json" "$snapshots/$stamp/metadata.json"
|
||||
(cd "$snapshots/$stamp" && sha256sum -c payload.sha256)
|
||||
rm -rf -- "${stage:?}"
|
||||
stage=''
|
||||
else
|
||||
chown -R root:root "$stage"
|
||||
chmod 0700 "$stage"
|
||||
chmod 0600 "$stage/payload.tar" "$stage/payload.sha256" "$stage/metadata.json"
|
||||
mv -- "$stage" "$snapshots/$stamp"
|
||||
stage=''
|
||||
fi
|
||||
latest_link=$(readlink "$backup_root/latest" 2>/dev/null || true)
|
||||
latest_stamp=${latest_link##*/}
|
||||
if [[ -z "$latest_stamp" || "$stamp" > "$latest_stamp" ]]; then
|
||||
ln -s "snapshots/$stamp" "$backup_root/.latest.new"
|
||||
mv -Tf -- "$backup_root/.latest.new" "$backup_root/latest"
|
||||
fi
|
||||
python3 /usr/local/libexec/atlas-prometheus-prune "$snapshots" \
|
||||
{{ atlas_prometheus_pull_keep_daily }} {{ atlas_prometheus_pull_keep_weekly }} {{ atlas_prometheus_pull_keep_monthly }}
|
||||
echo "Verified and published Prometheus backup $stamp"
|
||||
@@ -0,0 +1,10 @@
|
||||
[Unit]
|
||||
Description=Schedule Atlas pull of prepared Prometheus backups
|
||||
|
||||
[Timer]
|
||||
OnCalendar={{ atlas_prometheus_pull_calendar }}
|
||||
Persistent=true
|
||||
Unit=atlas-prometheus-pull.service
|
||||
|
||||
[Install]
|
||||
WantedBy=timers.target
|
||||
@@ -30,3 +30,7 @@ backend_phase1_timezone: Europe/Rome
|
||||
backend_phase1_services:
|
||||
- atlas-navidrome.service
|
||||
- atlas-syncthing.service
|
||||
backend_phase1_music_sync_enabled: false
|
||||
backend_phase1_music_source_dir: "{{ backend_phase1_archive_dir }}/Music"
|
||||
backend_phase1_music_sync_calendar: "*-*-* 00:45:00 Europe/Rome"
|
||||
backend_phase1_user_systemd_dir: "{{ backend_phase1_user_home }}/.config/systemd/user"
|
||||
|
||||
@@ -18,21 +18,28 @@
|
||||
- backend_phase1_app_data_root.startswith('/')
|
||||
- backend_phase1_navidrome_data_dir.startswith(backend_phase1_app_data_root + '/')
|
||||
- backend_phase1_syncthing_root.startswith(backend_phase1_app_data_root + '/')
|
||||
- >-
|
||||
not (backend_phase1_music_sync_enabled | bool) or
|
||||
(backend_phase1_music_source_dir.startswith(backend_phase1_archive_dir + '/')
|
||||
and backend_phase1_music_sync_calendar | length > 0)
|
||||
fail_msg: >-
|
||||
Disable the rootful media-stack gate and provide the Atlas LAN bind
|
||||
address, firewall sources, and absolute ZFS-backed paths before
|
||||
enabling phase one. This role does not manage Prometheus or migrate
|
||||
application data.
|
||||
tags: [music_sync]
|
||||
|
||||
- name: Read the rootless service account
|
||||
ansible.builtin.getent:
|
||||
database: passwd
|
||||
key: "{{ backend_phase1_username }}"
|
||||
tags: [music_sync]
|
||||
|
||||
- name: Record rootless service account IDs
|
||||
ansible.builtin.set_fact:
|
||||
backend_phase1_uid: "{{ ansible_facts['getent_passwd'][backend_phase1_username][1] }}"
|
||||
backend_phase1_gid: "{{ ansible_facts['getent_passwd'][backend_phase1_username][2] }}"
|
||||
tags: [music_sync]
|
||||
|
||||
- name: Read system service state before starting rootless Syncthing
|
||||
ansible.builtin.service_facts:
|
||||
@@ -65,6 +72,7 @@
|
||||
loop_control:
|
||||
label: "{{ item.dataset }}"
|
||||
register: backend_phase1_zfs_facts
|
||||
tags: [music_sync]
|
||||
|
||||
- name: Require mounted datasets at the declared paths
|
||||
ansible.builtin.assert:
|
||||
@@ -79,6 +87,23 @@
|
||||
loop: "{{ backend_phase1_zfs_facts.results }}"
|
||||
loop_control:
|
||||
label: "{{ item.item.dataset }}"
|
||||
tags: [music_sync]
|
||||
|
||||
- name: Inspect the music copy source
|
||||
ansible.builtin.stat:
|
||||
path: "{{ backend_phase1_music_source_dir }}"
|
||||
register: backend_phase1_music_source_stat
|
||||
when: backend_phase1_music_sync_enabled | bool
|
||||
tags: [music_sync]
|
||||
|
||||
- name: Require an existing music source directory
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- backend_phase1_music_source_stat.stat.isdir | default(false)
|
||||
fail_msg: >-
|
||||
{{ backend_phase1_music_source_dir }} must exist before enabling the daily music copy.
|
||||
when: backend_phase1_music_sync_enabled | bool
|
||||
tags: [music_sync]
|
||||
|
||||
- name: Enable lingering for the rootless service account
|
||||
ansible.builtin.command:
|
||||
@@ -87,12 +112,14 @@
|
||||
- enable-linger
|
||||
- "{{ backend_phase1_username }}"
|
||||
creates: "/var/lib/systemd/linger/{{ backend_phase1_username }}"
|
||||
tags: [music_sync]
|
||||
|
||||
- name: Start the rootless user systemd manager
|
||||
ansible.builtin.systemd:
|
||||
name: "user@{{ backend_phase1_uid }}.service"
|
||||
state: started
|
||||
when: not ansible_check_mode
|
||||
tags: [music_sync]
|
||||
|
||||
- name: Create rootless Quadlet and application directories
|
||||
ansible.builtin.file:
|
||||
@@ -113,6 +140,43 @@
|
||||
loop_control:
|
||||
label: "{{ item.path }}"
|
||||
|
||||
- name: Install rsync for the daily music copy
|
||||
ansible.builtin.dnf:
|
||||
name: rsync
|
||||
state: present
|
||||
when: backend_phase1_music_sync_enabled | bool
|
||||
tags: [music_sync]
|
||||
|
||||
- name: Create the rootless user systemd directory
|
||||
ansible.builtin.file:
|
||||
path: "{{ backend_phase1_user_systemd_dir }}"
|
||||
state: directory
|
||||
owner: "{{ backend_phase1_username }}"
|
||||
group: "{{ backend_phase1_user_group }}"
|
||||
mode: "0700"
|
||||
when: backend_phase1_music_sync_enabled | bool
|
||||
tags: [music_sync]
|
||||
|
||||
- name: Install the daily music copy service
|
||||
ansible.builtin.template:
|
||||
src: atlas-music-sync.service.j2
|
||||
dest: "{{ backend_phase1_user_systemd_dir }}/atlas-music-sync.service"
|
||||
owner: "{{ backend_phase1_username }}"
|
||||
group: "{{ backend_phase1_user_group }}"
|
||||
mode: "0644"
|
||||
when: backend_phase1_music_sync_enabled | bool
|
||||
tags: [music_sync]
|
||||
|
||||
- name: Install the daily music copy timer
|
||||
ansible.builtin.template:
|
||||
src: atlas-music-sync.timer.j2
|
||||
dest: "{{ backend_phase1_user_systemd_dir }}/atlas-music-sync.timer"
|
||||
owner: "{{ backend_phase1_username }}"
|
||||
group: "{{ backend_phase1_user_group }}"
|
||||
mode: "0644"
|
||||
when: backend_phase1_music_sync_enabled | bool
|
||||
tags: [music_sync]
|
||||
|
||||
- name: Render the rootless Navidrome Quadlet
|
||||
ansible.builtin.template:
|
||||
src: atlas-navidrome.container.j2
|
||||
@@ -140,6 +204,7 @@
|
||||
XDG_RUNTIME_DIR: "/run/user/{{ backend_phase1_uid }}"
|
||||
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ backend_phase1_uid }}/bus"
|
||||
when: not ansible_check_mode
|
||||
tags: [music_sync]
|
||||
|
||||
- name: Permit NPM access to phase-one web interfaces through Aegis
|
||||
ansible.posix.firewalld:
|
||||
@@ -186,3 +251,19 @@
|
||||
when:
|
||||
- backend_phase1_start_services | bool
|
||||
- not ansible_check_mode
|
||||
|
||||
- name: Enable the daily music copy timer
|
||||
become_user: "{{ backend_phase1_username }}"
|
||||
ansible.builtin.systemd:
|
||||
name: atlas-music-sync.timer
|
||||
scope: user
|
||||
state: started
|
||||
enabled: true
|
||||
daemon_reload: true
|
||||
environment:
|
||||
XDG_RUNTIME_DIR: "/run/user/{{ backend_phase1_uid }}"
|
||||
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ backend_phase1_uid }}/bus"
|
||||
when:
|
||||
- backend_phase1_music_sync_enabled | bool
|
||||
- not ansible_check_mode
|
||||
tags: [music_sync]
|
||||
|
||||
@@ -0,0 +1,10 @@
|
||||
# Managed by Ansible. Do not edit manually.
|
||||
[Unit]
|
||||
Description=Copy Atlas Archive music to the Navidrome library
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
ExecStartPre=/usr/bin/mountpoint -q {{ backend_phase1_archive_dir }}
|
||||
ExecStartPre=/usr/bin/mountpoint -q {{ backend_phase1_music_dir }}
|
||||
ExecStartPre=/usr/bin/test -d {{ backend_phase1_music_source_dir }}
|
||||
ExecStart=/usr/bin/rsync -aH --no-perms --no-owner --no-group --delay-updates --stats -- {{ backend_phase1_music_source_dir }}/ {{ backend_phase1_music_dir }}/
|
||||
@@ -0,0 +1,11 @@
|
||||
# Managed by Ansible. Do not edit manually.
|
||||
[Unit]
|
||||
Description=Schedule the daily Atlas Navidrome music copy
|
||||
|
||||
[Timer]
|
||||
OnCalendar={{ backend_phase1_music_sync_calendar }}
|
||||
Persistent=true
|
||||
Unit=atlas-music-sync.service
|
||||
|
||||
[Install]
|
||||
WantedBy=timers.target
|
||||
103
ansible/roles/profile_server/tasks/backup_export_identity.yml
Normal file
103
ansible/roles/profile_server/tasks/backup_export_identity.yml
Normal file
@@ -0,0 +1,103 @@
|
||||
---
|
||||
- name: Validate Prometheus backup export identity inputs
|
||||
tags: [services, backup, prometheus_backup]
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- inventory_hostname == 'prometheus'
|
||||
- server_backup_username is match('^[a-z_][a-z0-9_-]*$')
|
||||
- server_backup_username not in ['root', server_username]
|
||||
- server_backup_export_root.startswith('/var/lib/')
|
||||
- server_backup_public_key_name is match('^[a-z0-9_-]+$')
|
||||
- hostvars['atlas'].atlas_manage_prometheus_backup_pull | default(false) | bool
|
||||
fail_msg: Enable Atlas and Prometheus backup roles together with dedicated identity settings.
|
||||
when: server_backup_export_enabled | bool
|
||||
|
||||
- name: Create dedicated Prometheus backup export group
|
||||
tags: [services, backup, prometheus_backup]
|
||||
ansible.builtin.group:
|
||||
name: "{{ server_backup_username }}"
|
||||
system: true
|
||||
state: present
|
||||
when: server_backup_export_enabled | bool
|
||||
|
||||
- name: Create locked Prometheus backup export account
|
||||
tags: [services, backup, prometheus_backup]
|
||||
ansible.builtin.user:
|
||||
name: "{{ server_backup_username }}"
|
||||
group: "{{ server_backup_username }}"
|
||||
groups: []
|
||||
append: false
|
||||
comment: Read-only prepared backup export for Atlas
|
||||
home: "{{ server_backup_export_root }}"
|
||||
create_home: false
|
||||
shell: /bin/bash
|
||||
password_lock: true
|
||||
system: true
|
||||
state: present
|
||||
when: server_backup_export_enabled | bool
|
||||
|
||||
- name: Require restricted rrsync helper on Prometheus
|
||||
tags: [services, backup, prometheus_backup]
|
||||
ansible.builtin.stat:
|
||||
path: "{{ server_backup_rrsync_path }}"
|
||||
register: server_backup_rrsync_file
|
||||
when: server_backup_export_enabled | bool
|
||||
|
||||
- name: Validate restricted rrsync helper
|
||||
tags: [services, backup, prometheus_backup]
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- server_backup_rrsync_file.stat.exists
|
||||
- server_backup_rrsync_file.stat.isreg
|
||||
- server_backup_rrsync_file.stat.pw_name == 'root'
|
||||
fail_msg: Rocky rsync must provide the root-owned rrsync support script.
|
||||
when: server_backup_export_enabled | bool
|
||||
|
||||
- name: Create prepared backup export root
|
||||
tags: [services, backup, prometheus_backup]
|
||||
ansible.builtin.file:
|
||||
path: "{{ server_backup_export_root }}"
|
||||
state: directory
|
||||
owner: root
|
||||
group: "{{ server_backup_username }}"
|
||||
mode: "0750"
|
||||
when: server_backup_export_enabled | bool
|
||||
|
||||
- name: Create restricted Prometheus backup SSH directories
|
||||
tags: [services, backup, prometheus_backup]
|
||||
ansible.builtin.file:
|
||||
path: "{{ item }}"
|
||||
state: directory
|
||||
owner: root
|
||||
group: "{{ server_backup_username }}"
|
||||
mode: "0750"
|
||||
loop:
|
||||
- "{{ server_backup_export_root }}/.ssh"
|
||||
- "{{ server_backup_export_root }}/.ssh/authorized_keys.d"
|
||||
when: server_backup_export_enabled | bool
|
||||
|
||||
- name: Read Atlas public key for Prometheus backup pull
|
||||
tags: [services, backup, prometheus_backup]
|
||||
ansible.builtin.slurp:
|
||||
src: "{{ hostvars['atlas'].atlas_prometheus_pull_private_key_path | default('/etc/atlas-prometheus-pull/id_ed25519') }}.pub"
|
||||
delegate_to: atlas
|
||||
become: true
|
||||
register: server_backup_atlas_public_key
|
||||
when:
|
||||
- server_backup_export_enabled | bool
|
||||
- not ansible_check_mode
|
||||
|
||||
- name: Authorize only restricted read-only backup access from Atlas
|
||||
tags: [services, backup, prometheus_backup]
|
||||
ansible.builtin.copy:
|
||||
content: >-
|
||||
{{ 'command="/usr/bin/python3 ' ~ server_backup_rrsync_path ~ ' -ro '
|
||||
~ server_backup_export_root ~ '/versions",restrict '
|
||||
~ (server_backup_atlas_public_key.content | b64decode | trim) ~ '\n' }}
|
||||
dest: "{{ server_backup_export_root }}/.ssh/authorized_keys.d/{{ server_backup_public_key_name }}"
|
||||
owner: root
|
||||
group: "{{ server_backup_username }}"
|
||||
mode: "0640"
|
||||
when:
|
||||
- server_backup_export_enabled | bool
|
||||
- not ansible_check_mode
|
||||
91
ansible/roles/profile_server/tasks/backup_export_job.yml
Normal file
91
ansible/roles/profile_server/tasks/backup_export_job.yml
Normal file
@@ -0,0 +1,91 @@
|
||||
---
|
||||
- name: Validate Prometheus backup export job inputs
|
||||
tags: [services, backup, prometheus_backup]
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- server_backup_export_source_keep | int >= 2
|
||||
- server_backup_export_paths | length > 0
|
||||
- server_backup_export_paths | unique | length == server_backup_export_paths | length
|
||||
- >-
|
||||
server_backup_export_paths
|
||||
| select('match', '^[a-zA-Z0-9][a-zA-Z0-9._/-]*$') | list | length
|
||||
== server_backup_export_paths | length
|
||||
- >-
|
||||
server_backup_export_paths
|
||||
| reject('search', '(^|/)\.\.(/|$)') | list | length
|
||||
== server_backup_export_paths | length
|
||||
- >-
|
||||
server_backup_export_excludes
|
||||
| select('match', '^[a-zA-Z0-9][a-zA-Z0-9._/-]*$') | list | length
|
||||
== server_backup_export_excludes | length
|
||||
- >-
|
||||
server_backup_export_excludes
|
||||
| reject('search', '(^|/)\.\.(/|$)') | list | length
|
||||
== server_backup_export_excludes | length
|
||||
fail_msg: Define safe relative paths and at least two prepared export versions.
|
||||
when: server_backup_export_enabled | bool
|
||||
|
||||
- name: Validate Prometheus backup export calendar
|
||||
tags: [services, backup, prometheus_backup]
|
||||
ansible.builtin.command:
|
||||
argv: [systemd-analyze, calendar, "{{ server_backup_export_calendar }}"]
|
||||
changed_when: false
|
||||
check_mode: false
|
||||
when: server_backup_export_enabled | bool
|
||||
|
||||
- name: Ensure prepared Prometheus backup versions directory exists
|
||||
tags: [services, backup, prometheus_backup]
|
||||
ansible.builtin.file:
|
||||
path: "{{ server_backup_export_root }}/versions"
|
||||
state: directory
|
||||
owner: root
|
||||
group: "{{ server_backup_username }}"
|
||||
mode: "0750"
|
||||
when: server_backup_export_enabled | bool
|
||||
|
||||
- name: Install Prometheus backup export helper
|
||||
tags: [services, backup, prometheus_backup, gitea_cutover, npm_quadlet_backup]
|
||||
ansible.builtin.template:
|
||||
src: prometheus-backup-export.sh.j2
|
||||
dest: /usr/local/sbin/prometheus-backup-export
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0750"
|
||||
validate: "bash -n %s"
|
||||
when: server_backup_export_enabled | bool
|
||||
|
||||
- name: Install Prometheus backup export systemd units
|
||||
tags: [services, backup, prometheus_backup]
|
||||
ansible.builtin.template:
|
||||
src: "{{ item }}.j2"
|
||||
dest: "/etc/systemd/system/{{ item }}"
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0644"
|
||||
loop:
|
||||
- prometheus-backup-export.service
|
||||
- prometheus-backup-export.timer
|
||||
loop_control:
|
||||
label: "{{ item }}"
|
||||
register: server_backup_export_units
|
||||
when: server_backup_export_enabled | bool
|
||||
|
||||
- name: Reload systemd after Prometheus backup export unit changes
|
||||
tags: [services, backup, prometheus_backup]
|
||||
ansible.builtin.systemd:
|
||||
daemon_reload: true
|
||||
when:
|
||||
- server_backup_export_enabled | bool
|
||||
- server_backup_export_units is changed
|
||||
- not ansible_check_mode
|
||||
|
||||
- name: Enable Prometheus backup export timer only after explicit activation
|
||||
tags: [services, backup, prometheus_backup]
|
||||
ansible.builtin.systemd:
|
||||
name: prometheus-backup-export.timer
|
||||
enabled: true
|
||||
state: started
|
||||
when:
|
||||
- server_backup_export_enabled | bool
|
||||
- server_backup_export_start_timer | bool
|
||||
- not ansible_check_mode
|
||||
38
ansible/roles/profile_server/tasks/gitea_final_export.yml
Normal file
38
ansible/roles/profile_server/tasks/gitea_final_export.yml
Normal file
@@ -0,0 +1,38 @@
|
||||
---
|
||||
- name: Install the explicit Gitea final-export helper
|
||||
tags: [services, gitea_final_export]
|
||||
ansible.builtin.template:
|
||||
src: prometheus-gitea-final-export.sh.j2
|
||||
dest: /usr/local/sbin/prometheus-gitea-final-export
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0750"
|
||||
when:
|
||||
- server_gitea_cutover_tools_enabled | bool
|
||||
- not server_legacy_stack_retired | bool
|
||||
|
||||
- name: Require the prepared source and explicit final-export approval
|
||||
tags: [services, gitea_final_export]
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- server_gitea_cutover_tools_enabled | bool
|
||||
- server_backup_export_enabled | bool
|
||||
- not server_legacy_stack_retired | bool
|
||||
- not ansible_check_mode
|
||||
fail_msg: >-
|
||||
Install the cutover helper and perform an explicit non-check-mode run
|
||||
only after the Gitea outage gate has been approved.
|
||||
when: server_gitea_final_export | bool
|
||||
|
||||
- name: Stop source Gitea and publish the final consistent export
|
||||
tags: [services, gitea_final_export]
|
||||
ansible.builtin.command:
|
||||
argv:
|
||||
- /usr/local/sbin/prometheus-gitea-final-export
|
||||
register: server_gitea_final_export_result
|
||||
changed_when: server_gitea_final_export_result.rc == 0
|
||||
no_log: true
|
||||
when:
|
||||
- server_gitea_final_export | bool
|
||||
- not server_legacy_stack_retired | bool
|
||||
- not ansible_check_mode
|
||||
59
ansible/roles/profile_server/tasks/gitea_npm_proxy.yml
Normal file
59
ansible/roles/profile_server/tasks/gitea_npm_proxy.yml
Normal file
@@ -0,0 +1,59 @@
|
||||
---
|
||||
- name: Validate the NPM Gitea cutover override
|
||||
tags: [services, gitea_cutover]
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- server_gitea_cutover_tools_enabled | bool
|
||||
- server_gitea_npm_domains | length > 0
|
||||
- server_gitea_npm_domains | select('match', '^[a-zA-Z0-9.-]+$') | list | length == server_gitea_npm_domains | length
|
||||
fail_msg: Declare the exact NPM Gitea hostnames before enabling the Atlas upstream.
|
||||
when: server_gitea_on_atlas | bool
|
||||
|
||||
- name: Ensure the NPM custom configuration directory exists
|
||||
tags: [services, gitea_cutover]
|
||||
ansible.builtin.file:
|
||||
path: /opt/npm/data/nginx/custom
|
||||
state: directory
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0755"
|
||||
when: server_gitea_cutover_tools_enabled | bool
|
||||
|
||||
- name: Render the Gitea-only NPM runtime upstream override
|
||||
tags: [services, gitea_cutover]
|
||||
ansible.builtin.template:
|
||||
src: prometheus-gitea-npm-proxy.conf.j2
|
||||
dest: /opt/npm/data/nginx/custom/server_proxy.conf
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0644"
|
||||
register: server_gitea_npm_override
|
||||
when: server_gitea_on_atlas | bool
|
||||
|
||||
- name: Remove the Gitea NPM override when source routing is selected
|
||||
tags: [services, gitea_cutover]
|
||||
ansible.builtin.file:
|
||||
path: /opt/npm/data/nginx/custom/server_proxy.conf
|
||||
state: absent
|
||||
when:
|
||||
- server_gitea_cutover_tools_enabled | bool
|
||||
- not server_gitea_on_atlas | bool
|
||||
|
||||
- name: Validate NPM configuration after a Gitea upstream change
|
||||
tags: [services, gitea_cutover]
|
||||
ansible.builtin.command:
|
||||
argv: [podman, exec, nginx-proxy-manager, nginx, -t]
|
||||
changed_when: false
|
||||
when:
|
||||
- server_gitea_on_atlas | bool
|
||||
- server_gitea_npm_override is changed
|
||||
- not ansible_check_mode
|
||||
|
||||
- name: Reload NPM after validating the Gitea upstream change
|
||||
tags: [services, gitea_cutover]
|
||||
ansible.builtin.command:
|
||||
argv: [podman, exec, nginx-proxy-manager, nginx, -s, reload]
|
||||
when:
|
||||
- server_gitea_on_atlas | bool
|
||||
- server_gitea_npm_override is changed
|
||||
- not ansible_check_mode
|
||||
61
ansible/roles/profile_server/tasks/gitea_ssh_proxy.yml
Normal file
61
ansible/roles/profile_server/tasks/gitea_ssh_proxy.yml
Normal file
@@ -0,0 +1,61 @@
|
||||
---
|
||||
- name: Validate the Prometheus Gitea SSH cutover inputs
|
||||
tags: [services, gitea_cutover]
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- server_gitea_cutover_tools_enabled | bool
|
||||
- server_gitea_atlas_address is match('^[0-9]{1,3}(\.[0-9]{1,3}){3}$')
|
||||
- server_gitea_ssh_public_port | int > 1024
|
||||
- server_gitea_ssh_public_port | int < 65536
|
||||
- server_gitea_ssh_target_port | int > 1024
|
||||
- server_gitea_ssh_target_port | int < 65536
|
||||
- server_gitea_ssh_public_port | int != 22
|
||||
fail_msg: Keep administrative SSH on 22 and provide the Atlas rootless Gitea SSH endpoint.
|
||||
when: server_gitea_on_atlas | bool
|
||||
|
||||
- name: Install the Gitea SSH socket proxy units without activating them
|
||||
tags: [services, gitea_cutover]
|
||||
ansible.builtin.template:
|
||||
src: "{{ item }}.j2"
|
||||
dest: "/etc/systemd/system/{{ item }}"
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0644"
|
||||
loop:
|
||||
- prometheus-gitea-ssh-proxy.socket
|
||||
- prometheus-gitea-ssh-proxy.service
|
||||
loop_control:
|
||||
label: "{{ item }}"
|
||||
register: server_gitea_ssh_proxy_units
|
||||
when: server_gitea_cutover_tools_enabled | bool
|
||||
|
||||
- name: Reload systemd after Gitea SSH proxy unit changes
|
||||
tags: [services, gitea_cutover]
|
||||
ansible.builtin.systemd:
|
||||
daemon_reload: true
|
||||
when:
|
||||
- server_gitea_cutover_tools_enabled | bool
|
||||
- server_gitea_ssh_proxy_units is changed
|
||||
- not ansible_check_mode
|
||||
|
||||
- name: Manage the public Gitea SSH socket separately from administrative SSH
|
||||
tags: [services, gitea_cutover]
|
||||
ansible.builtin.systemd:
|
||||
name: prometheus-gitea-ssh-proxy.socket
|
||||
state: "{{ 'started' if server_gitea_on_atlas | bool else 'stopped' }}"
|
||||
enabled: "{{ server_gitea_on_atlas | bool }}"
|
||||
when:
|
||||
- server_gitea_cutover_tools_enabled | bool
|
||||
- not ansible_check_mode
|
||||
|
||||
- name: Open only the public Gitea SSH port after cutover
|
||||
tags: [services, gitea_cutover]
|
||||
ansible.posix.firewalld:
|
||||
port: "{{ server_gitea_ssh_public_port }}/tcp"
|
||||
zone: "{{ server_firewalld_zone }}"
|
||||
state: "{{ 'enabled' if server_gitea_on_atlas | bool else 'disabled' }}"
|
||||
permanent: true
|
||||
immediate: true
|
||||
when:
|
||||
- server_gitea_cutover_tools_enabled | bool
|
||||
- server_firewall_backend == 'firewalld'
|
||||
112
ansible/roles/profile_server/tasks/legacy_cleanup.yml
Normal file
112
ansible/roles/profile_server/tasks/legacy_cleanup.yml
Normal file
@@ -0,0 +1,112 @@
|
||||
---
|
||||
- name: Require explicit retirement of the migrated Prometheus source
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- inventory_hostname == 'prometheus'
|
||||
- server_legacy_stack_retired | bool
|
||||
- server_gitea_on_atlas | bool
|
||||
- server_npm_quadlet_cutover | bool
|
||||
- server_backup_export_enabled | bool
|
||||
|
||||
- name: Verify legacy paths have no mounts or container users
|
||||
ansible.builtin.command:
|
||||
argv:
|
||||
- python3
|
||||
- -c
|
||||
- |
|
||||
import json, os, pathlib, subprocess
|
||||
def run(*args):
|
||||
return subprocess.check_output(args, text=True).strip()
|
||||
paths = ['/opt/gitea', '/home/git/.ssh', '/opt/navidrome',
|
||||
'/opt/postgres', '/opt/music', '/opt/containerd', '/opt/docker']
|
||||
mounts = json.loads(run('findmnt', '--json', '--list', '-o', 'TARGET'))['filesystems']
|
||||
for path in paths:
|
||||
assert os.path.realpath(path) == path, 'Symlink in cleanup path: ' + path
|
||||
for mount in mounts:
|
||||
target = mount['target']
|
||||
assert target != path and not target.startswith(path + '/'), 'Mounted cleanup path: ' + path
|
||||
ids = run('podman', 'ps', '-aq').split()
|
||||
containers = json.loads(run('podman', 'inspect', *ids)) if ids else []
|
||||
for container in containers:
|
||||
assert container['Name'].lstrip('/') == 'nginx-proxy-manager', 'Unexpected container; review before cleanup'
|
||||
for mount in container.get('Mounts', []):
|
||||
source = os.path.realpath(mount['Source'])
|
||||
for path in paths:
|
||||
assert source != path and not source.startswith(path + '/'), 'Container uses cleanup path: ' + path
|
||||
for path in ['/opt/music', '/opt/containerd']:
|
||||
if os.path.isdir(path):
|
||||
for entry in pathlib.Path(path).rglob('*'):
|
||||
assert entry.is_dir() and not entry.is_symlink(), 'Unexpected file in empty legacy path: ' + str(entry)
|
||||
if os.path.isdir('/opt/docker'):
|
||||
allowed = {'/opt/docker/server', '/opt/docker/server/docker-compose.yml'}
|
||||
for entry in pathlib.Path('/opt/docker').rglob('*'):
|
||||
assert str(entry) in allowed and not entry.is_symlink(), 'Unexpected legacy Docker content: ' + str(entry)
|
||||
assert run('systemctl', 'is-active', 'prometheus-npm.service') == 'active'
|
||||
assert subprocess.run(['systemctl', 'is-active', '--quiet', 'podman-compose-server.service']).returncode != 0
|
||||
assert subprocess.run(['systemctl', 'is-active', '--quiet', 'prometheus-backup-export.service']).returncode != 0
|
||||
print('Legacy cleanup preflight passed')
|
||||
changed_when: false
|
||||
check_mode: false
|
||||
|
||||
- name: Require the updated backup configuration before deleting fallback files
|
||||
ansible.builtin.command:
|
||||
argv:
|
||||
- python3
|
||||
- -c
|
||||
- |
|
||||
import pathlib, subprocess
|
||||
unit = subprocess.check_output(['systemctl', 'show', 'prometheus-backup-export.service',
|
||||
'-p', 'RequiresMountsFor', '--value'], text=True)
|
||||
assert '/opt/gitea' not in unit, 'Backup unit still depends on legacy Gitea'
|
||||
helper = pathlib.Path('/usr/local/sbin/prometheus-backup-export').read_text()
|
||||
assert 'podman-compose-server' not in helper and 'opt/docker/server' not in helper
|
||||
subprocess.run(['bash', '-n', '/usr/local/sbin/prometheus-backup-export'], check=True)
|
||||
changed_when: false
|
||||
when: not ansible_check_mode
|
||||
|
||||
- name: Delete only the explicitly approved legacy data and fallback files
|
||||
ansible.builtin.file:
|
||||
path: "{{ item }}"
|
||||
state: absent
|
||||
loop:
|
||||
- /opt/gitea
|
||||
- /home/git/.ssh
|
||||
- /opt/navidrome
|
||||
- /opt/postgres
|
||||
- /opt/music
|
||||
- /opt/containerd
|
||||
- /opt/docker
|
||||
- /usr/local/sbin/prometheus-gitea-final-export
|
||||
- /etc/systemd/system/podman-compose-server.service
|
||||
register: server_legacy_deleted
|
||||
diff: false
|
||||
|
||||
- name: Reload systemd after removing the inactive legacy unit
|
||||
ansible.builtin.systemd:
|
||||
daemon_reload: true
|
||||
when:
|
||||
- server_legacy_deleted is changed
|
||||
- not ansible_check_mode
|
||||
|
||||
- name: Inspect the obsolete Git home without following symlinks
|
||||
ansible.builtin.stat:
|
||||
path: /home/git
|
||||
follow: false
|
||||
register: server_legacy_git_home
|
||||
|
||||
- name: Require the obsolete Git account to be absent before removing its empty home
|
||||
ansible.builtin.command:
|
||||
argv: [getent, passwd, git]
|
||||
register: server_legacy_git_account
|
||||
changed_when: false
|
||||
failed_when: server_legacy_git_account.rc != 2
|
||||
check_mode: false
|
||||
when: server_legacy_git_home.stat.exists
|
||||
|
||||
# rmdir refuses any nonempty directory; never recursively delete this parent.
|
||||
- name: Remove only the empty obsolete Git home
|
||||
ansible.builtin.command:
|
||||
argv: [rmdir, /home/git]
|
||||
register: server_legacy_git_home_removed
|
||||
changed_when: server_legacy_git_home_removed.rc == 0
|
||||
when: server_legacy_git_home.stat.exists
|
||||
33
ansible/roles/profile_server/tasks/legacy_image_cleanup.yml
Normal file
33
ansible/roles/profile_server/tasks/legacy_image_cleanup.yml
Normal file
@@ -0,0 +1,33 @@
|
||||
---
|
||||
- name: Require the migrated Prometheus topology for image cleanup
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- inventory_hostname == 'prometheus'
|
||||
- server_gitea_on_atlas | bool
|
||||
- server_npm_quadlet_cutover | bool
|
||||
- server_legacy_images | default([]) | length > 0
|
||||
- >-
|
||||
server_legacy_images | difference([
|
||||
'docker.gitea.com/gitea:1.25.2',
|
||||
'docker.io/deluan/navidrome:latest',
|
||||
'docker.io/library/postgres:13']) | length == 0
|
||||
|
||||
- name: Check whether the explicitly selected legacy images exist
|
||||
ansible.builtin.command:
|
||||
argv: [podman, image, exists, "{{ item }}"]
|
||||
loop: "{{ server_legacy_images }}"
|
||||
register: server_legacy_image_presence
|
||||
changed_when: false
|
||||
failed_when: server_legacy_image_presence.rc not in [0, 1]
|
||||
check_mode: false
|
||||
|
||||
# No --force: Podman must refuse images referenced by any existing container.
|
||||
- name: Remove only unused explicitly selected legacy images
|
||||
ansible.builtin.command:
|
||||
argv: [podman, image, rm, "{{ item.item }}"]
|
||||
loop: "{{ server_legacy_image_presence.results }}"
|
||||
loop_control:
|
||||
label: "{{ item.item }}"
|
||||
when: item.rc == 0
|
||||
register: server_legacy_image_removal
|
||||
changed_when: server_legacy_image_removal.rc == 0
|
||||
@@ -11,6 +11,7 @@
|
||||
- name: Configure DuckDNS updater
|
||||
tags: [dotfiles, dotfiles:server, duckdns]
|
||||
ansible.builtin.import_tasks: duckdns.yml
|
||||
when: server_duckdns_enabled | bool
|
||||
|
||||
- name: Ensure server directories exist
|
||||
tags: [dotfiles, services]
|
||||
@@ -23,6 +24,9 @@
|
||||
loop: "{{ server_directories | default([]) }}"
|
||||
loop_control:
|
||||
label: "{{ item.path }}"
|
||||
when:
|
||||
- item.path != '/opt/gitea/data' or not server_gitea_on_atlas | bool
|
||||
- item.path != server_container_stack_dir or not server_legacy_stack_retired | bool
|
||||
|
||||
- name: Copy server dotfiles
|
||||
tags: [dotfiles, dotfiles:server]
|
||||
@@ -37,7 +41,7 @@
|
||||
label: "{{ item.dest }}"
|
||||
|
||||
- name: Render server templates
|
||||
tags: [dotfiles, dotfiles:server]
|
||||
tags: [dotfiles, dotfiles:server, gitea_cutover]
|
||||
ansible.builtin.template:
|
||||
src: "{{ item.src }}"
|
||||
dest: "{{ item.dest if item.dest.startswith('/') else server_user_home ~ '/' ~ item.dest }}"
|
||||
@@ -48,10 +52,41 @@
|
||||
loop_control:
|
||||
label: "{{ item.dest }}"
|
||||
no_log: "{{ item.no_log | default(false) }}"
|
||||
when: item.src != 'server/docker-compose.yml.j2' or not server_legacy_stack_retired | bool
|
||||
|
||||
- name: Manage Podman Compose stack
|
||||
tags: [services, podman]
|
||||
ansible.builtin.include_tasks: podman-compose.yml
|
||||
when: not server_legacy_stack_retired | bool
|
||||
|
||||
- name: Import staged NPM Quadlet tasks
|
||||
ansible.builtin.import_tasks: npm_quadlet.yml
|
||||
|
||||
- name: Import explicit legacy server image cleanup
|
||||
ansible.builtin.import_tasks: legacy_image_cleanup.yml
|
||||
tags: [never, server_image_cleanup]
|
||||
when: server_legacy_image_cleanup | default(false) | bool
|
||||
|
||||
- name: Import Prometheus backup export identity tasks
|
||||
ansible.builtin.import_tasks: backup_export_identity.yml
|
||||
|
||||
- name: Import Prometheus backup export job tasks
|
||||
ansible.builtin.import_tasks: backup_export_job.yml
|
||||
tags: [server_legacy_cleanup]
|
||||
|
||||
- name: Import explicitly approved legacy server data cleanup
|
||||
ansible.builtin.import_tasks: legacy_cleanup.yml
|
||||
tags: [never, server_legacy_cleanup]
|
||||
when: server_legacy_cleanup | bool
|
||||
|
||||
- name: Import explicit Prometheus Gitea final-export tasks
|
||||
ansible.builtin.import_tasks: gitea_final_export.yml
|
||||
|
||||
- name: Import Prometheus Gitea SSH proxy tasks
|
||||
ansible.builtin.import_tasks: gitea_ssh_proxy.yml
|
||||
|
||||
- name: Import Prometheus Gitea NPM proxy override tasks
|
||||
ansible.builtin.import_tasks: gitea_npm_proxy.yml
|
||||
|
||||
- name: Ensure server SSH authorized key fragments directory exists
|
||||
tags: [services, ssh]
|
||||
@@ -77,13 +112,17 @@
|
||||
when: server_ssh_authorized_keys | length > 0
|
||||
|
||||
- name: Configure server SSH authorized key fragments
|
||||
tags: [services, ssh]
|
||||
tags: [services, ssh, prometheus_backup]
|
||||
ansible.builtin.lineinfile:
|
||||
path: /etc/ssh/sshd_config
|
||||
regexp: '^\s*AuthorizedKeysFile\s+'
|
||||
line: >-
|
||||
AuthorizedKeysFile {{ server_ssh_authorized_keys | map(attribute='name')
|
||||
| map('regex_replace', '^', '%h/.ssh/authorized_keys.d/') | join(' ') }}
|
||||
AuthorizedKeysFile {{
|
||||
((server_ssh_authorized_keys | map(attribute='name')
|
||||
| map('regex_replace', '^', '%h/.ssh/authorized_keys.d/') | list)
|
||||
+ (['%h/.ssh/authorized_keys.d/' ~ server_backup_public_key_name]
|
||||
if server_backup_export_enabled | bool else [])) | join(' ')
|
||||
}}
|
||||
state: present
|
||||
validate: "sshd -t -f %s"
|
||||
notify: Reload SSH service
|
||||
@@ -100,11 +139,13 @@
|
||||
notify: Reload SSH service
|
||||
|
||||
- name: Restrict SSH login to allowed users on server
|
||||
tags: [services]
|
||||
tags: [services, prometheus_backup]
|
||||
ansible.builtin.lineinfile:
|
||||
path: /etc/ssh/sshd_config
|
||||
regexp: '^\s*AllowUsers\s+'
|
||||
line: "AllowUsers {{ server_sshd_allow_users | join(' ') }}"
|
||||
line: >-
|
||||
AllowUsers {{ (server_sshd_allow_users
|
||||
+ ([server_backup_username] if server_backup_export_enabled | bool else [])) | join(' ') }}
|
||||
state: present
|
||||
validate: "sshd -t -f %s"
|
||||
notify: Reload SSH service
|
||||
|
||||
71
ansible/roles/profile_server/tasks/npm_quadlet.yml
Normal file
71
ansible/roles/profile_server/tasks/npm_quadlet.yml
Normal file
@@ -0,0 +1,71 @@
|
||||
---
|
||||
- name: Require staged NPM Quadlet for an active cutover
|
||||
tags: [services, npm_quadlet]
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- not server_npm_quadlet_cutover | bool or server_npm_quadlet_stage | bool
|
||||
fail_msg: The NPM Quadlet cutover requires the staged container and network.
|
||||
|
||||
- name: Validate staged NPM Quadlet inputs
|
||||
tags: [services, npm_quadlet]
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- server_npm_quadlet_image is defined
|
||||
- server_npm_quadlet_image is match('^docker\.io/jc21/nginx-proxy-manager@sha256:[a-f0-9]{64}$')
|
||||
- server_gitea_on_atlas | bool
|
||||
fail_msg: Stage the exact running NPM image only after Gitea has left Compose.
|
||||
when: server_npm_quadlet_stage | bool
|
||||
|
||||
- name: Ensure rootful Quadlet directory exists for NPM
|
||||
tags: [services, npm_quadlet]
|
||||
ansible.builtin.file:
|
||||
path: /etc/containers/systemd
|
||||
state: directory
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0755"
|
||||
when: server_npm_quadlet_stage | bool
|
||||
|
||||
- name: Render staged NPM container and network Quadlets
|
||||
tags: [services, npm_quadlet]
|
||||
ansible.builtin.template:
|
||||
src: "{{ item }}.j2"
|
||||
dest: "/etc/containers/systemd/{{ item }}"
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0644"
|
||||
loop:
|
||||
- prometheus-npm.container
|
||||
- server-web.network
|
||||
loop_control:
|
||||
label: "{{ item }}"
|
||||
register: server_npm_quadlet_units
|
||||
when: server_npm_quadlet_stage | bool
|
||||
|
||||
- name: Reload systemd after staging NPM Quadlets
|
||||
tags: [services, npm_quadlet]
|
||||
ansible.builtin.systemd:
|
||||
daemon_reload: true
|
||||
when:
|
||||
- server_npm_quadlet_stage | bool
|
||||
- server_npm_quadlet_units is changed
|
||||
- not ansible_check_mode
|
||||
|
||||
- name: Verify the staged NPM Quadlet was generated
|
||||
tags: [services, npm_quadlet]
|
||||
ansible.builtin.command:
|
||||
argv: [systemctl, show, prometheus-npm.service, --property=LoadState, --value]
|
||||
register: server_npm_quadlet_load_state
|
||||
changed_when: false
|
||||
when:
|
||||
- server_npm_quadlet_stage | bool
|
||||
- not ansible_check_mode
|
||||
|
||||
- name: Reject an invalid staged NPM Quadlet
|
||||
tags: [services, npm_quadlet]
|
||||
ansible.builtin.assert:
|
||||
that: server_npm_quadlet_load_state.stdout == 'loaded'
|
||||
fail_msg: Quadlet generator did not produce prometheus-npm.service.
|
||||
when:
|
||||
- server_npm_quadlet_stage | bool
|
||||
- not ansible_check_mode
|
||||
@@ -0,0 +1,15 @@
|
||||
[Unit]
|
||||
Description=Prepare a read-only Prometheus application backup for Atlas
|
||||
RequiresMountsFor=/opt/npm {% if not server_gitea_on_atlas | bool %}/opt/gitea {% endif %}{{ server_backup_export_root }}
|
||||
ConditionFileIsExecutable=/usr/local/sbin/prometheus-backup-export
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
ExecStart=/usr/local/sbin/prometheus-backup-export
|
||||
User=root
|
||||
Group=root
|
||||
UMask=0077
|
||||
TimeoutStartSec=infinity
|
||||
Nice=10
|
||||
IOSchedulingClass=best-effort
|
||||
IOSchedulingPriority=7
|
||||
@@ -0,0 +1,118 @@
|
||||
#!/usr/bin/env bash
|
||||
set -Eeuo pipefail
|
||||
umask 077
|
||||
|
||||
export_root={{ server_backup_export_root | quote }}
|
||||
versions="$export_root/versions"
|
||||
{% if server_legacy_stack_retired | bool %}
|
||||
stack_unit=prometheus-npm.service
|
||||
{% else %}
|
||||
stack_unit=''
|
||||
compose_active=false
|
||||
quadlet_active=false
|
||||
systemctl is-active --quiet podman-compose-server.service && compose_active=true
|
||||
systemctl is-active --quiet prometheus-npm.service && quadlet_active=true
|
||||
if [[ "$compose_active" == "$quadlet_active" ]]; then
|
||||
echo 'Expected exactly one active NPM service (Compose or Quadlet)' >&2
|
||||
exit 1
|
||||
fi
|
||||
if "$quadlet_active"; then
|
||||
stack_unit=prometheus-npm.service
|
||||
else
|
||||
stack_unit=podman-compose-server.service
|
||||
fi
|
||||
{% endif %}
|
||||
stamp=$(date -u +%Y%m%dT%H%M%SZ)
|
||||
stage=''
|
||||
stack_stopped=false
|
||||
|
||||
exec 9>/run/lock/prometheus-backup-export.lock
|
||||
flock -n 9 || { echo 'A backup export is already running' >&2; exit 1; }
|
||||
|
||||
cleanup() {
|
||||
local rc=$?
|
||||
trap - EXIT
|
||||
if "$stack_stopped"; then
|
||||
if systemctl is-active --quiet "$stack_unit"; then
|
||||
systemctl restart "$stack_unit" || rc=1
|
||||
else
|
||||
systemctl start "$stack_unit" || rc=1
|
||||
fi
|
||||
fi
|
||||
if (( rc != 0 )) && [[ -n "$stage" && -d "$stage" ]]; then
|
||||
rm -rf -- "$stage"
|
||||
fi
|
||||
exit "$rc"
|
||||
}
|
||||
trap cleanup EXIT
|
||||
trap 'exit 129' HUP
|
||||
trap 'exit 130' INT
|
||||
trap 'exit 143' TERM
|
||||
|
||||
systemctl is-active --quiet "$stack_unit" || {
|
||||
echo "The managed NPM unit $stack_unit must be active before preparing a backup" >&2
|
||||
exit 1
|
||||
}
|
||||
|
||||
paths=(
|
||||
{% for path in server_backup_export_paths %}
|
||||
{{ path | quote }}
|
||||
{% endfor %}
|
||||
)
|
||||
excludes=(
|
||||
{% for path in server_backup_export_excludes %}
|
||||
--exclude={{ path | quote }}
|
||||
{% endfor %}
|
||||
)
|
||||
for path in "${paths[@]}"; do
|
||||
[[ -e "/$path" ]] || { echo "Required backup path missing: /$path" >&2; exit 1; }
|
||||
done
|
||||
[[ ! -e "$versions/$stamp" ]] || { echo "Export version already exists: $stamp" >&2; exit 1; }
|
||||
stage=$(mktemp -d "$export_root/.staging.XXXXXXXX")
|
||||
|
||||
# SQLite databases and their accompanying files are copied while both
|
||||
# managed containers are stopped. The EXIT trap restarts the stack on error.
|
||||
stack_stopped=true
|
||||
systemctl stop "$stack_unit"
|
||||
tar --acls --xattrs --selinux "${excludes[@]}" -C / -cf "$stage/payload.tar" "${paths[@]}"
|
||||
systemctl start "$stack_unit"
|
||||
for container in nginx-proxy-manager{% if not server_gitea_on_atlas | bool %} gitea{% endif %}; do
|
||||
running=false
|
||||
for _ in {1..30}; do
|
||||
if [[ $(podman inspect --format '{{ '{{.State.Running}}' }}' "$container" 2>/dev/null) == true ]]; then
|
||||
running=true
|
||||
break
|
||||
fi
|
||||
sleep 2
|
||||
done
|
||||
"$running" || { echo "Container did not restart: $container" >&2; exit 1; }
|
||||
done
|
||||
ready=false
|
||||
for _ in {1..60}; do
|
||||
if curl -fsS --connect-timeout 2 --max-time 3 -o /dev/null http://127.0.0.1:81/; then
|
||||
ready=true
|
||||
break
|
||||
fi
|
||||
sleep 2
|
||||
done
|
||||
"$ready" || { echo 'NPM administration did not become ready after backup' >&2; exit 1; }
|
||||
stack_stopped=false
|
||||
|
||||
tar -tf "$stage/payload.tar" >/dev/null
|
||||
(cd "$stage" && sha256sum payload.tar >payload.sha256)
|
||||
printf '{"schema":1,"host":"prometheus","created_utc":"%s"}\n' "$stamp" >"$stage/metadata.json"
|
||||
chown root:{{ server_backup_username }} "$stage" "$stage/payload.tar" "$stage/payload.sha256" "$stage/metadata.json"
|
||||
chmod 0750 "$stage"
|
||||
chmod 0640 "$stage/payload.tar" "$stage/payload.sha256" "$stage/metadata.json"
|
||||
mv -- "$stage" "$versions/$stamp"
|
||||
stage=''
|
||||
ln -s "$stamp" "$versions/.current.new"
|
||||
mv -Tf -- "$versions/.current.new" "$versions/current"
|
||||
|
||||
# Keep a small source-side safety window; Atlas owns long-term retention.
|
||||
mapfile -t old_versions < <(find "$versions" -mindepth 1 -maxdepth 1 -type d \
|
||||
-printf '%f\n' | grep -E '^[0-9]{8}T[0-9]{6}Z$' | sort -r | tail -n +{{ server_backup_export_source_keep + 1 }})
|
||||
for old in "${old_versions[@]}"; do
|
||||
rm -rf -- "${versions:?}/$old"
|
||||
done
|
||||
echo "Prepared Prometheus backup export $stamp"
|
||||
@@ -0,0 +1,10 @@
|
||||
[Unit]
|
||||
Description=Prepare daily Prometheus application backup for Atlas
|
||||
|
||||
[Timer]
|
||||
OnCalendar={{ server_backup_export_calendar }}
|
||||
Persistent=false
|
||||
Unit=prometheus-backup-export.service
|
||||
|
||||
[Install]
|
||||
WantedBy=timers.target
|
||||
@@ -0,0 +1,80 @@
|
||||
#!/usr/bin/env bash
|
||||
set -Eeuo pipefail
|
||||
umask 077
|
||||
|
||||
export_root={{ server_backup_export_root | quote }}
|
||||
versions="$export_root/versions"
|
||||
stamp=$(date -u +%Y%m%dT%H%M%SZ)
|
||||
stage=''
|
||||
gitea_stopped=false
|
||||
|
||||
exec 9>/run/lock/prometheus-backup-export.lock
|
||||
flock -n 9 || { echo 'A Prometheus backup export is already running' >&2; exit 1; }
|
||||
|
||||
cleanup() {
|
||||
local rc=$?
|
||||
trap - EXIT
|
||||
if (( rc != 0 )) && "$gitea_stopped"; then
|
||||
podman start gitea >/dev/null || rc=1
|
||||
fi
|
||||
if (( rc != 0 )) && [[ -n "$stage" && -d "$stage" ]]; then
|
||||
rm -rf -- "$stage"
|
||||
fi
|
||||
exit "$rc"
|
||||
}
|
||||
trap cleanup EXIT
|
||||
trap 'exit 129' HUP
|
||||
trap 'exit 130' INT
|
||||
trap 'exit 143' TERM
|
||||
|
||||
systemctl is-active --quiet podman-compose-server.service || {
|
||||
echo 'Prometheus Compose stack is not active' >&2; exit 1;
|
||||
}
|
||||
if systemctl is-active --quiet prometheus-backup-export.timer; then
|
||||
echo 'Stop the scheduled export timer for the cutover first' >&2
|
||||
exit 1
|
||||
fi
|
||||
[[ $(podman inspect --format '{{ '{{.State.Running}}' }}' gitea) == true ]] || {
|
||||
echo 'Source Gitea must be running before the final export' >&2; exit 1;
|
||||
}
|
||||
[[ -d /opt/gitea/data && -d /home/git/.ssh ]] || {
|
||||
echo 'Required source Gitea paths are missing' >&2; exit 1;
|
||||
}
|
||||
[[ ! -e "$versions/$stamp" ]] || {
|
||||
echo 'Final export timestamp already exists' >&2; exit 1;
|
||||
}
|
||||
|
||||
gitea_stopped=true
|
||||
podman stop --time 30 gitea >/dev/null
|
||||
[[ $(podman inspect --format '{{ '{{.State.Running}}' }}' gitea) == false ]] || {
|
||||
echo 'Source Gitea did not stop' >&2; exit 1;
|
||||
}
|
||||
python3 - <<'PY'
|
||||
import sqlite3
|
||||
path = '/opt/gitea/data/gitea/gitea.db'
|
||||
with sqlite3.connect(f'file:{path}?mode=ro', uri=True) as database:
|
||||
if database.execute('PRAGMA quick_check').fetchone()[0] != 'ok':
|
||||
raise SystemExit('Source Gitea SQLite quick_check failed')
|
||||
PY
|
||||
|
||||
stage=$(mktemp -d "$export_root/.staging.XXXXXXXX")
|
||||
tar --acls --xattrs --selinux -C / -cf "$stage/payload.tar" \
|
||||
opt/gitea/data home/git/.ssh
|
||||
tar -tf "$stage/payload.tar" >/dev/null
|
||||
(cd "$stage" && sha256sum payload.tar >payload.sha256)
|
||||
printf '{"schema":1,"host":"prometheus","purpose":"gitea-cutover","created_utc":"%s"}\n' \
|
||||
"$stamp" >"$stage/metadata.json"
|
||||
|
||||
[[ $(podman inspect --format '{{ '{{.State.Running}}' }}' gitea) == false ]] || {
|
||||
echo 'Source Gitea restarted during final export' >&2; exit 1;
|
||||
}
|
||||
chown root:{{ server_backup_username }} "$stage" "$stage/payload.tar" \
|
||||
"$stage/payload.sha256" "$stage/metadata.json"
|
||||
chmod 0750 "$stage"
|
||||
chmod 0640 "$stage/payload.tar" "$stage/payload.sha256" "$stage/metadata.json"
|
||||
mv -- "$stage" "$versions/$stamp"
|
||||
stage=''
|
||||
ln -s "$stamp" "$versions/.current.new"
|
||||
mv -Tf -- "$versions/.current.new" "$versions/current"
|
||||
|
||||
echo "Prepared final Gitea export $stamp; source Gitea remains stopped"
|
||||
@@ -0,0 +1,6 @@
|
||||
# Managed by Ansible. NPM's variable proxy upstream uses Nginx DNS, not /etc/hosts.
|
||||
{% for domain in server_gitea_npm_domains %}
|
||||
if ($host = {{ domain }}) {
|
||||
set $server {{ server_gitea_atlas_address }};
|
||||
}
|
||||
{% endfor %}
|
||||
@@ -0,0 +1,12 @@
|
||||
[Unit]
|
||||
Description=Forward public Gitea SSH to Atlas through Aegis
|
||||
Requires=prometheus-gitea-ssh-proxy.socket
|
||||
After=network-online.target wg-quick@wg0.service
|
||||
|
||||
[Service]
|
||||
ExecStart=/usr/lib/systemd/systemd-socket-proxyd {{ server_gitea_atlas_address }}:{{ server_gitea_ssh_target_port }}
|
||||
DynamicUser=true
|
||||
NoNewPrivileges=true
|
||||
ProtectSystem=strict
|
||||
ProtectHome=true
|
||||
PrivateTmp=true
|
||||
@@ -0,0 +1,9 @@
|
||||
[Unit]
|
||||
Description=Public Gitea SSH socket on Prometheus
|
||||
|
||||
[Socket]
|
||||
ListenStream=0.0.0.0:{{ server_gitea_ssh_public_port }}
|
||||
NoDelay=true
|
||||
|
||||
[Install]
|
||||
WantedBy=sockets.target
|
||||
@@ -0,0 +1,26 @@
|
||||
[Unit]
|
||||
Description=Nginx Proxy Manager on Prometheus
|
||||
RequiresMountsFor=/opt/npm/data /opt/npm/letsencrypt
|
||||
|
||||
[Container]
|
||||
Image={{ server_npm_quadlet_image }}
|
||||
ContainerName=nginx-proxy-manager
|
||||
Network=server-web.network
|
||||
NetworkAlias=nginx-proxy-manager
|
||||
AddHost=host.containers.internal:host-gateway
|
||||
PublishPort=80:80
|
||||
PublishPort=443:443
|
||||
PublishPort=127.0.0.1:81:81
|
||||
Volume=/opt/npm/data:/data
|
||||
Volume=/opt/npm/letsencrypt:/etc/letsencrypt
|
||||
Pull=missing
|
||||
|
||||
[Service]
|
||||
Restart=always
|
||||
TimeoutStartSec=180
|
||||
TimeoutStopSec=120
|
||||
|
||||
{% if server_npm_quadlet_cutover | bool %}
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
{% endif %}
|
||||
@@ -0,0 +1,5 @@
|
||||
[Network]
|
||||
NetworkName=server_web
|
||||
Driver=bridge
|
||||
Subnet=10.89.0.0/24
|
||||
Gateway=10.89.0.1
|
||||
@@ -4,7 +4,7 @@ name: server
|
||||
|
||||
services:
|
||||
nginx-proxy-manager:
|
||||
image: docker.io/jc21/nginx-proxy-manager:latest
|
||||
image: {{ server_npm_quadlet_image if server_npm_quadlet_stage | bool else 'docker.io/jc21/nginx-proxy-manager:latest' }}
|
||||
container_name: nginx-proxy-manager
|
||||
restart: unless-stopped
|
||||
ports:
|
||||
@@ -38,6 +38,7 @@ services:
|
||||
# networks:
|
||||
# - web
|
||||
|
||||
{% if not server_gitea_on_atlas | bool %}
|
||||
gitea:
|
||||
image: docker.gitea.com/gitea:1.25.2
|
||||
container_name: gitea
|
||||
@@ -55,6 +56,7 @@ services:
|
||||
ports:
|
||||
- "3000:3000"
|
||||
- "127.0.0.1:222:22"
|
||||
{% endif %}
|
||||
|
||||
|
||||
networks:
|
||||
|
||||
74
docs/atlas-dr-lab.md
Normal file
74
docs/atlas-dr-lab.md
Normal file
@@ -0,0 +1,74 @@
|
||||
# Isolated Atlas DR lab
|
||||
|
||||
This is a **scaled rehearsal**, not a substitute for a full-data restore. The
|
||||
`atlas-dr-lab` libvirt VM on Ikaros was left **shut off** on 2026-09-30. Its
|
||||
persistent volumes are in the default libvirt pool: the current 30 GiB OS
|
||||
volume `atlas-dr-lab-os-rebuild2.qcow2`, the pre-rebuild OS volume
|
||||
`atlas-dr-lab-os.qcow2`, and four independent 4 GiB
|
||||
`atlas-dr-lab-data{1,2,3,4}.qcow2` volumes. The VM uses libvirt's `default`
|
||||
NAT network (last DHCP address `192.168.122.168`), 2 vCPU, and 4 GiB RAM.
|
||||
The data disks have `virtio-atlasdrdata{1,2,3,4}` serials. No physical disk or
|
||||
production Atlas storage is attached. The VM has no autostart.
|
||||
|
||||
## Rebuild inputs and isolation
|
||||
|
||||
- Use Rocky's **9.8 GenericCloud Base x86_64** image
|
||||
`Rocky-9-GenericCloud-Base-9.8-20260525.0.x86_64.qcow2` from
|
||||
`https://download.rockylinux.org/pub/rocky/9.8/images/x86_64/`.
|
||||
Verify its `.CHECKSUM` file; the observed SHA-256 was
|
||||
`92c206cc6f790c61583247eefe87890f8828420662c17cacf247cec78ab4eec8`.
|
||||
- Use a dedicated lab-only inventory merged **after** the repository
|
||||
inventory, and always `--limit atlas_dr_lab`. The temporary 2026-09-30
|
||||
inventory/playbook and logs are in `/tmp/atlas-dr-lab-image/`; copy a
|
||||
sanitized inventory to durable private storage before `/tmp` is cleared if
|
||||
the lab will be repeated. Never reuse `host_vars/atlas.yml`, production
|
||||
Vault secrets, or production disk by-id paths for the lab.
|
||||
- The lab host belongs to `platform_rocky` and `atlas`. It uses `dradmin`
|
||||
(UID/GID 1000) with the operator's **public** SSH key and a random,
|
||||
unknown password hash, the libvirt DHCP address, pool `zpool`, mount root
|
||||
`/zpool`, the four `virtio-atlasdrdata*` by-id paths, a 1 GiB backup
|
||||
reservation, and `rocky_manage_openzfs_repo: true` with only `zfs` in
|
||||
`host_packages`. The following gates remain false: sharing, firewall,
|
||||
media stack, ZFS timers, Borg, USB, monitoring, and Prometheus pull.
|
||||
`atlas_manage_storage` is true. Set `atlas_create_pool: true` **only for the
|
||||
first disposable pool creation**, then set it false before any later run.
|
||||
- A minimal lab playbook selects `atlas_dr_lab`, `become: true`, and the
|
||||
existing `packages_rocky` and `profile_atlas` roles. Use a separate
|
||||
`ANSIBLE_CONFIG` without the production Vault password script, and keep
|
||||
host-key checking on with a lab-specific known-hosts file. The 2026-09-30
|
||||
runs used `-i ansible/inventory/hosts.yml -i <lab-inventory.yml>` and
|
||||
`--limit atlas_dr_lab` throughout.
|
||||
|
||||
## Rehearsal and narrow checks
|
||||
|
||||
1. Before any pool operation, compare `virsh -c qemu:///system domblklist
|
||||
atlas-dr-lab` with the four intended qcow2 paths, and in the guest compare
|
||||
`/dev/disk/by-id/virtio-atlasdrdata*` with `lsblk`. Do not proceed if a
|
||||
physical disk or production identity appears.
|
||||
2. For a first-time disposable build only, run the lab playbook with
|
||||
`--tags pool` and `atlas_create_pool: true`, then immediately set the gate
|
||||
false. Run the full lab playbook and check `zpool status -P zpool`,
|
||||
`zfs list -r zpool`, SELinux, and failed systemd units.
|
||||
3. Write a non-sensitive canary under the lab `/zpool/archive` and snapshot
|
||||
it. Record the pool GUID and canary SHA-256. Export the lab pool cleanly,
|
||||
shut down the VM, and replace **only the OS volume** with a fresh verified
|
||||
Rocky image. Preserve all four data volumes. Reconfigure cloud-init for a
|
||||
new instance; the seed CD-ROM must use **SATA**. The SCSI seed attachment
|
||||
tried during this rehearsal was not detected by cloud-init and was
|
||||
replaced with a SATA attachment before proceeding.
|
||||
4. On the new OS, apply `packages_rocky` to reinstall OpenZFS. First run
|
||||
`zpool import -d /dev/disk/by-id` **without importing**, compare GUID and
|
||||
vdev membership, then use ordinary `zpool import -d /dev/disk/by-id zpool`.
|
||||
Do not use `-f`, `-F`, `-X`, rollback, or pool creation.
|
||||
5. Reapply `profile_atlas` with the lab gates and `atlas_create_pool: false`.
|
||||
Verify the canary, restored snapshot file in an empty temporary directory,
|
||||
dataset hierarchy, SELinux, and pool health. A second full playbook run
|
||||
should report `changed=0`. Remove temporary restored files and shut down
|
||||
the VM after testing.
|
||||
|
||||
The observed 2026-09-30 pool GUID was `8880368391795119587`; the canary
|
||||
SHA-256 was `949701c7a95fadae1fddc21abe846c4312212dbfeb7477948f3188fc3ec34a78`.
|
||||
The post-rebuild Ansible run succeeded, a repeat run reported `changed=0`,
|
||||
12 datasets and the original snapshot were present, and the pool was healthy.
|
||||
The snapshot-restored file matched content and basic metadata. See
|
||||
[`atlas-recovery.md`](atlas-recovery.md) for the production runbook and limits.
|
||||
241
docs/atlas-gitea-migration.md
Normal file
241
docs/atlas-gitea-migration.md
Normal file
@@ -0,0 +1,241 @@
|
||||
# Gitea migration from Prometheus to Atlas
|
||||
|
||||
The later 2026-10-03 canonical-domain change to `git.fscotto.co` is recorded
|
||||
in `docs/domain-fscotto-co.md`. Public SSH remains on TCP/2222; the old
|
||||
DuckDNS Proxy Host was observed disabled. Earlier domain references below
|
||||
describe migration evidence, not the current canonical URL.
|
||||
|
||||
This records the staged migration and its observed partial cutover. Gitea is
|
||||
temporary on Atlas until Uranus; NPM remains on Prometheus. On 2026-10-03
|
||||
the operator explicitly approved removal of the old Prometheus Gitea data,
|
||||
SSH fragment and final-export helper. NPM now uses a rootful Quadlet with no
|
||||
installed Compose fallback. The source-retention and rollback steps below
|
||||
are historical migration gates, not current recovery instructions.
|
||||
Existing backup archives were preserved; use current Atlas data and verified
|
||||
backups for recovery. Do not recreate or restart stale source Gitea.
|
||||
|
||||
## Observed source before cutover and chosen topology (2026-10-01)
|
||||
|
||||
- Prometheus runs the rootful `docker.gitea.com/gitea:1.25.2` image in its
|
||||
managed Compose stack. `/opt/gitea/data` is about 280 MiB, uses SQLite,
|
||||
and contains 33 repositories. A live read-only SQLite `quick_check` passed.
|
||||
`/home/git/.ssh` is a separate small bind mount; `/opt/gitea/data/ssh`
|
||||
contains the existing SSH host keys. Neither tree may be discarded.
|
||||
- Gitea answers HTTP 200 on Prometheus port 3000. NPM currently forwards
|
||||
`git.fscotto.duckdns.org` and `git.ov-ad3410.infomaniak.ch` to the Compose
|
||||
hostname `gitea:3000`. Public DNS resolves to Prometheus. The container's
|
||||
SSH port is bound only to `127.0.0.1:222`; this is not a public Gitea SSH
|
||||
listener. Prometheus' public port 22 remains administrative SSH.
|
||||
- Atlas has a healthy pool and a verified, private Prometheus backup under
|
||||
`/zpool/backup/hosts/prometheus/latest`. The 2026-10-01 scheduled export
|
||||
and pull succeeded. The intended target is a separate
|
||||
`/zpool/services/data/gitea` dataset, not `Archive` or the backup dataset.
|
||||
- The approved cutover keeps NPM on Prometheus, changes the two HTTP Proxy
|
||||
Hosts' effective upstream to Atlas over the Prometheus--Aegis gateway, and offers public Gitea
|
||||
SSH on port 2222 via the same gateway. Prometheus port 22 is unchanged.
|
||||
HTTPS and SSH must be validated together before declaring cutover.
|
||||
- The initial staging ran as a **rootless user Quadlet** under a dedicated,
|
||||
non-login Atlas account, using the pinned `1.25.2-rootless` image. This was an explicit
|
||||
rootful-to-rootless **data-layout conversion**, not a drop-in image swap:
|
||||
the target mounts `/var/lib/gitea` and `/etc/gitea`, and uses Gitea's
|
||||
built-in SSH server instead of the source image's OpenSSH daemon. Keep the
|
||||
application version unchanged until the conversion has passed an isolated
|
||||
restore test. The host's rootful Quadlet directory must not be used.
|
||||
|
||||
## Phase 1: prepare without traffic changes
|
||||
|
||||
Preparation completed on 2026-10-01: Ansible created
|
||||
`zpool/services/data/gitea`, a dedicated non-login `gitea` account (UID/GID
|
||||
1101), separate subordinate IDs, parent-dataset traverse ACLs, and an inactive
|
||||
user Quadlet under `/var/lib/atlas-gitea/.config/containers/systemd/`. The
|
||||
Quadlet has no `[Install]` section and, until the final cutover, binds only
|
||||
loopback staging ports 3001/2223 if started manually. A second targeted
|
||||
Ansible run changed nothing; the generated service was inactive and neither
|
||||
staging port listened.
|
||||
|
||||
The explicit rehearsal is managed by:
|
||||
|
||||
```bash
|
||||
ansible-playbook ansible/site.yml --limit atlas --tags gitea_restore \
|
||||
-e atlas_gitea_restore_test=true
|
||||
```
|
||||
|
||||
On 2026-10-01 this selected the latest verified Prometheus backup, checked its
|
||||
SHA-256, extracted only `opt/gitea/data`, moved `app.ini` into the rootless
|
||||
config mount, rewrote `/data/` paths, enabled built-in SSH on internal port
|
||||
2222, and retained the three source SSH host-key pairs. SQLite `quick_check`
|
||||
passed, all 33 restored repositories passed `git fsck`, and each source/target
|
||||
public host-key fingerprint matched. A temporary `1.25.2-rootless` container
|
||||
with `--network none` answered HTTP internally and listened on internal
|
||||
SSH/2222. The container was removed; the user Quadlet remains inactive, with
|
||||
no staging listener. The second restore run changed nothing. This copy is
|
||||
deliberately stale once new source writes occur and **must not** be used as the
|
||||
final cutover copy.
|
||||
|
||||
Target backup checks on 2026-10-01: the managed recursive hourly ZFS snapshot
|
||||
`atlas-auto-hourly-20261001T193401Z` contains the new dataset. The managed
|
||||
Borg service completed archive `atlas-20261001T193420Z`, whose contents list
|
||||
includes the staged Gitea database. A separate one-file restore from each
|
||||
source into private `/var/tmp` directories matched the live staged database
|
||||
and passed SQLite `quick_check`. Temporary files and the on-demand snapshot
|
||||
mount were removed; the Borg temporary snapshot was cleaned up and the pool
|
||||
remained healthy. This is file-level proof, **not** a full Gitea recovery.
|
||||
The operator's UUID-bound offline USB run published version
|
||||
`20261001T201220Z-254397` on 2026-10-02. A separate read-only mount and
|
||||
temporary restore of `services/data/gitea/data/gitea/gitea.db` matched
|
||||
contents, owner, group, mode, size, mtime and POSIX ACL; SQLite
|
||||
`quick_check` returned `ok`. The temporary mount and copy were removed,
|
||||
LUKS was closed, and the pool was healthy. This is a file-level restore test,
|
||||
not a complete Gitea recovery rehearsal from USB.
|
||||
|
||||
1. Provision a dedicated target dataset and non-login service identity via
|
||||
Ansible, keeping UID/GID distinct from Atlas' reserved Immich `1100`.
|
||||
Install the user Quadlet in that identity's
|
||||
`~/.config/containers/systemd/`, **without** an `[Install]` section;
|
||||
do not enable, start, or expose it yet.
|
||||
2. Verify the selected Atlas backup SHA-256 and metadata, then extract **only**
|
||||
`opt/gitea/data` to private staging. Keep `home/git/.ssh` in the source
|
||||
backup for rollback; the rootless image does not consume its OpenSSH mount.
|
||||
Never unpack NPM,
|
||||
WireGuard, or other host configuration from this sensitive tarball into a
|
||||
live namespace. Convert the rootful `/data` tree on a disposable copy:
|
||||
place application data under `/var/lib/gitea`, move `app.ini` to
|
||||
`/etc/gitea`, and rewrite every absolute `/data/...` path for the new
|
||||
layout. Enable `START_SSH_SERVER`, use internal SSH port 2222, and retain
|
||||
the source host-key pairs for the built-in server only after verifying
|
||||
their fingerprints and compatibility. Do not rely on the old
|
||||
`/home/git/.ssh` OpenSSH mount in the rootless image. Set only the target
|
||||
copy's ownership and path-scoped SELinux labels.
|
||||
3. Validate SQLite integrity, repository count and representative `git fsck`,
|
||||
LFS/attachment presence, permissions, and an isolated rootless test
|
||||
container with no production ingress or outbound network. Because the
|
||||
source stays active, this is a rehearsal copy, not the final cutover copy.
|
||||
Regenerate Git hooks if the changed installation path requires it.
|
||||
4. ZFS, Borg and UUID-bound offline USB inclusion and one-file restores have
|
||||
passed. These do not replace the final consistent source copy.
|
||||
|
||||
## Phase 2: explicit final cutover
|
||||
|
||||
The opt-in `/usr/local/sbin/prometheus-gitea-final-export` helper was installed
|
||||
on 2026-10-01 and passed `bash -n`. It refuses to
|
||||
run while the scheduled Prometheus export timer is active. When explicitly
|
||||
triggered, it stops only the source Gitea container, checks SQLite, publishes
|
||||
a checksum-verified Gitea-only version for Atlas' existing pull, and leaves
|
||||
the source stopped on success. NPM remains running. A failure before
|
||||
completion restarts source Gitea. Its Ansible gate is
|
||||
`--tags gitea_final_export -e server_gitea_final_export=true`.
|
||||
After Atlas pulls that version, its separate
|
||||
`--tags gitea_final_restore -e atlas_gitea_final_restore=true` gate accepts
|
||||
only metadata marked `gitea-cutover`, validates a private staged replacement,
|
||||
and swaps it for the marked rehearsal. The swap and its rollback path passed
|
||||
synthetic tests on 2026-10-01; the live gate succeeded on 2026-10-02.
|
||||
|
||||
On 2026-10-02 the operator approved the outage. The final stopped-source
|
||||
export `20261002T071525Z` passed the Atlas pull checksum; the guarded restore
|
||||
replaced the rehearsal. SQLite `quick_check`, all 33 repository `git fsck`
|
||||
checks, and the source/target SSH host-key comparison passed. The rootless
|
||||
Atlas Quadlet serves LAN HTTP/3000 and SSH/2222, reachable from Prometheus
|
||||
through Aegis; its firewall admits only Aegis. The final marker gates startup.
|
||||
|
||||
Prometheus now runs the NPM-only Compose stack. Both NPM database records still
|
||||
say `gitea:3000`, but Nginx evaluates this variable upstream through its
|
||||
runtime DNS resolver, which **does not** use a Compose `extra_hosts` alias.
|
||||
The initial alias attempt returned 502. A managed `server_proxy.conf` override
|
||||
sets `$server` to Atlas' IP for only the two declared Gitea domains; it passed
|
||||
`nginx -t` and primary HTTPS/API returned 200 after a clean NPM restart
|
||||
without the alias; a representative public `git ls-remote` also succeeded.
|
||||
Navidrome and Syncthing Proxy Hosts still responded. No NPM SQLite records
|
||||
or credentials were changed. The
|
||||
secondary hostname `git.ov-ad3410.infomaniak.ch` did not resolve from Ikaros
|
||||
and had no generated NPM config file at the time of inspection.
|
||||
|
||||
On 2026-10-03 the operator retired this unused secondary hostname. Its NPM
|
||||
Proxy Host was already soft-deleted; Ansible now declares only
|
||||
`git.fscotto.duckdns.org` and removes the secondary runtime override.
|
||||
|
||||
Prometheus' public TCP/2222 socket proxies to Atlas without changing admin
|
||||
SSH/22. The local socket presents the preserved Gitea ED25519 host key, but
|
||||
an external TCP/2222 connection from Ikaros initially timed out. During that
|
||||
test no SYN reached Prometheus `eth0`; its socket and firewalld port were active.
|
||||
After the VPS firewall was opened later on 2026-10-02, the public port connected,
|
||||
its ED25519 host-key fingerprint matched Atlas, Gitea authenticated the `ikaros`
|
||||
key as `fscotto`, and a public SSH `git ls-remote` for `fscotto/infra.git`
|
||||
returned HEAD. The operator subsequently reported successful authenticated
|
||||
SSH pull and push; the agent did not perform a write test. HTTPS write/login
|
||||
remain untested. Do not
|
||||
restart the stale source after public HTTPS has accepted target writes.
|
||||
|
||||
The Prometheus export timer resumed with NPM-only paths. A recursive ZFS
|
||||
snapshot at `20261002T073032Z` and encrypted Borg archive
|
||||
`atlas-20261002T073044Z` captured the Atlas target after cutover; Borg exited
|
||||
successfully, cleaned its temporary snapshot, and the pool was healthy.
|
||||
|
||||
## Corrected Atlas service owner (2026-10-02)
|
||||
|
||||
The operator required the host Quadlet to belong to `admin`, while the Unix
|
||||
user **inside** the container must be named `gitea`. The pinned derived
|
||||
`Containerfile.gitea-rootless` changes only the base image's UID/GID 1000
|
||||
passwd/group names from `git` to `gitea`; it retains the rootless image's
|
||||
paths and entrypoint. Gitea's `RUN_USER` is `gitea`, while its built-in SSH
|
||||
user and advertised clone user remain `git`, preserving `git@` URLs. The
|
||||
selective restore helper now generates the same three settings for any future
|
||||
explicit restore, instead of recreating a `RUN_USER = git` target.
|
||||
|
||||
A disposable, loopback-only container using a copy of a Gitea ZFS snapshot
|
||||
passed HTTP, SQLite, internal-user and SSH host-key checks without touching
|
||||
live data. After explicit outage approval, the opt-in
|
||||
`--tags gitea_owner_migration -e atlas_gitea_owner_migration=true` run stopped
|
||||
the old user service, took safety snapshot
|
||||
`zpool/services/data/gitea@gitea-owner-migration-20261002T100104`, transferred
|
||||
only the Gitea dataset to `admin`, tested an `admin` staging Quadlet on
|
||||
loopback, then promoted it to the production LAN ports. The old Atlas Quadlet
|
||||
was removed. The old host `gitea` account and its sub-ID range are retained
|
||||
for a deliberate rollback; they must not restart stale Gitea. The parent
|
||||
traverse ACL is removed by the normal Gitea role once the new owner is live.
|
||||
|
||||
The new service returned HTTP 200 locally and through public primary HTTPS;
|
||||
Navidrome and Syncthing remained active under `admin`, the pool was healthy,
|
||||
and a second normal Gitea Ansible run was idempotent. This does **not** close
|
||||
the separate external TCP/2222 or authenticated clone/push validation gap.
|
||||
|
||||
1. Agree on an outage and record source/target versions, pool health, the
|
||||
latest backups, SSH host-key fingerprints, and both current NPM routes.
|
||||
Stop the Prometheus export timer for the change window so it cannot
|
||||
restart the old Compose stack unexpectedly.
|
||||
2. Quiesce source writes with the final-export helper: it stops Gitea before
|
||||
the consistent export and leaves it stopped after success. Pull that export
|
||||
to Atlas and verify checksum and timestamp. Keep
|
||||
`/opt/gitea/data` and `/home/git/.ssh` intact for rollback. Do not allow
|
||||
source Gitea to restart after accepting writes on Atlas.
|
||||
3. Restore the final Gitea-only payload to the target and repeat integrity
|
||||
checks. Verify its advertised SSH port is 2222, its existing HTTPS
|
||||
`ROOT_URL`, repositories, LFS/attachments, and SSH host-key identity. Enable
|
||||
the production Atlas Quadlet only after the final-restore marker exists;
|
||||
its firewall permits only Aegis to reach HTTP and SSH. Validate local HTTP
|
||||
and the target service before switching NPM.
|
||||
4. Enable the public TCP/2222 socket proxy on Prometheus to Atlas over Aegis
|
||||
without changing administrative TCP/22. Switch Prometheus to the desired
|
||||
NPM-only Compose stack and use the managed Gitea-only NPM runtime upstream
|
||||
override. Do not use Compose `extra_hosts`: Nginx bypasses it for the
|
||||
variable upstream. The old Gitea data stays intact. Do not change public DNS.
|
||||
5. Test HTTPS login, representative clone/push, LFS, and public SSH clone/push
|
||||
on port 2222 from outside the Atlas LAN. Record the last source write and
|
||||
first healthy target service times; do not claim RPO/RTO without measuring.
|
||||
6. Resume the Prometheus NPM-only backup export timer after the desired stack
|
||||
is active and verify its next result. Verify the next Atlas snapshot/Borg
|
||||
run covers Gitea and test a restored target copy. Do not delete old source
|
||||
data.
|
||||
|
||||
## Rollback gate
|
||||
|
||||
Before Atlas accepts writes, restore the old Compose definition and remove the
|
||||
NPM override, disable the public 2222 proxy, and restart the unchanged source
|
||||
Gitea if target validation fails. **After Atlas accepts writes, do not blindly restart the source:** its
|
||||
SQLite database and repositories are stale. Quiesce Atlas, capture its new
|
||||
data, and decide a reverse migration or an extended outage explicitly.
|
||||
|
||||
Upstream references: [rootful container layout](https://docs.gitea.com/1.25/installation/install-with-docker/),
|
||||
[rootless image layout and incompatibility](https://docs.gitea.com/installation/install-with-docker-rootless/),
|
||||
[rootless Podman Quadlet](https://docs.gitea.com/installation/install-with-podman-quadlet/),
|
||||
[standard-image conversion](https://docs.gitea.com/1.24/installation/install-with-docker-rootless/),
|
||||
and [restore and hook regeneration](https://docs.gitea.com/1.26/administration/backup-and-restore/).
|
||||
173
docs/atlas-icloudpd-migration.md
Normal file
173
docs/atlas-icloudpd-migration.md
Normal file
@@ -0,0 +1,173 @@
|
||||
# iCloudPD: Aegis to Atlas
|
||||
|
||||
Atlas is the temporary ingestion host until Uranus. Aegis iCloudPD and its
|
||||
state were retired. Ansible declares Atlas storage, the rootless Quadlet,
|
||||
and a private `icloudpd.conf` with the Apple ID from the existing Vault key.
|
||||
The password, keyring and MFA cookies remain application-managed; initialization
|
||||
is interactive.
|
||||
Do not place cookies, keyring files, passwords, or the Apple ID in this document,
|
||||
unencrypted repository content, or a terminal transcript.
|
||||
|
||||
## Historical source and current destination (2026-10-02)
|
||||
|
||||
- Before retirement, Aegis' rootful `icloudpd.service` was active (no reported restarts, running
|
||||
since 2026-07-25), but its declared data bind `/var/lib/icloudpd/data`
|
||||
has **zero top-level entries** and is 4 KiB as observed on 2026-10-02.
|
||||
Its persistent config has two top-level entries. `pi` cannot run passwordless
|
||||
sudo, so the container's internal filesystem and root-only state have **not**
|
||||
been audited. Do not conclude there are no photos to preserve: they could be
|
||||
inside the container overlay because the declared bind targets the wrong
|
||||
home. The current
|
||||
Quadlet mounts that data directory at `/home/root/iCloud`; the image's
|
||||
documented default is `/home/user/iCloud` with its default `user=user`.
|
||||
- The non-secret `folder_structure` value in the persisted Aegis config is a
|
||||
systemd generator path, **not** `{:%Y/%m/%d}`. The Quadlet passes percent
|
||||
characters in `Environment=` without systemd escaping; that is the likely
|
||||
cause. A running unit therefore does not prove that Aegis ingests photos.
|
||||
Do not copy this config or assume that its MFA state is usable on Atlas.
|
||||
- Atlas' `zpool` is healthy. `/zpool/archive/Pictures` already contains about
|
||||
25 GiB of unrelated data; iCloudPD gets only a new managed
|
||||
`/zpool/archive/Pictures/iCloudPD` subtree. Both that subtree and
|
||||
`zpool/services/data/icloudpd` were created on 2026-10-02. Never rsync with `--delete` into
|
||||
Pictures or adopt its existing contents. `/zpool/media/photobook` is reserved
|
||||
for Immich and remains untouched, including its Aegis-only NFS export.
|
||||
|
||||
The upstream image documents `/config/icloudpd.conf` as its primary
|
||||
configuration (environment configuration is deprecated), an exact
|
||||
`/home/${user}/iCloud/.mounted` failsafe, and an interactive `--Initialise`
|
||||
step for keyring and MFA cookies. The configuration must use the same download
|
||||
path, user/UID, and folder format as the bind mounts. References:
|
||||
[image configuration](https://github.com/boredazfcuk/docker-icloudpd/blob/master/CONFIGURATION.md),
|
||||
[Podman user namespaces](https://docs.podman.io/en/latest/markdown/podman-pod.unit.5.html).
|
||||
|
||||
## Declared Atlas target
|
||||
|
||||
| Item | Location or policy |
|
||||
| --- | --- |
|
||||
| Downloaded photos | `/zpool/archive/Pictures/iCloudPD`, a new managed subtree of the SMB `Archive` dataset |
|
||||
| Config, keyring, MFA cookies | `zpool/services/data/icloudpd` at `/zpool/services/data/icloudpd/config`, outside Archive |
|
||||
| Host service owner | `admin` rootless user manager; no rootful Quadlet or published port |
|
||||
| Container identity | Entry process root in its user namespace; downloader UID/GID 1000 maps to host `admin` |
|
||||
| Image | Digest-pinned `docker.io/boredazfcuk/icloudpd`, with no registry auto-update |
|
||||
| SELinux | Private `:Z` config bind; shared `:z` photo bind because Archive is also exposed through SMB and used by Syncthing. The label and SMB behavior require runtime testing. |
|
||||
| Access | The new subtree is `admin:admin` mode 0750. No Photobook ownership, ACL, or export changes. |
|
||||
| Sync policy | Daily interval; explicit directory/file modes 750/640; no iCloud deletion and no deletion of destination-only files |
|
||||
|
||||
The photo subtree receives a managed marker and the image's `.mounted` file.
|
||||
An existing unmarked path is refused rather than taken over. The existing
|
||||
Pictures tree is not chowned or emptied. The Quadlet now has `[Install]` with
|
||||
`WantedBy=default.target`, so the lingering admin user manager starts it at boot.
|
||||
Ansible keeps the service running. Ansible renders a mode-0600
|
||||
`icloudpd.conf` with `no_log` and no diff, but does not pull the image,
|
||||
initialize MFA, or run a cutover task. Boot startup was approved on 2026-10-03
|
||||
after a reboot left the previously manual-started service inactive.
|
||||
|
||||
The previous gated check-mode tests and isolated Quadlet-generator test proved
|
||||
only the proposed layout; they predate the simplified declarative role. They
|
||||
were not a production deployment or an authentication test.
|
||||
|
||||
## Evidence already gathered without production writes
|
||||
|
||||
The digest-pinned image was pulled into **admin's** Atlas Podman store. An
|
||||
isolated `/var/tmp` test ran with no network, a fake Apple ID, private temporary
|
||||
config/photo mounts, `keep-id:uid=1000,gid=1000`, and no new privileges. Both
|
||||
container root and UID 1000 wrote to the mounts; UID
|
||||
1000's files mapped to host `admin`. A short-lived container remained running,
|
||||
retained the intended `/home/user/iCloud` and literal `{:%Y/%m/%d}` config,
|
||||
and saw an admin-owned `.mounted` marker. The container and temporary files
|
||||
were removed. A second isolated test showed that dropping **all** container
|
||||
capabilities prevents its root entrypoint from reading an admin-owned 0600
|
||||
config; with the default rootless user-namespace capabilities it could read
|
||||
and write that file. The Quadlet retains `NoNewPrivileges=true` but does not
|
||||
drop every capability. This proves only the container layout and namespace mapping,
|
||||
**not** Apple authentication, a real download, SMB visibility, scheduled
|
||||
operation, backup coverage, or recovery.
|
||||
|
||||
The earlier disposable Photobook ACL test is superseded by the operator's
|
||||
clarification that Photobook belongs to Immich. It is not evidence for the
|
||||
current Archive destination, and the proposed Photobook ACL change was never
|
||||
deployed.
|
||||
|
||||
Backup path review on 2026-10-02: the managed Borg and USB scripts snapshot
|
||||
the pool recursively and bind every mounted child dataset, so both
|
||||
`archive` and the proposed `services/data/icloudpd` fall within their
|
||||
declared source scope. Borg's runner switches to the dedicated `borg` account
|
||||
with only `CAP_DAC_READ_SEARCH`; a read-only check using those exact `setpriv`
|
||||
capability flags could traverse/read Archive, whereas plain
|
||||
`sudo -u borg` could not. USB copies as root and preserves POSIX ACLs, but not
|
||||
generic xattrs/SELinux labels. **This was scope and permission evidence, not a
|
||||
completed backup or restore of iCloudPD data**, which did not exist at the time.
|
||||
|
||||
## Validation status and remaining checks
|
||||
|
||||
- Aegis retirement is complete: `icloudpd.service` is `not-found`/`inactive`,
|
||||
the rootful Quadlet and `/var/lib/icloudpd` are absent, and AdGuard is active.
|
||||
The temporary retirement tasks are no longer in the Aegis role. The Podman
|
||||
image cache may remain; it is not service data.
|
||||
- Atlas storage and the `admin` Quadlet are deployed. The second Ansible
|
||||
run changed nothing and did not start the service; a later manual start
|
||||
generated the config. `/zpool/media/photobook` was unchanged.
|
||||
- The image generated `/zpool/services/data/icloudpd/config/icloudpd.conf`
|
||||
on first start. Ansible replaced that default file with a private template
|
||||
using the Apple ID already in Vault. The operator initialized password
|
||||
and MFA interactively; never put credentials or codes in the repository,
|
||||
chat, or Ansible extra-vars. Automatic boot startup was separately approved
|
||||
on 2026-10-03; this does not change the interactive MFA procedure.
|
||||
- Initial ingestion completed on 2026-10-03. Still check folder structure,
|
||||
ownership, SELinux and SMB access, no unintended deletions, the next daily
|
||||
cycle, completed Borg and USB versions, and isolated restore of photos and
|
||||
private state. A recursive hourly `zpool/archive` snapshot exists after
|
||||
ingestion, but no iCloudPD-specific backup restore has passed. The first
|
||||
real scrub and measured recovery targets are separate open items.
|
||||
|
||||
On 2026-10-02 Atlas storage and the inactive Quadlet were deployed; a second
|
||||
Ansible run made zero changes. The generated service was inactive, and no
|
||||
`icloudpd.conf` existed. Two interactive-sudo Aegis runs removed its service,
|
||||
Quadlet and `/var/lib/icloudpd`, then cleared the failed-unit record left by a
|
||||
SIGKILL during shutdown. Read-only verification found `LoadState=not-found`,
|
||||
`ActiveState=inactive`, both paths absent, and AdGuard active.
|
||||
|
||||
On 2026-10-02 the operator requested the first manual start. The rootless
|
||||
service stayed active, and the image generated `icloudpd.conf` under the
|
||||
private config dataset. Its mode was tightened from 0644 to 0600. The generated
|
||||
`apple_id` field is empty; no MFA or download is verified. The service has no
|
||||
boot-time install target, so it is not configured for automatic startup.
|
||||
|
||||
The 2026-10-02 Atlas `icloudpd` run rendered the Vault-backed template without
|
||||
printing its contents; the second run made zero changes. File owner is
|
||||
`admin:admin`, mode 0600, and the Apple ID field is nonempty. The rootless
|
||||
service remained active with zero restarts. At that point keyring initialization,
|
||||
cookie creation and a real download were unverified. The template now reads
|
||||
`vault_atlas_icloudpd_apple_id`, which is already present in the encrypted
|
||||
Vault; no password or MFA code was added to the template.
|
||||
|
||||
The attempted interactive initialization then lost its container. Diagnosis
|
||||
found that the image launcher requires `traceroute` to pass its iCloud
|
||||
reachability check. Rootless Podman without `NET_RAW` returned `Operation not
|
||||
permitted` despite working Atlas/container DNS and host HTTPS. An isolated
|
||||
container with only `CAP_NET_RAW` passed the same check. The Quadlet now grants
|
||||
that single capability while keeping `NoNewPrivileges=true`; a manual restart
|
||||
passed `traceroute`, and the app stayed running. Logs then showed only the missing
|
||||
keyring and a wait for `--Initialise` again. The app expanded the generated config
|
||||
on startup, so Ansible now seeds it only when absent and idempotently maintains
|
||||
only its declared options. A second live Ansible run made zero changes. At
|
||||
that point MFA, actual ingestion, and backup/restore were unverified.
|
||||
|
||||
On 2026-10-03, after interactive initialization, the rootless service was
|
||||
active and the previous 24h of logs showed download activity with no
|
||||
authentication failures or errors. At 02:16 the application reported `All
|
||||
photos and videos have been downloaded` and `Download complete for user`.
|
||||
The destination contained 11,658 files totaling 86,020,430,015 bytes; this
|
||||
is a filesystem file count, not a count of distinct iCloud assets. A later
|
||||
read-only check found the service still active. This closes initial
|
||||
authentication and ingestion only: a subsequent daily cycle and end-to-end
|
||||
recovery of the new photos and private state remain untested.
|
||||
|
||||
On 2026-10-03 Atlas rebooted at 10:17 CEST; iCloudPD stayed inactive because
|
||||
its Quadlet had no install target. A manual start restored the running service
|
||||
and the application began listing iCloud files. The operator then approved
|
||||
persistent boot startup. The managed Quadlet now declares
|
||||
`WantedBy=default.target`; the live generator created
|
||||
`default.target.wants/atlas-icloudpd.service`, admin has `Linger=yes`, and the
|
||||
service remained active with zero restarts. No NAS reboot was performed to
|
||||
test this change; actual post-reboot startup remains untested.
|
||||
141
docs/atlas-recovery.md
Normal file
141
docs/atlas-recovery.md
Normal file
@@ -0,0 +1,141 @@
|
||||
# Atlas recovery runbook
|
||||
|
||||
This runbook is for a **replacement Rocky Linux 9 installation**, not a normal
|
||||
playbook run. A scaled whole-OS rebuild with a disposable pool passed in an
|
||||
isolated VM on 2026-09-30, but no production-size whole-host recovery has been
|
||||
tested. The existing production pool must be imported, never created or
|
||||
rewritten. The provisional targets are **RPO 24 hours**
|
||||
and **RTO 72 hours**, for Archive and Atlas services alike. They are planning
|
||||
objectives, not demonstrated recovery times. The manual USB cadence may leave
|
||||
an older copy; a recent Borg archive is needed to meet the RPO after total
|
||||
pool loss.
|
||||
|
||||
## Before an incident
|
||||
|
||||
- Keep an offline copy of the encrypted Ansible Vault, its unlock material,
|
||||
the exported Borg repository key, and the Borg passphrase. Do not store
|
||||
unlock material in this repository or in a recovery command line.
|
||||
On 2026-09-30 the operator confirmed these are available independently of
|
||||
Atlas and the Ansible controller; their usability has not been tested here.
|
||||
- Keep the Atlas installation media and a reproducible checkout of this
|
||||
repository available independently of Atlas. Record the exact Git revision
|
||||
used for a successful deployment.
|
||||
- Record the pool's current disk identities with `zpool status -P zpool` and
|
||||
`lsblk -o NAME,SIZE,MODEL,SERIAL,FSTYPE,UUID`. Compare these with
|
||||
`atlas_zpool_disks` before touching a replacement host. The `host_vars`
|
||||
values are historical identifiers, not evidence that a newly attached disk
|
||||
is the same device.
|
||||
- Verify that the latest hourly/daily snapshots, Borg archive, and offline USB
|
||||
version exist and note their timestamps. A timer being enabled is not proof
|
||||
that a backup completed.
|
||||
|
||||
## Incident gate
|
||||
|
||||
1. Identify whether the fault is the OS disk, one or more pool disks, accidental
|
||||
deletion, or an unavailable host. Preserve failed media when possible.
|
||||
2. Stop writes to affected services and capture the last known good backup
|
||||
timestamps. Do not run `zpool create`, `zpool destroy`, `zfs rollback`,
|
||||
`zpool import -F`, `zpool import -X`, `zpool import -f`, or disk formatting
|
||||
as a diagnostic shortcut.
|
||||
3. Choose one recovery source below. Do not merge several sources into the
|
||||
production namespace without comparing their timestamps and content.
|
||||
|
||||
## Rebuild the OS and import the existing pool
|
||||
|
||||
1. Install Rocky Linux 9 on a **separate system disk**. Configure basic network,
|
||||
SSH, a temporary sudo administrator, SELinux enforcing, and the current
|
||||
OpenZFS kmod repository. Keep the pool drives untouched.
|
||||
2. Run read-only identification: `lsblk -f`, `zpool import`, and
|
||||
`zpool import -d /dev/disk/by-id`. Check the pool GUID, vdev layout, and
|
||||
stable drive identities against the incident record. If any differ, stop.
|
||||
3. Import only after matching the expected pool and host ownership. A pool
|
||||
cleanly exported from the old host can be imported with
|
||||
`zpool import -d /dev/disk/by-id zpool`. If it reports that the pool is
|
||||
active elsewhere or needs a rewind/force, stop and investigate rather than
|
||||
adding flags. Verify with `zpool status -v zpool`, `zfs list -r zpool`,
|
||||
`zfs get -r mountpoint,canmount zpool`, and `findmnt -R /zpool`.
|
||||
4. Leave `atlas_create_pool: false`. Ensure `host_vars/atlas.yml` reflects the
|
||||
replacement host's actual SSH address and disk identities before running
|
||||
Ansible. Apply `ansible/site.yml --limit atlas` with the bootstrap admin
|
||||
connection override as documented in the Atlas setup section of README.
|
||||
This may start shares/services, so keep clients disconnected or services
|
||||
gated until data and permissions are verified.
|
||||
5. Check `getenforce`, `zpool status -v zpool`, `systemctl --failed`, SSH,
|
||||
firewalld, Cockpit, NFS, SMB, and the backup/monitoring timers. Do not
|
||||
report recovery complete on the basis of Ansible success alone.
|
||||
|
||||
## Choose the data source
|
||||
|
||||
- **Local snapshot, pool intact:** inspect `zfs list -t snapshot -r zpool`.
|
||||
Mount/access the chosen snapshot read-only and copy selected files to an
|
||||
empty staging directory; compare content, owner, mode, mtime, and POSIX ACL.
|
||||
Move into the live namespace only after an operator-approved scope review.
|
||||
Do not use an automatic rollback: it can discard newer changes in the
|
||||
dataset and descendants.
|
||||
- **Offline USB:** verify the configured LUKS and ext4 UUIDs from
|
||||
`host_vars/atlas.yml` before unlocking. Mount ext4 read-only with `ro,noload`,
|
||||
use only a published `atlas/latest` version, and restore to an empty staging
|
||||
directory. Compare checksums and metadata. The USB copy intentionally omits
|
||||
generic xattrs and SELinux labels; relabel only the restored destination.
|
||||
Never run the backup service to perform a restore.
|
||||
- **Hetzner Borg:** use the dedicated pinned host key, repository path,
|
||||
offline exported recovery key, and Vault-backed passphrase. List archives
|
||||
and extract a selected archive into an empty staging directory, never the
|
||||
live `/zpool` tree. A repository check and sample restore were previously
|
||||
performed; that does not prove this incident's archive is complete. Compare
|
||||
content and metadata before publication. Avoid `borg break-lock` while any
|
||||
backup/check job may still be active.
|
||||
|
||||
After publishing restored files, run the explicit Ansible `restorecon` tag only
|
||||
for the paths actually restored, for example:
|
||||
|
||||
```bash
|
||||
ansible-playbook ansible/site.yml --limit atlas --tags restorecon \
|
||||
-e '{"atlas_restorecon_paths":["/zpool/archive"]}'
|
||||
```
|
||||
|
||||
Then check ownership/ACLs, application-specific integrity, SMB/NFS client
|
||||
access, backup service health, and `zpool status -v zpool`. Reconnect clients
|
||||
only after these checks pass. Record the last recoverable timestamp (actual
|
||||
RPO) and elapsed service outage (actual RTO) in the incident log.
|
||||
|
||||
## Scaled isolated rehearsal (2026-09-30)
|
||||
|
||||
The lab setup, repeatable checks, and preserved VM state are recorded in
|
||||
[`atlas-dr-lab.md`](atlas-dr-lab.md).
|
||||
|
||||
On Ikaros, a local libvirt `atlas-dr-lab` VM used a 30 GiB Rocky 9.8 system
|
||||
disk and four separate, disposable 4 GiB virtio data disks with stable
|
||||
`/dev/disk/by-id` identities. The official Rocky cloud image matched its
|
||||
published SHA-256. The lab inventory was separate from production, used a
|
||||
fresh lab-only password hash and the operator's public SSH key, and disabled
|
||||
sharing, Borg, USB backup, monitoring, media services, and the Prometheus pull.
|
||||
No production disk, Vault secret, or production data was attached or copied.
|
||||
|
||||
1. The existing `packages_rocky` and `profile_atlas` roles installed OpenZFS,
|
||||
created a RAIDZ2 `zpool` through the explicit one-time pool gate, and built
|
||||
all 12 declared datasets with a lab-sized 1 GiB backup reservation. The
|
||||
pool creation gate was set false immediately afterward.
|
||||
2. A 4 MiB canary file was written under the lab `archive` dataset and a ZFS
|
||||
snapshot created. The pool was cleanly exported and the VM shut down.
|
||||
3. Only the system-disk volume was replaced by a fresh Rocky cloud image;
|
||||
the four virtio data volumes were retained. Ansible reinstalled OpenZFS.
|
||||
Read-only `zpool import -d /dev/disk/by-id` showed the expected RAIDZ2
|
||||
topology and pool GUID `8880368391795119587` before an ordinary import
|
||||
without `-f`, rewind, or rollback.
|
||||
4. The imported pool was healthy. The canary SHA-256 matched its pre-rebuild
|
||||
value. `profile_atlas` completed against the imported pool and a second
|
||||
run reported `changed=0`. A file restored from the preserved snapshot into
|
||||
`/var/tmp` matched SHA-256, owner, group, mode, size, and mtime; the temporary
|
||||
copy was removed. Final checks found SELinux Enforcing, 12 datasets, the
|
||||
snapshot, no failed units, and a healthy pool. The VM was shut down while
|
||||
retaining its disposable volumes for a future rehearsal.
|
||||
|
||||
This proves the **sequence** for a cleanly exported, small pool and the tested
|
||||
Ansible subset, not recovery duration or capacity at 2 TB. The earlier
|
||||
2026-09-25 independent production ZFS/USB file restores and the earlier Borg
|
||||
temporary-directory restore remain separate evidence. The VM did not restore
|
||||
production USB/Borg archives, exercise services with production data, test an
|
||||
unclean import, or prove the provisional RPO/RTO. Before relying on 24h/72h,
|
||||
measure a representative full restore and service cutover in a suitably sized
|
||||
future change window. Never use the production Atlas pool for a rehearsal.
|
||||
20
docs/atlas-sharing-decision.md
Normal file
20
docs/atlas-sharing-decision.md
Normal file
@@ -0,0 +1,20 @@
|
||||
# Atlas SMB/NFS namespace decision
|
||||
|
||||
Decision date: 2026-09-30. Keep the current namespaces **separate**.
|
||||
|
||||
- `/zpool/archive` is the SMB3 `Archive` share for authorized Samba accounts.
|
||||
- `/zpool/media/photobook` is the Aegis-only NFSv4 export, `all_squash`-mapped
|
||||
to UID/GID `1100`.
|
||||
- No new dual-protocol namespace, broad export, group, or ACL model is needed.
|
||||
Existing permissions and client access remain unchanged.
|
||||
|
||||
The two paths serve different ownership and exposure needs. A common namespace
|
||||
would expand the permissions design and require same-file SMB/NFS interoperability
|
||||
testing without a present requirement. Revisit only when a specific workflow
|
||||
needs both protocols on the same files; then decide UID/GID, group, POSIX ACL,
|
||||
SELinux policy and client behavior before changing exports or permissions.
|
||||
|
||||
Read-only Atlas verification on 2026-09-30 confirmed that Samba `Archive` points
|
||||
to `/zpool/archive`, NFS exports `/zpool/media/photobook` only to
|
||||
`192.168.178.54` with `all_squash` and anonymous UID/GID `1100`, both datasets
|
||||
are distinct, and `zpool` is healthy. No sharing configuration was changed.
|
||||
63
docs/atlas-updates.md
Normal file
63
docs/atlas-updates.md
Normal file
@@ -0,0 +1,63 @@
|
||||
# Atlas Rocky/OpenZFS update and reboot procedure (draft)
|
||||
|
||||
This is an operator-controlled maintenance procedure. The playbook does not
|
||||
reboot Atlas, replace a pool device, or perform a pool feature upgrade.
|
||||
|
||||
## Preflight
|
||||
|
||||
1. Schedule an outage and confirm no Borg, USB, snapshot, scrub, or resilver
|
||||
job is active. A service in `activating` is still active; do not interrupt it.
|
||||
2. Check `zpool status -v zpool` (including scrub status), `zfs list -r zpool`,
|
||||
`systemctl --failed`, and `systemctl list-timers --all`. Resolve pool errors
|
||||
first. Record current `uname -r`, `modinfo zfs | grep '^version:'`,
|
||||
`rpm -q kernel-core kmod-zfs zfs`, and the current boot entry.
|
||||
3. Confirm a recent successful Borg archive and a usable snapshot. Confirm
|
||||
the latest published offline USB version and its physical availability;
|
||||
do not start a USB backup merely to satisfy a checklist without capacity,
|
||||
UUID, and operator checks. Record timestamps, not just timer state.
|
||||
4. Ensure console/KVM or another independent recovery route is available.
|
||||
Check free space in `/boot` and the root filesystem. Review proposed DNF
|
||||
transactions before consenting to package changes.
|
||||
|
||||
## Change window
|
||||
|
||||
1. Stop client writes and quiesce stateful applications deliberately. Record
|
||||
which services were stopped; do not assume `ansible-playbook --check` does
|
||||
this. Avoid updating during a running scrub or backup.
|
||||
2. Use `dnf upgrade --assumeno` first to review the kernel, `kmod-zfs`, `zfs`,
|
||||
and dependencies. Confirm a matching kmod will be available for the target
|
||||
kernel. If compatibility is uncertain, defer the update.
|
||||
3. Apply the approved DNF transaction. Do not run `zpool upgrade` or enable
|
||||
new pool feature flags as part of ordinary OS maintenance; that can remove
|
||||
downgrade options. Preserve at least one known-good boot entry.
|
||||
4. Reboot **manually** during the agreed outage. Ansible must not trigger it.
|
||||
|
||||
## Post-boot gate
|
||||
|
||||
1. Verify `uname -r`, `modinfo zfs`, `rpm -q kernel-core kmod-zfs zfs`,
|
||||
`zpool status -v zpool`, `zfs list -r zpool`, and `findmnt -R /zpool`.
|
||||
2. Verify SELinux remains enforcing; inspect `systemctl --failed` and the
|
||||
journal for ZFS, mount, SSH, NFS, SMB, Cockpit, Podman, and backup errors.
|
||||
3. Validate a read-only file listing through SMB and an NFS client access
|
||||
check before reopening writes. Check the rootless temporary services and
|
||||
all backup/monitoring timers. Run the Atlas health monitor in `--dry-run`
|
||||
mode, then a real check after inspection.
|
||||
4. Re-enable clients and record versions, downtime, anomalies, and next
|
||||
successful snapshot/Borg run. A green boot alone is not a completed update.
|
||||
|
||||
## Failure response
|
||||
|
||||
If the new kernel cannot load ZFS, boot the previous known-good kernel from
|
||||
the console and inspect package/kmod matching before trying another reboot.
|
||||
Do not force-import, rewind, clear errors, or upgrade pool features to make a
|
||||
failed OS update appear successful. Preserve logs and stop for a recovery
|
||||
decision if the pool does not import cleanly.
|
||||
|
||||
The procedure-definition item is complete, but the procedure is **not yet
|
||||
rehearsed** on a replacement host or during a real Atlas update. Record the
|
||||
first controlled execution and its post-boot evidence separately.
|
||||
|
||||
Read-only preflight on 2026-09-30 observed kernel
|
||||
`5.14.0-687.52.1.el9_8.x86_64`, ZFS module/package `2.2.11-1`, a healthy
|
||||
`zpool`, enforcing SELinux, and no failed systemd units. This did not review
|
||||
an upgrade transaction, stop services, or reboot the host.
|
||||
82
docs/domain-fscotto-co.md
Normal file
82
docs/domain-fscotto-co.md
Normal file
@@ -0,0 +1,82 @@
|
||||
# fscotto.co domain transition
|
||||
|
||||
## Observed state (2026-10-03)
|
||||
|
||||
Namecheap remains the DNS provider. The operator moved GitHub Pages to
|
||||
`blog.fscotto.co` in `fscotto/fscotto.github.io`, aligned Hugo and Pages
|
||||
settings, and changed the apex A record to `179.237.102.172`. The blog
|
||||
remains a CNAME to `fscotto.github.io`; mail records were left unchanged.
|
||||
A new Hugo deployment and cache clearing resolved the initial stale DNS
|
||||
and generated URLs. Blog HTTPS returned 200 with valid TLS.
|
||||
|
||||
The `git`, `music` and `syncthing` subdomains are CNAMEs to `fscotto.co`.
|
||||
The operator added NPM Proxy Hosts with certificates, WebSocket support
|
||||
and Force SSL:
|
||||
|
||||
| Hostname | HTTP upstream |
|
||||
| --- | --- |
|
||||
| git.fscotto.co | 192.168.178.55:3000 |
|
||||
| music.fscotto.co | 192.168.178.55:4533 |
|
||||
| syncthing.fscotto.co | 192.168.178.55:8384 |
|
||||
|
||||
All three redirected HTTP to HTTPS and returned final HTTPS 200 with valid
|
||||
TLS. Only the Syncthing GUI uses NPM; native synchronization is unchanged.
|
||||
NPM administration remains loopback-only on port 81 via SSH tunnel.
|
||||
|
||||
## Gitea canonical hostname
|
||||
|
||||
Atlas declares `atlas_gitea_public_domain: git.fscotto.co`. Ansible manages
|
||||
only `[server] DOMAIN`, `ROOT_URL` and `SSH_DOMAIN` in the existing private
|
||||
app.ini, preserving unrelated settings and mode 0600. Private configuration
|
||||
backups are created; diffs and secret-bearing results are suppressed.
|
||||
Only Gitea restarts when these fields change; a repeat run changed nothing.
|
||||
|
||||
HTTPS uses `https://git.fscotto.co/`; public SSH remains TCP/2222.
|
||||
Agent read-only checks returned the same HEAD from `fscotto/infra.git`
|
||||
over HTTPS and authenticated SSH. SSH host identity was checked against
|
||||
the already-trusted old endpoint key. No test push or user-authenticated
|
||||
web login was performed by the agent.
|
||||
|
||||
```bash
|
||||
ansible-playbook ansible/site.yml --limit atlas --tags gitea_public_domain --check --diff
|
||||
```
|
||||
|
||||
Client remotes do not update automatically. Update them deliberately after
|
||||
checking repository paths; integrations and webhooks are separate operations.
|
||||
For the verified infrastructure repository only:
|
||||
|
||||
```bash
|
||||
git remote set-url origin ssh://git@git.fscotto.co:2222/fscotto/infra.git
|
||||
```
|
||||
|
||||
Do not copy this path into unrelated clones. Verify Gitea's known SSH key
|
||||
before accepting the new hostname's identity.
|
||||
|
||||
## Local DuckDNS retirement
|
||||
|
||||
Prometheus declares `server_duckdns_enabled: false`. On 2026-10-03 the explicit
|
||||
Ansible cleanup removed the five-minute rocky cron entry and the private
|
||||
`~/duckdns` directory containing only `duck.sh` and `duck.log`. The temporary
|
||||
cleanup tasks and flag were subsequently removed from the playbook at the
|
||||
operator's request. Only the disabled provisioning state remains; ordinary
|
||||
provisioning cannot recreate the updater.
|
||||
The external DuckDNS name, Vault token, disabled NPM hosts and certificates
|
||||
remain untouched for a separate future decision.
|
||||
The repeat cleanup changed nothing; ordinary DuckDNS provisioning was skipped.
|
||||
The cron table had no remaining entries, NPM and the export timer were active,
|
||||
and NPM administration still listened only on `127.0.0.1:81`.
|
||||
|
||||
## Operator-confirmed transition completion
|
||||
|
||||
On 2026-10-03 the operator confirmed completion of:
|
||||
|
||||
- Web login on the new Gitea hostname.
|
||||
- Updates to remaining Git remotes, webhooks and integrations.
|
||||
- Removal of obsolete DuckDNS NPM Proxy Hosts, unused certificates and the old upstream override.
|
||||
- Review and removal of completed one-time procedures from the playbook.
|
||||
|
||||
These are operator confirmations, not new agent runtime checks or a test push.
|
||||
At the earlier inspection the three old DuckDNS Proxy Hosts were disabled,
|
||||
not deleted; that observation predates the confirmed cleanup. Existing backup
|
||||
archives remain preserved. DNS/Pages/NPM changes were operator actions;
|
||||
the Gitea application configuration change was deployed through Ansible.
|
||||
131
docs/prometheus-backup.md
Normal file
131
docs/prometheus-backup.md
Normal file
@@ -0,0 +1,131 @@
|
||||
# Prometheus to Atlas backup pull
|
||||
|
||||
The playbook and both hosts have the dedicated identity, restricted SSH
|
||||
access, helpers, and systemd units. A manual export, pull, and temporary
|
||||
restore passed on 2026-09-30. The first scheduled export and pull passed on
|
||||
2026-10-01. After NPM moved to its Quadlet, another manual export, pull, and
|
||||
isolated restore passed on 2026-10-03. The first scheduled cycle after that
|
||||
cutover is still pending. See `docs/prometheus-npm-quadlet.md`.
|
||||
|
||||
## Declared design
|
||||
|
||||
- Prometheus prepares a tar archive of Nginx Proxy Manager data and certificates,
|
||||
its active Quadlet and network definitions,
|
||||
and SSH/firewalld/WireGuard configuration. Gitea now runs on Atlas and is no
|
||||
longer included in new Prometheus exports. NPM access logs are excluded.
|
||||
The archive contains credentials, certificates, and the WireGuard private
|
||||
key: protect both copies accordingly.
|
||||
- The approved consistency mode stops the NPM Quadlet for local tar creation
|
||||
at 02:00 Europe/Rome, then restarts it even if archiving fails. After the
|
||||
approved legacy cleanup, the helper requires the Quadlet active and has
|
||||
no Compose dependency. A manual test outside that window requires separate approval.
|
||||
- Prometheus publishes the archive with its checksum as a versioned, read-only
|
||||
source under `/var/lib/prometheus-backup-export`. A locked service account
|
||||
has no sudo or supplementary groups. Its only authorized SSH key is forced
|
||||
through Rocky's `rrsync -ro`; root owns the key file and export directories,
|
||||
so the account cannot add an unrestricted key or change prepared data.
|
||||
- Atlas generates and retains the private Ed25519 identity under
|
||||
`/etc/atlas-prometheus-pull`. Its pinned Prometheus host key came through
|
||||
the controller's already strict SSH trust; the observed fingerprint was
|
||||
`SHA256:rfedk7DHI9mLB3UHk/4F3HHlSIiswtCAFsAXvfh6iXk` on 2026-09-30.
|
||||
Atlas pulls only the prepared `current/` version, verifies SHA-256, tar
|
||||
readability, metadata, and source freshness, then publishes atomically
|
||||
below `/zpool/backup/hosts/prometheus/snapshots`. Long-term retention runs
|
||||
only after publication. A local `rrsync` fixture verified the in-tree
|
||||
`current` symlink. A live Atlas-to-Prometheus SSH test verified that the
|
||||
account could list only the prepared versions directory,
|
||||
cannot obtain a shell, and cannot write to the export. The key is restricted
|
||||
to `/var/lib/prometheus-backup-export/versions`, not the account's `.ssh`.
|
||||
- Approved source preparation is 02:00 Europe/Rome, pull 03:00, three source
|
||||
versions, and 30 daily/8 weekly/12 monthly Atlas versions. The source
|
||||
timer is non-persistent to avoid an unexpected outage after a missed run.
|
||||
Atlas rejects a prepared source older than 24 hours.
|
||||
- The Atlas pull joins the existing health monitor's timer/failure checks
|
||||
only when enabled. Its failure hook uses 45Drives Alerts; email delivery
|
||||
is not claimed. A failed source preparation should produce a stale-source
|
||||
pull failure, not a silently successful reuse of an old archive.
|
||||
|
||||
## Activation and verification
|
||||
|
||||
1. The user confirmed downtime/consistency mode, schedule, retention, and
|
||||
targeted configuration scope. Review the tar path list and exclusions
|
||||
against the actual containers.
|
||||
2. The identity and units are deployed. Re-run the targeted
|
||||
check, confirm the Atlas public key remains only the restricted Prometheus
|
||||
account's key, and verify `sshd -T -C user=prometheus-backup,...` plus
|
||||
read-only SSH denial tests after any SSH configuration change.
|
||||
3. During an agreed window, start the Prometheus export service manually.
|
||||
Confirm the active NPM service is healthy afterward, inspect the archive
|
||||
without exposing file contents, and verify the checksum/metadata.
|
||||
4. Start the Atlas pull service manually. Confirm the SSH host pin, source
|
||||
freshness, checksum, tar listing, published `latest`, retention behavior,
|
||||
clean temporary directories, and healthy pool.
|
||||
5. Independently restore the selected archive to an empty staging directory
|
||||
(never `/`) and compare NPM SQLite, data, active Quadlet files, certificates,
|
||||
permissions, and representative files. Historical pre-Gitea-cutover
|
||||
versions also include Gitea repositories; current versions do not. Test
|
||||
application startup only in an isolated environment or an approved restore
|
||||
window.
|
||||
6. Both timers are enabled. Verify their calendars and the next actual run
|
||||
after any service-ownership change. A successful manual test is not proof
|
||||
of a later scheduled cycle.
|
||||
|
||||
Narrow static validation:
|
||||
|
||||
```bash
|
||||
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
|
||||
ansible-playbook ansible/site.yml --syntax-check
|
||||
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
|
||||
ansible-playbook ansible/site.yml --limit prometheus,atlas \
|
||||
--tags prometheus_backup --check --diff
|
||||
```
|
||||
|
||||
Do not run the export service as part of a routine playbook deployment. The
|
||||
service restart and any restore/cutover require separate operator decisions.
|
||||
|
||||
On 2026-09-30 the initial targeted `--check --diff` run ended `changed=0`
|
||||
with gates false. After enabling **implementation only**, a targeted real run
|
||||
installed the identities and units; both timers were confirmed `disabled` and
|
||||
`inactive`, the Compose stack stayed active, and the new account was locked
|
||||
with no supplementary groups. No application was stopped.
|
||||
The rendered shell helpers passed `bash -n` and ShellCheck; the retention
|
||||
helper passed an isolated 400-version fixture. These static/isolated checks
|
||||
were followed by live SSH, export, pull, and temporary restore checks.
|
||||
Read-only preflight on 2026-09-30 found the Compose service active, all
|
||||
declared source paths present, both timers inactive, and no prepared versions.
|
||||
The source filesystem had about 6.0 GB free. Of the 2.1 GB NPM data tree,
|
||||
2.1 GB was excluded access logs, so the expected archive is much smaller than
|
||||
the raw tree size; capacity still needs verification after actual exports.
|
||||
The manual export produced a 285,777,920-byte tar (273 MiB allocated at the
|
||||
source), and Prometheus retained about 5.8 GB free. NPM and Gitea restarted;
|
||||
both containers were running and their local HTTP endpoints returned 200.
|
||||
Atlas pulled the same version, verified SHA-256, published `latest`, and kept
|
||||
the pool healthy. A full extract to `/var/tmp` yielded 4,747 files; both
|
||||
SQLite databases passed `PRAGMA integrity_check`, and one restored Gitea Git
|
||||
repository passed `git fsck`. The temporary restore directory was removed.
|
||||
This did not test application startup on an isolated host.
|
||||
|
||||
After these checks, Ansible enabled the Prometheus 02:00 Europe/Rome export
|
||||
timer and Atlas 03:00 Europe/Rome pull timer. Their first scheduled run passed
|
||||
on 2026-10-01; Atlas verified and published `20261001T000001Z` as `latest`.
|
||||
Atlas' health monitor includes the pull timer.
|
||||
|
||||
On 2026-10-03 the stopped-source version `20261003T091009Z` was verified and
|
||||
pulled before the NPM cutover. The post-cutover version `20261003T091633Z`
|
||||
was exported by the Quadlet-aware helper, checksum-verified, pulled to Atlas,
|
||||
and restored to an isolated temporary directory. NPM SQLite `quick_check`
|
||||
passed with ten proxy hosts and six certificate records. The archive contains
|
||||
both Quadlet definitions. A manifest of all 70 regular Let's Encrypt files
|
||||
and 12 symlinks, including content hashes and link targets, matched the live
|
||||
Prometheus tree. No private key or secret content was printed. The next
|
||||
scheduled export/pull is still pending observation.
|
||||
|
||||
## Post-cleanup validation (2026-10-03)
|
||||
|
||||
The operator-approved removal of legacy data and Compose fallback also
|
||||
removed those backup input paths and the obsolete Gitea mount dependency.
|
||||
A separately approved export and Atlas pull published `20261003T112906Z`.
|
||||
Both SHA-256 checks passed; an isolated SQLite restore passed `quick_check`
|
||||
and contained ten proxy hosts. Both active Quadlet definitions were present;
|
||||
retired paths were absent. Existing backup archives were not deleted by cleanup.
|
||||
The first scheduled cycle after these changes remains unverified.
|
||||
145
docs/prometheus-npm-quadlet.md
Normal file
145
docs/prometheus-npm-quadlet.md
Normal file
@@ -0,0 +1,145 @@
|
||||
# Prometheus NPM Quadlet cutover
|
||||
|
||||
## Current state (2026-10-03)
|
||||
|
||||
Nginx Proxy Manager runs as the **rootful** generated
|
||||
`prometheus-npm.service` on Prometheus. The Quadlet files are
|
||||
`/etc/containers/systemd/prometheus-npm.container` and
|
||||
`/etc/containers/systemd/server-web.network`; the image is pinned by digest
|
||||
in `ansible/inventory/host_vars/prometheus.yml`. The generated service is
|
||||
wanted by `multi-user.target` and requires the generated network service.
|
||||
The old Compose unit, Compose file and Gitea final-export helper were
|
||||
removed by the operator-approved cleanup on 2026-10-03. The retired
|
||||
application data and empty legacy directories were also removed.
|
||||
Prometheus host vars set `server_legacy_stack_retired: true` so normal runs
|
||||
do not recreate those files. Destructive deletion still requires a separate
|
||||
cleanup tag and explicit extra-var.
|
||||
|
||||
There was **no data copy** in this cutover. The Quadlet reuses the existing
|
||||
`/opt/npm/data:/data` and `/opt/npm/letsencrypt:/etc/letsencrypt` bind mounts
|
||||
with the same container name and `server_web` bridge (`10.89.0.0/24`). Ports
|
||||
80 and 443 remain public; administration port 81 remains bound to
|
||||
`127.0.0.1`. Gitea stays on Atlas, and NPM remains on Prometheus. The
|
||||
Compose fallback is no longer installed. The Quadlet uses `Pull=missing`,
|
||||
not an automatic floating-tag update.
|
||||
|
||||
## Cutover and recovery boundaries
|
||||
|
||||
The separate `scripts/cutover_prometheus_npm_quadlet.sh` was run **once** in
|
||||
the approved outage window, after source backup version
|
||||
`20261003T091009Z` was checksum-verified and pulled to Atlas. Its preflight
|
||||
required exactly the Compose owner, an inactive generated Quadlet, the
|
||||
expected image, and the current backup version. The execution held the
|
||||
backup-export lock, stopped the export timer, stopped and disabled Compose,
|
||||
started the Quadlet, checked the exact image ID, SQLite database counts,
|
||||
certificate content, Nginx configuration, and local Gitea/Syncthing HTTPS,
|
||||
then restarted the timer. Its failure trap would have restarted Compose.
|
||||
**Do not rerun that forward-cutover script after success**: its preconditions
|
||||
intentionally reject an active Quadlet.
|
||||
|
||||
Recovery is now a Quadlet rebuild and restoration from a verified Atlas
|
||||
backup, with an explicit outage decision before replacing live NPM state.
|
||||
The old Compose owner is no longer installed; reintroducing it would require
|
||||
a separately reviewed configuration and outage plan. The historical
|
||||
in-window rollback trap is not a supported post-cleanup rollback procedure.
|
||||
Do not restore an old database over a live instance or remove NPM bind mounts.
|
||||
|
||||
## Verified evidence
|
||||
|
||||
- Immediately after cutover, `prometheus-npm.service` was active with zero
|
||||
recorded restarts; Compose was inactive/disabled. The generated
|
||||
`multi-user.target.wants` link and network dependency were present. An
|
||||
actual reboot has not been performed solely for this test.
|
||||
- The running image ID matched the prior Compose image. Podman showed the
|
||||
original two bind mounts, `server_web`, public 80/443, and loopback-only 81.
|
||||
External HTTPS to Gitea and Syncthing returned 200 with TLS verification
|
||||
result 0. External access to TCP/81 timed out.
|
||||
- The first **manual post-cutover** export `20261003T091633Z` succeeded with
|
||||
the Quadlet as its active owner. The Atlas pull published that version;
|
||||
its SHA-256 payload check passed. An isolated restore passed NPM SQLite
|
||||
`quick_check` with ten proxy hosts and six certificate records. Both
|
||||
Quadlet definitions were present in the tar archive.
|
||||
- A path/content manifest of all 70 regular Let's Encrypt files and the
|
||||
path/target manifest of all 12 symlinks in the Atlas archive exactly
|
||||
matched the live Prometheus tree (aggregate SHA-256
|
||||
`ce0965fbd3ff44bb8502ed9f314e0131edd86d822039de115b39f6a2273c2da8`).
|
||||
The earlier apparent 70-vs-82 count was only a regular-file-versus-symlink
|
||||
counting difference, not missing certificate data. No certificate key
|
||||
contents were exposed during comparison.
|
||||
- The targeted `--tags npm_quadlet` normal Ansible run completed with
|
||||
`changed=0`, and the backup export timer remained active/enabled.
|
||||
|
||||
The first unattended 02:00 Europe/Rome export and 03:00 Atlas pull **after**
|
||||
this cutover have not yet occurred. Check their service results and the
|
||||
published version after the next cycle; the successful manual cycle proves
|
||||
the new path works but not its next scheduled execution.
|
||||
|
||||
```bash
|
||||
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
|
||||
ansible-playbook ansible/site.yml --limit prometheus --tags npm_quadlet --check --diff
|
||||
sudo systemctl status prometheus-npm.service prometheus-backup-export.timer
|
||||
sudo systemctl show podman-compose-server.service -p LoadState # expected: not-found
|
||||
```
|
||||
|
||||
The backup archive includes credentials, certificates, and WireGuard
|
||||
configuration. Do not publish it or print its contents in diagnostics; see
|
||||
`docs/prometheus-backup.md` for the restricted pull and restore procedure.
|
||||
|
||||
## Selective legacy image cleanup
|
||||
|
||||
On 2026-10-03 opt-in Ansible tasks removed only the unused Gitea 1.25.2,
|
||||
Navidrome latest and PostgreSQL 13 rootful images, without force or global
|
||||
prune. Podman refuses images referenced by existing containers. The second
|
||||
run changed nothing. NPM remained active with zero restarts; local admin
|
||||
and public Gitea HTTPS returned 200. Backup timer and SSH proxy stayed active.
|
||||
|
||||
Validation:
|
||||
```bash
|
||||
ansible-playbook ansible/site.yml --limit prometheus --tags server_image_cleanup --check --diff -e server_legacy_image_cleanup=true
|
||||
```
|
||||
|
||||
The image cleanup defaults to disabled and carries the `never` tag.
|
||||
Check mode probes image presence but skips removal; it does not prove
|
||||
Podman would accept deletion. It never removes NPM resources.
|
||||
|
||||
## Approved legacy data and fallback cleanup
|
||||
|
||||
The operator explicitly approved deletion on 2026-10-03. The separate
|
||||
`server_legacy_cleanup` tasks removed `/opt/gitea`, `/home/git/.ssh`,
|
||||
`/opt/navidrome`, `/opt/postgres`, `/opt/music`, `/opt/containerd`,
|
||||
`/opt/docker`, the old Compose unit and the final Gitea export helper.
|
||||
The empty `/home/git` parent is removed only with `rmdir`, after confirming
|
||||
the Git account is absent. Guards reject symlinked paths, nested mounts,
|
||||
unexpected containers, container users of these paths, unexpected content
|
||||
in the empty legacy trees, and an active Compose or export service.
|
||||
The second cleanup run changed nothing.
|
||||
|
||||
Before deletion, Ansible removed obsolete backup input paths and the
|
||||
Gitea mount dependency. Normal Compose/template/final-export task checks
|
||||
changed nothing and did not recreate the retired files. Deletion is opt-in:
|
||||
|
||||
```bash
|
||||
ansible-playbook ansible/site.yml --limit prometheus --tags server_legacy_cleanup --check --diff -e server_legacy_cleanup=true
|
||||
```
|
||||
|
||||
Remove check mode only for approved deletion. No active NPM data, certificate,
|
||||
image, network, volume, SSH proxy, WireGuard configuration or backup archive
|
||||
is removed. No services were restarted by the cleanup.
|
||||
|
||||
After separate approval for the brief managed NPM pause, the new export
|
||||
`20261003T112906Z` completed successfully and was pulled to Atlas. SHA-256
|
||||
passed on both hosts; an isolated SQLite restore passed `quick_check` and
|
||||
contained ten proxy hosts. Both Quadlet definitions were present, and
|
||||
retired paths were absent. Temporary restore files were removed.
|
||||
NPM was active with zero automatic restarts; primary public Gitea HTTPS
|
||||
returned 200 with valid TLS. Backup timer, SSH proxy and WireGuard stayed active.
|
||||
The first scheduled post-cleanup cycle remains unverified.
|
||||
|
||||
After separate operator approval on 2026-10-03, the unused secondary hostname
|
||||
`git.ov-ad3410.infomaniak.ch` was removed from the declared domains and
|
||||
the managed NPM runtime override. Its Proxy Host (id 10) was already
|
||||
soft-deleted, with no generated config or associated certificate. Historical
|
||||
deleted records and backup archives are preserved; no DNS changes were made.
|
||||
Only `git.fscotto.duckdns.org` remains declared for the Gitea override.
|
||||
Nginx validation and reload passed without restarting NPM; the primary
|
||||
public HTTPS endpoint returned 200 with valid TLS.
|
||||
106
scripts/cutover_prometheus_npm_quadlet.sh
Normal file
106
scripts/cutover_prometheus_npm_quadlet.sh
Normal file
@@ -0,0 +1,106 @@
|
||||
#!/usr/bin/env bash
|
||||
# Run on Prometheus as root with the exact verified source-export version.
|
||||
set -Eeuo pipefail
|
||||
|
||||
expected_export=${1:?Pass the verified Prometheus backup export version}
|
||||
mode=${2:---preflight}
|
||||
[[ $expected_export =~ ^[0-9]{8}T[0-9]{6}Z$ ]] || exit 2
|
||||
[[ $mode == --preflight || $mode == --execute ]] || exit 2
|
||||
[[ $EUID -eq 0 ]] || { echo 'Run as root on Prometheus' >&2; exit 2; }
|
||||
|
||||
compose_unit=podman-compose-server.service
|
||||
quadlet_unit=prometheus-npm.service
|
||||
backup_timer=prometheus-backup-export.timer
|
||||
versions=/var/lib/prometheus-backup-export/versions
|
||||
quadlet_file=/etc/containers/systemd/prometheus-npm.container
|
||||
|
||||
exec 9>/run/lock/prometheus-backup-export.lock
|
||||
flock -n 9 || { echo 'Backup/export lock is busy' >&2; exit 1; }
|
||||
|
||||
systemctl is-active --quiet "$compose_unit"
|
||||
if systemctl is-active --quiet "$quadlet_unit"; then
|
||||
echo 'NPM Quadlet is already active; refusing overlapping cutover' >&2
|
||||
exit 1
|
||||
fi
|
||||
[[ $(systemctl show "$quadlet_unit" -p LoadState --value) == loaded ]]
|
||||
[[ $(systemctl is-enabled "$compose_unit") == enabled ]]
|
||||
[[ $(readlink "$versions/current") == "$expected_export" ]]
|
||||
image=$(sed -n 's/^Image=//p' "$quadlet_file")
|
||||
[[ $image =~ ^docker\.io/jc21/nginx-proxy-manager@sha256:[a-f0-9]{64}$ ]]
|
||||
podman image exists "$image"
|
||||
(cd "$versions/current" && sha256sum -c payload.sha256 && tar -tf payload.tar >/dev/null)
|
||||
curl -fsS --connect-timeout 2 --max-time 5 -o /dev/null http://127.0.0.1:81/
|
||||
old_image=$(podman inspect nginx-proxy-manager --format '{{.Image}}')
|
||||
|
||||
data_signature() {
|
||||
python3 - <<'PY'
|
||||
import hashlib, os, sqlite3
|
||||
db = sqlite3.connect('file:/opt/npm/data/database.sqlite?mode=ro', uri=True)
|
||||
assert db.execute('pragma quick_check').fetchone()[0] == 'ok'
|
||||
counts = [db.execute('select count(*) from ' + table).fetchone()[0]
|
||||
for table in ('proxy_host', 'certificate', 'user')]
|
||||
db.close()
|
||||
digest = hashlib.sha256()
|
||||
for root, dirs, files in os.walk('/opt/npm/letsencrypt'):
|
||||
dirs.sort()
|
||||
for name in sorted(files):
|
||||
path = os.path.join(root, name)
|
||||
with open(path, 'rb') as stream:
|
||||
digest.update(path.encode() + b'\0' + stream.read())
|
||||
print(*counts, digest.hexdigest())
|
||||
PY
|
||||
}
|
||||
before=$(data_signature)
|
||||
if [[ $mode == --preflight ]]; then
|
||||
echo 'NPM Quadlet cutover preflight passed; no service was changed'
|
||||
exit 0
|
||||
fi
|
||||
|
||||
stopped_old=false
|
||||
rollback() {
|
||||
rc=$?
|
||||
trap - EXIT
|
||||
if (( rc != 0 )) && "$stopped_old"; then
|
||||
echo 'NPM Quadlet cutover failed; restoring Compose' >&2
|
||||
systemctl stop "$quadlet_unit" || true
|
||||
systemctl enable "$compose_unit" || true
|
||||
systemctl start "$compose_unit" || true
|
||||
systemctl start "$backup_timer" || true
|
||||
curl -fsS --connect-timeout 2 --max-time 10 -o /dev/null http://127.0.0.1:81/ || true
|
||||
fi
|
||||
exit "$rc"
|
||||
}
|
||||
trap rollback EXIT
|
||||
|
||||
stopped_old=true
|
||||
systemctl stop "$backup_timer"
|
||||
systemctl stop "$compose_unit"
|
||||
if podman container exists nginx-proxy-manager; then
|
||||
echo 'Compose left the NPM container behind; refusing duplicate ownership' >&2
|
||||
exit 1
|
||||
fi
|
||||
systemctl disable "$compose_unit"
|
||||
systemctl start "$quadlet_unit"
|
||||
|
||||
ready=false
|
||||
for _ in {1..60}; do
|
||||
if curl -fsS --connect-timeout 2 --max-time 3 -o /dev/null http://127.0.0.1:81/; then
|
||||
ready=true
|
||||
break
|
||||
fi
|
||||
sleep 2
|
||||
done
|
||||
"$ready"
|
||||
systemctl is-active --quiet "$quadlet_unit"
|
||||
[[ $(podman inspect nginx-proxy-manager --format '{{.Image}}') == "$old_image" ]]
|
||||
podman exec nginx-proxy-manager nginx -t
|
||||
[[ $(data_signature) == "$before" ]]
|
||||
for hostname in git.fscotto.duckdns.org syncthing.fscotto.duckdns.org; do
|
||||
status=$(curl -ksS --connect-timeout 3 --max-time 10 \
|
||||
--resolve "$hostname:443:127.0.0.1" -o /dev/null -w '%{http_code}' \
|
||||
"https://$hostname/")
|
||||
[[ $status == 200 ]]
|
||||
done
|
||||
systemctl start "$backup_timer"
|
||||
stopped_old=false
|
||||
echo 'NPM Quadlet cutover passed local application and data checks'
|
||||
@@ -1,83 +1,71 @@
|
||||
$ANSIBLE_VAULT;1.1;AES256
|
||||
61353065386233646137323235306631353635663530363237636231316265643562353465323430
|
||||
6165646466623962313835313537633137633766373930380a316335323962616265643136346666
|
||||
63336133336131346336383534356637623831363138323165633262386333363535393365383233
|
||||
6234393835653439370a313963313365373633323464343263383661383336363662633133643232
|
||||
34366634383862363635653034313531623330396639616462343630326162316535643465653532
|
||||
36326534333637376462353561343964633636366331363833313263353133383636623537303663
|
||||
35393032316439336666343161653439643638376134363535656262343963393365623432336433
|
||||
35383934313762313037326430316666363731666231336534326661353034333063643364343230
|
||||
65333739303566366263333565333465613136646237623937393733623438613832393634663463
|
||||
39376131313234333039633735613233373931613232653036663665316636303961653834366339
|
||||
36353730316132316233303964303839363161346564396163336137663134353062363733656430
|
||||
37643339326661653031376265646132623162373562393437373437313732396537383939333666
|
||||
62353036316633306666313461663033303830393765396131643035353730383931646239663935
|
||||
32626461316364386135303761383837613063336466363162323332663764616464373565383231
|
||||
61346463336566346533326535376439643133613762383633396131323632356533636139336365
|
||||
62393838316634623932643034376631333539343965383436613364643962363834346337353334
|
||||
32656439366439313734353963343133333533653839613632323338336131373566613835393536
|
||||
31663433616334373432376531346435336530303936356461303163646463613661643161313661
|
||||
66663866343565616631616338353737356164353562366164383736346131666662623132333466
|
||||
39383865653631373232393433663430643961646265386166333137643966303834363262373636
|
||||
62396434373363353636376133666133663162653265313139313732353639336232333862643036
|
||||
64386231336561396537326139346566306434633934343038663165396665363032383466633662
|
||||
62336163633964363435386630343966333162333730336138333239646631633132663931376462
|
||||
33663139356261313065376636613930353735396131306538306664646135636336643032623131
|
||||
38346264333331353633326535326431626563323036313665643337353563333339646430386564
|
||||
31613435383036313430316366323636663735326336393338353835323861333564363832656462
|
||||
35336435623261326363633033316130393062616339353263643062633331646137376135656365
|
||||
35636139336564346164616235616431326531333433646330386134323932373339646536356464
|
||||
66343533326534326165323564663533653666633035343163633832393361336462343937623165
|
||||
62383931326630363036396333313931393836366439653433623165666166356338653364336534
|
||||
35333936653833386163633738326164386166613561333530633937343230363366333662666539
|
||||
39333361633933663735303438663239303536363433313962643137386533633539326365383765
|
||||
37636538386339333935386132353265353031643662616330316463623661663738353433313830
|
||||
36373963633166333464653338343830373063323536383364393033393235326639613662343737
|
||||
38663362636331343061646465313237313431373433353361353265333766633463353632646536
|
||||
31323231306138323031396630656538363930373439336234343963616334363632653738316465
|
||||
63653938373830336362313238656266613362636634616537653863336132343931616262396130
|
||||
66393239303866656232653832343132366537333537343635666563343639323433383163613335
|
||||
39613533376634316133633430303535306266656333626264343733666335393661666561396633
|
||||
39346265316137326465326635396362333565393133623637633132616232326263663662343137
|
||||
33363733306135363361643031306265363733656362386666306334333035393839636533343363
|
||||
35396638616636633639343930373136376339346162393061393765363837646365383866636131
|
||||
33653465666239393133616232636231333332396138376332393664343364643835306530393238
|
||||
34663237303530303837663535646263393931373531393039356336316561653130356262636562
|
||||
38336362326639653237626634376334666565653036353236313634376364626338646538386536
|
||||
38626636386466373566646166393963643164343536373236396138303532393161363335386638
|
||||
32633032393737626363613463323366366637616361313537356136626661626633613739323338
|
||||
35383963666431343566356562333234663936376562616638636261303466633539376334303331
|
||||
39303834663234663063356233313962326664383839393832303462643636393034383434303465
|
||||
64333635376135326333356435373734643430623736373234643335343130383066326436356664
|
||||
63346663326364343634303930343338336139313864316165366232643537366635653764353763
|
||||
31363863633261643263303433373330366161323166366462336332313135366338393334653764
|
||||
66353733653137663835663731373364613030373334663061313433373861613665363236633130
|
||||
65613965366636343465336533613438373466383737373366653965633437323562643966396431
|
||||
39303033643438633762633263326132663466643438656366363431616237633031333936313831
|
||||
30323930383233313032323638356333626230333764363662313662646536643839353032353462
|
||||
30326166653937353130623133303533343934633565393831623033303234316330353432313266
|
||||
30636536633933376365623665616262663236383731633633346232613366333137396139306363
|
||||
35633336643266326335303261666666653536666630613639376336373237646134306462616537
|
||||
33343561373162666332613634643837343566646161373065366637653135613632353334636363
|
||||
63363232303963646530333366663862323264326536643337323266396566316233613630303637
|
||||
66646366376466373931613734363931316230323063373666653062373364396433633762633762
|
||||
38613933323733653238383935623230383562646563363833653838636165626365646537383639
|
||||
33666535656363393562316336633439636138373365623431393965653765306138646234663938
|
||||
65653133663663393731646337386535333261643932336132396237323930306136643534353930
|
||||
65636438396432623034626561613137336138623265393064383034623863303166356138393564
|
||||
37373164626634653662326234333539663735323464613334616130643937373730363263633366
|
||||
31393937326432386165343338313031376565313866363731643534313233303064373935303538
|
||||
31343832336230393636653432653162336361383963633766343461653466316337353931333363
|
||||
63313137303564336630343937356564643763383764613362366634373362666465626334336539
|
||||
64366533376165306532343461613265366266383862323032333465336161663161376630316465
|
||||
30306562666163646235656664653635366461366435663961623635383437663564356563346462
|
||||
31636234663765623838333237393239373564366262613637363938653463396530613963643837
|
||||
38636634376637366332623035313465393762653865623130336263343663303066366135616639
|
||||
63333964356466613038303263366462346261353030646532366361393965306435613131316463
|
||||
65366266376637323764643239323730366565633335666638666334663635373961303637383861
|
||||
35313431646434656562333937663837393038386361616630626532636339306432353434656165
|
||||
33663261383166386432383465666136376237346565303164363461666663346130346162316338
|
||||
62373061353034316234303835663439396434343738303764376665336239626238386436386234
|
||||
61306166383637366266393730323732386163366261393630336431633862353761343763363665
|
||||
61323039396234393835303633363339373633653334343766653032313230343464326664356566
|
||||
3462623830666664626633373966363866333337383730313066
|
||||
37646664613266633436346262613633613830623366383138613432366365373765353230333134
|
||||
3332333764313337396637323133623937343738373133370a333930356365653034323235643230
|
||||
36633864343161653833356636373931383761663864663334336236373733326266386639366335
|
||||
3131313661313637320a313937633361646333333962303335333233346166343831373039663964
|
||||
33366532386135663463643965363766643063616436316463666232666138323236346231303537
|
||||
34303535333866376430363063623934623761373865656231656661383935393866353566346430
|
||||
64333434613432376436343438343561383235366631623730653533633535326237666265653439
|
||||
34366264653665643063663361313339663034323932326233366636326336323432303434373765
|
||||
36316532316265343434383438623239666232373633626330333464303361643630303635643834
|
||||
39313136623830303762313462343637633763626333393033346637663931663238653734626131
|
||||
38393963646563333732353531653239643330326539643538323164343934356166343034316565
|
||||
33346431333735636537613930383331393265313962626234363237373562313231393061326439
|
||||
64363765323935316661353531366165343139633963336139313737306332613364643031666161
|
||||
30386362643930316265616564306336633133303166363665333462316265313364393939306162
|
||||
31303639313933356337386134623934663461643161306666633261653538633232343036653833
|
||||
66316466636233343136393765636333353230353738313833333265663238303730313936326664
|
||||
38373239353162363438323964333030666563346161643437326335666162356264396135393532
|
||||
63363862373136346532653734336335616132386237303031363433663132343861633937386130
|
||||
30633938616364303462303030303966303939633066393264303462393730363233373937356439
|
||||
36663533376232663737613734653532313136343939663539373866333638396266666163383864
|
||||
63613532393334373539346338616163383637633237666234613437663966653733616361353830
|
||||
61656666376133363330633863346637376266343134633037313132313361366638616261363839
|
||||
39393062396237666333303937363536346561343763663133323236393037383532396465336138
|
||||
35613463356532376534386433626337613030343266353332306462306463336336343830666138
|
||||
33656138363837633337393865643633623261613335366263643162663637623636666162653632
|
||||
30626238616266323332616234393838343330663662393433366630393566316336636530303165
|
||||
38343665623437356636643236393734396264356632326133623264633862633333626330336663
|
||||
61376263656665653731636133316161653635323138303866623862303065366232633736623336
|
||||
39656666386435343062656138313061616661313966326432663236626631316162623961616636
|
||||
35343939613262303066626537396164616666316265643065373638663436643961336138313862
|
||||
39666163646538356338356631346534633139643636393866646462646533363265663234633761
|
||||
65363661336138353239656165393836386134666331663036653132306433343764643666306333
|
||||
33623661626565633333306337303263633335386632386330353730316436313931326164363862
|
||||
37616265653161633632353865346639653961653836353962303762336535666266386535363165
|
||||
39653138646663376634323131613463333035326639313266613830616431316131383464353533
|
||||
62656634346637636164626461613137303461633761336232373133653532323566303136663030
|
||||
30633337346534636566343934306662356238396365306563336666623435353731613136333036
|
||||
36343436373932323265306639363761353364383635333136366231373166613861633032343233
|
||||
61376338616630343639333964356162613332323835333730333135356665383431626138643534
|
||||
66393966666465303763316230386538393863303063386564303165303962346139373338303436
|
||||
39373032663538323532323766353864643338326561313564373562616430326264386362666532
|
||||
36613132306462336631363035343732636465343562643430343035373961366566383130656165
|
||||
64613938393265343037633161653937323933646637653036306532366237313838346361333932
|
||||
34663565653264626137323239336532643262356166633665313761336162303635346666383863
|
||||
61383930383033626337383366353766393536653135383062656639323361353539356232613736
|
||||
63646235663363333333623463313961326533653236363938383765663439613832653039386436
|
||||
38393734633536313731323437336332353564363564333736663037386530333639326338656561
|
||||
66336637353238383231613666313261383234336531666132396230373931623363323832633064
|
||||
39613036346166393936613939363865616135653830366435643538336365353333613831353962
|
||||
36623537363434373137633063373934383439333462646361613737303239643834303535366138
|
||||
34656465316431656461373737643537303936636539383934373831616438343965373765373535
|
||||
63653164363731653030303466646539636361383664343763646163663238383435653035653666
|
||||
63383165626365653261303834333234626534396333353231303261396361616233363334383336
|
||||
33316462636133336132656364613439396131613565646565396365316238323962353462653736
|
||||
64386139313266663963643962363133386133393166306163626632646463333363323830306164
|
||||
35353936653137383761326132373739306163613764386531613032313235373331303530383633
|
||||
64646533313434653734366233633535323564386431306538633666383661303038613330653832
|
||||
34643463396137643034353439653334653836333161396130363637326339383363303037306330
|
||||
61353635633334343432646461396439393439383639336139316161373737333961653731393333
|
||||
61636164343838346365373736356161386430356533303331333838333732363233613931613863
|
||||
66633662383466306332366563373865323861323833353238356563363635313463366333653432
|
||||
65303839653963376566383737346231343663363363313332383365646363373737323839613564
|
||||
34613362303335316363363661653639386538326337386537333765643161613961316531613563
|
||||
38386564636637643762643830666138383361396233303339643665343261356462393830376662
|
||||
32656334346536636536343263336565333234353831616565366538393661353561376538346334
|
||||
61396135623230366433303932396130636331333263316333643861626564343330386636613063
|
||||
32383061616435643736653264313839363232346332343565336464353138396339623533393237
|
||||
38353632646565323735643462626239663736643033643231613464663866663262366632353434
|
||||
37363866343239363131633464316133396462353336613962306332343563333962333934616330
|
||||
3536356634376131633039373834376533633065303533653333
|
||||
|
||||
@@ -8,7 +8,7 @@ vault_git_work_email: "REPLACE_ME"
|
||||
vault_git_work_gpg: "REPLACE_ME"
|
||||
vault_ikaros_authorized_ssh_keys:
|
||||
- "ssh-ed25519 REPLACE_ME"
|
||||
vault_aegis_icloudpd_apple_id: "REPLACE_ME"
|
||||
vault_atlas_icloudpd_apple_id: "REPLACE_ME"
|
||||
vault_atlas_admin_password_hash: "REPLACE_WITH_A_SHADOW_COMPATIBLE_HASH"
|
||||
vault_atlas_samba_password: "REPLACE_ME"
|
||||
vault_atlas_immich_db_password: "REPLACE_ME"
|
||||
|
||||
Reference in New Issue
Block a user