Files
infra/AGENTS.md
Fabio Scotto di Santolo 7bc7f0e645 Feature/prometheus npm quadlet (#15)
* Stage Prometheus NPM Quadlet with backup-safe cutover

* Complete Prometheus NPM Quadlet cutover
2026-10-03 11:53:12 +02:00

495 lines
43 KiB
Markdown

# AGENTS.md
Ansible-driven personal infrastructure repo for Fedora and Void desktops, Fedora IoT, WSL, a Rocky Linux 9 server, and an Atlas NAS.
## Source Of Truth
- Main orchestration: `ansible/site.yml`
- Inventory and layering inputs: `ansible/inventory/hosts.yml`, `ansible/inventory/group_vars/*.yml`, `ansible/inventory/host_vars/*.yml`
- Dotfiles live under `dotfiles/`
- AI agent instructions (bootstrap, rules, knowledge) are centralized in `dotfiles/common/.config/ai/` and shared between OpenCode, Codex, and Gemini CLI.
- OpenCode loads its entrypoint configuration from `dotfiles/common/.config/opencode/opencode.json`.
- Codex config is rendered from `dotfiles/common/.codex/config.toml.j2` so `model_instructions_file` points to the deployed `~/.config/ai/bootstrap.md`.
## Topology
- Current personal desktop: `ikaros = platform_fedora + role_personal_workstation + graphical_desktop + desktop_gnome`
- Current laptop: `nymph = platform_fedora + graphical_desktop + desktop_gnome`
- Void desktop profile is also the base for other future/reference hosts via `platform_void + graphical_desktop`
- Workstation: `deadalus` is Windows + Fedora WSL.
- Rocky server: `prometheus` belongs to `rocky_server`.
- NAS: `atlas` (Rocky Linux 9, reached through SSH)
- Always-on LAN node: `aegis` (Fedora IoT on Raspberry Pi 4, reached through SSH)
- Hosts intentionally belong to multiple groups; trust `ansible/site.yml` over hostname assumptions.
- Inventory axes are independent: `platform_*`, `role_*`, and `desktop_*`. Legacy `void` and `desktop` remain compatibility parents.
## Working Rules
- Preserve layering `all -> platform -> role -> desktop -> host`.
- Keep `ansible/site.yml` small; orchestration belongs there, implementation belongs in roles.
- Prefer minimal, targeted edits. Preserve idempotency and existing ordering.
- Use Git Flow branch prefixes: `feature/` for new functionality, `bugfix/` for non-urgent fixes,
`hotfix/` for urgent production fixes, `release/` for release preparation, and `support/` for
maintained release lines. Do not use abbreviated prefixes such as `feat/`.
- Desktop and WSL hosts use `ansible_connection: local`; remote infrastructure hosts use SSH.
- Treat `secrets/` as sensitive. Never print secret values.
- Tmux plugins are bootstrapped by TPM on the host; the repo only keeps tmux config and custom helper scripts.
- Read the relevant role tasks, templates, vars, and deployed dotfiles before editing.
## Validation
- Default minimum:
- `ansible-playbook ansible/site.yml --syntax-check`
- Repo-wide checks:
- `ansible-lint ansible/site.yml`
- `ansible-lint ansible/roles`
- `yamllint ansible/`
- Host-focused dry runs:
- Fedora desktop work: `ansible-playbook ansible/site.yml --limit ikaros --check --diff`
- Fedora laptop work: `ansible-playbook ansible/site.yml --limit nymph --check --diff`
- WSL workstation dev: `ansible-playbook ansible/site.yml --limit deadalus --check --diff`
- Server: `ansible-playbook ansible/site.yml --limit prometheus --check --diff`
- Rocky server after activation: `ansible-playbook ansible/site.yml --limit <host> --check --diff`
- Atlas NAS: `ansible-playbook ansible/site.yml --limit atlas --check --diff`
- Aegis IoT: `ansible-playbook ansible/site.yml --limit aegis --check --diff`
- Aegis NFS client layer: `ansible-playbook ansible/site.yml --limit aegis --tags nfs --list-tasks`
- Aegis host DNS: `ansible-playbook ansible/site.yml --limit aegis --tags dns --check --diff`
- Focused checks:
- Emacs is disabled by default; temporary Emacs check: `ansible-playbook ansible/site.yml --limit <host> --tags emacs --check --diff -e emacs_enabled=true`
- AI coding agents: `ansible-playbook ansible/site.yml --limit <host> --tags ai_agents --check --diff`
- Mail bootstrap: `sh -n scripts/bootstrap_mail.sh` and `shellcheck scripts/bootstrap_mail.sh`
- Server NPM Quadlet: `systemctl status prometheus-npm.service`; disabled Compose fallback render:
`podman-compose -f /opt/docker/server/docker-compose.yml config`
- Atlas media stack:
`ansible-playbook ansible/site.yml --limit atlas --tags storage,sharing,containers --check --diff`
- Atlas rootless Gitea staging (does not start Gitea):
`ansible-playbook ansible/site.yml --limit atlas --tags gitea --check --diff`
- Atlas iCloudPD storage and inactive Quadlet (does not start it):
`ansible-playbook ansible/site.yml --limit atlas --tags icloudpd --check --diff`
- Atlas explicit Gitea host-owner migration (live outage; never a normal run):
`ansible-playbook ansible/site.yml --limit atlas --tags gitea_owner_migration -e atlas_gitea_owner_migration=true`
- Atlas explicit isolated Gitea restore rehearsal (not part of normal runs):
`ansible-playbook ansible/site.yml --limit atlas --tags gitea_restore -e atlas_gitea_restore_test=true`
- Atlas final Gitea replacement gate (dry-run only until a stopped-source export is pulled):
`ansible-playbook ansible/site.yml --limit atlas --tags gitea_final_restore --check --diff -e atlas_gitea_final_restore=true`
- Prometheus final Gitea export helper (dry-run installs only; outage action remains opt-in):
`ansible-playbook ansible/site.yml --limit prometheus --tags gitea_final_export --check --diff`
- Gitea cutover network configuration before activation:
`ansible-playbook ansible/site.yml --limit prometheus --tags gitea_cutover,prometheus_backup --check --diff -e server_gitea_on_atlas=true`
and `ansible-playbook ansible/site.yml --limit atlas --tags gitea --check --diff`
- Atlas daily Navidrome music copy:
`ansible-playbook ansible/site.yml --limit atlas --tags music_sync --check --diff`
- Atlas network/share hardening:
`ansible-playbook ansible/site.yml --limit atlas --tags hardening,sharing --check --diff`
- Atlas ZFS snapshot retention and scrub timers:
`ansible-playbook ansible/site.yml --limit atlas --tags snapshots,scrub --check --diff`
- Atlas encrypted Borg backup:
`ansible-playbook ansible/site.yml --limit atlas --tags packages,borg --check --diff`
- Atlas Borg progress logging only:
`ansible-playbook ansible/site.yml --limit atlas --tags borg_logging --check --diff`
- Atlas manual offline USB backup and 45Drives Alerts reminder:
`ansible-playbook ansible/site.yml --limit atlas --tags usb_backup,usb_reminder --check --diff`
- Atlas pool, disk, capacity, temperature, and job monitoring:
`ansible-playbook ansible/site.yml --limit atlas --tags monitoring --check --diff`
- Atlas explicit post-restore SELinux relabeling:
`ansible-playbook ansible/site.yml --limit atlas --tags restorecon --check -e '{"atlas_restorecon_paths":["/zpool/archive"]}'`
- Prometheus/Aegis WireGuard gateway:
`ansible-playbook ansible/site.yml --limit prometheus,aegis --tags wireguard --check --diff`
- Prometheus NPM Quadlet steady state (does not perform a cutover):
`ansible-playbook ansible/site.yml --limit prometheus --tags npm_quadlet --check --diff`
- DuckDNS config only: `ansible-playbook ansible/site.yml --limit prometheus --tags duckdns --check --diff`
## Conventions
- Use FQCN Ansible modules.
- Prefer declarative modules over `command`/`shell`; when `shell` is required, make idempotency and failure behavior explicit.
- Start YAML files with `---`, use 2-space indentation, and keep file modes quoted like `"0644"`.
- Keep booleans as booleans and structured vars as YAML lists/maps.
- Put host-specific overrides in `host_vars`, not shared `group_vars`.
- Use `no_log: true` for secret-bearing task inputs or outputs.
## Desktop Notes
- `desktop_profile` names independently selectable desktop groups such as `desktop_gnome`, `desktop_sway`, and `desktop_niri`. Keep platform-specific session bootstrap in platform-specific roles.
- `desktop_environment` is fixed to `minimal` for Void desktops. `profile_desktop_common` owns shared Void bootstrap; `profile_desktop_sway` and `profile_desktop_niri` manage the enabled sessions, while `profile_desktop_gnome` copies shared desktop dotfiles for Fedora/GNOME without managing GNOME settings. `desktop_sessions_enabled` and `desktop_default_session` apply to the minimal mode.
- Emacs has one authoring-oriented `.emacs.d`, deployed by `dotfiles_common` when `emacs_enabled` is true. Fedora/GNOME desktops and workstation profiles enable it; keep platform dependencies in package group vars rather than branching in Emacs Lisp.
- NTFS filesystem support is provided by `ntfs-3g` in `ansible/inventory/group_vars/void.yml`.
- Void user services are managed by `turnstile` and live under `dotfiles/desktop/.config/service/`.
- `ssh-agent` keeps the stable socket `~/.local/state/ssh-agent/socket`.
- Critical session entrypoints:
- `dotfiles/desktop/.config/sway/config` plus `host.conf` and `session-env` deployed via `host_sway_dotfiles` (sway / Wayland)
- `dotfiles/desktop/.config/niri/config.kdl` and `session-env` deployed via `desktop_niri_dotfiles` (Niri / Wayland)
- Void Niri lives in `profile_desktop_niri`, gated on `'niri' in desktop_sessions_enabled`; it installs the `emptty` `niri.desktop` session, the `/usr/local/bin/start-niri` launcher, and the xdg-desktop-portal config, mirroring `profile_desktop_sway`.
- Fedora GNOME (`desktop_gnome`) assumes GNOME comes from the Fedora Workstation base install; Ansible deploys shared desktop dotfiles and git/GPG config for `ikaros` and `nymph`, not GNOME settings.
- Do not switch or restart the display manager during a playbook run from an active graphical session.
- `nymph` is the Fedora/GNOME laptop target; keep GNOME settings unmanaged for now and add host-specific tuning only after real use.
## Void Package And Dotfile Bucket Rules
`platform_void` is the reusable Void platform selection. The legacy `void` group remains a compatibility parent so existing `group_vars/void.yml` and `when: "'void' in group_names"` checks keep working during the transition.
The Void desktop package lists in `ansible/inventory/group_vars/void.yml` are kept disjoint by role:
- `void_packages_base` — system runtime only (init/services, kernel, audio core, networking, filesystem, firewall, hardware daemons, runit logging).
- `desktop_common_packages` — GUI infrastructure shared by the minimal desktop mode.
- `desktop_minimal_packages` — applications, integration components, and the `emptty` display manager.
- `desktop_sway_packages` — binaries specific to the Sway session.
`profile_packages` remains the shared package bucket for Void and Fedora profiles. Rocky uses
`rocky_profile_packages` so RPM-specific names do not leak back into the other platforms; do not move
desktop-specific Void entries through either bucket.
The dotfile vars follow the same split: `desktop_common_dotfiles` carries mode-independent content and `desktop_minimal_dotfiles` carries Thunar, Udiskie, and MIME defaults. `desktop_void_dotfiles` remains reserved for files that need the Void runtime.
## Workstation Notes
- `deadalus` is modeled as Windows + Fedora WSL and is the sole workstation target.
- Fedora WSL belongs to `platform_fedora`, `workstation_dev_fedora`, and the shared WSL layer. It must not receive Flatpak or Snap runtimes.
- Fedora WSL installs Mise from the official `jdxcode/mise` COPR and uses its pinned Temurin Java 11 JDK; update the declared Mise version deliberately.
- Windows applications are installed manually and are not managed from the WSL profile.
## Rocky Server Notes
- DuckDNS is rendered by `profile_server` from host-local `server_duckdns_domain` and
`vault_duckdns_token`. Keep the rotated token in encrypted Vault or untracked local vars, never in
dotfiles. The private `~/duckdns/duck.sh` keeps the existing entrypoint; rendering uses `no_log`
and disables diffs. Provisioning does not execute the updater or change its external schedule.
- `rocky_server` is a child of both `platform_rocky` and `server`; `prometheus` is its active target.
- The target must already provide `server_username` with local sudo access before the profile runs.
- The Rocky profile installs Podman and podman-compose and renders the disabled legacy
`podman-compose-server` unit for rollback. On Prometheus, Nginx Proxy Manager is now the rootful
`prometheus-npm.service` Quadlet with a pinned image digest and the existing `/opt/npm/data` and
`/opt/npm/letsencrypt` bind mounts. The rootful `server_web` bridge remains `10.89.0.0/24`.
Gitea runs on Atlas; PostgreSQL and Navidrome are absent from the desired Prometheus stack.
The profile does not delete legacy data, update DNS, or perform an implicit cutover.
- Firewalld enables SSH, Cockpit (`9090/tcp`), HTTP and HTTPS. Nginx Proxy Manager publishes `80/tcp` and
`443/tcp`; bind its administration interface only to `127.0.0.1:81` and use `npm-tunnel` from Ikaros or Nymph.
Nextcloud remains disabled; do not provision `/srv/nextcloud` directories.
- `scripts/migrate_prometheus_data.sh` is the separate, source-host-run NPM/Gitea migration path. It dry-runs by
default and requires explicit source-stack quiescing before copying persistent Docker data with rsync.
- Atlas-only OpenZFS, NFS, Samba, and Syncthing stay selected through Atlas host variables and must not
leak into `rocky_server`. Cockpit plus its Navigator and Podman extensions are selected explicitly for
Prometheus through its host variables.
## Atlas NAS Notes
- `atlas` is a remote Rocky Linux 9 NAS. Keep its connection, LAN, pool and mountpoint values in
`host_vars/atlas.yml`. Bootstrap it once with `-e atlas_connection_username=<existing-admin>`;
subsequent runs use the dedicated Atlas account.
- The pool is normally pre-existing. A one-time bootstrap may create it only when `atlas_create_pool=true`
is explicitly supplied and `atlas_zpool_disks` contains exactly four real `/dev/disk/by-id/...` paths.
Never partition, force, destroy, roll back, or modify the vdev layout of an existing pool.
- `atlas_manage_storage`, `atlas_manage_sharing`, and `atlas_manage_firewall` are enabled in Atlas host vars as
the declared steady state; set one false only for a deliberate suspension. `atlas_manage_media_stack` remains false
until the future rootful Immich stack has its required Vault inputs and target validation.
- Atlas requires `vault_atlas_admin_password_hash` for Cockpit and, while sharing is enabled,
`vault_atlas_samba_password`. The future rootful media stack also requires
`vault_atlas_immich_db_password`. Never print these values.
- Atlas creates the complete declared hierarchy only under the verified existing or explicitly bootstrapped pool: `archive`,
`services`, `services/data`, `services/data/navidrome`, `services/data/syncthing`, `media`, `media/music`,
`media/photobook`, `backup`, `backup/hosts`, and `backup/hosts/prometheus`. `backup` has a `500G`
reservation covering its descendants. `archive` is the SMB-shared raw-data namespace; container state is never beneath it.
- The `immich` system account is fixed to UID/GID `1100`, has no login shell or `wheel` membership, and receives only
the `video` and `render` supplementary groups. Immich's rootful Quadlets run as `1100:1100`; Server and ML receive
`/dev/dri`, while the Photobook external library is read-only at `/external/photobook`.
- Atlas applies persistent kernel network hardening: redirects and source routes are rejected, martians logged, reverse-path filtering remains loose for WireGuard, and IPv4 forwarding is disabled. SSH permits only the declared administrator using public-key authentication; root login, passwords, agent and remote forwarding
are disabled, while local forwarding remains available for private administrative tunnels. Photobook is exported only to the configured Aegis IP with all access squashed to UID/GID
`1100`. Targeted SELinux is enforced persistently; a required reboot is reported but never initiated automatically. The primary LAN interface is assigned explicitly to the managed firewalld zone, and firewall rules are applied before NFS or SMB are started; their service state and TCP listeners are then verified. SMB3 exposes `Archive` to Vault-backed authorized accounts on mandatory encrypted, signed SMB3 over TCP/445 only and admits the configured LAN without host-specific exclusions.
- Atlas NPM and Immich share a rootful Podman network. NPM publishes HTTP/HTTPS, but its administration port remains
bound to `127.0.0.1:81`; do not expose it directly to the LAN or Internet.
- `profile_backend_phase1` temporarily runs rootless Navidrome and Syncthing on Atlas until Uranus replaces
them. It binds only to Atlas' LAN IP, never `wg0`; Navidrome and the Syncthing GUI admit only Aegis as
the source-NAT gateway, while native Syncthing ports admit the configured LAN. It initializes fresh
state only and never migrates or deletes source application data. The enabled rootless
`atlas-music-sync.timer` copies `/zpool/archive/Music` to `/zpool/media/music` daily at 00:45
Europe/Rome without deleting destination files; it requires both datasets to be mounted.
- `wireguard_overlay` manages `wg0` between Prometheus (`10.0.0.1`) and Aegis (`10.0.0.2`). It persists private
keys only on their respective hosts, exchanges only derived public keys through Ansible, and verifies a real peer
handshake. Prometheus opens `51820/udp`; Aegis is the LAN gateway. Its persistent IPv4 forwarding, narrowly scoped
WireGuard-to-LAN firewalld policy, and source masquerading permit Prometheus to reach LAN services without a static
route on the router. Prometheus includes `192.168.178.0/24` in Aegis' peer `AllowedIPs`; add the Uranus VIP there
when it is assigned. After a firewalld reload, restore Prometheus' rootful Podman networking with
`podman network reload --all` so the existing proxy stack retains container DNS.
## Atlas NAS TODO
Completed validation: the existing RAIDZ2 pool and datasets, SELinux, LAN firewall, SSH, Cockpit with
the selected 45Drives plugins, encrypted SMB3 `Archive`, the Aegis-only NFSv4 `photobook` export, and the
Prometheus--Aegis WireGuard gateway are operational. The gateway handshake, forwarding, source masquerading,
and TCP reachability to Atlas were verified. Temporary Navidrome and Syncthing are available through their
manual NPM Proxy Hosts; Syncthing uses `/data/Org` backed by the SMB-shared Archive dataset. Aegis has also
validated NFSv4.2 read, write, delete, and `all_squash` mapping to UID/GID `1100` end-to-end. The ZFS
snapshot timers are active; a recursive hourly snapshot and scheduled retention prune completed
successfully. The first monthly scrub remains a runtime check.
### Priority 1 - Data protection
- [x] Deploy Ansible-managed recursive ZFS snapshots with 24 hourly, 30 daily, 8 weekly, and 12 monthly
generations, plus a monthly scrub on the first Sunday at 03:00. The timers and first hourly snapshot were
verified on Atlas. Cockpit Scheduler is for visibility or manual operations only, and snapshot
rollback is never automated.
- [ ] Verify the first monthly ZFS scrub from its actual service result. Scheduled retention pruning
was observed on 2026-09-30; timer activation alone does not establish a successful scrub.
- [x] Activate and validate the encrypted offsite Borg backup to the Hetzner Storage Box. Atlas uses the
dedicated SSH identity, pinned ED25519 host key, Vault-backed `repokey` encryption, and a locked
non-login `borg` account with no sudo or supplementary groups. The initial snapshot-consistent backup,
Borg repository check, and temporary-directory restore completed successfully; the restored `Archive`
tree matched the live data, and temporary snapshots and mounts were removed. The exported recovery key
was copied offline. Daily backup retries and logging, 30 daily, 8 weekly and 12 monthly archives,
compaction, and monthly repository checks are enabled. Runs report a ZFS-based estimated percentage.
On 2026-09-30 a successful incremental run also removed the stale 2026-09-29 snapshot and its own
temporary snapshot after exit; the earlier `RuntimeDirectory` cleanup failure is resolved.
- [x] Populate `/zpool/archive` with the currently available data so offsite and offline backup tests run
against a representative load.
- [x] Evaluate Borg against the populated pool. The 2026-09-29 archive took 1 h 32 min for 2.18 TB
original / 2.04 TB compressed data, with 13.49 GB deduplicated size; retention and compaction
succeeded. On 2026-09-30 a subsequent incremental archive completed in about 22 seconds with
successful cleanup. The monitor reported 37% Storage Box quota used. These are observed runs, not
a guarantee of future duration or compression ratio.
- [x] Add the UUID-bound offline USB backup with versioned rsync, locking, capacity checks, verification,
safe unmounting and a tested restore procedure; never trigger it for an arbitrary USB disk. The
LUKS/ext4 identities were read-only verified; the manual service and 45Drives Alerts reminder timer were
deployed on Atlas. Interactive LUKS unlock is part of the manual service; only the reminder is
scheduled for the first Saturday of each month at 10:00 Europe/Rome via the existing 45Drives
notifier. A manual test produced an Alerts notification, not an email. The first USB attempt failed
on a `security.selinux` xattr and was interrupted; the xattr filter is deployed and the temporary
recursive snapshot and open LUKS mapper were cleaned up. A later run reported checksum verification
and published the USB version, but failed while removing host-namespace ZFS snapshot mounts. Those
exact mounts and snapshots were cleaned up. An `ExecStopPost` helper now removes only the named
temporary snapshot after the backup process exits. A new full run checksum-verified and published a
USB version; the service ended successfully, the mapper closed, no temporary USB snapshot remained,
and the pool was healthy. On 2026-09-25 an independent, read-only USB restore test copied one file from
the published `atlas/latest` version into `/var/tmp` and matched contents, owner, mode, size, mtime and
POSIX ACL. The temporary copy and mount were removed, the mapper closed, and the pool remained healthy.
- [x] Test restores independently from a ZFS snapshot, Borg, and the offline USB backup before relying on
any backup path. The earlier Borg temporary-directory restore passed. On 2026-09-25 a separate,
read-only ZFS snapshot test restored one file to `/var/tmp`, confirmed matching contents, ownership,
mode, mtime and ACL, then removed its temporary copy and on-demand mount. This is a file-level smoke
test, not full dataset recovery. An independent USB file restore passed on 2026-09-25 with matching
content and metadata; the later scaled OS-rebuild rehearsal is documented under Priority 2.
- [x] Add monitoring and alerting for pool health, scrub/resilver, SMART data, temperatures, free space,
snapshot/local-backup growth, Hetzner Storage Box quota, and failed maintenance or backup timers.
The half-hourly Atlas health monitor and systemd final-failure hooks are deployed. A live probe
found no issues; the service and timer succeeded, and a labelled 45Drives Alerts test notification
was submitted. Alerts are deduplicated; email delivery is not claimed. The Storage Box quota probe
runs `df -m` over the dedicated pinned-key SSH identity and does not open the Borg repository.
The failed-job hook was corrected to pass the literal systemd unit name; its expansion was verified
on Atlas, but a new real failure notification has not been deliberately triggered.
### Priority 2 - NAS operability and recovery
- [x] Document and test disaster recovery in `docs/atlas-recovery.md`: the operator confirmed Vault
and Borg recovery material is available offline; provisional targets are RPO 24h/RTO 72h. On
2026-09-30 an isolated small Rocky VM was rebuilt with the Atlas Ansible roles, imported its
preserved RAIDZ2 pool without force/rewind, and restored a file from the preserved snapshot;
the second Ansible run was idempotent. Earlier independent production ZFS, USB, and Borg file
restore tests remain separate evidence. A production-size full restore, unclean import, and
measured 24h/72h compliance are not claimed.
- [x] Define a controlled Rocky kernel/OpenZFS update and reboot procedure in
`docs/atlas-updates.md`. The first real change-window execution is not yet
validated; the procedure never reboots automatically or upgrades pool features.
- [x] Add the Atlas-initiated least-privilege Prometheus backup pull: Prometheus exposes only prepared
read-only dumps through a dedicated account and Atlas retains the private SSH key, pinned host key,
atomic pull, verification, retention and systemd service/timer. The dedicated key/account and unit
files are deployed; live read-only SSH, shell denial, and write denial were verified. On 2026-09-30
a manual export, Atlas pull, checksum verification, and temporary restore passed; both SQLite
databases passed integrity checks and a restored Git repository passed `git fsck`. Both daily
timers are enabled for 02:00/03:00 Europe/Rome. On 2026-10-01 their first scheduled export and
pull succeeded: Atlas verified the payload checksum and published `20261001T000001Z` as `latest`.
- [x] Decide whether a common SMB/NFS namespace is required: no. `Archive` (SMB) and `photobook` (NFS)
remain intentionally distinct; `docs/atlas-sharing-decision.md` records the decision. No ACL or export
change is authorized by this decision.
### Priority 3 - Service expansion
- [x] Populate `/zpool/media/music` and validate Navidrome. On 2026-09-30, 21,158 files
(93,937,810,350 regular-file bytes) were copied from `/zpool/archive/Music` using a temporary
ZFS snapshot; a checksum-based rsync dry run found no differences or extra files. Navidrome saw
all files through its read-only mount, completed a scan, indexed 18,168 tracks, and responded
over HTTP. Some imported playlists still reference obsolete Windows paths. The source was left
intact and the temporary snapshot was removed.
- [x] Schedule a daily, non-deleting copy from `Archive/Music` to the separate Navidrome music
dataset. The rootless `atlas-music-sync.timer` is enabled for 00:45 Europe/Rome; a manual
idempotent service run succeeded on 2026-10-01. The first scheduled run triggered at
00:45 CEST on 2026-10-02 and exited successfully (`Result=success`, status 0); the next
run is scheduled for 2026-10-03 00:45 CEST.
- [x] Design the staged Prometheus-to-Atlas Gitea migration in `docs/atlas-gitea-migration.md`.
The approved topology keeps NPM on Prometheus and moves HTTPS and public SSH (TCP/2222) together;
Gitea runs as an `admin`-owned rootless user Quadlet on Atlas with an internal `gitea` user.
The rootful-to-rootless data-layout
conversion passed an isolated restore rehearsal. The later partial cutover is tracked below.
- [x] Prepare the dedicated Atlas Gitea dataset, non-login UID/GID 1101 with a separate rootless Podman
sub-ID range, and disabled user Quadlet. On 2026-10-01 the targeted Ansible run and a second idempotent
run passed; the generated unit was inactive, with no staging HTTP/SSH listener. POSIX ACLs on only the
service-namespace parents grant this account traversal without access to sibling datasets.
- [x] Perform an isolated rootless restore rehearsal from the verified Prometheus backup. On 2026-10-01
the SHA-256-checked selective extraction and path/SSH conversion succeeded; SQLite `quick_check`
passed, all 33 repositories passed `git fsck`, and source/target public SSH host-key fingerprints
matched. The pinned rootless image answered HTTP and listened on internal SSH/2222 with
`--network none`; the temporary container was removed and the Quadlet stayed inactive. A second
restore run made no changes. This is a rehearsal copy, not the final consistent cutover copy.
- [x] Verify ZFS and Borg coverage of the staged Gitea dataset. On 2026-10-01 the managed recursive
hourly snapshot `atlas-auto-hourly-20261001T193401Z` included it, and the managed incremental
Borg archive `atlas-20261001T193420Z` included its database. A private one-file restore from
each independently matched the staged database and passed SQLite `quick_check`; temporary files
and snapshot mounts were removed, the Borg service ended successfully, and the pool was healthy.
- [x] Include the new Gitea dataset in a UUID-bound offline USB version and test a file restore
before accepting production writes. The operator's 2026-10-01 manual run published version
`20261001T201220Z-254397` successfully on 2026-10-02. Its Gitea database was restored to a
temporary directory from a read-only mount: contents, owner, group, mode, size, mtime and POSIX
ACL matched, and SQLite `quick_check` passed. Temporary files and mounts were removed, LUKS
was closed, and the pool remained healthy. A redundant run was stopped during verification;
its temporary snapshot was cleaned up and the service's resulting failed state was reset.
- [x] Install a separate opt-in final Gitea export helper on Prometheus. Its 2026-10-01 targeted
deployment and `bash -n` passed while Gitea and NPM stayed running. It refuses an active export
timer, stops only Gitea, verifies SQLite, publishes a checksum-verified Gitea-only version for
Atlas' existing pull, and leaves the source stopped on success. It was invoked on 2026-10-02
after the export timer was stopped; version `20261002T071525Z` was pulled and verified on Atlas.
- [x] Prepare the Atlas final-restore gate without replacing the rehearsal: it accepts only a
checksum-verified `gitea-cutover` export, refuses a running target, stages and validates the new
layout before replacing the marked rehearsal, and rolls back a failed swap. Synthetic success
and rollback tests passed on 2026-10-01. On 2026-10-02 the final gate replaced the rehearsal;
SQLite `quick_check`, all 33 repository `git fsck` checks, checksum and SSH host-key comparison passed.
- [x] Start the rootless Atlas Gitea Quadlet and move the primary HTTPS route. On 2026-10-02 Atlas
answered HTTP 200 through the Aegis gateway. NPM stayed on Prometheus; its variable upstream
required a managed Nginx `server_proxy.conf` override because runtime DNS ignores Compose
`extra_hosts`. The primary public HTTPS page and API returned 200, and `git ls-remote` succeeded
for a representative repository after NPM restart; the Navidrome and Syncthing Proxy Hosts also
responded. The source
Gitea container was removed from the desired Compose stack without deleting its data; the
Prometheus backup export timer resumed for NPM only. A post-cutover recursive ZFS snapshot and
encrypted Borg archive `atlas-20261002T073044Z` completed successfully.
- [x] Move the live Gitea Quadlet and dataset from the legacy host `gitea` account to `admin`
after a disposable snapshot-copy test of the pinned derived image. On 2026-10-02 the explicit
outage run stopped only legacy Gitea, made safety snapshot
`zpool/services/data/gitea@gitea-owner-migration-20261002T100104`, changed dataset ownership,
and validated loopback staging (HTTP 200, internal `gitea` UID/GID 1000, SQLite `quick_check`)
before promoting the `admin` Quadlet. Production LAN and public HTTPS returned 200; Navidrome
and Syncthing remained active, the pool was healthy, and the normal Gitea run changed nothing.
The old host account and data on Prometheus remain preserved; the old Atlas Quadlet and its
parent-dataset traverse ACL were removed. A subsequent normal run changed nothing.
- [x] Validate public Gitea SSH/2222 and an authenticated read from Ikaros. After the VPS
firewall was opened on 2026-10-02, TCP/2222 connected, the public ED25519 host-key
fingerprint matched Atlas, Gitea authenticated `fscotto` using the `ikaros` key, and
`git ls-remote` returned HEAD for `fscotto/infra.git` over public SSH.
- [x] Validate authenticated SSH pull and push. On 2026-10-02 the operator reported both
operations working through the public SSH endpoint; the earlier agent-run `git ls-remote`
remains the independent read-only check. The agent did not perform a test push.
- [x] Validate Gitea login and write via HTTPS. On 2026-10-03 the operator confirmed
authenticated web login and Git clone/pull/push through the public HTTPS endpoint. Do not
restart the stale source Gitea after Atlas has accepted writes.
- [ ] Design and deploy Nextcloud as another explicitly temporary Atlas service before Uranus. Give it
separate persistent application, database, and cache storage; keep credentials in Vault; publish it only
through NPM over the Prometheus--Aegis gateway; and define backup, upgrade, and eventual Uranus-migration
procedures before exposing user data. Do not deploy Nextcloud before the data-protection checklist is complete.
- [ ] Keep `atlas_manage_media_stack` disabled until the future Immich deployment has validated `/dev/dri`,
container paths, and the required Vault database secret.
### Priority 4 - Optional workflows
- [x] Deploy the declared Atlas iCloudPD state dataset and inactive rootless `admin` Quadlet.
Photos belong under `/zpool/archive/Pictures/iCloudPD`; private config/MFA state belongs in
`zpool/services/data/icloudpd`. Photobook remains reserved for Immich. Ansible now renders
`icloudpd.conf` with the Apple ID from the existing Vault key, but does not store the password,
manage MFA, or enable automatic startup. The isolated no-network layout test is documented in
`docs/atlas-icloudpd-migration.md`. On 2026-10-02 Atlas deployment and a second idempotent run
passed; no app config existed at deployment. A manual first start on 2026-10-02 generated
`icloudpd.conf`; an Ansible run then replaced it with a private mode-0600 Vault-backed template
and an idempotent second run. The image later expanded the config, so Ansible now seeds it
only when absent and maintains the declared fields. Its launcher requires `traceroute`; the
rootless Quadlet grants only `NET_RAW`, tested in isolation and after restart. The service
was subsequently initialized interactively; initial ingestion is tracked below.
- [x] Retire Aegis iCloudPD completely. The operator authorized deleting its Quadlet,
`/var/lib/icloudpd` data, and MFA state despite an unaudited container overlay. After two
interactive-sudo runs on 2026-10-02, the unit is `not-found`/`inactive`, the Quadlet and state
directory are absent, and AdGuard remains active. The temporary retirement tasks have since
been removed from the Aegis role; it no longer manages iCloudPD.
- [x] Validate Atlas iCloudPD authentication and initial ingestion. On 2026-10-03 the active
rootless service logged `All photos and videos have been downloaded` at 02:16 and reported
completion for the user. The destination held 11,658 files (86,020,430,015 bytes); the preceding 24h
logs showed download activity without authentication failures or errors. A later read-only check
found the service still active. This confirms the initial download, not the next daily cycle.
- [x] Declare HEIC decoding for Fedora graphical desktops without converting the originals on Atlas.
The Fedora role installs RPM Fusion Free with a pinned signing-key fingerprint and
`libheif-freeworld` on Ikaros and Nymph. The package was confirmed installed on Ikaros on
2026-10-03; Nymph deployment and an actual image-opening test were not observed.
- [ ] Validate Atlas iCloudPD filesystem/SELinux/SMB access, the next daily sync, ZFS/Borg/USB
backup inclusion, and isolated restore of photos and private state. A recursive hourly snapshot
of `zpool/archive` exists after ingestion, but no iCloudPD-specific backup version or restore
has been verified. The first monthly scrub remains a separate open data-protection check.
## Prometheus NPM Quadlet cutover
- [x] Stage a rootful NPM Quadlet using the exact running image and the existing data/certificate
mounts, bridge subnet, public HTTP/HTTPS ports, and loopback-only administration port.
The generated service depends on `server-web-network.service` and is wanted by `multi-user.target`.
- [x] Take and verify the stopped-source export before switching owners. Version
`20261003T091009Z` was pulled to Atlas and its NPM SQLite database checked in isolation.
- [x] Cut over NPM to `prometheus-npm.service` on 2026-10-03. The legacy Compose unit is inactive
and disabled; the Quadlet is active with zero recorded restarts. Public Gitea and Syncthing
HTTPS returned 200 with valid TLS, while public TCP/81 remained unreachable.
- [x] Validate the post-cutover backup path. The export and Atlas pull published
`20261003T091633Z`; checksum, SQLite `quick_check`, ten proxy hosts, six certificate records,
both Quadlet files were present, and the complete Let's Encrypt tree (70 regular files plus
12 symlinks) matched the live data. A targeted normal Ansible run changed nothing. Details and rollback
boundaries are in `docs/prometheus-npm-quadlet.md`.
- [ ] Observe the first scheduled export and Atlas pull after the cutover; the manual end-to-end
cycle passed, but the next unattended cycle has not yet occurred.
## Cerberus Management Node (Deferred)
`cerberus` is postponed until the office in the new house is physically set up. It is not an inventory
host and this section is a design and implementation backlog, not authorization to provision it early.
The planned node is a Lenovo ThinkCentre M700 Tiny with an Intel Core i3-6100T, 8 GB RAM, a 256 GB SSD,
and native 1 Gbps Ethernet. It will connect to a multi-input KVM switch using a passive DisplayPort-to-HDMI
cable, sharing the monitor and peripherals with Ikaros. Fedora Sericea (immutable Fedora with the Sway
Wayland compositor) is the intended OS. Cerberus is an isolated management plane: a dedicated Toolbox
environment will run Ansible for future `uranus` cluster provisioning. Rootless Podman will host Grafana,
Prometheus, and Loki. The 256 GB local SSD is the hot tier retaining metrics and logs for 30 days; scheduled,
validated exports of older historical data will use a dedicated Atlas NFS dataset as cold storage.
### Implementation plan
- [ ] Confirm the office, KVM switch, passive DisplayPort-to-HDMI path, shared monitor/peripherals, and native
1 Gbps Ethernet are physically operational before adding Cerberus to inventory.
- [ ] Install and update Fedora Sericea with Sway; document the immutable-host lifecycle and keep host changes
declarative rather than treating the base OS as a mutable workstation.
- [ ] Model Cerberus as its own host with independent platform, role, desktop, network, and storage inputs;
do not repurpose Ikaros variables or make it a Uranus cluster member.
- [ ] Provision an isolated Toolbox-based Ansible controller with the required collections and a reproducible
project checkout; define its least-privilege SSH access, known-host handling, and Vault workflow without
storing secrets in the image or repository.
- [ ] Define the explicit Uranus provisioning workflow from Cerberus, including inventory boundaries,
validation-only runs, and separate approval for any destructive cluster operation.
- [ ] Design rootless Podman/Quadlet services for Grafana, Prometheus, and Loki, including persistent local
state, service ownership, LAN exposure/authentication, resource limits, updates, and backups.
- [ ] Size and enforce a 30-day local hot-retention policy for metrics and logs on the 256 GB SSD; validate
actual disk growth and alert before capacity exhaustion.
- [ ] Create and validate a dedicated Atlas NFS cold-storage dataset and least-privilege export for Cerberus;
do not use a broad existing share or couple it to unrelated Atlas application state.
- [ ] Implement scheduled, idempotent exports of data older than 30 days to the Atlas NFS cold tier, with
locking, capacity checks, integrity verification, retention rules, failure monitoring, and a tested restore.
- [ ] Validate management-plane recovery: rebuild Cerberus, restore observability history from Atlas, and
confirm that Uranus provisioning can resume without depending on unreproducible local state.
## Coding Agent Notes
- Shared agent definitions and lifecycle flags live in `ai_agents` in `ansible/inventory/group_vars/all.yml`.
- Shared agent dotfiles live in `ai_agents_dotfiles`; rendered configs live in `ai_agents_templates`.
- Every `ai_agents.<agent>` entry has independent `install_enabled`, `deploy_enabled`, and `uninstall_enabled` flags. Installation and removal must not both be true for the same agent; the common pre-task fails before changes when they conflict.
- Fedora, Void desktop, and WSL workstation profiles consume the shared agent definitions; do not duplicate package entries in profile-specific vars. IBM Bob on the workstation follows its own flags.
- `dotfiles_common` deploys `ai_agents_dotfiles` and renders `ai_agents_templates` only when deployment is enabled.
- Removal is limited to the managed npm packages and `/usr/local/bin/bob`; never remove agent dotfiles, instructions, credentials, or user data.
- Keep `.config/ai/` as the common instruction source; update agent-specific entrypoints to reference it rather than duplicating instruction text.
## Tooling Notes
- Install local tooling with:
- `python3 -m pip install ansible ansible-lint yamllint shellcheck-py`
- `ansible-galaxy collection install -r ansible/collections/requirements.yml`
- Required collections currently include `ansible.posix` and `community.general`.
- `.yamllint` treats `line-length` as a warning at 120 chars and disables `document-start` and `comments-indentation`.
## When Updating Docs
- Keep `README.md` and `AGENTS.md` aligned when workflows materially change.
- If you add a new operational area, also add the narrowest validation command for it.
- Call out checks you could not run and any follow-up verification needed.
## Aegis Fedora IoT Notes
- `aegis` is a remote Fedora IoT Raspberry Pi 4 node. Bootstrap it once with
`ansible/bootstrap/aegis.bu`; the remaining configuration is applied by `profile_aegis` over SSH.
- Fedora IoT is immutable. Do not add it to mutable Fedora package or shared dotfile roles.
- `profile_aegis` owns the `nfs-utils` and `wireguard-tools` rpm-ostree layers and reports the required reboot
without initiating it. `wireguard_overlay` then configures Aegis as the WireGuard LAN gateway with persistent IPv4
forwarding, a scoped inter-zone policy, and source masquerading. It also owns rootful Podman Quadlets, persistent container
state under `/var/lib`, the Podman auto-update timer, LAN-restricted firewalld rules, and SSH hardening. Keep
`aegis_lan_subnet`, `aegis_adguard_web_port`, and `aegis_network_connection_uuid` host-specific;
SSH permits only the declared
key-authenticated users, never root or password authentication. Keep Apple IDs and other
credentials in Vault and use `no_log` for their rendering.
- `aegis_adguard_web_port` defaults to `80`. The initial AdGuard Home wizard port `3000` is intentionally unmanaged: open and close it manually only while
completing initial setup. Disable the local systemd-resolved stub through `profile_aegis` before
AdGuard binds port 53; keep
`/etc/resolv.conf` linked to `/run/systemd/resolve/resolv.conf`. LAN clients may use AdGuard, but
Aegis must use the independent upstream DNS declared by `aegis_host_dns_servers` so Greenboot does
not depend on the AdGuard container during startup.
- Aegis iCloudPD has been retired and is no longer managed by this role. Its service, Quadlet,
data, and MFA state were removed with the operator's explicit authorization.