47 KiB
AGENTS.md
Ansible-driven personal infrastructure repo for Fedora and Void desktops, Fedora IoT, WSL, a Rocky Linux 9 server, and an Atlas NAS.
Source Of Truth
- Main orchestration:
ansible/site.yml - Inventory and layering inputs:
ansible/inventory/hosts.yml,ansible/inventory/group_vars/*.yml,ansible/inventory/host_vars/*.yml - Dotfiles live under
dotfiles/ - AI agent instructions (bootstrap, rules, knowledge) are centralized in
dotfiles/common/.config/ai/and shared between OpenCode, Codex, and Gemini CLI. - OpenCode loads its entrypoint configuration from
dotfiles/common/.config/opencode/opencode.json. - Codex config is rendered from
dotfiles/common/.codex/config.toml.j2somodel_instructions_filepoints to the deployed~/.config/ai/bootstrap.md.
Topology
- Current personal desktop:
ikaros = platform_fedora + role_personal_workstation + graphical_desktop + desktop_gnome - Current laptop:
nymph = platform_fedora + graphical_desktop + desktop_gnome - Void desktop profile is also the base for other future/reference hosts via
platform_void + graphical_desktop - Workstation:
deadalusis Windows + Fedora WSL. - Rocky server:
prometheusbelongs torocky_server. - NAS:
atlas(Rocky Linux 9, reached through SSH) - Always-on LAN node:
aegis(Fedora IoT on Raspberry Pi 4, reached through SSH) - Hosts intentionally belong to multiple groups; trust
ansible/site.ymlover hostname assumptions. - Inventory axes are independent:
platform_*,role_*, anddesktop_*. Legacyvoidanddesktopremain compatibility parents.
Working Rules
- Preserve layering
all -> platform -> role -> desktop -> host. - Keep
ansible/site.ymlsmall; orchestration belongs there, implementation belongs in roles. - Prefer minimal, targeted edits. Preserve idempotency and existing ordering.
- Keep completed one-time cleanup operations out of the playbook. Execute them directly with explicit authorization; retain only the ongoing desired-state configuration and historical documentation, not permanent cleanup flags or tasks.
- Use Git Flow branch prefixes:
feature/for new functionality,bugfix/for non-urgent fixes,hotfix/for urgent production fixes,release/for release preparation, andsupport/for maintained release lines. Do not use abbreviated prefixes such asfeat/. - Desktop and WSL hosts use
ansible_connection: local; remote infrastructure hosts use SSH. - Treat
secrets/as sensitive. Never print secret values. - Tmux plugins are bootstrapped by TPM on the host; the repo only keeps tmux config and custom helper scripts.
- Read the relevant role tasks, templates, vars, and deployed dotfiles before editing.
Validation
- Default minimum:
ansible-playbook ansible/site.yml --syntax-check
- Repo-wide checks:
ansible-lint ansible/site.ymlansible-lint ansible/rolesyamllint ansible/
- Host-focused dry runs:
- Fedora desktop work:
ansible-playbook ansible/site.yml --limit ikaros --check --diff - Fedora laptop work:
ansible-playbook ansible/site.yml --limit nymph --check --diff - WSL workstation dev:
ansible-playbook ansible/site.yml --limit deadalus --check --diff - Server:
ansible-playbook ansible/site.yml --limit prometheus --check --diff - Rocky server after activation:
ansible-playbook ansible/site.yml --limit <host> --check --diff - Atlas NAS:
ansible-playbook ansible/site.yml --limit atlas --check --diff - Aegis IoT:
ansible-playbook ansible/site.yml --limit aegis --check --diff - Aegis NFS client layer:
ansible-playbook ansible/site.yml --limit aegis --tags nfs --list-tasks - Aegis host DNS:
ansible-playbook ansible/site.yml --limit aegis --tags dns --check --diff
- Fedora desktop work:
- Focused checks:
- Emacs is disabled by default; temporary Emacs check:
ansible-playbook ansible/site.yml --limit <host> --tags emacs --check --diff -e emacs_enabled=true - AI coding agents:
ansible-playbook ansible/site.yml --limit <host> --tags ai_agents --check --diff - Mail bootstrap:
sh -n scripts/bootstrap_mail.shandshellcheck scripts/bootstrap_mail.sh - Server NPM Quadlet:
systemctl status prometheus-npm.service; the Compose fallback is retired. - Explicit Prometheus legacy cleanup (destructive only without check mode):
ansible-playbook ansible/site.yml --limit prometheus --tags server_legacy_cleanup --check --diff -e server_legacy_cleanup=true - Atlas media stack:
ansible-playbook ansible/site.yml --limit atlas --tags storage,sharing,containers --check --diff - Atlas rootless Gitea staging (does not start Gitea):
ansible-playbook ansible/site.yml --limit atlas --tags gitea --check --diff - Atlas canonical Gitea domain (restarts only Gitea on a real configuration change):
ansible-playbook ansible/site.yml --limit atlas --tags gitea_public_domain --check --diff - Atlas iCloudPD storage and boot-started Quadlet:
ansible-playbook ansible/site.yml --limit atlas --tags icloudpd --check --diff - Atlas explicit Gitea host-owner migration (live outage; never a normal run):
ansible-playbook ansible/site.yml --limit atlas --tags gitea_owner_migration -e atlas_gitea_owner_migration=true - Atlas explicit isolated Gitea restore rehearsal (not part of normal runs):
ansible-playbook ansible/site.yml --limit atlas --tags gitea_restore -e atlas_gitea_restore_test=true - Atlas final Gitea replacement gate (dry-run only until a stopped-source export is pulled):
ansible-playbook ansible/site.yml --limit atlas --tags gitea_final_restore --check --diff -e atlas_gitea_final_restore=true - Prometheus final Gitea export helper (dry-run installs only; outage action remains opt-in):
ansible-playbook ansible/site.yml --limit prometheus --tags gitea_final_export --check --diff - Gitea cutover network configuration before activation:
ansible-playbook ansible/site.yml --limit prometheus --tags gitea_cutover,prometheus_backup --check --diff -e server_gitea_on_atlas=trueandansible-playbook ansible/site.yml --limit atlas --tags gitea --check --diff - Atlas daily Navidrome music copy:
ansible-playbook ansible/site.yml --limit atlas --tags music_sync --check --diff - Atlas network/share hardening:
ansible-playbook ansible/site.yml --limit atlas --tags hardening,sharing --check --diff - Atlas ZFS snapshot retention and scrub timers:
ansible-playbook ansible/site.yml --limit atlas --tags snapshots,scrub --check --diff - Atlas encrypted Borg backup:
ansible-playbook ansible/site.yml --limit atlas --tags packages,borg --check --diff - Atlas Borg progress logging only:
ansible-playbook ansible/site.yml --limit atlas --tags borg_logging --check --diff - Atlas manual offline USB backup and 45Drives Alerts reminder:
ansible-playbook ansible/site.yml --limit atlas --tags usb_backup,usb_reminder --check --diff - Atlas pool, disk, capacity, temperature, and job monitoring:
ansible-playbook ansible/site.yml --limit atlas --tags monitoring --check --diff - Atlas explicit post-restore SELinux relabeling:
ansible-playbook ansible/site.yml --limit atlas --tags restorecon --check -e '{"atlas_restorecon_paths":["/zpool/archive"]}' - Prometheus/Aegis WireGuard gateway:
ansible-playbook ansible/site.yml --limit prometheus,aegis --tags wireguard --check --diff - Prometheus NPM Quadlet steady state (does not perform a cutover):
ansible-playbook ansible/site.yml --limit prometheus --tags npm_quadlet --check --diff - DuckDNS config only (skipped on Prometheus):
ansible-playbook ansible/site.yml --limit prometheus --tags duckdns --check --diff
- Emacs is disabled by default; temporary Emacs check:
Conventions
- Use FQCN Ansible modules.
- Prefer declarative modules over
command/shell; whenshellis required, make idempotency and failure behavior explicit. - Start YAML files with
---, use 2-space indentation, and keep file modes quoted like"0644". - Keep booleans as booleans and structured vars as YAML lists/maps.
- Put host-specific overrides in
host_vars, not sharedgroup_vars. - Use
no_log: truefor secret-bearing task inputs or outputs.
Desktop Notes
desktop_profilenames independently selectable desktop groups such asdesktop_gnome,desktop_sway, anddesktop_niri. Keep platform-specific session bootstrap in platform-specific roles.desktop_environmentis fixed tominimalfor Void desktops.profile_desktop_commonowns shared Void bootstrap;profile_desktop_swayandprofile_desktop_nirimanage the enabled sessions, whileprofile_desktop_gnomecopies shared desktop dotfiles for Fedora/GNOME without managing GNOME settings.desktop_sessions_enabledanddesktop_default_sessionapply to the minimal mode.- Emacs has one authoring-oriented
.emacs.d, deployed bydotfiles_commonwhenemacs_enabledis true. Fedora/GNOME desktops and workstation profiles enable it; keep platform dependencies in package group vars rather than branching in Emacs Lisp. - NTFS filesystem support is provided by
ntfs-3ginansible/inventory/group_vars/void.yml. - Void user services are managed by
turnstileand live underdotfiles/desktop/.config/service/. ssh-agentkeeps the stable socket~/.local/state/ssh-agent/socket.- Critical session entrypoints:
dotfiles/desktop/.config/sway/configplushost.confandsession-envdeployed viahost_sway_dotfiles(sway / Wayland)dotfiles/desktop/.config/niri/config.kdlandsession-envdeployed viadesktop_niri_dotfiles(Niri / Wayland)
- Void Niri lives in
profile_desktop_niri, gated on'niri' in desktop_sessions_enabled; it installs theempttyniri.desktopsession, the/usr/local/bin/start-nirilauncher, and the xdg-desktop-portal config, mirroringprofile_desktop_sway. - Fedora GNOME (
desktop_gnome) assumes GNOME comes from the Fedora Workstation base install; Ansible deploys shared desktop dotfiles and git/GPG config forikarosandnymph, not GNOME settings. - Do not switch or restart the display manager during a playbook run from an active graphical session.
nymphis the Fedora/GNOME laptop target; keep GNOME settings unmanaged for now and add host-specific tuning only after real use.
Void Package And Dotfile Bucket Rules
platform_void is the reusable Void platform selection. The legacy void group remains a compatibility parent so existing group_vars/void.yml and when: "'void' in group_names" checks keep working during the transition.
The Void desktop package lists in ansible/inventory/group_vars/void.yml are kept disjoint by role:
void_packages_base— system runtime only (init/services, kernel, audio core, networking, filesystem, firewall, hardware daemons, runit logging).desktop_common_packages— GUI infrastructure shared by the minimal desktop mode.desktop_minimal_packages— applications, integration components, and theempttydisplay manager.desktop_sway_packages— binaries specific to the Sway session.profile_packagesremains the shared package bucket for Void and Fedora profiles. Rocky usesrocky_profile_packagesso RPM-specific names do not leak back into the other platforms; do not move desktop-specific Void entries through either bucket. The dotfile vars follow the same split:desktop_common_dotfilescarries mode-independent content anddesktop_minimal_dotfilescarries Thunar, Udiskie, and MIME defaults.desktop_void_dotfilesremains reserved for files that need the Void runtime.
Workstation Notes
deadalusis modeled as Windows + Fedora WSL and is the sole workstation target.- Fedora WSL belongs to
platform_fedora,workstation_dev_fedora, and the shared WSL layer. It must not receive Flatpak or Snap runtimes. - Fedora WSL installs Mise from the official
jdxcode/miseCOPR and uses its pinned Temurin Java 11 JDK; update the declared Mise version deliberately. - Windows applications are installed manually and are not managed from the WSL profile.
Rocky Server Notes
- Prometheus disables DuckDNS provisioning with
server_duckdns_enabled: false. Its updater, log and five-minute cron entry were explicitly retired; the external DuckDNS name and Vault token remain untouched. The completed one-time cleanup has no remaining playbook tasks. - When enabled, DuckDNS is rendered by
profile_serverfrom host-localserver_duckdns_domainandvault_duckdns_token. Keep the rotated token in encrypted Vault or untracked local vars, never in dotfiles. The private~/duckdns/duck.shkeeps the existing entrypoint; rendering usesno_logand disables diffs. Provisioning does not execute the updater or change its external schedule. rocky_serveris a child of bothplatform_rockyandserver;prometheusis its active target.- The target must already provide
server_usernamewith local sudo access before the profile runs. - The Rocky profile installs Podman and podman-compose. Prometheus explicitly retires the legacy
Compose unit, files and final-export helper with
server_legacy_stack_retired: true. Its approved opt-in cleanup removed old application data on 2026-10-03; normal runs do not delete data or recreate the retired files. On Prometheus, Nginx Proxy Manager is now the rootfulprometheus-npm.serviceQuadlet with a pinned image digest and the existing/opt/npm/dataand/opt/npm/letsencryptbind mounts. The rootfulserver_webbridge remains10.89.0.0/24. Gitea runs on Atlas; PostgreSQL and Navidrome are absent from the desired Prometheus stack. Normal runs do not delete legacy data, update DNS, or perform an implicit cutover; destructive cleanup requires its explicit tag and opt-in extra-var. - Firewalld enables SSH, Cockpit (
9090/tcp), HTTP and HTTPS. Nginx Proxy Manager publishes80/tcpand443/tcp; bind its administration interface only to127.0.0.1:81and usenpm-tunnelfrom Ikaros or Nymph. Nextcloud remains disabled; do not provision/srv/nextclouddirectories. scripts/migrate_prometheus_data.shis the separate, source-host-run NPM/Gitea migration path. It dry-runs by default and requires explicit source-stack quiescing before copying persistent Docker data with rsync.- Atlas-only OpenZFS, NFS, Samba, and Syncthing stay selected through Atlas host variables and must not
leak into
rocky_server. Cockpit plus its Navigator and Podman extensions are selected explicitly for Prometheus through its host variables.
Atlas NAS Notes
atlasis a remote Rocky Linux 9 NAS. Keep its connection, LAN, pool and mountpoint values inhost_vars/atlas.yml. Bootstrap it once with-e atlas_connection_username=<existing-admin>; subsequent runs use the dedicated Atlas account.- The pool is normally pre-existing. A one-time bootstrap may create it only when
atlas_create_pool=trueis explicitly supplied andatlas_zpool_diskscontains exactly four real/dev/disk/by-id/...paths. Never partition, force, destroy, roll back, or modify the vdev layout of an existing pool. atlas_manage_storage,atlas_manage_sharing, andatlas_manage_firewallare enabled in Atlas host vars as the declared steady state; set one false only for a deliberate suspension.atlas_manage_media_stackremains false until the future rootful Immich stack has its required Vault inputs and target validation.- Atlas requires
vault_atlas_admin_password_hashfor Cockpit and, while sharing is enabled,vault_atlas_samba_password. The future rootful media stack also requiresvault_atlas_immich_db_password. Never print these values. - Atlas creates the complete declared hierarchy only under the verified existing or explicitly bootstrapped pool:
archive,services,services/data,services/data/navidrome,services/data/syncthing,media,media/music,media/photobook,backup,backup/hosts, andbackup/hosts/prometheus.backuphas a500Greservation covering its descendants.archiveis the SMB-shared raw-data namespace; container state is never beneath it. - The
immichsystem account is fixed to UID/GID1100, has no login shell orwheelmembership, and receives only thevideoandrendersupplementary groups. Immich's rootful Quadlets run as1100:1100; Server and ML receive/dev/dri, while the Photobook external library is read-only at/external/photobook. - Atlas applies persistent kernel network hardening: redirects and source routes are rejected, martians logged, reverse-path filtering remains loose for WireGuard, and IPv4 forwarding is disabled. SSH permits only the declared administrator using public-key authentication; root login, passwords, agent and remote forwarding
are disabled, while local forwarding remains available for private administrative tunnels. Photobook is exported only to the configured Aegis IP with all access squashed to UID/GID
1100. Targeted SELinux is enforced persistently; a required reboot is reported but never initiated automatically. The primary LAN interface is assigned explicitly to the managed firewalld zone, and firewall rules are applied before NFS or SMB are started; their service state and TCP listeners are then verified. SMB3 exposesArchiveto Vault-backed authorized accounts on mandatory encrypted, signed SMB3 over TCP/445 only and admits the configured LAN without host-specific exclusions. - Atlas NPM and Immich share a rootful Podman network. NPM publishes HTTP/HTTPS, but its administration port remains
bound to
127.0.0.1:81; do not expose it directly to the LAN or Internet. profile_backend_phase1temporarily runs rootless Navidrome and Syncthing on Atlas until Uranus replaces them. It binds only to Atlas' LAN IP, neverwg0; Navidrome and the Syncthing GUI admit only Aegis as the source-NAT gateway, while native Syncthing ports admit the configured LAN. It initializes fresh state only and never migrates or deletes source application data. The enabled rootlessatlas-music-sync.timercopies/zpool/archive/Musicto/zpool/media/musicdaily at 00:45 Europe/Rome without deleting destination files; it requires both datasets to be mounted.wireguard_overlaymanageswg0between Prometheus (10.0.0.1) and Aegis (10.0.0.2). It persists private keys only on their respective hosts, exchanges only derived public keys through Ansible, and verifies a real peer handshake. Prometheus opens51820/udp; Aegis is the LAN gateway. Its persistent IPv4 forwarding, narrowly scoped WireGuard-to-LAN firewalld policy, and source masquerading permit Prometheus to reach LAN services without a static route on the router. Prometheus includes192.168.178.0/24in Aegis' peerAllowedIPs; add the Uranus VIP there when it is assigned. After a firewalld reload, restore Prometheus' rootful Podman networking withpodman network reload --allso the existing proxy stack retains container DNS.
Atlas NAS TODO
Completed validation: the existing RAIDZ2 pool and datasets, SELinux, LAN firewall, SSH, Cockpit with
the selected 45Drives plugins, encrypted SMB3 Archive, the Aegis-only NFSv4 photobook export, and the
Prometheus--Aegis WireGuard gateway are operational. The gateway handshake, forwarding, source masquerading,
and TCP reachability to Atlas were verified. Temporary Navidrome and Syncthing are available through their
manual NPM Proxy Hosts; Syncthing uses /data/Org backed by the SMB-shared Archive dataset. Aegis has also
validated NFSv4.2 read, write, delete, and all_squash mapping to UID/GID 1100 end-to-end. The ZFS
snapshot timers are active; a recursive hourly snapshot and scheduled retention prune completed
successfully. The first monthly scrub remains a runtime check.
Priority 1 - Data protection
- Deploy Ansible-managed recursive ZFS snapshots with 24 hourly, 30 daily, 8 weekly, and 12 monthly generations, plus a monthly scrub on the first Sunday at 03:00. The timers and first hourly snapshot were verified on Atlas. Cockpit Scheduler is for visibility or manual operations only, and snapshot rollback is never automated.
- Verify the first monthly ZFS scrub from its actual service result. Scheduled retention pruning was observed on 2026-09-30; timer activation alone does not establish a successful scrub.
- Activate and validate the encrypted offsite Borg backup to the Hetzner Storage Box. Atlas uses the
dedicated SSH identity, pinned ED25519 host key, Vault-backed
repokeyencryption, and a locked non-loginborgaccount with no sudo or supplementary groups. The initial snapshot-consistent backup, Borg repository check, and temporary-directory restore completed successfully; the restoredArchivetree matched the live data, and temporary snapshots and mounts were removed. The exported recovery key was copied offline. Daily backup retries and logging, 30 daily, 8 weekly and 12 monthly archives, compaction, and monthly repository checks are enabled. Runs report a ZFS-based estimated percentage. On 2026-09-30 a successful incremental run also removed the stale 2026-09-29 snapshot and its own temporary snapshot after exit; the earlierRuntimeDirectorycleanup failure is resolved. - Populate
/zpool/archivewith the currently available data so offsite and offline backup tests run against a representative load. - Evaluate Borg against the populated pool. The 2026-09-29 archive took 1 h 32 min for 2.18 TB original / 2.04 TB compressed data, with 13.49 GB deduplicated size; retention and compaction succeeded. On 2026-09-30 a subsequent incremental archive completed in about 22 seconds with successful cleanup. The monitor reported 37% Storage Box quota used. These are observed runs, not a guarantee of future duration or compression ratio.
- Add the UUID-bound offline USB backup with versioned rsync, locking, capacity checks, verification,
safe unmounting and a tested restore procedure; never trigger it for an arbitrary USB disk. The
LUKS/ext4 identities were read-only verified; the manual service and 45Drives Alerts reminder timer were
deployed on Atlas. Interactive LUKS unlock is part of the manual service; only the reminder is
scheduled for the first Saturday of each month at 10:00 Europe/Rome via the existing 45Drives
notifier. A manual test produced an Alerts notification, not an email. The first USB attempt failed
on a
security.selinuxxattr and was interrupted; the xattr filter is deployed and the temporary recursive snapshot and open LUKS mapper were cleaned up. A later run reported checksum verification and published the USB version, but failed while removing host-namespace ZFS snapshot mounts. Those exact mounts and snapshots were cleaned up. AnExecStopPosthelper now removes only the named temporary snapshot after the backup process exits. A new full run checksum-verified and published a USB version; the service ended successfully, the mapper closed, no temporary USB snapshot remained, and the pool was healthy. On 2026-09-25 an independent, read-only USB restore test copied one file from the publishedatlas/latestversion into/var/tmpand matched contents, owner, mode, size, mtime and POSIX ACL. The temporary copy and mount were removed, the mapper closed, and the pool remained healthy. - Test restores independently from a ZFS snapshot, Borg, and the offline USB backup before relying on
any backup path. The earlier Borg temporary-directory restore passed. On 2026-09-25 a separate,
read-only ZFS snapshot test restored one file to
/var/tmp, confirmed matching contents, ownership, mode, mtime and ACL, then removed its temporary copy and on-demand mount. This is a file-level smoke test, not full dataset recovery. An independent USB file restore passed on 2026-09-25 with matching content and metadata; the later scaled OS-rebuild rehearsal is documented under Priority 2. - Add monitoring and alerting for pool health, scrub/resilver, SMART data, temperatures, free space,
snapshot/local-backup growth, Hetzner Storage Box quota, and failed maintenance or backup timers.
The half-hourly Atlas health monitor and systemd final-failure hooks are deployed. A live probe
found no issues; the service and timer succeeded, and a labelled 45Drives Alerts test notification
was submitted. Alerts are deduplicated; email delivery is not claimed. The Storage Box quota probe
runs
df -mover the dedicated pinned-key SSH identity and does not open the Borg repository. The failed-job hook was corrected to pass the literal systemd unit name; its expansion was verified on Atlas, but a new real failure notification has not been deliberately triggered.
Priority 2 - NAS operability and recovery
- Document and test disaster recovery in
docs/atlas-recovery.md: the operator confirmed Vault and Borg recovery material is available offline; provisional targets are RPO 24h/RTO 72h. On 2026-09-30 an isolated small Rocky VM was rebuilt with the Atlas Ansible roles, imported its preserved RAIDZ2 pool without force/rewind, and restored a file from the preserved snapshot; the second Ansible run was idempotent. Earlier independent production ZFS, USB, and Borg file restore tests remain separate evidence. A production-size full restore, unclean import, and measured 24h/72h compliance are not claimed. - Define a controlled Rocky kernel/OpenZFS update and reboot procedure in
docs/atlas-updates.md. The first real change-window execution is not yet validated; the procedure never reboots automatically or upgrades pool features. - Add the Atlas-initiated least-privilege Prometheus backup pull: Prometheus exposes only prepared
read-only dumps through a dedicated account and Atlas retains the private SSH key, pinned host key,
atomic pull, verification, retention and systemd service/timer. The dedicated key/account and unit
files are deployed; live read-only SSH, shell denial, and write denial were verified. On 2026-09-30
a manual export, Atlas pull, checksum verification, and temporary restore passed; both SQLite
databases passed integrity checks and a restored Git repository passed
git fsck. Both daily timers are enabled for 02:00/03:00 Europe/Rome. On 2026-10-01 their first scheduled export and pull succeeded: Atlas verified the payload checksum and published20261001T000001Zaslatest. - Decide whether a common SMB/NFS namespace is required: no.
Archive(SMB) andphotobook(NFS) remain intentionally distinct;docs/atlas-sharing-decision.mdrecords the decision. No ACL or export change is authorized by this decision.
Priority 3 - Service expansion
- Populate
/zpool/media/musicand validate Navidrome. On 2026-09-30, 21,158 files (93,937,810,350 regular-file bytes) were copied from/zpool/archive/Musicusing a temporary ZFS snapshot; a checksum-based rsync dry run found no differences or extra files. Navidrome saw all files through its read-only mount, completed a scan, indexed 18,168 tracks, and responded over HTTP. Some imported playlists still reference obsolete Windows paths. The source was left intact and the temporary snapshot was removed. - Schedule a daily, non-deleting copy from
Archive/Musicto the separate Navidrome music dataset. The rootlessatlas-music-sync.timeris enabled for 00:45 Europe/Rome; a manual idempotent service run succeeded on 2026-10-01. The first scheduled run triggered at 00:45 CEST on 2026-10-02 and exited successfully (Result=success, status 0); the next run is scheduled for 2026-10-03 00:45 CEST. - Design the staged Prometheus-to-Atlas Gitea migration in
docs/atlas-gitea-migration.md. The approved topology keeps NPM on Prometheus and moves HTTPS and public SSH (TCP/2222) together; Gitea runs as anadmin-owned rootless user Quadlet on Atlas with an internalgiteauser. The rootful-to-rootless data-layout conversion passed an isolated restore rehearsal. The later partial cutover is tracked below. - Prepare the dedicated Atlas Gitea dataset, non-login UID/GID 1101 with a separate rootless Podman sub-ID range, and disabled user Quadlet. On 2026-10-01 the targeted Ansible run and a second idempotent run passed; the generated unit was inactive, with no staging HTTP/SSH listener. POSIX ACLs on only the service-namespace parents grant this account traversal without access to sibling datasets.
- Perform an isolated rootless restore rehearsal from the verified Prometheus backup. On 2026-10-01
the SHA-256-checked selective extraction and path/SSH conversion succeeded; SQLite
quick_checkpassed, all 33 repositories passedgit fsck, and source/target public SSH host-key fingerprints matched. The pinned rootless image answered HTTP and listened on internal SSH/2222 with--network none; the temporary container was removed and the Quadlet stayed inactive. A second restore run made no changes. This is a rehearsal copy, not the final consistent cutover copy. - Verify ZFS and Borg coverage of the staged Gitea dataset. On 2026-10-01 the managed recursive
hourly snapshot
atlas-auto-hourly-20261001T193401Zincluded it, and the managed incremental Borg archiveatlas-20261001T193420Zincluded its database. A private one-file restore from each independently matched the staged database and passed SQLitequick_check; temporary files and snapshot mounts were removed, the Borg service ended successfully, and the pool was healthy. - Include the new Gitea dataset in a UUID-bound offline USB version and test a file restore
before accepting production writes. The operator's 2026-10-01 manual run published version
20261001T201220Z-254397successfully on 2026-10-02. Its Gitea database was restored to a temporary directory from a read-only mount: contents, owner, group, mode, size, mtime and POSIX ACL matched, and SQLitequick_checkpassed. Temporary files and mounts were removed, LUKS was closed, and the pool remained healthy. A redundant run was stopped during verification; its temporary snapshot was cleaned up and the service's resulting failed state was reset. - Install a separate opt-in final Gitea export helper on Prometheus. Its 2026-10-01 targeted
deployment and
bash -npassed while Gitea and NPM stayed running. It refuses an active export timer, stops only Gitea, verifies SQLite, publishes a checksum-verified Gitea-only version for Atlas' existing pull, and leaves the source stopped on success. It was invoked on 2026-10-02 after the export timer was stopped; version20261002T071525Zwas pulled and verified on Atlas. - Prepare the Atlas final-restore gate without replacing the rehearsal: it accepts only a
checksum-verified
gitea-cutoverexport, refuses a running target, stages and validates the new layout before replacing the marked rehearsal, and rolls back a failed swap. Synthetic success and rollback tests passed on 2026-10-01. On 2026-10-02 the final gate replaced the rehearsal; SQLitequick_check, all 33 repositorygit fsckchecks, checksum and SSH host-key comparison passed. - Start the rootless Atlas Gitea Quadlet and move the primary HTTPS route. On 2026-10-02 Atlas
answered HTTP 200 through the Aegis gateway. NPM stayed on Prometheus; its variable upstream
required a managed Nginx
server_proxy.confoverride because runtime DNS ignores Composeextra_hosts. The primary public HTTPS page and API returned 200, andgit ls-remotesucceeded for a representative repository after NPM restart; the Navidrome and Syncthing Proxy Hosts also responded. The source Gitea container was removed from the desired Compose stack without deleting its data; the Prometheus backup export timer resumed for NPM only. A post-cutover recursive ZFS snapshot and encrypted Borg archiveatlas-20261002T073044Zcompleted successfully. - Move the live Gitea Quadlet and dataset from the legacy host
giteaaccount toadminafter a disposable snapshot-copy test of the pinned derived image. On 2026-10-02 the explicit outage run stopped only legacy Gitea, made safety snapshotzpool/services/data/gitea@gitea-owner-migration-20261002T100104, changed dataset ownership, and validated loopback staging (HTTP 200, internalgiteaUID/GID 1000, SQLitequick_check) before promoting theadminQuadlet. Production LAN and public HTTPS returned 200; Navidrome and Syncthing remained active, the pool was healthy, and the normal Gitea run changed nothing. The old host account and data on Prometheus remain preserved; the old Atlas Quadlet and its parent-dataset traverse ACL were removed. A subsequent normal run changed nothing. - Validate public Gitea SSH/2222 and an authenticated read from Ikaros. After the VPS
firewall was opened on 2026-10-02, TCP/2222 connected, the public ED25519 host-key
fingerprint matched Atlas, Gitea authenticated
fscottousing theikaroskey, andgit ls-remotereturned HEAD forfscotto/infra.gitover public SSH. - Validate authenticated SSH pull and push. On 2026-10-02 the operator reported both
operations working through the public SSH endpoint; the earlier agent-run
git ls-remoteremains the independent read-only check. The agent did not perform a test push. - Validate Gitea login and write via HTTPS. On 2026-10-03 the operator confirmed authenticated web login and Git clone/pull/push through the public HTTPS endpoint. Do not restart the stale source Gitea after Atlas has accepted writes.
- Design and deploy Nextcloud as another explicitly temporary Atlas service before Uranus. Give it separate persistent application, database, and cache storage; keep credentials in Vault; publish it only through NPM over the Prometheus--Aegis gateway; and define backup, upgrade, and eventual Uranus-migration procedures before exposing user data. Do not deploy Nextcloud before the data-protection checklist is complete.
- Move Gitea canonical HTTPS and SSH hostname to
git.fscotto.coon 2026-10-03 through Ansible. Only Gitea restarted; second run changed nothing. HTTPS and authenticated SSH reads returned the same repository HEAD. The new NPM hostnames passed TLS/HTTP checks; old DuckDNS Proxy Hosts were observed disabled. Details are indocs/domain-fscotto-co.md. - Confirm login on the new Gitea hostname and update remaining client remotes/integrations. The operator confirmed completion on 2026-10-03; the agent did not perform a test push.
- Remove obsolete DuckDNS NPM Proxy Hosts, unused certificates and the old upstream override. The operator confirmed completion on 2026-10-03; no new agent runtime check was performed.
- Review and remove completed one-time procedures from the playbook. The operator confirmed completion on 2026-10-03.
- Retire Prometheus' local DuckDNS updater on 2026-10-03 through Ansible: the five-minute cron entry and private updater/log directory were removed. Provisioning is disabled; repeat cleanup changed nothing. HTTPS services, private NPM administration and the export timer stayed healthy. The external name and Vault token remain untouched for possible future use on a local host.
- Keep
atlas_manage_media_stackdisabled until the future Immich deployment has validated/dev/dri, container paths, and the required Vault database secret.
Priority 4 - Optional workflows
- Deploy the declared Atlas iCloudPD state dataset and inactive rootless
adminQuadlet. Photos belong under/zpool/archive/Pictures/iCloudPD; private config/MFA state belongs inzpool/services/data/icloudpd. Photobook remains reserved for Immich. Ansible now rendersicloudpd.confwith the Apple ID from the existing Vault key, but does not store the password or manage MFA. Automatic startup was approved on 2026-10-03; the Quadlet now usesWantedBy=default.targetand Ansible keeps the service running. The isolated no-network layout test is documented indocs/atlas-icloudpd-migration.md. On 2026-10-02 Atlas deployment and a second idempotent run passed; no app config existed at deployment. A manual first start on 2026-10-02 generatedicloudpd.conf; an Ansible run then replaced it with a private mode-0600 Vault-backed template and an idempotent second run. The image later expanded the config, so Ansible now seeds it only when absent and maintains the declared fields. Its launcher requirestraceroute; the rootless Quadlet grants onlyNET_RAW, tested in isolation and after restart. The service was subsequently initialized interactively; initial ingestion is tracked below. - Retire Aegis iCloudPD completely. The operator authorized deleting its Quadlet,
/var/lib/icloudpddata, and MFA state despite an unaudited container overlay. After two interactive-sudo runs on 2026-10-02, the unit isnot-found/inactive, the Quadlet and state directory are absent, and AdGuard remains active. The temporary retirement tasks have since been removed from the Aegis role; it no longer manages iCloudPD. - Validate Atlas iCloudPD authentication and initial ingestion. On 2026-10-03 the active
rootless service logged
All photos and videos have been downloadedat 02:16 and reported completion for the user. The destination held 11,658 files (86,020,430,015 bytes); the preceding 24h logs showed download activity without authentication failures or errors. A later read-only check found the service still active. This confirms the initial download, not the next daily cycle. - Declare HEIC decoding for Fedora graphical desktops without converting the originals on Atlas.
The Fedora role installs RPM Fusion Free with a pinned signing-key fingerprint and
libheif-freeworldon Ikaros and Nymph. The package was confirmed installed on Ikaros on 2026-10-03; Nymph deployment and an actual image-opening test were not observed. - Validate Atlas iCloudPD filesystem/SELinux/SMB access, the next daily sync, ZFS/Borg/USB
backup inclusion, and isolated restore of photos and private state. A recursive hourly snapshot
of
zpool/archiveexists after ingestion, but no iCloudPD-specific backup version or restore has been verified. The first monthly scrub remains a separate open data-protection check.
Prometheus NPM Quadlet cutover
- Stage a rootful NPM Quadlet using the exact running image and the existing data/certificate
mounts, bridge subnet, public HTTP/HTTPS ports, and loopback-only administration port.
The generated service depends on
server-web-network.serviceand is wanted bymulti-user.target. - Take and verify the stopped-source export before switching owners. Version
20261003T091009Zwas pulled to Atlas and its NPM SQLite database checked in isolation. - Cut over NPM to
prometheus-npm.serviceon 2026-10-03. The legacy Compose unit is inactive and disabled; the Quadlet is active with zero recorded restarts. Public Gitea and Syncthing HTTPS returned 200 with valid TLS, while public TCP/81 remained unreachable. - Validate the post-cutover backup path. The export and Atlas pull published
20261003T091633Z; checksum, SQLitequick_check, ten proxy hosts, six certificate records, both Quadlet files were present, and the complete Let's Encrypt tree (70 regular files plus 12 symlinks) matched the live data. A targeted normal Ansible run changed nothing. Details and rollback boundaries are indocs/prometheus-npm-quadlet.md. - Remove only unused Gitea, Navidrome and PostgreSQL images with opt-in
Ansible tasks on 2026-10-03. Second run changed nothing; NPM stayed active
with zero restarts, HTTP/HTTPS passed, backup timer and SSH proxy stayed active.
This image-only step preserved data and fallback; the later approved deletion is tracked below. Validation:
ansible-playbook ansible/site.yml --limit prometheus --tags server_image_cleanup --check --diff -e server_legacy_image_cleanup=true - Complete explicitly approved old-data and Compose fallback removal on 2026-10-03.
Backup paths and mount dependencies were reconciled before deletion; repeat cleanup changed
nothing. The normal Compose/template/helper check did not recreate retired files.
A separately approved manual export/pull published
20261003T112906Z; checksum and isolated SQLite restore passed with ten proxy hosts and both Quadlet definitions. NPM, primary HTTPS, WireGuard, SSH proxy and backup timer remained healthy; existing backup archives were preserved. - Retire the unused secondary Gitea hostname
git.ov-ad3410.infomaniak.chon 2026-10-03. Its NPM Proxy Host was already soft-deleted and had no associated certificate. Its Ansible domain and runtime override were removed; nginx -t and reload passed without restarting NPM. Primary HTTPS returned 200 with valid TLS. At that step onlygit.fscotto.duckdns.orgremained declared; the subsequent domain transition and operator-confirmed cleanup are tracked above. - Observe the first scheduled export and Atlas pull after the cutover; the manual end-to-end cycle passed, but the next unattended cycle has not yet occurred.
Cerberus Management Node (Deferred)
cerberus is postponed until the office in the new house is physically set up. It is not an inventory
host and this section is a design and implementation backlog, not authorization to provision it early.
The planned node is a Lenovo ThinkCentre M700 Tiny with an Intel Core i3-6100T, 8 GB RAM, a 256 GB SSD,
and native 1 Gbps Ethernet. It will connect to a multi-input KVM switch using a passive DisplayPort-to-HDMI
cable, sharing the monitor and peripherals with Ikaros. Fedora Sericea (immutable Fedora with the Sway
Wayland compositor) is the intended OS. Cerberus is an isolated management plane: a dedicated Toolbox
environment will run Ansible for future uranus cluster provisioning. Rootless Podman will host Grafana,
Prometheus, and Loki. The 256 GB local SSD is the hot tier retaining metrics and logs for 30 days; scheduled,
validated exports of older historical data will use a dedicated Atlas NFS dataset as cold storage.
Implementation plan
- Confirm the office, KVM switch, passive DisplayPort-to-HDMI path, shared monitor/peripherals, and native 1 Gbps Ethernet are physically operational before adding Cerberus to inventory.
- Install and update Fedora Sericea with Sway; document the immutable-host lifecycle and keep host changes declarative rather than treating the base OS as a mutable workstation.
- Model Cerberus as its own host with independent platform, role, desktop, network, and storage inputs; do not repurpose Ikaros variables or make it a Uranus cluster member.
- Provision an isolated Toolbox-based Ansible controller with the required collections and a reproducible project checkout; define its least-privilege SSH access, known-host handling, and Vault workflow without storing secrets in the image or repository.
- Define the explicit Uranus provisioning workflow from Cerberus, including inventory boundaries, validation-only runs, and separate approval for any destructive cluster operation.
- Design rootless Podman/Quadlet services for Grafana, Prometheus, and Loki, including persistent local state, service ownership, LAN exposure/authentication, resource limits, updates, and backups.
- Size and enforce a 30-day local hot-retention policy for metrics and logs on the 256 GB SSD; validate actual disk growth and alert before capacity exhaustion.
- Create and validate a dedicated Atlas NFS cold-storage dataset and least-privilege export for Cerberus; do not use a broad existing share or couple it to unrelated Atlas application state.
- Implement scheduled, idempotent exports of data older than 30 days to the Atlas NFS cold tier, with locking, capacity checks, integrity verification, retention rules, failure monitoring, and a tested restore.
- Validate management-plane recovery: rebuild Cerberus, restore observability history from Atlas, and confirm that Uranus provisioning can resume without depending on unreproducible local state.
Coding Agent Notes
- Shared agent definitions and lifecycle flags live in
ai_agentsinansible/inventory/group_vars/all.yml. - Shared agent dotfiles live in
ai_agents_dotfiles; rendered configs live inai_agents_templates. - Every
ai_agents.<agent>entry has independentinstall_enabled,deploy_enabled, anduninstall_enabledflags. Installation and removal must not both be true for the same agent; the common pre-task fails before changes when they conflict. - Fedora, Void desktop, and WSL workstation profiles consume the shared agent definitions; do not duplicate package entries in profile-specific vars. IBM Bob on the workstation follows its own flags.
dotfiles_commondeploysai_agents_dotfilesand rendersai_agents_templatesonly when deployment is enabled.- Removal is limited to the managed npm packages and
/usr/local/bin/bob; never remove agent dotfiles, instructions, credentials, or user data. - Keep
.config/ai/as the common instruction source; update agent-specific entrypoints to reference it rather than duplicating instruction text.
Tooling Notes
- Install local tooling with:
python3 -m pip install ansible ansible-lint yamllint shellcheck-pyansible-galaxy collection install -r ansible/collections/requirements.yml
- Required collections currently include
ansible.posixandcommunity.general. .yamllinttreatsline-lengthas a warning at 120 chars and disablesdocument-startandcomments-indentation.
When Updating Docs
- Keep
README.mdandAGENTS.mdaligned when workflows materially change. - If you add a new operational area, also add the narrowest validation command for it.
- Call out checks you could not run and any follow-up verification needed.
Aegis Fedora IoT Notes
aegisis a remote Fedora IoT Raspberry Pi 4 node. Bootstrap it once withansible/bootstrap/aegis.bu; the remaining configuration is applied byprofile_aegisover SSH.- Fedora IoT is immutable. Do not add it to mutable Fedora package or shared dotfile roles.
profile_aegisowns thenfs-utilsandwireguard-toolsrpm-ostree layers and reports the required reboot without initiating it.wireguard_overlaythen configures Aegis as the WireGuard LAN gateway with persistent IPv4 forwarding, a scoped inter-zone policy, and source masquerading. It also owns rootful Podman Quadlets, persistent container state under/var/lib, the Podman auto-update timer, LAN-restricted firewalld rules, and SSH hardening. Keepaegis_lan_subnet,aegis_adguard_web_port, andaegis_network_connection_uuidhost-specific; SSH permits only the declared key-authenticated users, never root or password authentication. Keep Apple IDs and other credentials in Vault and useno_logfor their rendering.aegis_adguard_web_portdefaults to80. The initial AdGuard Home wizard port3000is intentionally unmanaged: open and close it manually only while completing initial setup. Disable the local systemd-resolved stub throughprofile_aegisbefore AdGuard binds port 53; keep/etc/resolv.conflinked to/run/systemd/resolve/resolv.conf. LAN clients may use AdGuard, but Aegis must use the independent upstream DNS declared byaegis_host_dns_serversso Greenboot does not depend on the AdGuard container during startup.- Aegis iCloudPD has been retired and is no longer managed by this role. Its service, Quadlet, data, and MFA state were removed with the operator's explicit authorization.