# Gitea migration from Prometheus to Atlas This is a staged migration plan, not a cutover authorization. Keep the source Gitea, its data, both NPM Proxy Hosts, and public DNS unchanged until the target and rollback have been tested. Gitea is temporary on Atlas until Uranus; NPM remains on Prometheus. ## Observed source and chosen topology (2026-10-01) - Prometheus runs the rootful `docker.gitea.com/gitea:1.25.2` image in its managed Compose stack. `/opt/gitea/data` is about 280 MiB, uses SQLite, and contains 33 repositories. A live read-only SQLite `quick_check` passed. `/home/git/.ssh` is a separate small bind mount; `/opt/gitea/data/ssh` contains the existing SSH host keys. Neither tree may be discarded. - Gitea answers HTTP 200 on Prometheus port 3000. NPM currently forwards `git.fscotto.duckdns.org` and `git.ov-ad3410.infomaniak.ch` to the Compose hostname `gitea:3000`. Public DNS resolves to Prometheus. The container's SSH port is bound only to `127.0.0.1:222`; this is not a public Gitea SSH listener. Prometheus' public port 22 remains administrative SSH. - Atlas has a healthy pool and a verified, private Prometheus backup under `/zpool/backup/hosts/prometheus/latest`. The 2026-10-01 scheduled export and pull succeeded. The intended target is a separate `/zpool/services/data/gitea` dataset, not `Archive` or the backup dataset. - The approved cutover keeps NPM on Prometheus, changes the two HTTP Proxy Hosts' effective upstream to Atlas over the Prometheus--Aegis gateway, and offers public Gitea SSH on port 2222 via the same gateway. Prometheus port 22 is unchanged. HTTPS and SSH must be validated together before declaring cutover. - Run Gitea as a **rootless user Quadlet** under a dedicated, non-login Atlas account, using the pinned `1.25.2-rootless` image. This is an explicit rootful-to-rootless **data-layout conversion**, not a drop-in image swap: the target mounts `/var/lib/gitea` and `/etc/gitea`, and uses Gitea's built-in SSH server instead of the source image's OpenSSH daemon. Keep the application version unchanged until the conversion has passed an isolated restore test. The host's rootful Quadlet directory must not be used. ## Phase 1: prepare without traffic changes Preparation completed on 2026-10-01: Ansible created `zpool/services/data/gitea`, a dedicated non-login `gitea` account (UID/GID 1101), separate subordinate IDs, parent-dataset traverse ACLs, and an inactive user Quadlet under `/var/lib/atlas-gitea/.config/containers/systemd/`. The Quadlet has no `[Install]` section and, until the final cutover, binds only loopback staging ports 3001/2223 if started manually. A second targeted Ansible run changed nothing; the generated service was inactive and neither staging port listened. The explicit rehearsal is managed by: ```bash ansible-playbook ansible/site.yml --limit atlas --tags gitea_restore \ -e atlas_gitea_restore_test=true ``` On 2026-10-01 this selected the latest verified Prometheus backup, checked its SHA-256, extracted only `opt/gitea/data`, moved `app.ini` into the rootless config mount, rewrote `/data/` paths, enabled built-in SSH on internal port 2222, and retained the three source SSH host-key pairs. SQLite `quick_check` passed, all 33 restored repositories passed `git fsck`, and each source/target public host-key fingerprint matched. A temporary `1.25.2-rootless` container with `--network none` answered HTTP internally and listened on internal SSH/2222. The container was removed; the user Quadlet remains inactive, with no staging listener. The second restore run changed nothing. This copy is deliberately stale once new source writes occur and **must not** be used as the final cutover copy. Target backup checks on 2026-10-01: the managed recursive hourly ZFS snapshot `atlas-auto-hourly-20261001T193401Z` contains the new dataset. The managed Borg service completed archive `atlas-20261001T193420Z`, whose contents list includes the staged Gitea database. A separate one-file restore from each source into private `/var/tmp` directories matched the live staged database and passed SQLite `quick_check`. Temporary files and the on-demand snapshot mount were removed; the Borg temporary snapshot was cleaned up and the pool remained healthy. This is file-level proof, **not** a full Gitea recovery. The UUID-bound offline USB disk is connected but its LUKS mapper is closed; its manual backup requires interactive unlock. It has not yet captured or restored this new dataset. 1. Provision a dedicated target dataset and non-login service identity via Ansible, keeping UID/GID distinct from Atlas' reserved Immich `1100`. Install the user Quadlet in that identity's `~/.config/containers/systemd/`, **without** an `[Install]` section; do not enable, start, or expose it yet. 2. Verify the selected Atlas backup SHA-256 and metadata, then extract **only** `opt/gitea/data` to private staging. Keep `home/git/.ssh` in the source backup for rollback; the rootless image does not consume its OpenSSH mount. Never unpack NPM, WireGuard, or other host configuration from this sensitive tarball into a live namespace. Convert the rootful `/data` tree on a disposable copy: place application data under `/var/lib/gitea`, move `app.ini` to `/etc/gitea`, and rewrite every absolute `/data/...` path for the new layout. Enable `START_SSH_SERVER`, use internal SSH port 2222, and retain the source host-key pairs for the built-in server only after verifying their fingerprints and compatibility. Do not rely on the old `/home/git/.ssh` OpenSSH mount in the rootless image. Set only the target copy's ownership and path-scoped SELinux labels. 3. Validate SQLite integrity, repository count and representative `git fsck`, LFS/attachment presence, permissions, and an isolated rootless test container with no production ingress or outbound network. Because the source stays active, this is a rehearsal copy, not the final cutover copy. Regenerate Git hooks if the changed installation path requires it. 4. ZFS and Borg inclusion and one-file restores have passed. Complete a UUID-bound offline USB version and a one-file restore for the new dataset before accepting user traffic. ## Phase 2: explicit final cutover The opt-in `/usr/local/sbin/prometheus-gitea-final-export` helper was installed on 2026-10-01 and passed `bash -n`; it has **not** been invoked. It refuses to run while the scheduled Prometheus export timer is active. When explicitly triggered, it stops only the source Gitea container, checks SQLite, publishes a checksum-verified Gitea-only version for Atlas' existing pull, and leaves the source stopped on success. NPM remains running. A failure before completion restarts source Gitea. Its Ansible gate is `--tags gitea_final_export -e server_gitea_final_export=true`. After Atlas pulls that version, its separate `--tags gitea_final_restore -e atlas_gitea_final_restore=true` gate accepts only metadata marked `gitea-cutover`, validates a private staged replacement, and swaps it for the marked rehearsal. The swap and its rollback path passed synthetic tests on 2026-10-01; the gate has not been used on live Gitea data. The network change is also prepared but inactive. `server_gitea_on_atlas=true` removes the rootful Gitea service from the desired Prometheus Compose stack, adds `gitea:192.168.178.55` to NPM's container hosts file, and removes Gitea from future Prometheus backup exports. Both existing NPM Proxy Host records remain at `gitea:3000`, but that name then resolves to Atlas; no direct SQLite edit or NPM login is required. A separate systemd socket on public TCP/2222 proxies SSH to Atlas TCP/2222 over the gateway, leaving administrative TCP/22 unchanged. Atlas' production flag changes the user Quadlet from loopback staging ports to LAN ports 3000/2222, grants only Aegis access in firewalld, and starts it **only** after the `.final-sha256` marker exists. Neither flag is enabled yet. The future Prometheus configuration passed a check-run; the installed socket units passed `systemd-analyze verify` while remaining inactive. Source Gitea still answered HTTP 200 after preparation. The operator cannot unlock the UUID-bound USB disk now and chose to defer traffic activation until a new USB version covers Gitea and its file restore passes. This blocks the final export, production flags, and public cutover; the prepared configuration alone does not constitute a migration. 1. Agree on an outage and record source/target versions, pool health, the latest backups, SSH host-key fingerprints, and both current NPM routes. Stop the Prometheus export timer for the change window so it cannot restart the old Compose stack unexpectedly. 2. Quiesce source writes with the final-export helper: it stops Gitea before the consistent export and leaves it stopped after success. Pull that export to Atlas and verify checksum and timestamp. Keep `/opt/gitea/data` and `/home/git/.ssh` intact for rollback. Do not allow source Gitea to restart after accepting writes on Atlas. 3. Restore the final Gitea-only payload to the target and repeat integrity checks. Verify its advertised SSH port is 2222, its existing HTTPS `ROOT_URL`, repositories, LFS/attachments, and SSH host-key identity. Enable the production Atlas Quadlet only after the final-restore marker exists; its firewall permits only Aegis to reach HTTP and SSH. Validate local HTTP and the target service before switching NPM. 4. Enable the public TCP/2222 socket proxy on Prometheus to Atlas over Aegis without changing administrative TCP/22. Switch Prometheus to the desired NPM-only Compose stack and recreate NPM with the managed `gitea` host alias so **both** existing Proxy Hosts reach Atlas without changing their database records. The old Gitea data stays intact. Do not change public DNS. 5. Test HTTPS login, representative clone/push, LFS, and public SSH clone/push on port 2222 from outside the Atlas LAN. Record the last source write and first healthy target service times; do not claim RPO/RTO without measuring. 6. After successful traffic validation, resume the Prometheus NPM-only backup export timer and verify its next result. Verify the next Atlas snapshot/Borg run covers Gitea and test a restored target copy. Do not delete old source data. ## Rollback gate Before Atlas accepts writes, restore the old Compose definition (removing the NPM `gitea` host alias), disable the public 2222 proxy, and restart the unchanged source Gitea if target validation fails. **After Atlas accepts writes, do not blindly restart the source:** its SQLite database and repositories are stale. Quiesce Atlas, capture its new data, and decide a reverse migration or an extended outage explicitly. Upstream references: [rootful container layout](https://docs.gitea.com/1.25/installation/install-with-docker/), [rootless image layout and incompatibility](https://docs.gitea.com/installation/install-with-docker-rootless/), [rootless Podman Quadlet](https://docs.gitea.com/installation/install-with-podman-quadlet/), [standard-image conversion](https://docs.gitea.com/1.24/installation/install-with-docker-rootless/), and [restore and hook regeneration](https://docs.gitea.com/1.26/administration/backup-and-restore/).