Compare commits

..

15 Commits

Author SHA1 Message Date
Fabio Scotto di Santolo
9e76309833 Merge main and reconcile Atlas checklist 2026-10-02 17:54:33 +02:00
Fabio Scotto di Santolo
ed3fee06e8 Record operator-validated Gitea SSH pull and push 2026-10-02 17:47:40 +02:00
Fabio Scotto di Santolo
309d64b4ed Record successful public Gitea SSH authentication 2026-10-02 10:19:23 +02:00
Fabio Scotto di Santolo
dd33a4f55d Move Atlas Gitea Quadlet to admin with internal gitea user 2026-10-02 10:06:51 +02:00
Fabio Scotto di Santolo
0028fe8c4d Cut over Gitea HTTPS to Atlas with managed NPM upstream 2026-10-02 09:36:45 +02:00
Fabio Scotto di Santolo
12037fcc9a Enable restored rootless Gitea on Atlas 2026-10-02 09:35:44 +02:00
Fabio Scotto di Santolo
3f9a626759 Validate Gitea USB backup restore 2026-10-02 09:13:06 +02:00
Fabio Scotto di Santolo
31fedb8d44 Prepare gated Gitea HTTPS and SSH cutover 2026-10-01 22:09:57 +02:00
Fabio Scotto di Santolo
9b5ee77905 Prepare guarded final Gitea restore on Atlas 2026-10-01 22:00:41 +02:00
Fabio Scotto di Santolo
54fb7d46d7 Prepare consistent final Gitea export on Prometheus 2026-10-01 21:57:29 +02:00
Fabio Scotto di Santolo
a609e68f42 Record ZFS and Borg coverage for staged Gitea 2026-10-01 21:38:39 +02:00
Fabio Scotto di Santolo
06d3b175cb Rehearse rootless Gitea restore from verified backup 2026-10-01 21:33:18 +02:00
Fabio Scotto di Santolo
256d758b1a Prepare isolated rootless Gitea Quadlet on Atlas 2026-10-01 21:23:40 +02:00
Fabio Scotto di Santolo
9d0013769c Correct Atlas Gitea migration to rootless Quadlet 2026-10-01 21:15:42 +02:00
Fabio Scotto di Santolo
5b0f415163 Document staged Gitea migration to Atlas 2026-10-01 21:10:32 +02:00
27 changed files with 1553 additions and 23 deletions

View File

@@ -57,6 +57,19 @@ Ansible-driven personal infrastructure repo for Fedora and Void desktops, Fedora
- Server compose render: `podman-compose -f /opt/docker/server/docker-compose.yml config` and `systemctl status podman-compose-server`
- Atlas media stack:
`ansible-playbook ansible/site.yml --limit atlas --tags storage,sharing,containers --check --diff`
- Atlas rootless Gitea staging (does not start Gitea):
`ansible-playbook ansible/site.yml --limit atlas --tags gitea --check --diff`
- Atlas explicit Gitea host-owner migration (live outage; never a normal run):
`ansible-playbook ansible/site.yml --limit atlas --tags gitea_owner_migration -e atlas_gitea_owner_migration=true`
- Atlas explicit isolated Gitea restore rehearsal (not part of normal runs):
`ansible-playbook ansible/site.yml --limit atlas --tags gitea_restore -e atlas_gitea_restore_test=true`
- Atlas final Gitea replacement gate (dry-run only until a stopped-source export is pulled):
`ansible-playbook ansible/site.yml --limit atlas --tags gitea_final_restore --check --diff -e atlas_gitea_final_restore=true`
- Prometheus final Gitea export helper (dry-run installs only; outage action remains opt-in):
`ansible-playbook ansible/site.yml --limit prometheus --tags gitea_final_export --check --diff`
- Gitea cutover network configuration before activation:
`ansible-playbook ansible/site.yml --limit prometheus --tags gitea_cutover,prometheus_backup --check --diff -e server_gitea_on_atlas=true`
and `ansible-playbook ansible/site.yml --limit atlas --tags gitea --check --diff`
- Atlas daily Navidrome music copy:
`ansible-playbook ansible/site.yml --limit atlas --tags music_sync --check --diff`
- Atlas network/share hardening:
@@ -272,7 +285,75 @@ successfully. The first monthly scrub remains a runtime check.
intact and the temporary snapshot was removed.
- [x] Schedule a daily, non-deleting copy from `Archive/Music` to the separate Navidrome music
dataset. The rootless `atlas-music-sync.timer` is enabled for 00:45 Europe/Rome; a manual
idempotent service run succeeded on 2026-10-01. The first scheduled run remains to be verified.
idempotent service run succeeded on 2026-10-01. The first scheduled run triggered at
00:45 CEST on 2026-10-02 and exited successfully (`Result=success`, status 0); the next
run is scheduled for 2026-10-03 00:45 CEST.
- [x] Design the staged Prometheus-to-Atlas Gitea migration in `docs/atlas-gitea-migration.md`.
The approved topology keeps NPM on Prometheus and moves HTTPS and public SSH (TCP/2222) together;
Gitea runs as an `admin`-owned rootless user Quadlet on Atlas with an internal `gitea` user.
The rootful-to-rootless data-layout
conversion passed an isolated restore rehearsal. The later partial cutover is tracked below.
- [x] Prepare the dedicated Atlas Gitea dataset, non-login UID/GID 1101 with a separate rootless Podman
sub-ID range, and disabled user Quadlet. On 2026-10-01 the targeted Ansible run and a second idempotent
run passed; the generated unit was inactive, with no staging HTTP/SSH listener. POSIX ACLs on only the
service-namespace parents grant this account traversal without access to sibling datasets.
- [x] Perform an isolated rootless restore rehearsal from the verified Prometheus backup. On 2026-10-01
the SHA-256-checked selective extraction and path/SSH conversion succeeded; SQLite `quick_check`
passed, all 33 repositories passed `git fsck`, and source/target public SSH host-key fingerprints
matched. The pinned rootless image answered HTTP and listened on internal SSH/2222 with
`--network none`; the temporary container was removed and the Quadlet stayed inactive. A second
restore run made no changes. This is a rehearsal copy, not the final consistent cutover copy.
- [x] Verify ZFS and Borg coverage of the staged Gitea dataset. On 2026-10-01 the managed recursive
hourly snapshot `atlas-auto-hourly-20261001T193401Z` included it, and the managed incremental
Borg archive `atlas-20261001T193420Z` included its database. A private one-file restore from
each independently matched the staged database and passed SQLite `quick_check`; temporary files
and snapshot mounts were removed, the Borg service ended successfully, and the pool was healthy.
- [x] Include the new Gitea dataset in a UUID-bound offline USB version and test a file restore
before accepting production writes. The operator's 2026-10-01 manual run published version
`20261001T201220Z-254397` successfully on 2026-10-02. Its Gitea database was restored to a
temporary directory from a read-only mount: contents, owner, group, mode, size, mtime and POSIX
ACL matched, and SQLite `quick_check` passed. Temporary files and mounts were removed, LUKS
was closed, and the pool remained healthy. A redundant run was stopped during verification;
its temporary snapshot was cleaned up and the service's resulting failed state was reset.
- [x] Install a separate opt-in final Gitea export helper on Prometheus. Its 2026-10-01 targeted
deployment and `bash -n` passed while Gitea and NPM stayed running. It refuses an active export
timer, stops only Gitea, verifies SQLite, publishes a checksum-verified Gitea-only version for
Atlas' existing pull, and leaves the source stopped on success. It was invoked on 2026-10-02
after the export timer was stopped; version `20261002T071525Z` was pulled and verified on Atlas.
- [x] Prepare the Atlas final-restore gate without replacing the rehearsal: it accepts only a
checksum-verified `gitea-cutover` export, refuses a running target, stages and validates the new
layout before replacing the marked rehearsal, and rolls back a failed swap. Synthetic success
and rollback tests passed on 2026-10-01. On 2026-10-02 the final gate replaced the rehearsal;
SQLite `quick_check`, all 33 repository `git fsck` checks, checksum and SSH host-key comparison passed.
- [x] Start the rootless Atlas Gitea Quadlet and move the primary HTTPS route. On 2026-10-02 Atlas
answered HTTP 200 through the Aegis gateway. NPM stayed on Prometheus; its variable upstream
required a managed Nginx `server_proxy.conf` override because runtime DNS ignores Compose
`extra_hosts`. The primary public HTTPS page and API returned 200, and `git ls-remote` succeeded
for a representative repository after NPM restart; the Navidrome and Syncthing Proxy Hosts also
responded. The source
Gitea container was removed from the desired Compose stack without deleting its data; the
Prometheus backup export timer resumed for NPM only. A post-cutover recursive ZFS snapshot and
encrypted Borg archive `atlas-20261002T073044Z` completed successfully.
- [x] Move the live Gitea Quadlet and dataset from the legacy host `gitea` account to `admin`
after a disposable snapshot-copy test of the pinned derived image. On 2026-10-02 the explicit
outage run stopped only legacy Gitea, made safety snapshot
`zpool/services/data/gitea@gitea-owner-migration-20261002T100104`, changed dataset ownership,
and validated loopback staging (HTTP 200, internal `gitea` UID/GID 1000, SQLite `quick_check`)
before promoting the `admin` Quadlet. Production LAN and public HTTPS returned 200; Navidrome
and Syncthing remained active, the pool was healthy, and the normal Gitea run changed nothing.
The old host account and data on Prometheus remain preserved; the old Atlas Quadlet and its
parent-dataset traverse ACL were removed. A subsequent normal run changed nothing.
- [x] Validate public Gitea SSH/2222 and an authenticated read from Ikaros. After the VPS
firewall was opened on 2026-10-02, TCP/2222 connected, the public ED25519 host-key
fingerprint matched Atlas, Gitea authenticated `fscotto` using the `ikaros` key, and
`git ls-remote` returned HEAD for `fscotto/infra.git` over public SSH.
- [x] Validate authenticated SSH pull and push. On 2026-10-02 the operator reported both
operations working through the public SSH endpoint; the earlier agent-run `git ls-remote`
remains the independent read-only check. The agent did not perform a test push.
- [ ] Validate HTTPS write/login before declaring the full cutover complete. The
secondary NPM hostname `git.ov-ad3410.infomaniak.ch` did not resolve from Ikaros and had
no generated NPM config file at the previous inspection. Do not restart the stale source
Gitea after Atlas has accepted writes.
- [ ] Design and deploy Nextcloud as another explicitly temporary Atlas service before Uranus. Give it
separate persistent application, database, and cache storage; keep credentials in Vault; publish it only
through NPM over the Prometheus--Aegis gateway; and define backup, upgrade, and eventual Uranus-migration

View File

@@ -314,7 +314,8 @@ servizi sono inizializzati **ex novo**, senza migrare lo stato precedente, rispe
`/zpool/media/music` è stata popolata separatamente da `/zpool/archive/Music` il 2026-09-30;
Navidrome ha completato la scansione. Il timer rootless `atlas-music-sync.timer` copia i file nuovi
o modificati ogni giorno alle 00:45 Europe/Rome, senza eliminare quelli presenti solo nella
destinazione; entrambi i dataset ZFS devono essere montati. Alcune playlist originali contengono
destinazione; entrambi i dataset ZFS devono essere montati. La prima esecuzione schedulata è
riuscita il 2026-10-02. Alcune playlist originali contengono
ancora vecchi percorsi Windows. I servizi sono vincolati all'indirizzo LAN di Atlas
(`192.168.178.55`), mai a WireGuard. `wireguard_overlay` collega invece Prometheus (`10.0.0.1`)
e Aegis (`10.0.0.2`): le chiavi private restano sui rispettivi host e Ansible scambia solo le pubbliche.
@@ -326,6 +327,14 @@ alla LAN. Dopo la verifica dei servizi, configurare manualmente i Proxy Host NPM
negli `AllowedIPs`; aggiungere la VIP Uranus quando esisterà. Dopo il reload di firewalld, Ansible
ricarica le reti Podman rootful di Prometheus per conservare DNS e connettività del proxy.
La migrazione Gitea da Prometheus ad Atlas è descritta in
[`docs/atlas-gitea-migration.md`](docs/atlas-gitea-migration.md). Gitea usa un Quadlet rootless
di `admin` su un dataset dedicato; l'immagine derivata mantiene UID/GID 1000 ma chiama l'utente
interno `gitea`. NPM resta su Prometheus e l'HTTPS pubblico primario serve Atlas. L'SSH pubblico
su TCP/2222 autentica la chiave `ikaros` e un `git ls-remote` è riuscito; l'operatore ha
confermato pull e push SSH. Resta da provare la scrittura via HTTPS. I dati sorgente restano
conservati su Prometheus senza avviarne il vecchio container.
Validare il gateway con:
```bash

View File

@@ -299,7 +299,17 @@ and Aegis (`10.0.0.2`). Their state is initialized ex novo in `/zpool/services/d
`/zpool/media/music` was populated separately from `/zpool/archive/Music` on 2026-09-30;
Navidrome completed its library scan. The rootless `atlas-music-sync.timer` copies new and changed
files daily at 00:45 Europe/Rome, without deleting destination-only files. Both ZFS datasets must
be mounted. Some source playlists still contain obsolete Windows paths.
be mounted. Its first scheduled run succeeded on 2026-10-02. Some source playlists still contain
obsolete Windows paths.
The Gitea move from Prometheus to Atlas is tracked in
[`docs/atlas-gitea-migration.md`](docs/atlas-gitea-migration.md). The final consistent copy runs in
Atlas' dedicated dataset under `admin`'s rootless user Quadlet. Its pinned derived image uses an
internal Unix user named `gitea` (UID/GID 1000), while clone URLs keep `git@`. NPM remains on Prometheus and the primary
public HTTPS route serves Atlas. Public SSH/2222 now authenticates the `ikaros` key and serves
read-only `git ls-remote`; the operator also confirmed SSH pull and push. HTTPS writes remain
untested. The old Gitea data remains on Prometheus, but its container
is absent from the desired stack.
The separate `wireguard_overlay` role manages `wg0` between Prometheus (`10.0.0.1`) and Aegis
(`10.0.0.2`), generating private keys once on their respective hosts and exchanging only public keys

View File

@@ -87,23 +87,26 @@ server_backup_export_root: /var/lib/prometheus-backup-export
server_backup_rrsync_path: /usr/share/doc/rsync/support/rrsync
server_backup_export_calendar: "*-*-* 02:00:00 Europe/Rome"
server_backup_export_start_timer: false
# Explicit Gitea cutover helper: installed separately from any outage action.
server_gitea_cutover_tools_enabled: false
server_gitea_final_export: false
server_gitea_on_atlas: false
server_gitea_atlas_address: "{{ hostvars['atlas'].ansible_host }}"
server_gitea_npm_domains: []
server_gitea_ssh_public_port: 2222
server_gitea_ssh_target_port: 2222
server_backup_export_source_keep: 3
server_backup_export_paths:
- opt/npm/data
- opt/npm/letsencrypt
- opt/gitea/data
- home/git/.ssh
- opt/docker/server/docker-compose.yml
- etc/systemd/system/podman-compose-server.service
- etc/ssh/sshd_config
- etc/ssh/sshd_config.d
- etc/firewalld
- etc/wireguard/wg0.conf
server_backup_export_excludes:
- opt/npm/data/logs
- opt/gitea/data/gitea/log
- opt/gitea/data/gitea/tmp
- opt/gitea/data/gitea/sessions
- opt/gitea/data/gitea/indexers
server_backup_export_paths: >-
{{ ['opt/npm/data', 'opt/npm/letsencrypt']
+ ([] if server_gitea_on_atlas | bool else ['opt/gitea/data', 'home/git/.ssh'])
+ ['opt/docker/server/docker-compose.yml',
'etc/systemd/system/podman-compose-server.service',
'etc/ssh/sshd_config', 'etc/ssh/sshd_config.d',
'etc/firewalld', 'etc/wireguard/wg0.conf'] }}
server_backup_export_excludes: >-
{{ ['opt/npm/data/logs']
+ ([] if server_gitea_on_atlas | bool else
['opt/gitea/data/gitea/log', 'opt/gitea/data/gitea/tmp',
'opt/gitea/data/gitea/sessions', 'opt/gitea/data/gitea/indexers']) }}
server_ssh_authorized_keys: []
server_ssh_authorized_key_directory: "{{ server_user_home }}/.ssh/authorized_keys.d"

View File

@@ -49,6 +49,9 @@ atlas_zfs_backup_reservation: 500G
atlas_zfs_dataset_photobook: media/photobook
atlas_mount_root: /zpool
atlas_manage_storage: true
# Rootless Gitea was restored from the stopped-source export before production activation.
atlas_manage_gitea: true
atlas_gitea_production_enabled: true
atlas_prometheus_pull_start_timer: true
atlas_manage_zfs_snapshots: true
atlas_zfs_snapshot_prefix: atlas-auto

View File

@@ -8,6 +8,12 @@ ansible_ssh_private_key_file: /home/fscotto/.ssh/id_ed25519
server_username: rocky
server_backup_export_enabled: true
server_backup_export_start_timer: true
# Install the final-copy helper only; it is never run by a normal playbook invocation.
server_gitea_cutover_tools_enabled: true
server_gitea_on_atlas: true
server_gitea_npm_domains:
- git.fscotto.duckdns.org
- git.ov-ad3410.infomaniak.ch
server_duckdns_domain: fscotto
server_ssh_authorized_keys:
- name: ikaros

View File

@@ -165,6 +165,35 @@ atlas_prometheus_pull_keep_monthly: 12
atlas_prometheus_pull_max_age_hours: 24
atlas_photobook_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_photobook }}"
# Rootless Gitea runs in admin's user manager; the image maps internal gitea to UID/GID 1000.
atlas_manage_gitea: false
atlas_gitea_username: "{{ atlas_admin_username }}"
atlas_gitea_group: "{{ atlas_admin_group }}"
atlas_gitea_uid: "{{ atlas_admin_uid }}"
atlas_gitea_gid: "{{ atlas_admin_gid }}"
atlas_gitea_home: "{{ atlas_admin_home }}"
atlas_gitea_container_uid: 1000
atlas_gitea_container_gid: 1000
atlas_gitea_legacy_username: gitea
atlas_gitea_legacy_uid: 1101
atlas_gitea_legacy_home: /var/lib/atlas-gitea
atlas_gitea_owner_migration: false
atlas_gitea_dataset: "{{ atlas_zfs_pool }}/services/data/gitea"
atlas_gitea_mountpoint: "{{ atlas_app_data_mountpoint }}/gitea"
atlas_gitea_quadlet_dir: "{{ atlas_gitea_home }}/.config/containers/systemd"
atlas_gitea_image: localhost/atlas-gitea:1.25.2-user-gitea-v1
atlas_gitea_image_build_dir: "{{ atlas_gitea_home }}/.local/share/atlas-gitea-image"
atlas_gitea_production_enabled: false
atlas_gitea_bind_address: "{{ ansible_host }}"
atlas_gitea_http_port: 3000
atlas_gitea_ssh_port: 2222
atlas_gitea_staging_bind_address: 127.0.0.1
atlas_gitea_staging_http_port: 3001
atlas_gitea_staging_ssh_port: 2223
atlas_gitea_restore_test: false
atlas_gitea_final_restore: false
atlas_gitea_restore_helper: /usr/local/libexec/atlas-gitea-restore-test
atlas_45drives_repo_url: https://repo.45drives.com/repofiles/rocky/45drives-enterprise.repo
atlas_45drives_repo_file: /etc/yum.repos.d/45drives-enterprise.repo
atlas_45drives_packages:

View File

@@ -0,0 +1,10 @@
FROM docker.gitea.com/gitea@sha256:f1943db2d2f1e447e857b3f0aee4ebb7b184500f86e5b80eae110fd435435906
# Preserve the official image's UID/GID, paths and entrypoint; change only the
# internal Unix identity. The host-side rootless owner is Atlas admin.
USER 0
RUN sed -i 's/^git:x:1000:1000:/gitea:x:1000:1000:/' /etc/passwd \
&& sed -i 's/^git:x:1000:/gitea:x:1000:/' /etc/group \
&& grep -q '^gitea:x:1000:1000:' /etc/passwd \
&& grep -q '^gitea:x:1000:' /etc/group
USER 1000:1000

View File

@@ -0,0 +1,249 @@
#!/usr/bin/python3
"""Rehearse a selective rootful-to-rootless Gitea restore, never a cutover."""
import argparse
import hashlib
import json
import os
from pathlib import Path, PurePosixPath
import re
import shutil
import sqlite3
import tarfile
import tempfile
SOURCE_PREFIX = PurePosixPath("opt/gitea/data")
HOST_KEYS = (
"ssh_host_ed25519_key",
"ssh_host_rsa_key",
"ssh_host_ecdsa_key",
)
SERVER_SETTINGS = {
"START_SSH_SERVER": "true",
"BUILTIN_SSH_SERVER_USER": "git",
"SSH_USER": "git",
"SSH_PORT": "2222",
"SSH_LISTEN_PORT": "2222",
"SSH_SERVER_HOST_KEYS": ", ".join(
f"/var/lib/gitea/ssh/{key}" for key in HOST_KEYS
),
}
def sha256(path):
digest = hashlib.sha256()
with path.open("rb") as stream:
for chunk in iter(lambda: stream.read(1024 * 1024), b""):
digest.update(chunk)
return digest.hexdigest()
def expected_digest(backup):
checksum = (backup / "payload.sha256").read_text().strip().split()
if len(checksum) != 2 or checksum[1] != "payload.tar":
raise ValueError("Unexpected Prometheus backup checksum manifest")
if not re.fullmatch(r"[0-9a-f]{64}", checksum[0]):
raise ValueError("Invalid Prometheus backup SHA-256")
return checksum[0]
def convert_config(config):
original = config.read_text()
output = []
section = ""
server_seen = set()
server_found = False
run_user_seen = False
def append_missing_server_settings():
for key, value in SERVER_SETTINGS.items():
if key not in server_seen:
output.append(f"{key} = {value}\n")
for line in original.splitlines(keepends=True):
match = re.match(r"^\s*\[([^]]+)\]\s*$", line)
if match:
if not run_user_seen:
output.append("RUN_USER = gitea\n")
run_user_seen = True
if section == "server":
append_missing_server_settings()
section = match.group(1).lower()
server_found |= section == "server"
output.append(line)
continue
setting = re.match(r"^(\s*)([A-Z_]+)(\s*=\s*)(.*?)(\r?\n?)$", line)
if setting and section == "" and setting.group(2) == "RUN_USER":
run_user_seen = True
line = f"{setting.group(1)}RUN_USER{setting.group(3)}gitea{setting.group(5)}"
elif setting and section == "server" and setting.group(2) in SERVER_SETTINGS:
key = setting.group(2)
server_seen.add(key)
line = f"{setting.group(1)}{key}{setting.group(3)}{SERVER_SETTINGS[key]}{setting.group(5)}"
else:
line = line.replace("/data/", "/var/lib/gitea/")
output.append(line)
if section == "server":
append_missing_server_settings()
if not server_found:
raise ValueError("Gitea server configuration missing")
config.write_text("".join(output))
config.chmod(0o600)
def extract_gitea(tar_path, staged_data):
count = 0
with tarfile.open(tar_path, mode="r") as archive:
for member in archive:
name = PurePosixPath(member.name)
if name == SOURCE_PREFIX:
continue
if SOURCE_PREFIX not in name.parents:
continue
relative = name.relative_to(SOURCE_PREFIX)
if not relative.parts or any(part in (".", "..") for part in relative.parts):
raise ValueError("Unsafe Gitea backup path")
if not (member.isdir() or member.isfile()):
raise ValueError("Unexpected Gitea backup member type")
destination = staged_data.joinpath(*relative.parts)
if member.isdir():
destination.mkdir(parents=True, exist_ok=True)
destination.chmod(0o700)
continue
destination.parent.mkdir(parents=True, exist_ok=True)
with archive.extractfile(member) as source, destination.open("xb") as target:
shutil.copyfileobj(source, target)
destination.chmod(member.mode & 0o777)
count += 1
if count == 0:
raise ValueError("No Gitea files in backup")
def validate(staged_data, staged_config):
database = staged_data / "gitea/gitea.db"
repositories = staged_data / "git/repositories"
if not database.is_file() or not repositories.is_dir():
raise ValueError("Missing SQLite database or Git repositories")
with sqlite3.connect(f"file:{database}?mode=ro", uri=True) as connection:
if connection.execute("PRAGMA quick_check").fetchone()[0] != "ok":
raise ValueError("Gitea SQLite quick_check failed")
if connection.execute("SELECT count(*) FROM repository").fetchone()[0] < 1:
raise ValueError("Gitea backup contains no repository records")
if not any(repositories.rglob("*.git")):
raise ValueError("Gitea backup contains no Git repository directories")
if not (staged_config / "app.ini").is_file():
raise ValueError("Gitea app.ini missing")
for name in HOST_KEYS:
if not (staged_data / "ssh" / name).is_file():
raise ValueError("Gitea SSH host key missing")
def chown_tree(root, uid, gid):
for directory, dirs, files in os.walk(root):
os.chown(directory, uid, gid)
for name in dirs + files:
os.chown(os.path.join(directory, name), uid, gid)
def replace_rehearsal(target, stage, digest, uid, gid):
previous_data = target / ".previous-rehearsal-data"
previous_config = target / ".previous-rehearsal-config"
if previous_data.exists() or previous_config.exists():
raise ValueError("An interrupted Gitea replacement needs manual recovery")
os.rename(target / "data", previous_data)
try:
os.rename(target / "config", previous_config)
os.rename(stage / "data", target / "data")
os.rename(stage / "config", target / "config")
final_marker = target / ".final-sha256"
final_marker.write_text(digest + "\n")
final_marker.chmod(0o600)
os.chown(final_marker, uid, gid)
(target / ".rehearsal-sha256").unlink()
except Exception:
for name, previous in (("data", previous_data), ("config", previous_config)):
current = target / name
if previous.exists():
if current.exists():
shutil.rmtree(current)
os.rename(previous, current)
(target / ".final-sha256").unlink(missing_ok=True)
raise
shutil.rmtree(previous_data)
shutil.rmtree(previous_config)
def main():
parser = argparse.ArgumentParser()
parser.add_argument("--backup", type=Path, required=True)
parser.add_argument("--target", type=Path, required=True)
parser.add_argument("--uid", type=int, required=True)
parser.add_argument("--gid", type=int, required=True)
parser.add_argument("--replace-rehearsal", action="store_true")
args = parser.parse_args()
backup = args.backup.resolve(strict=True)
target = args.target.resolve(strict=True)
if not str(backup).startswith("/zpool/backup/hosts/prometheus/snapshots/"):
raise ValueError("Refusing backup outside the Atlas Prometheus snapshots")
if str(target) != "/zpool/services/data/gitea":
raise ValueError("Refusing target outside the dedicated Gitea dataset")
if args.uid != 1000 or args.gid != 1000:
raise ValueError("Unexpected admin-owned Gitea account IDs")
expected = expected_digest(backup)
if sha256(backup / "payload.tar") != expected:
raise ValueError("Prometheus backup SHA-256 mismatch")
marker = target / (".final-sha256" if args.replace_rehearsal else ".rehearsal-sha256")
if marker.exists():
if marker.read_text().strip() != expected:
raise ValueError("A different Gitea restore already occupies this dataset")
validate(target / "data", target / "config")
print("unchanged")
return
if args.replace_rehearsal:
metadata = json.loads((backup / "metadata.json").read_text())
if metadata.get("purpose") != "gitea-cutover":
raise ValueError("Final restore requires an explicit Gitea cutover export")
if not (target / ".rehearsal-sha256").is_file():
raise ValueError("Only a marked rehearsal may be replaced")
if not all((target / name).is_dir() for name in ("data", "config")):
raise ValueError("Prepared Gitea volume paths are missing")
else:
if (target / ".final-sha256").exists():
raise ValueError("Refusing a rehearsal restore over final Gitea data")
for name in ("data", "config"):
directory = target / name
if not directory.is_dir() or any(directory.iterdir()):
raise ValueError("Gitea target is not empty; refusing overwrite")
with tempfile.TemporaryDirectory(prefix=".rehearsal-", dir=target) as temporary:
stage = Path(temporary)
staged_data = stage / "data"
staged_config = stage / "config"
staged_data.mkdir()
staged_config.mkdir()
extract_gitea(backup / "payload.tar", staged_data)
source_config = staged_data / "gitea/conf/app.ini"
if not source_config.is_file():
raise ValueError("Source Gitea app.ini missing")
shutil.copy2(source_config, staged_config / "app.ini")
source_config.unlink()
convert_config(staged_config / "app.ini")
validate(staged_data, staged_config)
chown_tree(stage, args.uid, args.gid)
if args.replace_rehearsal:
replace_rehearsal(target, stage, expected, args.uid, args.gid)
else:
for name in ("data", "config"):
(target / name).rmdir()
os.rename(stage / name, target / name)
marker.write_text(expected + "\n")
marker.chmod(0o600)
os.chown(marker, args.uid, args.gid)
print("restored")
if __name__ == "__main__":
main()

View File

@@ -0,0 +1,159 @@
---
- name: Prepare the isolated rootless Atlas Gitea target
tags: [atlas, gitea]
when: atlas_manage_gitea | bool
block:
- name: Require the existing Atlas application-data dataset
ansible.builtin.assert:
that:
- atlas_manage_storage | bool
- atlas_gitea_dataset == atlas_zfs_pool ~ '/services/data/gitea'
- atlas_gitea_mountpoint == atlas_app_data_mountpoint ~ '/gitea'
- atlas_gitea_username == atlas_admin_username
- atlas_gitea_group == atlas_admin_group
- atlas_gitea_uid | int == atlas_admin_uid | int
- atlas_gitea_gid | int == atlas_admin_gid | int
- atlas_gitea_container_uid | int == 1000
- atlas_gitea_container_gid | int == 1000
- atlas_gitea_staging_bind_address == '127.0.0.1'
- not (atlas_gitea_production_enabled | bool) or atlas_manage_firewall | bool
- not (atlas_gitea_production_enabled | bool) or atlas_gitea_bind_address == ansible_host
fail_msg: >-
Rootless Gitea requires Atlas storage, the admin user manager, the
dedicated dataset, and loopback-only staging ports.
- name: Inspect the final-restore marker before production activation
ansible.builtin.stat:
path: "{{ atlas_gitea_mountpoint }}/.final-sha256"
register: atlas_gitea_final_marker
when: atlas_gitea_production_enabled | bool
- name: Refuse production activation without the final consistent restore
ansible.builtin.assert:
that:
- atlas_gitea_final_marker.stat.isreg | default(false)
fail_msg: Restore the final stopped-source Gitea export before enabling production.
when: atlas_gitea_production_enabled | bool
- name: Verify the production Gitea dataset belongs to admin
ansible.builtin.stat:
path: "{{ atlas_gitea_mountpoint }}"
register: atlas_gitea_dataset_owner
when: atlas_gitea_production_enabled | bool
- name: Refuse to overlap the legacy host-account service
ansible.builtin.assert:
that:
- atlas_gitea_dataset_owner.stat.uid | int == atlas_admin_uid | int
- atlas_gitea_dataset_owner.stat.gid | int == atlas_admin_gid | int
fail_msg: >-
Run the explicit Gitea owner migration before enabling the admin
Quadlet; never chown an active legacy service in a normal run.
when: atlas_gitea_production_enabled | bool
- name: Remove the retired account's parent-dataset traverse ACL
ansible.posix.acl:
path: "{{ item }}"
etype: user
entity: "{{ atlas_gitea_legacy_username }}"
state: absent
loop:
- "{{ atlas_services_mountpoint }}"
- "{{ atlas_app_data_mountpoint }}"
when: atlas_gitea_production_enabled | bool
- name: Enable POSIX ACLs only on the service-namespace parents
community.general.zfs:
name: "{{ item }}"
state: present
extra_zfs_properties:
acltype: posix
loop:
- "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_services }}"
- "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_app_data }}"
- name: Create the dedicated Gitea ZFS dataset
community.general.zfs:
name: "{{ atlas_gitea_dataset }}"
state: present
extra_zfs_properties:
compression: zstd
mountpoint: "{{ atlas_gitea_mountpoint }}"
- name: Restrict the Gitea dataset and create rootless volume paths
ansible.builtin.file:
path: "{{ item }}"
state: directory
owner: "{{ atlas_gitea_username }}"
group: "{{ atlas_gitea_group }}"
mode: "0700"
loop:
- "{{ atlas_gitea_mountpoint }}"
- "{{ atlas_gitea_mountpoint }}/data"
- "{{ atlas_gitea_mountpoint }}/config"
- "{{ atlas_gitea_home }}/.config"
- "{{ atlas_gitea_home }}/.config/containers"
- "{{ atlas_gitea_quadlet_dir }}"
- name: Ensure lingering for the admin rootless account
ansible.builtin.command:
argv:
- loginctl
- enable-linger
- "{{ atlas_gitea_username }}"
creates: "/var/lib/systemd/linger/{{ atlas_gitea_username }}"
- name: Start the admin rootless user manager
ansible.builtin.systemd:
name: "user@{{ atlas_gitea_uid }}.service"
state: started
when: not ansible_check_mode
- name: Prepare the admin-owned Gitea image
ansible.builtin.import_tasks: gitea_image.yml
- name: Render the rootless Gitea Quadlet
ansible.builtin.template:
src: atlas-gitea.container.j2
dest: "{{ atlas_gitea_quadlet_dir }}/atlas-gitea.container"
owner: "{{ atlas_gitea_username }}"
group: "{{ atlas_gitea_group }}"
mode: "0644"
- name: Permit only Aegis to reach production Gitea HTTP and SSH
ansible.posix.firewalld:
rich_rule: >-
rule family="ipv4" source address="{{ atlas_aegis_ip }}"
port port="{{ item }}" protocol="tcp" accept
zone: "{{ atlas_firewalld_zone }}"
state: "{{ 'enabled' if atlas_gitea_production_enabled | bool else 'disabled' }}"
permanent: true
immediate: true
loop:
- "{{ atlas_gitea_http_port }}"
- "{{ atlas_gitea_ssh_port }}"
when: atlas_manage_firewall | bool
- name: Reload the rootless Gitea user manager without starting Gitea
become_user: "{{ atlas_gitea_username }}"
ansible.builtin.systemd:
scope: user
daemon_reload: true
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_gitea_uid }}/bus"
when: not ansible_check_mode
- name: Start and enable the rootless Gitea user Quadlet after final restore
become_user: "{{ atlas_gitea_username }}"
ansible.builtin.systemd:
name: atlas-gitea.service
scope: user
state: started
enabled: true
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_gitea_uid }}/bus"
when:
- atlas_gitea_production_enabled | bool
- not ansible_check_mode

View File

@@ -0,0 +1,51 @@
---
- name: Create the admin-owned Gitea image build directory
ansible.builtin.file:
path: "{{ atlas_gitea_image_build_dir }}"
state: directory
owner: "{{ atlas_admin_username }}"
group: "{{ atlas_admin_group }}"
mode: "0700"
- name: Install the pinned rootless Gitea Containerfile
ansible.builtin.copy:
src: Containerfile.gitea-rootless
dest: "{{ atlas_gitea_image_build_dir }}/Containerfile"
owner: "{{ atlas_admin_username }}"
group: "{{ atlas_admin_group }}"
mode: "0644"
- name: Check the admin-owned Gitea image
become_user: "{{ atlas_admin_username }}"
ansible.builtin.command:
argv: [podman, image, exists, "{{ atlas_gitea_image }}"]
args:
chdir: "{{ atlas_gitea_image_build_dir }}"
environment:
HOME: "{{ atlas_admin_home }}"
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
register: atlas_gitea_image_present
changed_when: false
failed_when: false
check_mode: false
- name: Build the pinned Gitea image with the internal gitea identity
become_user: "{{ atlas_admin_username }}"
ansible.builtin.command:
argv:
- podman
- build
- --pull=always
- --tag
- "{{ atlas_gitea_image }}"
- --file
- Containerfile
- .
args:
chdir: "{{ atlas_gitea_image_build_dir }}"
environment:
HOME: "{{ atlas_admin_home }}"
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
when:
- atlas_gitea_image_present.rc != 0
- not ansible_check_mode

View File

@@ -0,0 +1,275 @@
---
# Run only in an approved outage with -e atlas_gitea_owner_migration=true.
- name: Move live Gitea from the legacy host account to admin
tags: [atlas, gitea_owner_migration]
when: atlas_gitea_owner_migration | bool
block:
- name: Refuse a check-mode owner migration
ansible.builtin.assert:
that: not ansible_check_mode
fail_msg: The owner migration requires an explicit live outage.
- name: Inspect the Gitea dataset owner
ansible.builtin.stat:
path: "{{ atlas_gitea_mountpoint }}"
register: atlas_gitea_migration_owner
- name: Require either the legacy owner or an already migrated dataset
ansible.builtin.assert:
that:
- atlas_gitea_migration_owner.stat.isdir | default(false)
- atlas_gitea_migration_owner.stat.uid | int in [atlas_gitea_legacy_uid | int, atlas_admin_uid | int]
fail_msg: Refusing to modify a Gitea dataset with an unexpected owner.
- name: Migrate only a legacy-owned Gitea dataset
when: atlas_gitea_migration_owner.stat.uid | int == atlas_gitea_legacy_uid | int
block:
- name: Require the final cutover marker and configuration
ansible.builtin.stat:
path: "{{ item }}"
loop:
- "{{ atlas_gitea_mountpoint }}/.final-sha256"
- "{{ atlas_gitea_mountpoint }}/config/app.ini"
register: atlas_gitea_migration_files
- name: Refuse migration without both final data and configuration
ansible.builtin.assert:
that: atlas_gitea_migration_files.results | map(attribute='stat.isreg') | min
- name: Check that admin has no existing Gitea Quadlet
ansible.builtin.stat:
path: "{{ atlas_gitea_quadlet_dir }}/atlas-gitea.container"
register: atlas_gitea_admin_quadlet
- name: Refuse to overwrite an existing admin Quadlet
ansible.builtin.assert:
that: not atlas_gitea_admin_quadlet.stat.exists
- name: Check pool health before the outage
ansible.builtin.command:
argv: [zpool, status, -x, "{{ atlas_zfs_pool }}"]
register: atlas_gitea_pool_before
changed_when: false
failed_when: "'is healthy' not in atlas_gitea_pool_before.stdout"
- name: Ensure the admin Gitea image is available before stopping the source
ansible.builtin.import_tasks: gitea_image.yml
- name: Stop, snapshot and test the admin-owned staging service
block:
- name: Stop and disable the legacy Gitea user service
become_user: "{{ atlas_gitea_legacy_username }}"
ansible.builtin.systemd:
name: atlas-gitea.service
scope: user
state: stopped
enabled: false
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_legacy_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_gitea_legacy_uid }}/bus"
- name: Record the migration snapshot name
ansible.builtin.set_fact:
atlas_gitea_migration_snapshot: >-
{{ atlas_gitea_dataset }}@gitea-owner-migration-{{ ansible_facts.date_time.iso8601_basic_short }}
- name: Snapshot the stopped Gitea dataset for manual recovery
ansible.builtin.command:
argv: [zfs, snapshot, "{{ atlas_gitea_migration_snapshot }}"]
- name: Transfer only the Gitea dataset to admin
ansible.builtin.file:
path: "{{ atlas_gitea_mountpoint }}"
state: directory
owner: "{{ atlas_admin_username }}"
group: "{{ atlas_admin_group }}"
recurse: true
- name: Set the actual internal Unix process user
ansible.builtin.lineinfile:
path: "{{ atlas_gitea_mountpoint }}/config/app.ini"
regexp: '^RUN_USER\s*='
line: RUN_USER = gitea
mode: "0600"
no_log: true
diff: false
- name: Preserve public git clone URLs independently of the Unix user
community.general.ini_file:
path: "{{ atlas_gitea_mountpoint }}/config/app.ini"
section: server
option: "{{ item }}"
value: git
mode: "0600"
no_extra_spaces: false
loop: [BUILTIN_SSH_SERVER_USER, SSH_USER]
no_log: true
diff: false
- name: Render admin's loopback-only staging Quadlet
ansible.builtin.template:
src: atlas-gitea.container.j2
dest: "{{ atlas_gitea_quadlet_dir }}/atlas-gitea.container"
owner: "{{ atlas_admin_username }}"
group: "{{ atlas_admin_group }}"
mode: "0644"
vars:
atlas_gitea_production_enabled: false
- name: Reload the admin user manager for staging
become_user: "{{ atlas_admin_username }}"
ansible.builtin.systemd:
scope: user
daemon_reload: true
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus"
- name: Start admin's loopback-only staging service
become_user: "{{ atlas_admin_username }}"
ansible.builtin.systemd:
name: atlas-gitea.service
scope: user
state: started
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus"
- name: Verify staging HTTP before promotion
ansible.builtin.uri:
url: "http://127.0.0.1:{{ atlas_gitea_staging_http_port }}/"
status_code: 200
register: atlas_gitea_staging_http
retries: 30
delay: 2
until: atlas_gitea_staging_http is succeeded
- name: Verify the container really runs as internal gitea
become_user: "{{ atlas_admin_username }}"
ansible.builtin.command:
argv: [podman, exec, atlas-gitea, id, -un]
environment:
HOME: "{{ atlas_admin_home }}"
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
register: atlas_gitea_internal_user
changed_when: false
failed_when: atlas_gitea_internal_user.stdout != 'gitea'
- name: Verify the migrated SQLite database
ansible.builtin.command:
argv:
- sqlite3
- "{{ atlas_gitea_mountpoint }}/data/gitea/gitea.db"
- PRAGMA quick_check;
register: atlas_gitea_migration_sqlite
changed_when: false
failed_when: atlas_gitea_migration_sqlite.stdout != 'ok'
rescue:
- name: Stop admin's failed staging service
become_user: "{{ atlas_admin_username }}"
ansible.builtin.systemd:
name: atlas-gitea.service
scope: user
state: stopped
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus"
failed_when: false
- name: Restore the original Gitea configuration from the safety snapshot
ansible.builtin.command:
argv:
- cp
- -a
- "{{ atlas_gitea_mountpoint }}/.zfs/snapshot/{{ atlas_gitea_migration_snapshot.split('@')[1] }}/config/app.ini"
- "{{ atlas_gitea_mountpoint }}/config/app.ini"
when: atlas_gitea_migration_snapshot is defined
- name: Return the Gitea dataset to the legacy account
ansible.builtin.file:
path: "{{ atlas_gitea_mountpoint }}"
state: directory
owner: "{{ atlas_gitea_legacy_username }}"
group: "{{ atlas_gitea_legacy_username }}"
recurse: true
- name: Restart the legacy Gitea service
become_user: "{{ atlas_gitea_legacy_username }}"
ansible.builtin.systemd:
name: atlas-gitea.service
scope: user
state: started
enabled: true
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_legacy_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_gitea_legacy_uid }}/bus"
- name: Report the failed migration and preserved snapshot
ansible.builtin.fail:
msg: >-
Admin staging failed; legacy Gitea was restarted. Inspect
{{ atlas_gitea_migration_snapshot | default('the host journal') }}.
- name: Stop admin's validated staging service
become_user: "{{ atlas_admin_username }}"
ansible.builtin.systemd:
name: atlas-gitea.service
scope: user
state: stopped
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus"
- name: Render admin's production Gitea Quadlet
ansible.builtin.template:
src: atlas-gitea.container.j2
dest: "{{ atlas_gitea_quadlet_dir }}/atlas-gitea.container"
owner: "{{ atlas_admin_username }}"
group: "{{ atlas_admin_group }}"
mode: "0644"
vars:
atlas_gitea_production_enabled: true
- name: Reload admin's production user manager
become_user: "{{ atlas_admin_username }}"
ansible.builtin.systemd:
scope: user
daemon_reload: true
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus"
- name: Enable and start admin's production Gitea
become_user: "{{ atlas_admin_username }}"
ansible.builtin.systemd:
name: atlas-gitea.service
scope: user
state: started
enabled: true
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus"
- name: Verify production HTTP before retiring the old Quadlet
ansible.builtin.uri:
url: "http://{{ atlas_gitea_bind_address }}:{{ atlas_gitea_http_port }}/"
status_code: 200
register: atlas_gitea_production_http
retries: 30
delay: 2
until: atlas_gitea_production_http is succeeded
- name: Remove only the disabled legacy Quadlet
ansible.builtin.file:
path: "{{ atlas_gitea_legacy_home }}/.config/containers/systemd/atlas-gitea.container"
state: absent
- name: Reload the legacy user manager after Quadlet removal
become_user: "{{ atlas_gitea_legacy_username }}"
ansible.builtin.systemd:
scope: user
daemon_reload: true
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_legacy_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_gitea_legacy_uid }}/bus"

View File

@@ -0,0 +1,107 @@
---
- name: Restore Gitea from a verified Prometheus backup only on explicit request
tags: [atlas, gitea_restore, gitea_final_restore]
when: atlas_gitea_restore_test | bool or atlas_gitea_final_restore | bool
block:
- name: Require the prepared rootless Gitea target
ansible.builtin.assert:
that:
- atlas_manage_gitea | bool
- not (atlas_gitea_restore_test | bool and atlas_gitea_final_restore | bool)
- atlas_gitea_staging_bind_address == '127.0.0.1'
- atlas_gitea_mountpoint == atlas_app_data_mountpoint ~ '/gitea'
fail_msg: Prepare the isolated, loopback-only rootless Gitea target first.
- name: Confirm the rootless Gitea service is inactive
become_user: "{{ atlas_gitea_username }}"
ansible.builtin.command:
argv:
- systemctl
- --user
- is-active
- atlas-gitea.service
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_uid }}"
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_gitea_uid }}/bus"
register: atlas_gitea_restore_service_state
changed_when: false
failed_when: false
when: not ansible_check_mode
- name: Refuse to overwrite an active rootless Gitea service
ansible.builtin.assert:
that:
- atlas_gitea_restore_service_state.stdout == 'inactive'
fail_msg: The rootless Gitea user service must be known and inactive before restoring data.
when: not ansible_check_mode
- name: Check for a manually running rootless Gitea container
become_user: "{{ atlas_gitea_username }}"
ansible.builtin.command:
argv:
- podman
- ps
- --quiet
- --filter
- name=atlas-gitea
args:
chdir: "{{ atlas_gitea_home }}"
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_uid }}"
register: atlas_gitea_restore_container_state
changed_when: false
when: not ansible_check_mode
- name: Refuse to overwrite a running rootless Gitea container
ansible.builtin.assert:
that:
- atlas_gitea_restore_container_state.stdout | length == 0
fail_msg: Stop every rootless Atlas Gitea container before restoring data.
when: not ansible_check_mode
- name: Install the selective rootless Gitea restore helper
ansible.builtin.copy:
src: atlas-gitea-restore-test.py
dest: "{{ atlas_gitea_restore_helper }}"
owner: root
group: root
mode: "0700"
- name: Restore only Gitea data into the isolated target
ansible.builtin.command:
argv:
- "{{ atlas_gitea_restore_helper }}"
- --backup
- "{{ atlas_backup_prometheus_mountpoint }}/latest"
- --target
- "{{ atlas_gitea_mountpoint }}"
- --uid
- "{{ atlas_gitea_uid | string }}"
- --gid
- "{{ atlas_gitea_gid | string }}"
register: atlas_gitea_restore_result
changed_when: atlas_gitea_restore_result.stdout == 'restored'
no_log: true
when:
- atlas_gitea_restore_test | bool
- not ansible_check_mode
- name: Replace the marked rehearsal with the final consistent Gitea export
ansible.builtin.command:
argv:
- "{{ atlas_gitea_restore_helper }}"
- --backup
- "{{ atlas_backup_prometheus_mountpoint }}/latest"
- --target
- "{{ atlas_gitea_mountpoint }}"
- --uid
- "{{ atlas_gitea_uid | string }}"
- --gid
- "{{ atlas_gitea_gid | string }}"
- --replace-rehearsal
register: atlas_gitea_final_restore_result
changed_when: atlas_gitea_final_restore_result.stdout == 'restored'
no_log: true
when:
- atlas_gitea_final_restore | bool
- not ansible_check_mode

View File

@@ -14,6 +14,15 @@
- name: Import Atlas storage tasks
ansible.builtin.import_tasks: storage.yml
- name: Import explicit Atlas Gitea owner migration
ansible.builtin.import_tasks: gitea_owner_migration.yml
- name: Import staged Atlas rootless Gitea tasks
ansible.builtin.import_tasks: gitea.yml
- name: Import explicit Atlas Gitea restore rehearsal tasks
ansible.builtin.import_tasks: gitea_restore.yml
- name: Import Atlas ZFS maintenance tasks
ansible.builtin.import_tasks: zfs_maintenance.yml

View File

@@ -0,0 +1,30 @@
# Managed by Ansible. Staging does not start automatically.
[Unit]
Description=Atlas rootless Gitea
RequiresMountsFor={{ atlas_gitea_mountpoint }}
[Container]
ContainerName=atlas-gitea
Image={{ atlas_gitea_image }}
UserNS=keep-id:uid={{ atlas_gitea_container_uid }},gid={{ atlas_gitea_container_gid }}
{% if atlas_gitea_production_enabled | bool %}
PublishPort={{ atlas_gitea_bind_address }}:{{ atlas_gitea_http_port }}:3000
PublishPort={{ atlas_gitea_bind_address }}:{{ atlas_gitea_ssh_port }}:2222
{% else %}
PublishPort={{ atlas_gitea_staging_bind_address }}:{{ atlas_gitea_staging_http_port }}:3000
PublishPort={{ atlas_gitea_staging_bind_address }}:{{ atlas_gitea_staging_ssh_port }}:2222
{% endif %}
Volume={{ atlas_gitea_mountpoint }}/data:/var/lib/gitea:Z
Volume={{ atlas_gitea_mountpoint }}/config:/etc/gitea:Z
NoNewPrivileges=true
DropCapability=all
[Service]
Restart=on-failure
RestartSec=10
TimeoutStartSec=900
{% if atlas_gitea_production_enabled | bool %}
[Install]
WantedBy=default.target
{% endif %}

View File

@@ -44,7 +44,7 @@
when: server_backup_export_enabled | bool
- name: Install Prometheus backup export helper
tags: [services, backup, prometheus_backup]
tags: [services, backup, prometheus_backup, gitea_cutover]
ansible.builtin.template:
src: prometheus-backup-export.sh.j2
dest: /usr/local/sbin/prometheus-backup-export

View File

@@ -0,0 +1,34 @@
---
- name: Install the explicit Gitea final-export helper
tags: [services, gitea_final_export]
ansible.builtin.template:
src: prometheus-gitea-final-export.sh.j2
dest: /usr/local/sbin/prometheus-gitea-final-export
owner: root
group: root
mode: "0750"
when: server_gitea_cutover_tools_enabled | bool
- name: Require the prepared source and explicit final-export approval
tags: [services, gitea_final_export]
ansible.builtin.assert:
that:
- server_gitea_cutover_tools_enabled | bool
- server_backup_export_enabled | bool
- not ansible_check_mode
fail_msg: >-
Install the cutover helper and perform an explicit non-check-mode run
only after the Gitea outage gate has been approved.
when: server_gitea_final_export | bool
- name: Stop source Gitea and publish the final consistent export
tags: [services, gitea_final_export]
ansible.builtin.command:
argv:
- /usr/local/sbin/prometheus-gitea-final-export
register: server_gitea_final_export_result
changed_when: server_gitea_final_export_result.rc == 0
no_log: true
when:
- server_gitea_final_export | bool
- not ansible_check_mode

View File

@@ -0,0 +1,59 @@
---
- name: Validate the NPM Gitea cutover override
tags: [services, gitea_cutover]
ansible.builtin.assert:
that:
- server_gitea_cutover_tools_enabled | bool
- server_gitea_npm_domains | length > 0
- server_gitea_npm_domains | select('match', '^[a-zA-Z0-9.-]+$') | list | length == server_gitea_npm_domains | length
fail_msg: Declare the exact NPM Gitea hostnames before enabling the Atlas upstream.
when: server_gitea_on_atlas | bool
- name: Ensure the NPM custom configuration directory exists
tags: [services, gitea_cutover]
ansible.builtin.file:
path: /opt/npm/data/nginx/custom
state: directory
owner: root
group: root
mode: "0755"
when: server_gitea_cutover_tools_enabled | bool
- name: Render the Gitea-only NPM runtime upstream override
tags: [services, gitea_cutover]
ansible.builtin.template:
src: prometheus-gitea-npm-proxy.conf.j2
dest: /opt/npm/data/nginx/custom/server_proxy.conf
owner: root
group: root
mode: "0644"
register: server_gitea_npm_override
when: server_gitea_on_atlas | bool
- name: Remove the Gitea NPM override when source routing is selected
tags: [services, gitea_cutover]
ansible.builtin.file:
path: /opt/npm/data/nginx/custom/server_proxy.conf
state: absent
when:
- server_gitea_cutover_tools_enabled | bool
- not server_gitea_on_atlas | bool
- name: Validate NPM configuration after a Gitea upstream change
tags: [services, gitea_cutover]
ansible.builtin.command:
argv: [podman, exec, nginx-proxy-manager, nginx, -t]
changed_when: false
when:
- server_gitea_on_atlas | bool
- server_gitea_npm_override is changed
- not ansible_check_mode
- name: Reload NPM after validating the Gitea upstream change
tags: [services, gitea_cutover]
ansible.builtin.command:
argv: [podman, exec, nginx-proxy-manager, nginx, -s, reload]
when:
- server_gitea_on_atlas | bool
- server_gitea_npm_override is changed
- not ansible_check_mode

View File

@@ -0,0 +1,61 @@
---
- name: Validate the Prometheus Gitea SSH cutover inputs
tags: [services, gitea_cutover]
ansible.builtin.assert:
that:
- server_gitea_cutover_tools_enabled | bool
- server_gitea_atlas_address is match('^[0-9]{1,3}(\.[0-9]{1,3}){3}$')
- server_gitea_ssh_public_port | int > 1024
- server_gitea_ssh_public_port | int < 65536
- server_gitea_ssh_target_port | int > 1024
- server_gitea_ssh_target_port | int < 65536
- server_gitea_ssh_public_port | int != 22
fail_msg: Keep administrative SSH on 22 and provide the Atlas rootless Gitea SSH endpoint.
when: server_gitea_on_atlas | bool
- name: Install the Gitea SSH socket proxy units without activating them
tags: [services, gitea_cutover]
ansible.builtin.template:
src: "{{ item }}.j2"
dest: "/etc/systemd/system/{{ item }}"
owner: root
group: root
mode: "0644"
loop:
- prometheus-gitea-ssh-proxy.socket
- prometheus-gitea-ssh-proxy.service
loop_control:
label: "{{ item }}"
register: server_gitea_ssh_proxy_units
when: server_gitea_cutover_tools_enabled | bool
- name: Reload systemd after Gitea SSH proxy unit changes
tags: [services, gitea_cutover]
ansible.builtin.systemd:
daemon_reload: true
when:
- server_gitea_cutover_tools_enabled | bool
- server_gitea_ssh_proxy_units is changed
- not ansible_check_mode
- name: Manage the public Gitea SSH socket separately from administrative SSH
tags: [services, gitea_cutover]
ansible.builtin.systemd:
name: prometheus-gitea-ssh-proxy.socket
state: "{{ 'started' if server_gitea_on_atlas | bool else 'stopped' }}"
enabled: "{{ server_gitea_on_atlas | bool }}"
when:
- server_gitea_cutover_tools_enabled | bool
- not ansible_check_mode
- name: Open only the public Gitea SSH port after cutover
tags: [services, gitea_cutover]
ansible.posix.firewalld:
port: "{{ server_gitea_ssh_public_port }}/tcp"
zone: "{{ server_firewalld_zone }}"
state: "{{ 'enabled' if server_gitea_on_atlas | bool else 'disabled' }}"
permanent: true
immediate: true
when:
- server_gitea_cutover_tools_enabled | bool
- server_firewall_backend == 'firewalld'

View File

@@ -37,7 +37,7 @@
label: "{{ item.dest }}"
- name: Render server templates
tags: [dotfiles, dotfiles:server]
tags: [dotfiles, dotfiles:server, gitea_cutover]
ansible.builtin.template:
src: "{{ item.src }}"
dest: "{{ item.dest if item.dest.startswith('/') else server_user_home ~ '/' ~ item.dest }}"
@@ -59,6 +59,15 @@
- name: Import Prometheus backup export job tasks
ansible.builtin.import_tasks: backup_export_job.yml
- name: Import explicit Prometheus Gitea final-export tasks
ansible.builtin.import_tasks: gitea_final_export.yml
- name: Import Prometheus Gitea SSH proxy tasks
ansible.builtin.import_tasks: gitea_ssh_proxy.yml
- name: Import Prometheus Gitea NPM proxy override tasks
ansible.builtin.import_tasks: gitea_npm_proxy.yml
- name: Ensure server SSH authorized key fragments directory exists
tags: [services, ssh]
ansible.builtin.file:

View File

@@ -59,7 +59,7 @@ stack_stopped=true
systemctl stop "$stack_unit"
tar --acls --xattrs --selinux "${excludes[@]}" -C / -cf "$stage/payload.tar" "${paths[@]}"
systemctl start "$stack_unit"
for container in nginx-proxy-manager gitea; do
for container in nginx-proxy-manager{% if not server_gitea_on_atlas | bool %} gitea{% endif %}; do
running=false
for _ in {1..30}; do
if [[ $(podman inspect --format '{{ '{{.State.Running}}' }}' "$container" 2>/dev/null) == true ]]; then

View File

@@ -0,0 +1,80 @@
#!/usr/bin/env bash
set -Eeuo pipefail
umask 077
export_root={{ server_backup_export_root | quote }}
versions="$export_root/versions"
stamp=$(date -u +%Y%m%dT%H%M%SZ)
stage=''
gitea_stopped=false
exec 9>/run/lock/prometheus-backup-export.lock
flock -n 9 || { echo 'A Prometheus backup export is already running' >&2; exit 1; }
cleanup() {
local rc=$?
trap - EXIT
if (( rc != 0 )) && "$gitea_stopped"; then
podman start gitea >/dev/null || rc=1
fi
if (( rc != 0 )) && [[ -n "$stage" && -d "$stage" ]]; then
rm -rf -- "$stage"
fi
exit "$rc"
}
trap cleanup EXIT
trap 'exit 129' HUP
trap 'exit 130' INT
trap 'exit 143' TERM
systemctl is-active --quiet podman-compose-server.service || {
echo 'Prometheus Compose stack is not active' >&2; exit 1;
}
if systemctl is-active --quiet prometheus-backup-export.timer; then
echo 'Stop the scheduled export timer for the cutover first' >&2
exit 1
fi
[[ $(podman inspect --format '{{ '{{.State.Running}}' }}' gitea) == true ]] || {
echo 'Source Gitea must be running before the final export' >&2; exit 1;
}
[[ -d /opt/gitea/data && -d /home/git/.ssh ]] || {
echo 'Required source Gitea paths are missing' >&2; exit 1;
}
[[ ! -e "$versions/$stamp" ]] || {
echo 'Final export timestamp already exists' >&2; exit 1;
}
gitea_stopped=true
podman stop --time 30 gitea >/dev/null
[[ $(podman inspect --format '{{ '{{.State.Running}}' }}' gitea) == false ]] || {
echo 'Source Gitea did not stop' >&2; exit 1;
}
python3 - <<'PY'
import sqlite3
path = '/opt/gitea/data/gitea/gitea.db'
with sqlite3.connect(f'file:{path}?mode=ro', uri=True) as database:
if database.execute('PRAGMA quick_check').fetchone()[0] != 'ok':
raise SystemExit('Source Gitea SQLite quick_check failed')
PY
stage=$(mktemp -d "$export_root/.staging.XXXXXXXX")
tar --acls --xattrs --selinux -C / -cf "$stage/payload.tar" \
opt/gitea/data home/git/.ssh
tar -tf "$stage/payload.tar" >/dev/null
(cd "$stage" && sha256sum payload.tar >payload.sha256)
printf '{"schema":1,"host":"prometheus","purpose":"gitea-cutover","created_utc":"%s"}\n' \
"$stamp" >"$stage/metadata.json"
[[ $(podman inspect --format '{{ '{{.State.Running}}' }}' gitea) == false ]] || {
echo 'Source Gitea restarted during final export' >&2; exit 1;
}
chown root:{{ server_backup_username }} "$stage" "$stage/payload.tar" \
"$stage/payload.sha256" "$stage/metadata.json"
chmod 0750 "$stage"
chmod 0640 "$stage/payload.tar" "$stage/payload.sha256" "$stage/metadata.json"
mv -- "$stage" "$versions/$stamp"
stage=''
ln -s "$stamp" "$versions/.current.new"
mv -Tf -- "$versions/.current.new" "$versions/current"
echo "Prepared final Gitea export $stamp; source Gitea remains stopped"

View File

@@ -0,0 +1,6 @@
# Managed by Ansible. NPM's variable proxy upstream uses Nginx DNS, not /etc/hosts.
{% for domain in server_gitea_npm_domains %}
if ($host = {{ domain }}) {
set $server {{ server_gitea_atlas_address }};
}
{% endfor %}

View File

@@ -0,0 +1,12 @@
[Unit]
Description=Forward public Gitea SSH to Atlas through Aegis
Requires=prometheus-gitea-ssh-proxy.socket
After=network-online.target wg-quick@wg0.service
[Service]
ExecStart=/usr/lib/systemd/systemd-socket-proxyd {{ server_gitea_atlas_address }}:{{ server_gitea_ssh_target_port }}
DynamicUser=true
NoNewPrivileges=true
ProtectSystem=strict
ProtectHome=true
PrivateTmp=true

View File

@@ -0,0 +1,9 @@
[Unit]
Description=Public Gitea SSH socket on Prometheus
[Socket]
ListenStream=0.0.0.0:{{ server_gitea_ssh_public_port }}
NoDelay=true
[Install]
WantedBy=sockets.target

View File

@@ -38,6 +38,7 @@ services:
# networks:
# - web
{% if not server_gitea_on_atlas | bool %}
gitea:
image: docker.gitea.com/gitea:1.25.2
container_name: gitea
@@ -55,6 +56,7 @@ services:
ports:
- "3000:3000"
- "127.0.0.1:222:22"
{% endif %}
networks:

View File

@@ -0,0 +1,227 @@
# Gitea migration from Prometheus to Atlas
This records the staged migration and its observed partial cutover. Gitea is
temporary on Atlas until Uranus; NPM remains on Prometheus. Preserve the old
Prometheus data, but do not restart its stale Gitea after Atlas accepts writes.
## Observed source before cutover and chosen topology (2026-10-01)
- Prometheus runs the rootful `docker.gitea.com/gitea:1.25.2` image in its
managed Compose stack. `/opt/gitea/data` is about 280 MiB, uses SQLite,
and contains 33 repositories. A live read-only SQLite `quick_check` passed.
`/home/git/.ssh` is a separate small bind mount; `/opt/gitea/data/ssh`
contains the existing SSH host keys. Neither tree may be discarded.
- Gitea answers HTTP 200 on Prometheus port 3000. NPM currently forwards
`git.fscotto.duckdns.org` and `git.ov-ad3410.infomaniak.ch` to the Compose
hostname `gitea:3000`. Public DNS resolves to Prometheus. The container's
SSH port is bound only to `127.0.0.1:222`; this is not a public Gitea SSH
listener. Prometheus' public port 22 remains administrative SSH.
- Atlas has a healthy pool and a verified, private Prometheus backup under
`/zpool/backup/hosts/prometheus/latest`. The 2026-10-01 scheduled export
and pull succeeded. The intended target is a separate
`/zpool/services/data/gitea` dataset, not `Archive` or the backup dataset.
- The approved cutover keeps NPM on Prometheus, changes the two HTTP Proxy
Hosts' effective upstream to Atlas over the Prometheus--Aegis gateway, and offers public Gitea
SSH on port 2222 via the same gateway. Prometheus port 22 is unchanged.
HTTPS and SSH must be validated together before declaring cutover.
- The initial staging ran as a **rootless user Quadlet** under a dedicated,
non-login Atlas account, using the pinned `1.25.2-rootless` image. This was an explicit
rootful-to-rootless **data-layout conversion**, not a drop-in image swap:
the target mounts `/var/lib/gitea` and `/etc/gitea`, and uses Gitea's
built-in SSH server instead of the source image's OpenSSH daemon. Keep the
application version unchanged until the conversion has passed an isolated
restore test. The host's rootful Quadlet directory must not be used.
## Phase 1: prepare without traffic changes
Preparation completed on 2026-10-01: Ansible created
`zpool/services/data/gitea`, a dedicated non-login `gitea` account (UID/GID
1101), separate subordinate IDs, parent-dataset traverse ACLs, and an inactive
user Quadlet under `/var/lib/atlas-gitea/.config/containers/systemd/`. The
Quadlet has no `[Install]` section and, until the final cutover, binds only
loopback staging ports 3001/2223 if started manually. A second targeted
Ansible run changed nothing; the generated service was inactive and neither
staging port listened.
The explicit rehearsal is managed by:
```bash
ansible-playbook ansible/site.yml --limit atlas --tags gitea_restore \
-e atlas_gitea_restore_test=true
```
On 2026-10-01 this selected the latest verified Prometheus backup, checked its
SHA-256, extracted only `opt/gitea/data`, moved `app.ini` into the rootless
config mount, rewrote `/data/` paths, enabled built-in SSH on internal port
2222, and retained the three source SSH host-key pairs. SQLite `quick_check`
passed, all 33 restored repositories passed `git fsck`, and each source/target
public host-key fingerprint matched. A temporary `1.25.2-rootless` container
with `--network none` answered HTTP internally and listened on internal
SSH/2222. The container was removed; the user Quadlet remains inactive, with
no staging listener. The second restore run changed nothing. This copy is
deliberately stale once new source writes occur and **must not** be used as the
final cutover copy.
Target backup checks on 2026-10-01: the managed recursive hourly ZFS snapshot
`atlas-auto-hourly-20261001T193401Z` contains the new dataset. The managed
Borg service completed archive `atlas-20261001T193420Z`, whose contents list
includes the staged Gitea database. A separate one-file restore from each
source into private `/var/tmp` directories matched the live staged database
and passed SQLite `quick_check`. Temporary files and the on-demand snapshot
mount were removed; the Borg temporary snapshot was cleaned up and the pool
remained healthy. This is file-level proof, **not** a full Gitea recovery.
The operator's UUID-bound offline USB run published version
`20261001T201220Z-254397` on 2026-10-02. A separate read-only mount and
temporary restore of `services/data/gitea/data/gitea/gitea.db` matched
contents, owner, group, mode, size, mtime and POSIX ACL; SQLite
`quick_check` returned `ok`. The temporary mount and copy were removed,
LUKS was closed, and the pool was healthy. This is a file-level restore test,
not a complete Gitea recovery rehearsal from USB.
1. Provision a dedicated target dataset and non-login service identity via
Ansible, keeping UID/GID distinct from Atlas' reserved Immich `1100`.
Install the user Quadlet in that identity's
`~/.config/containers/systemd/`, **without** an `[Install]` section;
do not enable, start, or expose it yet.
2. Verify the selected Atlas backup SHA-256 and metadata, then extract **only**
`opt/gitea/data` to private staging. Keep `home/git/.ssh` in the source
backup for rollback; the rootless image does not consume its OpenSSH mount.
Never unpack NPM,
WireGuard, or other host configuration from this sensitive tarball into a
live namespace. Convert the rootful `/data` tree on a disposable copy:
place application data under `/var/lib/gitea`, move `app.ini` to
`/etc/gitea`, and rewrite every absolute `/data/...` path for the new
layout. Enable `START_SSH_SERVER`, use internal SSH port 2222, and retain
the source host-key pairs for the built-in server only after verifying
their fingerprints and compatibility. Do not rely on the old
`/home/git/.ssh` OpenSSH mount in the rootless image. Set only the target
copy's ownership and path-scoped SELinux labels.
3. Validate SQLite integrity, repository count and representative `git fsck`,
LFS/attachment presence, permissions, and an isolated rootless test
container with no production ingress or outbound network. Because the
source stays active, this is a rehearsal copy, not the final cutover copy.
Regenerate Git hooks if the changed installation path requires it.
4. ZFS, Borg and UUID-bound offline USB inclusion and one-file restores have
passed. These do not replace the final consistent source copy.
## Phase 2: explicit final cutover
The opt-in `/usr/local/sbin/prometheus-gitea-final-export` helper was installed
on 2026-10-01 and passed `bash -n`. It refuses to
run while the scheduled Prometheus export timer is active. When explicitly
triggered, it stops only the source Gitea container, checks SQLite, publishes
a checksum-verified Gitea-only version for Atlas' existing pull, and leaves
the source stopped on success. NPM remains running. A failure before
completion restarts source Gitea. Its Ansible gate is
`--tags gitea_final_export -e server_gitea_final_export=true`.
After Atlas pulls that version, its separate
`--tags gitea_final_restore -e atlas_gitea_final_restore=true` gate accepts
only metadata marked `gitea-cutover`, validates a private staged replacement,
and swaps it for the marked rehearsal. The swap and its rollback path passed
synthetic tests on 2026-10-01; the live gate succeeded on 2026-10-02.
On 2026-10-02 the operator approved the outage. The final stopped-source
export `20261002T071525Z` passed the Atlas pull checksum; the guarded restore
replaced the rehearsal. SQLite `quick_check`, all 33 repository `git fsck`
checks, and the source/target SSH host-key comparison passed. The rootless
Atlas Quadlet serves LAN HTTP/3000 and SSH/2222, reachable from Prometheus
through Aegis; its firewall admits only Aegis. The final marker gates startup.
Prometheus now runs the NPM-only Compose stack. Both NPM database records still
say `gitea:3000`, but Nginx evaluates this variable upstream through its
runtime DNS resolver, which **does not** use a Compose `extra_hosts` alias.
The initial alias attempt returned 502. A managed `server_proxy.conf` override
sets `$server` to Atlas' IP for only the two declared Gitea domains; it passed
`nginx -t` and primary HTTPS/API returned 200 after a clean NPM restart
without the alias; a representative public `git ls-remote` also succeeded.
Navidrome and Syncthing Proxy Hosts still responded. No NPM SQLite records
or credentials were changed. The
secondary hostname `git.ov-ad3410.infomaniak.ch` did not resolve from Ikaros
and had no generated NPM config file at the time of inspection.
Prometheus' public TCP/2222 socket proxies to Atlas without changing admin
SSH/22. The local socket presents the preserved Gitea ED25519 host key, but
an external TCP/2222 connection from Ikaros initially timed out. During that
test no SYN reached Prometheus `eth0`; its socket and firewalld port were active.
After the VPS firewall was opened later on 2026-10-02, the public port connected,
its ED25519 host-key fingerprint matched Atlas, Gitea authenticated the `ikaros`
key as `fscotto`, and a public SSH `git ls-remote` for `fscotto/infra.git`
returned HEAD. The operator subsequently reported successful authenticated
SSH pull and push; the agent did not perform a write test. HTTPS write/login
remain untested. Do not
restart the stale source after public HTTPS has accepted target writes.
The Prometheus export timer resumed with NPM-only paths. A recursive ZFS
snapshot at `20261002T073032Z` and encrypted Borg archive
`atlas-20261002T073044Z` captured the Atlas target after cutover; Borg exited
successfully, cleaned its temporary snapshot, and the pool was healthy.
## Corrected Atlas service owner (2026-10-02)
The operator required the host Quadlet to belong to `admin`, while the Unix
user **inside** the container must be named `gitea`. The pinned derived
`Containerfile.gitea-rootless` changes only the base image's UID/GID 1000
passwd/group names from `git` to `gitea`; it retains the rootless image's
paths and entrypoint. Gitea's `RUN_USER` is `gitea`, while its built-in SSH
user and advertised clone user remain `git`, preserving `git@` URLs. The
selective restore helper now generates the same three settings for any future
explicit restore, instead of recreating a `RUN_USER = git` target.
A disposable, loopback-only container using a copy of a Gitea ZFS snapshot
passed HTTP, SQLite, internal-user and SSH host-key checks without touching
live data. After explicit outage approval, the opt-in
`--tags gitea_owner_migration -e atlas_gitea_owner_migration=true` run stopped
the old user service, took safety snapshot
`zpool/services/data/gitea@gitea-owner-migration-20261002T100104`, transferred
only the Gitea dataset to `admin`, tested an `admin` staging Quadlet on
loopback, then promoted it to the production LAN ports. The old Atlas Quadlet
was removed. The old host `gitea` account and its sub-ID range are retained
for a deliberate rollback; they must not restart stale Gitea. The parent
traverse ACL is removed by the normal Gitea role once the new owner is live.
The new service returned HTTP 200 locally and through public primary HTTPS;
Navidrome and Syncthing remained active under `admin`, the pool was healthy,
and a second normal Gitea Ansible run was idempotent. This does **not** close
the separate external TCP/2222 or authenticated clone/push validation gap.
1. Agree on an outage and record source/target versions, pool health, the
latest backups, SSH host-key fingerprints, and both current NPM routes.
Stop the Prometheus export timer for the change window so it cannot
restart the old Compose stack unexpectedly.
2. Quiesce source writes with the final-export helper: it stops Gitea before
the consistent export and leaves it stopped after success. Pull that export
to Atlas and verify checksum and timestamp. Keep
`/opt/gitea/data` and `/home/git/.ssh` intact for rollback. Do not allow
source Gitea to restart after accepting writes on Atlas.
3. Restore the final Gitea-only payload to the target and repeat integrity
checks. Verify its advertised SSH port is 2222, its existing HTTPS
`ROOT_URL`, repositories, LFS/attachments, and SSH host-key identity. Enable
the production Atlas Quadlet only after the final-restore marker exists;
its firewall permits only Aegis to reach HTTP and SSH. Validate local HTTP
and the target service before switching NPM.
4. Enable the public TCP/2222 socket proxy on Prometheus to Atlas over Aegis
without changing administrative TCP/22. Switch Prometheus to the desired
NPM-only Compose stack and use the managed Gitea-only NPM runtime upstream
override. Do not use Compose `extra_hosts`: Nginx bypasses it for the
variable upstream. The old Gitea data stays intact. Do not change public DNS.
5. Test HTTPS login, representative clone/push, LFS, and public SSH clone/push
on port 2222 from outside the Atlas LAN. Record the last source write and
first healthy target service times; do not claim RPO/RTO without measuring.
6. Resume the Prometheus NPM-only backup export timer after the desired stack
is active and verify its next result. Verify the next Atlas snapshot/Borg
run covers Gitea and test a restored target copy. Do not delete old source
data.
## Rollback gate
Before Atlas accepts writes, restore the old Compose definition and remove the
NPM override, disable the public 2222 proxy, and restart the unchanged source
Gitea if target validation fails. **After Atlas accepts writes, do not blindly restart the source:** its
SQLite database and repositories are stale. Quiesce Atlas, capture its new
data, and decide a reverse migration or an extended outage explicitly.
Upstream references: [rootful container layout](https://docs.gitea.com/1.25/installation/install-with-docker/),
[rootless image layout and incompatibility](https://docs.gitea.com/installation/install-with-docker-rootless/),
[rootless Podman Quadlet](https://docs.gitea.com/installation/install-with-podman-quadlet/),
[standard-image conversion](https://docs.gitea.com/1.24/installation/install-with-docker-rootless/),
and [restore and hook regeneration](https://docs.gitea.com/1.26/administration/backup-and-restore/).