From 06d3b175cbbb5ff5da31f430000ab91a00fcb440 Mon Sep 17 00:00:00 2001 From: Fabio Scotto di Santolo Date: Thu, 1 Oct 2026 21:33:18 +0200 Subject: [PATCH] Rehearse rootless Gitea restore from verified backup --- AGENTS.md | 14 +- README.it.md | 3 +- README.md | 7 +- ansible/roles/profile_atlas/defaults/main.yml | 2 + .../files/atlas-gitea-restore-test.py | 196 ++++++++++++++++++ .../profile_atlas/tasks/gitea_restore.yml | 84 ++++++++ ansible/roles/profile_atlas/tasks/main.yml | 3 + docs/atlas-gitea-migration.md | 26 ++- 8 files changed, 326 insertions(+), 9 deletions(-) create mode 100644 ansible/roles/profile_atlas/files/atlas-gitea-restore-test.py create mode 100644 ansible/roles/profile_atlas/tasks/gitea_restore.yml diff --git a/AGENTS.md b/AGENTS.md index 2cfcf91..b060986 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -59,6 +59,8 @@ Ansible-driven personal infrastructure repo for Fedora and Void desktops, Fedora `ansible-playbook ansible/site.yml --limit atlas --tags storage,sharing,containers --check --diff` - Atlas rootless Gitea staging (does not start Gitea): `ansible-playbook ansible/site.yml --limit atlas --tags gitea --check --diff` + - Atlas explicit isolated Gitea restore rehearsal (not part of normal runs): + `ansible-playbook ansible/site.yml --limit atlas --tags gitea_restore -e atlas_gitea_restore_test=true` - Atlas network/share hardening: `ansible-playbook ansible/site.yml --limit atlas --tags hardening,sharing --check --diff` - Atlas ZFS snapshot retention and scrub timers: @@ -265,13 +267,19 @@ successfully. The first monthly scrub remains a runtime check. - [x] Design the staged Prometheus-to-Atlas Gitea migration in `docs/atlas-gitea-migration.md`. The approved topology keeps NPM on Prometheus and moves HTTPS and public SSH (TCP/2222) together; Gitea must run as a dedicated rootless user Quadlet on Atlas. The rootful-to-rootless data-layout - conversion requires an isolated restore test. No Gitea data has been moved or traffic changed. + conversion passed an isolated restore rehearsal. No production traffic has changed. - [x] Prepare the dedicated Atlas Gitea dataset, non-login UID/GID 1101 with a separate rootless Podman sub-ID range, and disabled user Quadlet. On 2026-10-01 the targeted Ansible run and a second idempotent run passed; the generated unit was inactive, with no staging HTTP/SSH listener. POSIX ACLs on only the service-namespace parents grant this account traversal without access to sibling datasets. -- [ ] Perform an isolated rootless restore test from the verified Prometheus backup; validate SQLite, - repositories, SSH host keys, and target backups before any traffic cutover. +- [x] Perform an isolated rootless restore rehearsal from the verified Prometheus backup. On 2026-10-01 + the SHA-256-checked selective extraction and path/SSH conversion succeeded; SQLite `quick_check` + passed, all 33 repositories passed `git fsck`, and source/target public SSH host-key fingerprints + matched. The pinned rootless image answered HTTP and listened on internal SSH/2222 with + `--network none`; the temporary container was removed and the Quadlet stayed inactive. A second + restore run made no changes. This is a rehearsal copy, not the final consistent cutover copy. +- [ ] Verify target Gitea backup coverage with a new recursive ZFS snapshot, Borg archive, and an + independent restore of the staged dataset before accepting production writes. - [ ] After an explicit outage approval, perform the final consistent copy and HTTPS/SSH cutover, then remove Gitea from Prometheus' desired stack and backup export without deleting source data. - [ ] Design and deploy Nextcloud as another explicitly temporary Atlas service before Uranus. Give it diff --git a/README.it.md b/README.it.md index e987874..7b0229b 100644 --- a/README.it.md +++ b/README.it.md @@ -324,7 +324,8 @@ ricarica le reti Podman rootful di Prometheus per conservare DNS e connettività La migrazione Gitea da Prometheus ad Atlas è predisposta in [`docs/atlas-gitea-migration.md`](docs/atlas-gitea-migration.md). Atlas ha un dataset e un account -dedicati con Quadlet utente rootless inattivo. I dati Gitea non sono ancora stati ripristinati o spostati. +dedicati con Quadlet utente rootless inattivo. Una copia isolata del backup Prometheus ha superato +i controlli SQLite, Git e del container rootless senza rete; non è la copia finale per il cutover. NPM resta su Prometheus; stack sorgente e instradamento pubblico rimangono invariati fino a un cutover HTTPS e SSH separato e validato. diff --git a/README.md b/README.md index ba6228b..04453db 100644 --- a/README.md +++ b/README.md @@ -300,9 +300,10 @@ and Aegis (`10.0.0.2`). Their state is initialized ex novo in `/zpool/services/d The Gitea move from Prometheus to Atlas is staged in [`docs/atlas-gitea-migration.md`](docs/atlas-gitea-migration.md). Atlas has a dedicated dataset and -non-login account with an inactive rootless user Quadlet. No Gitea data has been restored or moved yet. -NPM remains on Prometheus; the source stack and public routes stay unchanged until a separately -validated HTTPS and SSH cutover. +non-login account with an inactive rootless user Quadlet. An isolated copy from the verified +Prometheus backup passed SQLite, Git, and network-disabled +rootless-container checks; it is not the final cutover copy. NPM remains on Prometheus; the source +stack and public routes stay unchanged until a separately validated HTTPS and SSH cutover. The separate `wireguard_overlay` role manages `wg0` between Prometheus (`10.0.0.1`) and Aegis (`10.0.0.2`), generating private keys once on their respective hosts and exchanging only public keys diff --git a/ansible/roles/profile_atlas/defaults/main.yml b/ansible/roles/profile_atlas/defaults/main.yml index 96397fd..d10b8e3 100644 --- a/ansible/roles/profile_atlas/defaults/main.yml +++ b/ansible/roles/profile_atlas/defaults/main.yml @@ -181,6 +181,8 @@ atlas_gitea_image: docker.gitea.com/gitea:1.25.2-rootless atlas_gitea_staging_bind_address: 127.0.0.1 atlas_gitea_staging_http_port: 3001 atlas_gitea_staging_ssh_port: 2223 +atlas_gitea_restore_test: false +atlas_gitea_restore_helper: /usr/local/libexec/atlas-gitea-restore-test atlas_45drives_repo_url: https://repo.45drives.com/repofiles/rocky/45drives-enterprise.repo atlas_45drives_repo_file: /etc/yum.repos.d/45drives-enterprise.repo diff --git a/ansible/roles/profile_atlas/files/atlas-gitea-restore-test.py b/ansible/roles/profile_atlas/files/atlas-gitea-restore-test.py new file mode 100644 index 0000000..7ad80bd --- /dev/null +++ b/ansible/roles/profile_atlas/files/atlas-gitea-restore-test.py @@ -0,0 +1,196 @@ +#!/usr/bin/python3 +"""Rehearse a selective rootful-to-rootless Gitea restore, never a cutover.""" + +import argparse +import hashlib +import os +from pathlib import Path, PurePosixPath +import re +import shutil +import sqlite3 +import tarfile +import tempfile + + +SOURCE_PREFIX = PurePosixPath("opt/gitea/data") +HOST_KEYS = ( + "ssh_host_ed25519_key", + "ssh_host_rsa_key", + "ssh_host_ecdsa_key", +) +SERVER_SETTINGS = { + "START_SSH_SERVER": "true", + "SSH_PORT": "2222", + "SSH_LISTEN_PORT": "2222", + "SSH_SERVER_HOST_KEYS": ", ".join( + f"/var/lib/gitea/ssh/{key}" for key in HOST_KEYS + ), +} + + +def sha256(path): + digest = hashlib.sha256() + with path.open("rb") as stream: + for chunk in iter(lambda: stream.read(1024 * 1024), b""): + digest.update(chunk) + return digest.hexdigest() + + +def expected_digest(backup): + checksum = (backup / "payload.sha256").read_text().strip().split() + if len(checksum) != 2 or checksum[1] != "payload.tar": + raise ValueError("Unexpected Prometheus backup checksum manifest") + if not re.fullmatch(r"[0-9a-f]{64}", checksum[0]): + raise ValueError("Invalid Prometheus backup SHA-256") + return checksum[0] + + +def convert_config(config): + original = config.read_text() + output = [] + section = "" + server_seen = set() + server_found = False + + def append_missing_server_settings(): + for key, value in SERVER_SETTINGS.items(): + if key not in server_seen: + output.append(f"{key} = {value}\n") + + for line in original.splitlines(keepends=True): + match = re.match(r"^\s*\[([^]]+)\]\s*$", line) + if match: + if section == "server": + append_missing_server_settings() + section = match.group(1).lower() + server_found |= section == "server" + output.append(line) + continue + setting = re.match(r"^(\s*)([A-Z_]+)(\s*=\s*)(.*?)(\r?\n?)$", line) + if setting and section == "server" and setting.group(2) in SERVER_SETTINGS: + key = setting.group(2) + server_seen.add(key) + line = f"{setting.group(1)}{key}{setting.group(3)}{SERVER_SETTINGS[key]}{setting.group(5)}" + else: + line = line.replace("/data/", "/var/lib/gitea/") + output.append(line) + if section == "server": + append_missing_server_settings() + if not server_found: + raise ValueError("Gitea server configuration missing") + config.write_text("".join(output)) + config.chmod(0o600) + + +def extract_gitea(tar_path, staged_data): + count = 0 + with tarfile.open(tar_path, mode="r") as archive: + for member in archive: + name = PurePosixPath(member.name) + if name == SOURCE_PREFIX: + continue + if SOURCE_PREFIX not in name.parents: + continue + relative = name.relative_to(SOURCE_PREFIX) + if not relative.parts or any(part in (".", "..") for part in relative.parts): + raise ValueError("Unsafe Gitea backup path") + if not (member.isdir() or member.isfile()): + raise ValueError("Unexpected Gitea backup member type") + destination = staged_data.joinpath(*relative.parts) + if member.isdir(): + destination.mkdir(parents=True, exist_ok=True) + destination.chmod(0o700) + continue + destination.parent.mkdir(parents=True, exist_ok=True) + with archive.extractfile(member) as source, destination.open("xb") as target: + shutil.copyfileobj(source, target) + destination.chmod(member.mode & 0o777) + count += 1 + if count == 0: + raise ValueError("No Gitea files in backup") + + +def validate(staged_data, staged_config): + database = staged_data / "gitea/gitea.db" + repositories = staged_data / "git/repositories" + if not database.is_file() or not repositories.is_dir(): + raise ValueError("Missing SQLite database or Git repositories") + with sqlite3.connect(f"file:{database}?mode=ro", uri=True) as connection: + if connection.execute("PRAGMA quick_check").fetchone()[0] != "ok": + raise ValueError("Gitea SQLite quick_check failed") + if connection.execute("SELECT count(*) FROM repository").fetchone()[0] < 1: + raise ValueError("Gitea backup contains no repository records") + if not any(repositories.rglob("*.git")): + raise ValueError("Gitea backup contains no Git repository directories") + if not (staged_config / "app.ini").is_file(): + raise ValueError("Gitea app.ini missing") + for name in HOST_KEYS: + if not (staged_data / "ssh" / name).is_file(): + raise ValueError("Gitea SSH host key missing") + + +def chown_tree(root, uid, gid): + for directory, dirs, files in os.walk(root): + os.chown(directory, uid, gid) + for name in dirs + files: + os.chown(os.path.join(directory, name), uid, gid) + + +def main(): + parser = argparse.ArgumentParser() + parser.add_argument("--backup", type=Path, required=True) + parser.add_argument("--target", type=Path, required=True) + parser.add_argument("--uid", type=int, required=True) + parser.add_argument("--gid", type=int, required=True) + args = parser.parse_args() + + backup = args.backup.resolve(strict=True) + target = args.target.resolve(strict=True) + if not str(backup).startswith("/zpool/backup/hosts/prometheus/snapshots/"): + raise ValueError("Refusing backup outside the Atlas Prometheus snapshots") + if str(target) != "/zpool/services/data/gitea": + raise ValueError("Refusing target outside the dedicated Gitea dataset") + if args.uid != 1101 or args.gid != 1101: + raise ValueError("Unexpected dedicated Gitea account IDs") + expected = expected_digest(backup) + if sha256(backup / "payload.tar") != expected: + raise ValueError("Prometheus backup SHA-256 mismatch") + + marker = target / ".rehearsal-sha256" + if marker.exists(): + if marker.read_text().strip() != expected: + raise ValueError("A different Gitea rehearsal already occupies this dataset") + validate(target / "data", target / "config") + print("unchanged") + return + for name in ("data", "config"): + directory = target / name + if not directory.is_dir() or any(directory.iterdir()): + raise ValueError("Gitea target is not empty; refusing overwrite") + + with tempfile.TemporaryDirectory(prefix=".rehearsal-", dir=target) as temporary: + stage = Path(temporary) + staged_data = stage / "data" + staged_config = stage / "config" + staged_data.mkdir() + staged_config.mkdir() + extract_gitea(backup / "payload.tar", staged_data) + source_config = staged_data / "gitea/conf/app.ini" + if not source_config.is_file(): + raise ValueError("Source Gitea app.ini missing") + shutil.copy2(source_config, staged_config / "app.ini") + source_config.unlink() + convert_config(staged_config / "app.ini") + validate(staged_data, staged_config) + chown_tree(stage, args.uid, args.gid) + for name in ("data", "config"): + (target / name).rmdir() + os.rename(stage / name, target / name) + marker.write_text(expected + "\n") + marker.chmod(0o600) + os.chown(marker, args.uid, args.gid) + print("restored") + + +if __name__ == "__main__": + main() diff --git a/ansible/roles/profile_atlas/tasks/gitea_restore.yml b/ansible/roles/profile_atlas/tasks/gitea_restore.yml new file mode 100644 index 0000000..4a4a87c --- /dev/null +++ b/ansible/roles/profile_atlas/tasks/gitea_restore.yml @@ -0,0 +1,84 @@ +--- +- name: Rehearse an isolated rootless Gitea restore from the verified Prometheus backup + tags: [atlas, gitea_restore] + when: atlas_gitea_restore_test | bool + block: + - name: Require the prepared rootless Gitea target + ansible.builtin.assert: + that: + - atlas_manage_gitea | bool + - atlas_gitea_staging_bind_address == '127.0.0.1' + - atlas_gitea_mountpoint == atlas_app_data_mountpoint ~ '/gitea' + fail_msg: Prepare the isolated, loopback-only rootless Gitea target first. + + - name: Confirm the rootless Gitea service is inactive + become_user: "{{ atlas_gitea_username }}" + ansible.builtin.command: + argv: + - systemctl + - --user + - is-active + - atlas-gitea.service + environment: + XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_uid }}" + DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_gitea_uid }}/bus" + register: atlas_gitea_restore_service_state + changed_when: false + failed_when: false + when: not ansible_check_mode + + - name: Refuse to overwrite an active rootless Gitea service + ansible.builtin.assert: + that: + - atlas_gitea_restore_service_state.stdout == 'inactive' + fail_msg: The rootless Gitea user service must be known and inactive before restoring data. + when: not ansible_check_mode + + - name: Check for a manually running rootless Gitea container + become_user: "{{ atlas_gitea_username }}" + ansible.builtin.command: + argv: + - podman + - ps + - --quiet + - --filter + - name=atlas-gitea + args: + chdir: "{{ atlas_gitea_home }}" + environment: + XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_uid }}" + register: atlas_gitea_restore_container_state + changed_when: false + when: not ansible_check_mode + + - name: Refuse to overwrite a running rootless Gitea container + ansible.builtin.assert: + that: + - atlas_gitea_restore_container_state.stdout | length == 0 + fail_msg: Stop every rootless Atlas Gitea container before restoring data. + when: not ansible_check_mode + + - name: Install the selective rootless Gitea restore helper + ansible.builtin.copy: + src: atlas-gitea-restore-test.py + dest: "{{ atlas_gitea_restore_helper }}" + owner: root + group: root + mode: "0700" + + - name: Restore only Gitea data into the isolated target + ansible.builtin.command: + argv: + - "{{ atlas_gitea_restore_helper }}" + - --backup + - "{{ atlas_backup_prometheus_mountpoint }}/latest" + - --target + - "{{ atlas_gitea_mountpoint }}" + - --uid + - "{{ atlas_gitea_uid | string }}" + - --gid + - "{{ atlas_gitea_gid | string }}" + register: atlas_gitea_restore_result + changed_when: atlas_gitea_restore_result.stdout == 'restored' + no_log: true + when: not ansible_check_mode diff --git a/ansible/roles/profile_atlas/tasks/main.yml b/ansible/roles/profile_atlas/tasks/main.yml index 51dae83..6409b7e 100644 --- a/ansible/roles/profile_atlas/tasks/main.yml +++ b/ansible/roles/profile_atlas/tasks/main.yml @@ -17,6 +17,9 @@ - name: Import staged Atlas rootless Gitea tasks ansible.builtin.import_tasks: gitea.yml +- name: Import explicit Atlas Gitea restore rehearsal tasks + ansible.builtin.import_tasks: gitea_restore.yml + - name: Import Atlas ZFS maintenance tasks ansible.builtin.import_tasks: zfs_maintenance.yml diff --git a/docs/atlas-gitea-migration.md b/docs/atlas-gitea-migration.md index 8fc5061..9b599ce 100644 --- a/docs/atlas-gitea-migration.md +++ b/docs/atlas-gitea-migration.md @@ -42,7 +42,27 @@ user Quadlet under `/var/lib/atlas-gitea/.config/containers/systemd/`. The Quadlet has no `[Install]` section and, until the final cutover, binds only loopback staging ports 3001/2223 if started manually. A second targeted Ansible run changed nothing; the generated service was inactive and neither -staging port listened. **No Gitea payload has been restored to the target.** +staging port listened. + +The explicit rehearsal is managed by: + +```bash +ansible-playbook ansible/site.yml --limit atlas --tags gitea_restore \ + -e atlas_gitea_restore_test=true +``` + +On 2026-10-01 this selected the latest verified Prometheus backup, checked its +SHA-256, extracted only `opt/gitea/data`, moved `app.ini` into the rootless +config mount, rewrote `/data/` paths, enabled built-in SSH on internal port +2222, and retained the three source SSH host-key pairs. SQLite `quick_check` +passed, all 33 restored repositories passed `git fsck`, and each source/target +public host-key fingerprint matched. A temporary `1.25.2-rootless` container +with `--network none` answered HTTP internally and listened on internal +SSH/2222. The container was removed; the user Quadlet remains inactive, with +no staging listener. The second restore run changed nothing. This copy is +deliberately stale once new source writes occur and **must not** be used as the +final cutover copy. Target snapshot/Borg inclusion and an independent restore +are still pending. 1. Provision a dedicated target dataset and non-login service identity via Ansible, keeping UID/GID distinct from Atlas' reserved Immich `1100`. @@ -50,7 +70,9 @@ staging port listened. **No Gitea payload has been restored to the target.** `~/.config/containers/systemd/`, **without** an `[Install]` section; do not enable, start, or expose it yet. 2. Verify the selected Atlas backup SHA-256 and metadata, then extract **only** - `opt/gitea/data` and `home/git/.ssh` to private staging. Never unpack NPM, + `opt/gitea/data` to private staging. Keep `home/git/.ssh` in the source + backup for rollback; the rootless image does not consume its OpenSSH mount. + Never unpack NPM, WireGuard, or other host configuration from this sensitive tarball into a live namespace. Convert the rootful `/data` tree on a disposable copy: place application data under `/var/lib/gitea`, move `app.ini` to