diff --git a/AGENTS.md b/AGENTS.md index 627a45e..f642760 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -59,6 +59,8 @@ Ansible-driven personal infrastructure repo for Fedora and Void desktops, Fedora `ansible-playbook ansible/site.yml --limit atlas --tags storage,sharing,containers --check --diff` - Atlas rootless Gitea staging (does not start Gitea): `ansible-playbook ansible/site.yml --limit atlas --tags gitea --check --diff` + - Atlas explicit Gitea host-owner migration (live outage; never a normal run): + `ansible-playbook ansible/site.yml --limit atlas --tags gitea_owner_migration -e atlas_gitea_owner_migration=true` - Atlas explicit isolated Gitea restore rehearsal (not part of normal runs): `ansible-playbook ansible/site.yml --limit atlas --tags gitea_restore -e atlas_gitea_restore_test=true` - Atlas final Gitea replacement gate (dry-run only until a stopped-source export is pulled): @@ -273,7 +275,8 @@ successfully. The first monthly scrub remains a runtime check. - [ ] After data protection and recovery are validated, populate `/zpool/media/music` and validate Navidrome. - [x] Design the staged Prometheus-to-Atlas Gitea migration in `docs/atlas-gitea-migration.md`. The approved topology keeps NPM on Prometheus and moves HTTPS and public SSH (TCP/2222) together; - Gitea must run as a dedicated rootless user Quadlet on Atlas. The rootful-to-rootless data-layout + Gitea runs as an `admin`-owned rootless user Quadlet on Atlas with an internal `gitea` user. + The rootful-to-rootless data-layout conversion passed an isolated restore rehearsal. The later partial cutover is tracked below. - [x] Prepare the dedicated Atlas Gitea dataset, non-login UID/GID 1101 with a separate rootless Podman sub-ID range, and disabled user Quadlet. On 2026-10-01 the targeted Ansible run and a second idempotent @@ -316,6 +319,15 @@ successfully. The first monthly scrub remains a runtime check. Gitea container was removed from the desired Compose stack without deleting its data; the Prometheus backup export timer resumed for NPM only. A post-cutover recursive ZFS snapshot and encrypted Borg archive `atlas-20261002T073044Z` completed successfully. +- [x] Move the live Gitea Quadlet and dataset from the legacy host `gitea` account to `admin` + after a disposable snapshot-copy test of the pinned derived image. On 2026-10-02 the explicit + outage run stopped only legacy Gitea, made safety snapshot + `zpool/services/data/gitea@gitea-owner-migration-20261002T100104`, changed dataset ownership, + and validated loopback staging (HTTP 200, internal `gitea` UID/GID 1000, SQLite `quick_check`) + before promoting the `admin` Quadlet. Production LAN and public HTTPS returned 200; Navidrome + and Syncthing remained active, the pool was healthy, and the normal Gitea run changed nothing. + The old host account and data on Prometheus remain preserved; the old Atlas Quadlet and its + parent-dataset traverse ACL were removed. A subsequent normal run changed nothing. - [ ] Complete public SSH/2222 and representative authenticated HTTPS/SSH clone/push validation. Prometheus' TCP/2222 socket and firewalld rule are active and the local proxy presents the matching Atlas host key, but Ikaros' external TCP connection timed out and no SYN reached diff --git a/README.it.md b/README.it.md index 7b0229b..7882b90 100644 --- a/README.it.md +++ b/README.it.md @@ -322,12 +322,13 @@ alla LAN. Dopo la verifica dei servizi, configurare manualmente i Proxy Host NPM negli `AllowedIPs`; aggiungere la VIP Uranus quando esisterà. Dopo il reload di firewalld, Ansible ricarica le reti Podman rootful di Prometheus per conservare DNS e connettività del proxy. -La migrazione Gitea da Prometheus ad Atlas è predisposta in -[`docs/atlas-gitea-migration.md`](docs/atlas-gitea-migration.md). Atlas ha un dataset e un account -dedicati con Quadlet utente rootless inattivo. Una copia isolata del backup Prometheus ha superato -i controlli SQLite, Git e del container rootless senza rete; non è la copia finale per il cutover. -NPM resta su Prometheus; stack sorgente e instradamento pubblico rimangono invariati fino a un cutover -HTTPS e SSH separato e validato. +La migrazione Gitea da Prometheus ad Atlas è descritta in +[`docs/atlas-gitea-migration.md`](docs/atlas-gitea-migration.md). Gitea usa un Quadlet rootless +di `admin` su un dataset dedicato; l'immagine derivata mantiene UID/GID 1000 ma chiama l'utente +interno `gitea`. NPM resta su Prometheus e l'HTTPS pubblico primario serve Atlas. L'SSH pubblico +su TCP/2222 non era ancora raggiungibile dall'esterno il 2026-10-02; non considerare completo +il cutover HTTPS+SSH finché non sono validati clone/push autenticati. I dati sorgente restano +conservati su Prometheus senza avviarne il vecchio container. Validare il gateway con: diff --git a/README.md b/README.md index 76e50c7..e39ad93 100644 --- a/README.md +++ b/README.md @@ -300,7 +300,8 @@ and Aegis (`10.0.0.2`). Their state is initialized ex novo in `/zpool/services/d The Gitea move from Prometheus to Atlas is tracked in [`docs/atlas-gitea-migration.md`](docs/atlas-gitea-migration.md). The final consistent copy runs in -Atlas' dedicated dataset under a rootless user Quadlet. NPM remains on Prometheus and the primary +Atlas' dedicated dataset under `admin`'s rootless user Quadlet. Its pinned derived image uses an +internal Unix user named `gitea` (UID/GID 1000), while clone URLs keep `git@`. NPM remains on Prometheus and the primary public HTTPS route serves Atlas. The public SSH/2222 socket works locally on Prometheus, but an external connection did not reach its interface on 2026-10-02; check upstream filtering before declaring the HTTPS+SSH cutover complete. The old Gitea data remains on Prometheus, but its container diff --git a/ansible/inventory/host_vars/atlas.yml b/ansible/inventory/host_vars/atlas.yml index da7cc80..f2e0f2d 100644 --- a/ansible/inventory/host_vars/atlas.yml +++ b/ansible/inventory/host_vars/atlas.yml @@ -52,9 +52,6 @@ atlas_manage_storage: true # Rootless Gitea was restored from the stopped-source export before production activation. atlas_manage_gitea: true atlas_gitea_production_enabled: true -# Dedicated rootless Podman range; admin owns 100000-165535 on this host. -atlas_gitea_subid_start: 165536 -atlas_gitea_subid_count: 65536 atlas_prometheus_pull_start_timer: true atlas_manage_zfs_snapshots: true atlas_zfs_snapshot_prefix: atlas-auto diff --git a/ansible/roles/profile_atlas/defaults/main.yml b/ansible/roles/profile_atlas/defaults/main.yml index 5e7d354..a0e5975 100644 --- a/ansible/roles/profile_atlas/defaults/main.yml +++ b/ansible/roles/profile_atlas/defaults/main.yml @@ -165,19 +165,24 @@ atlas_prometheus_pull_keep_monthly: 12 atlas_prometheus_pull_max_age_hours: 24 atlas_photobook_mountpoint: "{{ atlas_mount_root }}/{{ atlas_zfs_dataset_photobook }}" -# Staged rootless Gitea target. Preparation never starts the user Quadlet or opens ingress. +# Rootless Gitea runs in admin's user manager; the image maps internal gitea to UID/GID 1000. atlas_manage_gitea: false -atlas_gitea_username: gitea -atlas_gitea_group: gitea -atlas_gitea_uid: 1101 -atlas_gitea_gid: 1101 -atlas_gitea_subid_start: 165536 -atlas_gitea_subid_count: 65536 -atlas_gitea_home: /var/lib/atlas-gitea +atlas_gitea_username: "{{ atlas_admin_username }}" +atlas_gitea_group: "{{ atlas_admin_group }}" +atlas_gitea_uid: "{{ atlas_admin_uid }}" +atlas_gitea_gid: "{{ atlas_admin_gid }}" +atlas_gitea_home: "{{ atlas_admin_home }}" +atlas_gitea_container_uid: 1000 +atlas_gitea_container_gid: 1000 +atlas_gitea_legacy_username: gitea +atlas_gitea_legacy_uid: 1101 +atlas_gitea_legacy_home: /var/lib/atlas-gitea +atlas_gitea_owner_migration: false atlas_gitea_dataset: "{{ atlas_zfs_pool }}/services/data/gitea" atlas_gitea_mountpoint: "{{ atlas_app_data_mountpoint }}/gitea" atlas_gitea_quadlet_dir: "{{ atlas_gitea_home }}/.config/containers/systemd" -atlas_gitea_image: docker.gitea.com/gitea:1.25.2-rootless +atlas_gitea_image: localhost/atlas-gitea:1.25.2-user-gitea-v1 +atlas_gitea_image_build_dir: "{{ atlas_gitea_home }}/.local/share/atlas-gitea-image" atlas_gitea_production_enabled: false atlas_gitea_bind_address: "{{ ansible_host }}" atlas_gitea_http_port: 3000 diff --git a/ansible/roles/profile_atlas/files/Containerfile.gitea-rootless b/ansible/roles/profile_atlas/files/Containerfile.gitea-rootless new file mode 100644 index 0000000..e4591a6 --- /dev/null +++ b/ansible/roles/profile_atlas/files/Containerfile.gitea-rootless @@ -0,0 +1,10 @@ +FROM docker.gitea.com/gitea@sha256:f1943db2d2f1e447e857b3f0aee4ebb7b184500f86e5b80eae110fd435435906 + +# Preserve the official image's UID/GID, paths and entrypoint; change only the +# internal Unix identity. The host-side rootless owner is Atlas admin. +USER 0 +RUN sed -i 's/^git:x:1000:1000:/gitea:x:1000:1000:/' /etc/passwd \ + && sed -i 's/^git:x:1000:/gitea:x:1000:/' /etc/group \ + && grep -q '^gitea:x:1000:1000:' /etc/passwd \ + && grep -q '^gitea:x:1000:' /etc/group +USER 1000:1000 diff --git a/ansible/roles/profile_atlas/files/atlas-gitea-restore-test.py b/ansible/roles/profile_atlas/files/atlas-gitea-restore-test.py index ea9efed..e478c06 100644 --- a/ansible/roles/profile_atlas/files/atlas-gitea-restore-test.py +++ b/ansible/roles/profile_atlas/files/atlas-gitea-restore-test.py @@ -21,6 +21,8 @@ HOST_KEYS = ( ) SERVER_SETTINGS = { "START_SSH_SERVER": "true", + "BUILTIN_SSH_SERVER_USER": "git", + "SSH_USER": "git", "SSH_PORT": "2222", "SSH_LISTEN_PORT": "2222", "SSH_SERVER_HOST_KEYS": ", ".join( @@ -52,6 +54,7 @@ def convert_config(config): section = "" server_seen = set() server_found = False + run_user_seen = False def append_missing_server_settings(): for key, value in SERVER_SETTINGS.items(): @@ -61,6 +64,9 @@ def convert_config(config): for line in original.splitlines(keepends=True): match = re.match(r"^\s*\[([^]]+)\]\s*$", line) if match: + if not run_user_seen: + output.append("RUN_USER = gitea\n") + run_user_seen = True if section == "server": append_missing_server_settings() section = match.group(1).lower() @@ -68,7 +74,10 @@ def convert_config(config): output.append(line) continue setting = re.match(r"^(\s*)([A-Z_]+)(\s*=\s*)(.*?)(\r?\n?)$", line) - if setting and section == "server" and setting.group(2) in SERVER_SETTINGS: + if setting and section == "" and setting.group(2) == "RUN_USER": + run_user_seen = True + line = f"{setting.group(1)}RUN_USER{setting.group(3)}gitea{setting.group(5)}" + elif setting and section == "server" and setting.group(2) in SERVER_SETTINGS: key = setting.group(2) server_seen.add(key) line = f"{setting.group(1)}{key}{setting.group(3)}{SERVER_SETTINGS[key]}{setting.group(5)}" @@ -180,8 +189,8 @@ def main(): raise ValueError("Refusing backup outside the Atlas Prometheus snapshots") if str(target) != "/zpool/services/data/gitea": raise ValueError("Refusing target outside the dedicated Gitea dataset") - if args.uid != 1101 or args.gid != 1101: - raise ValueError("Unexpected dedicated Gitea account IDs") + if args.uid != 1000 or args.gid != 1000: + raise ValueError("Unexpected admin-owned Gitea account IDs") expected = expected_digest(backup) if sha256(backup / "payload.tar") != expected: raise ValueError("Prometheus backup SHA-256 mismatch") diff --git a/ansible/roles/profile_atlas/tasks/gitea.yml b/ansible/roles/profile_atlas/tasks/gitea.yml index 5c3732b..74a9bf7 100644 --- a/ansible/roles/profile_atlas/tasks/gitea.yml +++ b/ansible/roles/profile_atlas/tasks/gitea.yml @@ -9,16 +9,18 @@ - atlas_manage_storage | bool - atlas_gitea_dataset == atlas_zfs_pool ~ '/services/data/gitea' - atlas_gitea_mountpoint == atlas_app_data_mountpoint ~ '/gitea' - - atlas_gitea_uid | int != atlas_admin_uid | int - - atlas_gitea_uid | int != atlas_immich_uid | int - - atlas_gitea_gid | int != atlas_admin_gid | int - - atlas_gitea_gid | int != atlas_immich_gid | int + - atlas_gitea_username == atlas_admin_username + - atlas_gitea_group == atlas_admin_group + - atlas_gitea_uid | int == atlas_admin_uid | int + - atlas_gitea_gid | int == atlas_admin_gid | int + - atlas_gitea_container_uid | int == 1000 + - atlas_gitea_container_gid | int == 1000 - atlas_gitea_staging_bind_address == '127.0.0.1' - not (atlas_gitea_production_enabled | bool) or atlas_manage_firewall | bool - not (atlas_gitea_production_enabled | bool) or atlas_gitea_bind_address == ansible_host fail_msg: >- - Rootless Gitea preparation requires Atlas storage, an isolated service - identity and dataset, and loopback-only staging ports. + Rootless Gitea requires Atlas storage, the admin user manager, the + dedicated dataset, and loopback-only staging ports. - name: Inspect the final-restore marker before production activation ansible.builtin.stat: @@ -33,43 +35,32 @@ fail_msg: Restore the final stopped-source Gitea export before enabling production. when: atlas_gitea_production_enabled | bool - - name: Create the dedicated Gitea group - ansible.builtin.group: - name: "{{ atlas_gitea_group }}" - gid: "{{ atlas_gitea_gid }}" - system: true - state: present + - name: Verify the production Gitea dataset belongs to admin + ansible.builtin.stat: + path: "{{ atlas_gitea_mountpoint }}" + register: atlas_gitea_dataset_owner + when: atlas_gitea_production_enabled | bool - - name: Create the non-login Gitea service account - ansible.builtin.user: - name: "{{ atlas_gitea_username }}" - uid: "{{ atlas_gitea_uid }}" - group: "{{ atlas_gitea_group }}" - home: "{{ atlas_gitea_home }}" - shell: /sbin/nologin - create_home: true - system: true - state: present + - name: Refuse to overlap the legacy host-account service + ansible.builtin.assert: + that: + - atlas_gitea_dataset_owner.stat.uid | int == atlas_admin_uid | int + - atlas_gitea_dataset_owner.stat.gid | int == atlas_admin_gid | int + fail_msg: >- + Run the explicit Gitea owner migration before enabling the admin + Quadlet; never chown an active legacy service in a normal run. + when: atlas_gitea_production_enabled | bool - - name: Restrict the Gitea service home - ansible.builtin.file: - path: "{{ atlas_gitea_home }}" - state: directory - owner: "{{ atlas_gitea_username }}" - group: "{{ atlas_gitea_group }}" - mode: "0700" - - - name: Reserve dedicated rootless UID and GID ranges for Gitea - ansible.builtin.lineinfile: + - name: Remove the retired account's parent-dataset traverse ACL + ansible.posix.acl: path: "{{ item }}" - regexp: '^{{ atlas_gitea_username }}:' - line: >- - {{ atlas_gitea_username }}:{{ atlas_gitea_subid_start }}:{{ atlas_gitea_subid_count }} - create: false - mode: "0644" + etype: user + entity: "{{ atlas_gitea_legacy_username }}" + state: absent loop: - - /etc/subuid - - /etc/subgid + - "{{ atlas_services_mountpoint }}" + - "{{ atlas_app_data_mountpoint }}" + when: atlas_gitea_production_enabled | bool - name: Enable POSIX ACLs only on the service-namespace parents community.general.zfs: @@ -81,17 +72,6 @@ - "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_services }}" - "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_app_data }}" - - name: Permit Gitea to traverse only the application-data parents - ansible.posix.acl: - path: "{{ item }}" - entity: "{{ atlas_gitea_username }}" - etype: user - permissions: x - state: present - loop: - - "{{ atlas_services_mountpoint }}" - - "{{ atlas_app_data_mountpoint }}" - - name: Create the dedicated Gitea ZFS dataset community.general.zfs: name: "{{ atlas_gitea_dataset }}" @@ -115,7 +95,7 @@ - "{{ atlas_gitea_home }}/.config/containers" - "{{ atlas_gitea_quadlet_dir }}" - - name: Enable lingering for the dedicated rootless account + - name: Ensure lingering for the admin rootless account ansible.builtin.command: argv: - loginctl @@ -123,12 +103,15 @@ - "{{ atlas_gitea_username }}" creates: "/var/lib/systemd/linger/{{ atlas_gitea_username }}" - - name: Start the dedicated rootless user manager + - name: Start the admin rootless user manager ansible.builtin.systemd: name: "user@{{ atlas_gitea_uid }}.service" state: started when: not ansible_check_mode + - name: Prepare the admin-owned Gitea image + ansible.builtin.import_tasks: gitea_image.yml + - name: Render the rootless Gitea Quadlet ansible.builtin.template: src: atlas-gitea.container.j2 diff --git a/ansible/roles/profile_atlas/tasks/gitea_image.yml b/ansible/roles/profile_atlas/tasks/gitea_image.yml new file mode 100644 index 0000000..fac48fa --- /dev/null +++ b/ansible/roles/profile_atlas/tasks/gitea_image.yml @@ -0,0 +1,51 @@ +--- +- name: Create the admin-owned Gitea image build directory + ansible.builtin.file: + path: "{{ atlas_gitea_image_build_dir }}" + state: directory + owner: "{{ atlas_admin_username }}" + group: "{{ atlas_admin_group }}" + mode: "0700" + +- name: Install the pinned rootless Gitea Containerfile + ansible.builtin.copy: + src: Containerfile.gitea-rootless + dest: "{{ atlas_gitea_image_build_dir }}/Containerfile" + owner: "{{ atlas_admin_username }}" + group: "{{ atlas_admin_group }}" + mode: "0644" + +- name: Check the admin-owned Gitea image + become_user: "{{ atlas_admin_username }}" + ansible.builtin.command: + argv: [podman, image, exists, "{{ atlas_gitea_image }}"] + args: + chdir: "{{ atlas_gitea_image_build_dir }}" + environment: + HOME: "{{ atlas_admin_home }}" + XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}" + register: atlas_gitea_image_present + changed_when: false + failed_when: false + check_mode: false + +- name: Build the pinned Gitea image with the internal gitea identity + become_user: "{{ atlas_admin_username }}" + ansible.builtin.command: + argv: + - podman + - build + - --pull=always + - --tag + - "{{ atlas_gitea_image }}" + - --file + - Containerfile + - . + args: + chdir: "{{ atlas_gitea_image_build_dir }}" + environment: + HOME: "{{ atlas_admin_home }}" + XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}" + when: + - atlas_gitea_image_present.rc != 0 + - not ansible_check_mode diff --git a/ansible/roles/profile_atlas/tasks/gitea_owner_migration.yml b/ansible/roles/profile_atlas/tasks/gitea_owner_migration.yml new file mode 100644 index 0000000..2d5352a --- /dev/null +++ b/ansible/roles/profile_atlas/tasks/gitea_owner_migration.yml @@ -0,0 +1,275 @@ +--- +# Run only in an approved outage with -e atlas_gitea_owner_migration=true. +- name: Move live Gitea from the legacy host account to admin + tags: [atlas, gitea_owner_migration] + when: atlas_gitea_owner_migration | bool + block: + - name: Refuse a check-mode owner migration + ansible.builtin.assert: + that: not ansible_check_mode + fail_msg: The owner migration requires an explicit live outage. + + - name: Inspect the Gitea dataset owner + ansible.builtin.stat: + path: "{{ atlas_gitea_mountpoint }}" + register: atlas_gitea_migration_owner + + - name: Require either the legacy owner or an already migrated dataset + ansible.builtin.assert: + that: + - atlas_gitea_migration_owner.stat.isdir | default(false) + - atlas_gitea_migration_owner.stat.uid | int in [atlas_gitea_legacy_uid | int, atlas_admin_uid | int] + fail_msg: Refusing to modify a Gitea dataset with an unexpected owner. + + - name: Migrate only a legacy-owned Gitea dataset + when: atlas_gitea_migration_owner.stat.uid | int == atlas_gitea_legacy_uid | int + block: + - name: Require the final cutover marker and configuration + ansible.builtin.stat: + path: "{{ item }}" + loop: + - "{{ atlas_gitea_mountpoint }}/.final-sha256" + - "{{ atlas_gitea_mountpoint }}/config/app.ini" + register: atlas_gitea_migration_files + + - name: Refuse migration without both final data and configuration + ansible.builtin.assert: + that: atlas_gitea_migration_files.results | map(attribute='stat.isreg') | min + + - name: Check that admin has no existing Gitea Quadlet + ansible.builtin.stat: + path: "{{ atlas_gitea_quadlet_dir }}/atlas-gitea.container" + register: atlas_gitea_admin_quadlet + + - name: Refuse to overwrite an existing admin Quadlet + ansible.builtin.assert: + that: not atlas_gitea_admin_quadlet.stat.exists + + - name: Check pool health before the outage + ansible.builtin.command: + argv: [zpool, status, -x, "{{ atlas_zfs_pool }}"] + register: atlas_gitea_pool_before + changed_when: false + failed_when: "'is healthy' not in atlas_gitea_pool_before.stdout" + + - name: Ensure the admin Gitea image is available before stopping the source + ansible.builtin.import_tasks: gitea_image.yml + + - name: Stop, snapshot and test the admin-owned staging service + block: + - name: Stop and disable the legacy Gitea user service + become_user: "{{ atlas_gitea_legacy_username }}" + ansible.builtin.systemd: + name: atlas-gitea.service + scope: user + state: stopped + enabled: false + environment: + XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_legacy_uid }}" + DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_gitea_legacy_uid }}/bus" + + - name: Record the migration snapshot name + ansible.builtin.set_fact: + atlas_gitea_migration_snapshot: >- + {{ atlas_gitea_dataset }}@gitea-owner-migration-{{ ansible_facts.date_time.iso8601_basic_short }} + + - name: Snapshot the stopped Gitea dataset for manual recovery + ansible.builtin.command: + argv: [zfs, snapshot, "{{ atlas_gitea_migration_snapshot }}"] + + - name: Transfer only the Gitea dataset to admin + ansible.builtin.file: + path: "{{ atlas_gitea_mountpoint }}" + state: directory + owner: "{{ atlas_admin_username }}" + group: "{{ atlas_admin_group }}" + recurse: true + + - name: Set the actual internal Unix process user + ansible.builtin.lineinfile: + path: "{{ atlas_gitea_mountpoint }}/config/app.ini" + regexp: '^RUN_USER\s*=' + line: RUN_USER = gitea + mode: "0600" + no_log: true + diff: false + + - name: Preserve public git clone URLs independently of the Unix user + community.general.ini_file: + path: "{{ atlas_gitea_mountpoint }}/config/app.ini" + section: server + option: "{{ item }}" + value: git + mode: "0600" + no_extra_spaces: false + loop: [BUILTIN_SSH_SERVER_USER, SSH_USER] + no_log: true + diff: false + + - name: Render admin's loopback-only staging Quadlet + ansible.builtin.template: + src: atlas-gitea.container.j2 + dest: "{{ atlas_gitea_quadlet_dir }}/atlas-gitea.container" + owner: "{{ atlas_admin_username }}" + group: "{{ atlas_admin_group }}" + mode: "0644" + vars: + atlas_gitea_production_enabled: false + + - name: Reload the admin user manager for staging + become_user: "{{ atlas_admin_username }}" + ansible.builtin.systemd: + scope: user + daemon_reload: true + environment: + XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}" + DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus" + + - name: Start admin's loopback-only staging service + become_user: "{{ atlas_admin_username }}" + ansible.builtin.systemd: + name: atlas-gitea.service + scope: user + state: started + environment: + XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}" + DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus" + + - name: Verify staging HTTP before promotion + ansible.builtin.uri: + url: "http://127.0.0.1:{{ atlas_gitea_staging_http_port }}/" + status_code: 200 + register: atlas_gitea_staging_http + retries: 30 + delay: 2 + until: atlas_gitea_staging_http is succeeded + + - name: Verify the container really runs as internal gitea + become_user: "{{ atlas_admin_username }}" + ansible.builtin.command: + argv: [podman, exec, atlas-gitea, id, -un] + environment: + HOME: "{{ atlas_admin_home }}" + XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}" + register: atlas_gitea_internal_user + changed_when: false + failed_when: atlas_gitea_internal_user.stdout != 'gitea' + + - name: Verify the migrated SQLite database + ansible.builtin.command: + argv: + - sqlite3 + - "{{ atlas_gitea_mountpoint }}/data/gitea/gitea.db" + - PRAGMA quick_check; + register: atlas_gitea_migration_sqlite + changed_when: false + failed_when: atlas_gitea_migration_sqlite.stdout != 'ok' + + rescue: + - name: Stop admin's failed staging service + become_user: "{{ atlas_admin_username }}" + ansible.builtin.systemd: + name: atlas-gitea.service + scope: user + state: stopped + environment: + XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}" + DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus" + failed_when: false + + - name: Restore the original Gitea configuration from the safety snapshot + ansible.builtin.command: + argv: + - cp + - -a + - "{{ atlas_gitea_mountpoint }}/.zfs/snapshot/{{ atlas_gitea_migration_snapshot.split('@')[1] }}/config/app.ini" + - "{{ atlas_gitea_mountpoint }}/config/app.ini" + when: atlas_gitea_migration_snapshot is defined + + - name: Return the Gitea dataset to the legacy account + ansible.builtin.file: + path: "{{ atlas_gitea_mountpoint }}" + state: directory + owner: "{{ atlas_gitea_legacy_username }}" + group: "{{ atlas_gitea_legacy_username }}" + recurse: true + + - name: Restart the legacy Gitea service + become_user: "{{ atlas_gitea_legacy_username }}" + ansible.builtin.systemd: + name: atlas-gitea.service + scope: user + state: started + enabled: true + environment: + XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_legacy_uid }}" + DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_gitea_legacy_uid }}/bus" + + - name: Report the failed migration and preserved snapshot + ansible.builtin.fail: + msg: >- + Admin staging failed; legacy Gitea was restarted. Inspect + {{ atlas_gitea_migration_snapshot | default('the host journal') }}. + + - name: Stop admin's validated staging service + become_user: "{{ atlas_admin_username }}" + ansible.builtin.systemd: + name: atlas-gitea.service + scope: user + state: stopped + environment: + XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}" + DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus" + + - name: Render admin's production Gitea Quadlet + ansible.builtin.template: + src: atlas-gitea.container.j2 + dest: "{{ atlas_gitea_quadlet_dir }}/atlas-gitea.container" + owner: "{{ atlas_admin_username }}" + group: "{{ atlas_admin_group }}" + mode: "0644" + vars: + atlas_gitea_production_enabled: true + + - name: Reload admin's production user manager + become_user: "{{ atlas_admin_username }}" + ansible.builtin.systemd: + scope: user + daemon_reload: true + environment: + XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}" + DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus" + + - name: Enable and start admin's production Gitea + become_user: "{{ atlas_admin_username }}" + ansible.builtin.systemd: + name: atlas-gitea.service + scope: user + state: started + enabled: true + environment: + XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}" + DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_admin_uid }}/bus" + + - name: Verify production HTTP before retiring the old Quadlet + ansible.builtin.uri: + url: "http://{{ atlas_gitea_bind_address }}:{{ atlas_gitea_http_port }}/" + status_code: 200 + register: atlas_gitea_production_http + retries: 30 + delay: 2 + until: atlas_gitea_production_http is succeeded + + - name: Remove only the disabled legacy Quadlet + ansible.builtin.file: + path: "{{ atlas_gitea_legacy_home }}/.config/containers/systemd/atlas-gitea.container" + state: absent + + - name: Reload the legacy user manager after Quadlet removal + become_user: "{{ atlas_gitea_legacy_username }}" + ansible.builtin.systemd: + scope: user + daemon_reload: true + environment: + XDG_RUNTIME_DIR: "/run/user/{{ atlas_gitea_legacy_uid }}" + DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ atlas_gitea_legacy_uid }}/bus" diff --git a/ansible/roles/profile_atlas/tasks/main.yml b/ansible/roles/profile_atlas/tasks/main.yml index 6409b7e..7f09b9d 100644 --- a/ansible/roles/profile_atlas/tasks/main.yml +++ b/ansible/roles/profile_atlas/tasks/main.yml @@ -14,6 +14,9 @@ - name: Import Atlas storage tasks ansible.builtin.import_tasks: storage.yml +- name: Import explicit Atlas Gitea owner migration + ansible.builtin.import_tasks: gitea_owner_migration.yml + - name: Import staged Atlas rootless Gitea tasks ansible.builtin.import_tasks: gitea.yml diff --git a/ansible/roles/profile_atlas/templates/atlas-gitea.container.j2 b/ansible/roles/profile_atlas/templates/atlas-gitea.container.j2 index 092327c..30f3190 100644 --- a/ansible/roles/profile_atlas/templates/atlas-gitea.container.j2 +++ b/ansible/roles/profile_atlas/templates/atlas-gitea.container.j2 @@ -6,7 +6,7 @@ RequiresMountsFor={{ atlas_gitea_mountpoint }} [Container] ContainerName=atlas-gitea Image={{ atlas_gitea_image }} -UserNS=keep-id:uid=1000,gid=1000 +UserNS=keep-id:uid={{ atlas_gitea_container_uid }},gid={{ atlas_gitea_container_gid }} {% if atlas_gitea_production_enabled | bool %} PublishPort={{ atlas_gitea_bind_address }}:{{ atlas_gitea_http_port }}:3000 PublishPort={{ atlas_gitea_bind_address }}:{{ atlas_gitea_ssh_port }}:2222 diff --git a/docs/atlas-gitea-migration.md b/docs/atlas-gitea-migration.md index 13d4d97..fcbf53e 100644 --- a/docs/atlas-gitea-migration.md +++ b/docs/atlas-gitea-migration.md @@ -24,8 +24,8 @@ Prometheus data, but do not restart its stale Gitea after Atlas accepts writes. Hosts' effective upstream to Atlas over the Prometheus--Aegis gateway, and offers public Gitea SSH on port 2222 via the same gateway. Prometheus port 22 is unchanged. HTTPS and SSH must be validated together before declaring cutover. -- Run Gitea as a **rootless user Quadlet** under a dedicated, non-login Atlas - account, using the pinned `1.25.2-rootless` image. This is an explicit +- The initial staging ran as a **rootless user Quadlet** under a dedicated, + non-login Atlas account, using the pinned `1.25.2-rootless` image. This was an explicit rootful-to-rootless **data-layout conversion**, not a drop-in image swap: the target mounts `/var/lib/gitea` and `/etc/gitea`, and uses Gitea's built-in SSH server instead of the source image's OpenSSH daemon. Keep the @@ -151,6 +151,34 @@ snapshot at `20261002T073032Z` and encrypted Borg archive `atlas-20261002T073044Z` captured the Atlas target after cutover; Borg exited successfully, cleaned its temporary snapshot, and the pool was healthy. +## Corrected Atlas service owner (2026-10-02) + +The operator required the host Quadlet to belong to `admin`, while the Unix +user **inside** the container must be named `gitea`. The pinned derived +`Containerfile.gitea-rootless` changes only the base image's UID/GID 1000 +passwd/group names from `git` to `gitea`; it retains the rootless image's +paths and entrypoint. Gitea's `RUN_USER` is `gitea`, while its built-in SSH +user and advertised clone user remain `git`, preserving `git@` URLs. The +selective restore helper now generates the same three settings for any future +explicit restore, instead of recreating a `RUN_USER = git` target. + +A disposable, loopback-only container using a copy of a Gitea ZFS snapshot +passed HTTP, SQLite, internal-user and SSH host-key checks without touching +live data. After explicit outage approval, the opt-in +`--tags gitea_owner_migration -e atlas_gitea_owner_migration=true` run stopped +the old user service, took safety snapshot +`zpool/services/data/gitea@gitea-owner-migration-20261002T100104`, transferred +only the Gitea dataset to `admin`, tested an `admin` staging Quadlet on +loopback, then promoted it to the production LAN ports. The old Atlas Quadlet +was removed. The old host `gitea` account and its sub-ID range are retained +for a deliberate rollback; they must not restart stale Gitea. The parent +traverse ACL is removed by the normal Gitea role once the new owner is live. + +The new service returned HTTP 200 locally and through public primary HTTPS; +Navidrome and Syncthing remained active under `admin`, the pool was healthy, +and a second normal Gitea Ansible run was idempotent. This does **not** close +the separate external TCP/2222 or authenticated clone/push validation gap. + 1. Agree on an outage and record source/target versions, pool health, the latest backups, SSH host-key fingerprints, and both current NPM routes. Stop the Prometheus export timer for the change window so it cannot