From 9f95e68190c003777713003e7d4bbe248eb07f70 Mon Sep 17 00:00:00 2001 From: Fabio Scotto di Santolo Date: Sun, 4 Oct 2026 17:00:48 +0200 Subject: [PATCH] Integrate consistent Nextcloud backups and recovery --- AGENTS.md | 27 +++- ansible/inventory/host_vars/atlas.yml | 13 ++ ansible/roles/profile_atlas/defaults/main.yml | 8 +- ansible/roles/profile_atlas/tasks/main.yml | 3 + .../roles/profile_atlas/tasks/nextcloud.yml | 23 ++- .../tasks/nextcloud_application.yml | 15 ++ .../profile_atlas/tasks/nextcloud_backup.yml | 55 ++++++++ .../tasks/nextcloud_external_access.yml | 100 +++++++++++++ .../tasks/nextcloud_external_mount.yml | 94 +++++++++++++ ansible/roles/profile_atlas/tasks/storage.yml | 1 + .../atlas-nextcloud-backup-dependency.conf.j2 | 3 + ...atlas-nextcloud-backup-recovery.service.j2 | 19 +++ .../atlas-nextcloud-backup.service.j2 | 21 +++ .../templates/atlas-nextcloud-backup.sh.j2 | 133 ++++++++++++++++++ .../atlas-nextcloud-external-scan.service.j2 | 12 ++ .../atlas-nextcloud-external-scan.timer.j2 | 10 ++ .../templates/atlas-nextcloud.container.j2 | 6 +- docs/atlas-nextcloud-design.md | 9 +- docs/atlas-nextcloud-recovery-test.md | 126 +++++++++++++++++ docs/atlas-nextcloud.md | 112 ++++++++++++++- 20 files changed, 772 insertions(+), 18 deletions(-) create mode 100644 ansible/roles/profile_atlas/tasks/nextcloud_backup.yml create mode 100644 ansible/roles/profile_atlas/tasks/nextcloud_external_access.yml create mode 100644 ansible/roles/profile_atlas/tasks/nextcloud_external_mount.yml create mode 100644 ansible/roles/profile_atlas/templates/atlas-nextcloud-backup-dependency.conf.j2 create mode 100644 ansible/roles/profile_atlas/templates/atlas-nextcloud-backup-recovery.service.j2 create mode 100644 ansible/roles/profile_atlas/templates/atlas-nextcloud-backup.service.j2 create mode 100644 ansible/roles/profile_atlas/templates/atlas-nextcloud-backup.sh.j2 create mode 100644 ansible/roles/profile_atlas/templates/atlas-nextcloud-external-scan.service.j2 create mode 100644 ansible/roles/profile_atlas/templates/atlas-nextcloud-external-scan.timer.j2 create mode 100644 docs/atlas-nextcloud-recovery-test.md diff --git a/AGENTS.md b/AGENTS.md index 25a5e43..dcf20ea 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -68,6 +68,8 @@ Ansible-driven personal infrastructure repo for Fedora and Void desktops, Fedora `ansible-playbook ansible/site.yml --limit atlas --tags gitea_public_domain --check --diff` - Atlas Nextcloud/ONLYOFFICE steady state: `ansible-playbook ansible/site.yml --limit atlas --tags nextcloud --check --diff` + - Atlas recurring consistent Nextcloud backup preparation: + `ansible-playbook ansible/site.yml --limit atlas --tags nextcloud_backup,monitoring --check --diff` - Atlas iCloudPD storage and boot-started Quadlet: `ansible-playbook ansible/site.yml --limit atlas --tags icloudpd --check --diff` - Ongoing Gitea proxy configuration: @@ -210,15 +212,17 @@ and TCP reachability to Atlas were verified. Temporary Navidrome and Syncthing a manual NPM Proxy Hosts; Syncthing uses `/data/Org` backed by the SMB-shared Archive dataset. Aegis has also validated NFSv4.2 read, write, delete, and `all_squash` mapping to UID/GID `1100` end-to-end. The ZFS snapshot timers are active; a recursive hourly snapshot and scheduled retention prune completed -successfully. The first monthly scrub remains a runtime check. +successfully. The first monthly scrub completed successfully on 2026-10-04; +the actual service result and pool scan were independently verified. ### Priority 1 - Data protection - [x] Deploy Ansible-managed recursive ZFS snapshots with 24 hourly, 30 daily, 8 weekly, and 12 monthly generations, plus a monthly scrub on the first Sunday at 03:00. The timers and first hourly snapshot were verified on Atlas. Cockpit Scheduler is for visibility or manual operations only, and snapshot rollback is never automated. -- [ ] Verify the first monthly ZFS scrub from its actual service result. Scheduled retention pruning - was observed on 2026-09-30; timer activation alone does not establish a successful scrub. +- [x] Verify the first monthly ZFS scrub from its actual service result. On 2026-10-04 + it completed at 05:02 CEST after 2:02:04, repairing 0 B with zero errors; + the service exited successfully and the pool reported no known data errors. - [x] Activate and validate the encrypted offsite Borg backup to the Hetzner Storage Box. Atlas uses the dedicated SSH identity, pinned ED25519 host key, Vault-backed `repokey` encryption, and a locked non-login `borg` account with no sudo or supplementary groups. The initial snapshot-consistent backup, @@ -376,9 +380,20 @@ successfully. The first monthly scrub remains a runtime check. web login, WebDAV, private-file isolation, Famiglia cross-user create/read/update/delete and CalDAV/CardDAV discovery passed. The Office connector and public health/API asset passed. Temporary test files were removed; no iCloud data was imported. -- [ ] Complete Nextcloud desktop/mobile editing and synchronization acceptance, and - application-consistent backup/restore validation. Close the first actual scrub and - protection checks before importing family data. +- [x] Test a manual consistent Nextcloud backup and isolated restore on 2026-10-04. + Paused application writers and cron, copied app/config/custom apps/themes and files, + dumped PostgreSQL, restored database roles and verified authenticated DAV contents, + account recovery and Famiglia permissions. Test containers had no external network, + published ports or live data mounts; they and the temporary restore copy were removed. + See `docs/atlas-nextcloud-recovery-test.md`; this is not recurring Borg/USB recovery evidence. +- [x] Integrate consistent Nextcloud bundles with recurring Borg and operator-started USB + backups, two-version local retention, interruption recovery and failure monitoring. + On 2026-10-04 Borg archive `atlas-20261004T095255Z` succeeded; its new bundle was + extracted from Hetzner and restored in isolation with checksums, accounts, Famiglia + permissions and authenticated DAV verified. No production database was replaced. +- [ ] Validate a new USB version and restore its consistent Nextcloud bundle after + operator connection/unlock. Dependency installation alone is not restore evidence. +- [ ] Complete Nextcloud desktop/mobile editing and synchronization acceptance before family import. iCloud migration and future Uranus transfer remain separate operations, not playbook flags. - [x] Move Gitea canonical HTTPS and SSH hostname to `git.fscotto.co` on 2026-10-03 through Ansible. Only Gitea restarted; second run changed nothing. diff --git a/ansible/inventory/host_vars/atlas.yml b/ansible/inventory/host_vars/atlas.yml index 6d60e54..f81736d 100644 --- a/ansible/inventory/host_vars/atlas.yml +++ b/ansible/inventory/host_vars/atlas.yml @@ -50,6 +50,8 @@ atlas_zfs_dataset_photobook: media/photobook atlas_mount_root: /zpool atlas_manage_storage: true atlas_manage_nextcloud: true +# Two local consistent bundles; long-term history stays in Borg/USB and ZFS. +atlas_nextcloud_backup_keep: 2 atlas_nextcloud_domain: cloud.fscotto.co atlas_onlyoffice_domain: office.fscotto.co # Resolved official amd64 images on 2026-10-03; updates are deliberate. @@ -64,6 +66,17 @@ atlas_nextcloud_users: - username: chiara display_name: Chiara password: "{{ vault_nextcloud_chiara_password }}" +atlas_nextcloud_external_mounts: + - name: Documenti + user: fabio + source: /zpool/archive/Documents + target: /mnt/archive-documents + readonly: false + - name: Foto iCloud + user: fabio + source: /zpool/archive/Pictures/iCloudPD + target: /mnt/archive-icloud + readonly: true atlas_nextcloud_apps: - id: groupfolders version: 21.0.9 diff --git a/ansible/roles/profile_atlas/defaults/main.yml b/ansible/roles/profile_atlas/defaults/main.yml index 69193d8..5af3437 100644 --- a/ansible/roles/profile_atlas/defaults/main.yml +++ b/ansible/roles/profile_atlas/defaults/main.yml @@ -16,9 +16,13 @@ atlas_nextcloud_image: "" atlas_nextcloud_postgres_image: "" atlas_nextcloud_redis_image: "" atlas_onlyoffice_image: "" +atlas_nextcloud_backup_root: "{{ atlas_mount_root }}/backup/nextcloud" +atlas_nextcloud_backup_keep: 2 atlas_nextcloud_admin: admin atlas_nextcloud_users: [] atlas_nextcloud_apps: [] +# Existing Archive directories; never import into the internal data namespace. +atlas_nextcloud_external_mounts: [] atlas_nextcloud_services: - atlas-nextcloud-db.service - atlas-nextcloud-redis.service @@ -147,7 +151,9 @@ atlas_monitor_effective_timers: >- atlas_monitor_effective_failure_units: >- {{ atlas_monitor_failure_units + (['atlas-prometheus-pull.service'] - if atlas_manage_prometheus_backup_pull | bool else []) }} + if atlas_manage_prometheus_backup_pull | bool else []) + + (['atlas-nextcloud-backup.service', 'atlas-nextcloud-backup-recovery.service'] + if atlas_manage_nextcloud | bool else []) }} atlas_monitor_remote_capacity: {} atlas_monitor_pool_warning_percent: 80 atlas_monitor_pool_critical_percent: 90 diff --git a/ansible/roles/profile_atlas/tasks/main.yml b/ansible/roles/profile_atlas/tasks/main.yml index 6e2810e..deb5d13 100644 --- a/ansible/roles/profile_atlas/tasks/main.yml +++ b/ansible/roles/profile_atlas/tasks/main.yml @@ -35,6 +35,9 @@ - name: Import Atlas offline USB backup tasks ansible.builtin.import_tasks: usb_backup.yml +- name: Import recurring Nextcloud backup preparation + ansible.builtin.import_tasks: nextcloud_backup.yml + - name: Import Atlas Prometheus backup pull identity tasks ansible.builtin.import_tasks: prometheus_pull_identity.yml diff --git a/ansible/roles/profile_atlas/tasks/nextcloud.yml b/ansible/roles/profile_atlas/tasks/nextcloud.yml index 73dcacf..4bb8f84 100644 --- a/ansible/roles/profile_atlas/tasks/nextcloud.yml +++ b/ansible/roles/profile_atlas/tasks/nextcloud.yml @@ -39,6 +39,10 @@ (atlas_nextcloud_users | map(attribute='password') | list) }} no_log: true + - name: Prepare access to declared existing Archive directories + ansible.builtin.include_tasks: nextcloud_external_access.yml + when: atlas_nextcloud_external_mounts | length > 0 + - name: Verify the existing application-data parent is mounted community.general.zfs_facts: name: "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_app_data }}" @@ -197,7 +201,8 @@ owner: "{{ atlas_admin_username }}" group: "{{ atlas_admin_group }}" mode: "0644" - loop: [atlas-nextcloud-cron.service, atlas-nextcloud-cron.timer] + loop: [atlas-nextcloud-cron.service, atlas-nextcloud-cron.timer, + atlas-nextcloud-external-scan.service, atlas-nextcloud-external-scan.timer] register: atlas_nextcloud_cron_units - name: Manage and verify rootless Nextcloud services @@ -227,7 +232,13 @@ scope: user name: "{{ item }}" state: >- - {{ 'restarted' if (atlas_nextcloud_quadlets is changed or + {{ 'restarted' if ( + atlas_nextcloud_quadlets.results | + selectattr('item', 'equalto', item | replace('.service', '.container')) | + selectattr('changed') | list | length > 0 or + atlas_nextcloud_quadlets.results | + selectattr('item', 'equalto', 'atlas-nextcloud.network') | + selectattr('changed') | list | length > 0 or atlas_nextcloud_private_configuration is changed or atlas_nextcloud_secret_files is changed) else 'started' }} loop: "{{ atlas_nextcloud_services }}" @@ -299,6 +310,14 @@ state: "{{ 'restarted' if atlas_nextcloud_cron_units is changed else 'started' }}" enabled: true + - name: Enable periodic targeted Archive discovery + ansible.builtin.systemd: + scope: user + name: atlas-nextcloud-external-scan.timer + state: "{{ 'restarted' if atlas_nextcloud_cron_units is changed else 'started' }}" + enabled: true + when: atlas_nextcloud_external_mounts | length > 0 + - name: Verify ONLYOFFICE local health without publishing the domain ansible.builtin.uri: url: "http://127.0.0.1:{{ atlas_onlyoffice_http_port }}/healthcheck" diff --git a/ansible/roles/profile_atlas/tasks/nextcloud_application.yml b/ansible/roles/profile_atlas/tasks/nextcloud_application.yml index 9fc41ea..42bea58 100644 --- a/ansible/roles/profile_atlas/tasks/nextcloud_application.yml +++ b/ansible/roles/profile_atlas/tasks/nextcloud_application.yml @@ -169,3 +169,18 @@ argv: [podman, exec, --user, '33', atlas-nextcloud, php, occ, background:cron] when: atlas_nextcloud_background_mode.stdout | trim != 'cron' changed_when: true + +- name: Enable shipped external storage support when required + ansible.builtin.command: + argv: [podman, exec, --user, '33', atlas-nextcloud, php, occ, app:enable, files_external] + when: + - atlas_nextcloud_external_mounts | length > 0 + - "'files_external' not in (atlas_nextcloud_current_apps.stdout | from_json).enabled" + changed_when: true + +- name: Maintain only the declared Archive mounts + ansible.builtin.include_tasks: nextcloud_external_mount.yml + loop: "{{ atlas_nextcloud_external_mounts }}" + loop_control: + loop_var: atlas_nextcloud_mount + label: "{{ atlas_nextcloud_mount.name }}" diff --git a/ansible/roles/profile_atlas/tasks/nextcloud_backup.yml b/ansible/roles/profile_atlas/tasks/nextcloud_backup.yml new file mode 100644 index 0000000..07711dd --- /dev/null +++ b/ansible/roles/profile_atlas/tasks/nextcloud_backup.yml @@ -0,0 +1,55 @@ +--- +- name: Manage recurring consistent Nextcloud backup preparation + tags: [atlas, nextcloud_backup] + when: atlas_manage_nextcloud | bool + block: + - name: Validate private backup scope and local bundle retention + ansible.builtin.assert: + that: + - atlas_nextcloud_backup_root == atlas_mount_root ~ '/backup/nextcloud' + - atlas_nextcloud_backup_keep | int >= 2 + - atlas_manage_borg_backup | bool + - atlas_manage_usb_backup | bool + + - name: Install recurring backup helper with shell syntax validation + ansible.builtin.template: + src: atlas-nextcloud-backup.sh.j2 + dest: /usr/local/sbin/atlas-nextcloud-backup + owner: root + group: root + mode: "0750" + validate: /bin/bash -n %s + + - name: Install Nextcloud backup preparation and boot recovery units + ansible.builtin.template: + src: "{{ item }}.j2" + dest: "/etc/systemd/system/{{ item }}" + owner: root + group: root + mode: "0644" + loop: [atlas-nextcloud-backup.service, atlas-nextcloud-backup-recovery.service] + + - name: Create backup dependency drop-in directories + ansible.builtin.file: + path: "/etc/systemd/system/{{ item }}.d" + state: directory + owner: root + group: root + mode: "0755" + loop: [atlas-borg-backup.service, atlas-usb-backup.service] + + - name: Require a fresh consistent bundle before offsite and manual USB backups + ansible.builtin.template: + src: atlas-nextcloud-backup-dependency.conf.j2 + dest: "/etc/systemd/system/{{ item }}.d/nextcloud.conf" + owner: root + group: root + mode: "0644" + loop: [atlas-borg-backup.service, atlas-usb-backup.service] + + - name: Reload systemd and enable interruption recovery without running a backup + ansible.builtin.systemd: + daemon_reload: true + name: atlas-nextcloud-backup-recovery.service + enabled: true + when: not ansible_check_mode diff --git a/ansible/roles/profile_atlas/tasks/nextcloud_external_access.yml b/ansible/roles/profile_atlas/tasks/nextcloud_external_access.yml new file mode 100644 index 0000000..4cf7ba3 --- /dev/null +++ b/ansible/roles/profile_atlas/tasks/nextcloud_external_access.yml @@ -0,0 +1,100 @@ +--- +- name: Restrict external storage to explicit Archive directories + ansible.builtin.assert: + that: + - item.source in [atlas_archive_mountpoint ~ '/Documents', atlas_icloudpd_photos_dir] + - item.target is match('^/mnt/archive-[a-z]+$') + - item.name is match('^[A-Za-z][A-Za-z ]+$') + - item.readonly is boolean + - item.user in (atlas_nextcloud_users | map(attribute='username') | list) + - item.source != atlas_icloudpd_photos_dir or item.readonly + loop: "{{ atlas_nextcloud_external_mounts }}" + +- name: Inspect existing sources without creating or moving data + ansible.builtin.stat: + path: "{{ item.source }}" + follow: false + loop: "{{ atlas_nextcloud_external_mounts }}" + register: atlas_nextcloud_external_sources + +- name: Refuse missing sources and symlinks + ansible.builtin.assert: + that: + - item.stat.isdir | default(false) + - not (item.stat.islnk | default(false)) + loop: "{{ atlas_nextcloud_external_sources.results }}" + loop_control: + label: "{{ item.item.source }}" + +- name: Verify the Archive dataset before modifying its ACL capability + community.general.zfs_facts: + name: "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_archive }}" + properties: name,mounted,mountpoint + register: atlas_nextcloud_external_dataset + +- name: Refuse an absent or unmounted Archive dataset + ansible.builtin.assert: + that: + - atlas_nextcloud_external_dataset.ansible_facts.ansible_zfs_datasets | length == 1 + - atlas_nextcloud_external_dataset.ansible_facts.ansible_zfs_datasets[0].mounted == 'yes' + - atlas_nextcloud_external_dataset.ansible_facts.ansible_zfs_datasets[0].mountpoint == atlas_archive_mountpoint + +- name: Enable persistent POSIX ACL support on the verified Archive dataset + community.general.zfs: + name: "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_archive }}" + state: present + extra_zfs_properties: + acltype: posix + +- name: Derive actual rootless web UID for narrowly scoped Archive ACLs + become_user: "{{ atlas_admin_username }}" + environment: + XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}" + ansible.builtin.command: + argv: + - podman + - unshare + - python3 + - -c + - >- + print(next(int(b)+33-int(a) for a,b,n in + (l.split() for l in open('/proc/self/uid_map')) if int(a)<=33- + {{ atlas_nextcloud_mount_list.stdout | from_json | + selectattr('mount_point', 'equalto', '/' ~ atlas_nextcloud_mount.name) | list }} + no_log: true + +- name: Refuse duplicates or repurposing of existing unrelated storage + ansible.builtin.assert: + that: + - atlas_nextcloud_matching_mounts | length <= 1 + - >- + atlas_nextcloud_matching_mounts | length == 0 or + (atlas_nextcloud_matching_mounts[0].configuration.datadir | default('') == atlas_nextcloud_mount.target + and atlas_nextcloud_matching_mounts[0].storage == '\\OC\\Files\\Storage\\Local') + fail_msg: Existing storage conflicts with the declared Archive mount; refusing an implicit replacement. + +- name: Create an absent local mount restricted to its declared user + ansible.builtin.command: + argv: + - podman + - exec + - --user + - '33' + - atlas-nextcloud + - php + - occ + - files_external:create + - "{{ atlas_nextcloud_mount.name }}" + - local + - null::null + - --config + - "datadir={{ atlas_nextcloud_mount.target }}" + - --applicable-user + - "{{ atlas_nextcloud_mount.user }}" + - --output=json + when: atlas_nextcloud_matching_mounts | length == 0 + register: atlas_nextcloud_mount_created + changed_when: true + +- name: Record the managed mount ID and options + ansible.builtin.set_fact: + atlas_nextcloud_mount_id: >- + {{ atlas_nextcloud_mount_created.stdout | trim if atlas_nextcloud_matching_mounts | length == 0 + else atlas_nextcloud_matching_mounts[0].mount_id }} + atlas_nextcloud_mount_options: >- + {{ {} if atlas_nextcloud_matching_mounts | length == 0 else atlas_nextcloud_matching_mounts[0].options }} + +- name: Restrict the managed mount to exactly its declared user + ansible.builtin.command: + argv: >- + {{ ['podman', 'exec', '--user', '33', 'atlas-nextcloud', 'php', 'occ', + 'files_external:applicable', atlas_nextcloud_mount_id | string, + '--add-user=' ~ atlas_nextcloud_mount.user] + + (atlas_nextcloud_matching_mounts[0].applicable_groups | + map('regex_replace', '^', '--remove-group=') | list) + + (atlas_nextcloud_matching_mounts[0].applicable_users | + reject('equalto', atlas_nextcloud_mount.user) | + map('regex_replace', '^', '--remove-user=') | list) }} + when: + - atlas_nextcloud_matching_mounts | length > 0 + - >- + atlas_nextcloud_matching_mounts[0].applicable_groups | length > 0 or + atlas_nextcloud_matching_mounts[0].applicable_users != [atlas_nextcloud_mount.user] + changed_when: true + +- name: Maintain read-only photos and external change detection + ansible.builtin.command: + argv: [podman, exec, --user, '33', atlas-nextcloud, php, occ, files_external:option, + "{{ atlas_nextcloud_mount_id }}", "{{ item.key }}", "{{ item.value | to_json }}"] + loop: + - {key: readonly, value: "{{ atlas_nextcloud_mount.readonly }}"} + - {key: filesystem_check_changes, value: 1} + - {key: enable_sharing, value: false} + # Nextcloud persists option values as strings ("1" / "" for booleans). + when: >- + item.key not in atlas_nextcloud_mount_options or + atlas_nextcloud_mount_options[item.key] | string != + (('1' if item.value else '') if item.value is boolean else item.value | string) + changed_when: true + +- name: Verify the managed local storage is accessible + ansible.builtin.command: + argv: [podman, exec, --user, '33', atlas-nextcloud, php, occ, files_external:verify, + "{{ atlas_nextcloud_mount_id }}"] + changed_when: false diff --git a/ansible/roles/profile_atlas/tasks/storage.yml b/ansible/roles/profile_atlas/tasks/storage.yml index 195eb58..5fb0eb1 100644 --- a/ansible/roles/profile_atlas/tasks/storage.yml +++ b/ansible/roles/profile_atlas/tasks/storage.yml @@ -10,6 +10,7 @@ properties: compression: zstd mountpoint: "{{ atlas_archive_mountpoint }}" + acltype: posix - name: "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_services }}" mountpoint: "{{ atlas_services_mountpoint }}" owner: "{{ atlas_admin_username }}" diff --git a/ansible/roles/profile_atlas/templates/atlas-nextcloud-backup-dependency.conf.j2 b/ansible/roles/profile_atlas/templates/atlas-nextcloud-backup-dependency.conf.j2 new file mode 100644 index 0000000..1e8277b --- /dev/null +++ b/ansible/roles/profile_atlas/templates/atlas-nextcloud-backup-dependency.conf.j2 @@ -0,0 +1,3 @@ +[Unit] +Requires=atlas-nextcloud-backup.service +After=atlas-nextcloud-backup.service diff --git a/ansible/roles/profile_atlas/templates/atlas-nextcloud-backup-recovery.service.j2 b/ansible/roles/profile_atlas/templates/atlas-nextcloud-backup-recovery.service.j2 new file mode 100644 index 0000000..f1c2267 --- /dev/null +++ b/ansible/roles/profile_atlas/templates/atlas-nextcloud-backup-recovery.service.j2 @@ -0,0 +1,19 @@ +[Unit] +Description=Recover interrupted Nextcloud backup preparation after boot +Requires=zfs.target user@{{ atlas_admin_uid }}.service +After=zfs.target user@{{ atlas_admin_uid }}.service +{% if atlas_manage_monitoring | bool %} +OnFailure=atlas-monitor-failure@%n.service +{% endif %} + +[Service] +Type=oneshot +User=root +UMask=0077 +StateDirectory=atlas-nextcloud-backup +StateDirectoryMode=0700 +ExecStart=/usr/local/sbin/atlas-nextcloud-backup --recover +TimeoutStartSec=5min + +[Install] +WantedBy=multi-user.target diff --git a/ansible/roles/profile_atlas/templates/atlas-nextcloud-backup.service.j2 b/ansible/roles/profile_atlas/templates/atlas-nextcloud-backup.service.j2 new file mode 100644 index 0000000..58ae0f8 --- /dev/null +++ b/ansible/roles/profile_atlas/templates/atlas-nextcloud-backup.service.j2 @@ -0,0 +1,21 @@ +[Unit] +Description=Prepare a consistent Nextcloud bundle before Atlas backups +Requires=zfs.target user@{{ atlas_admin_uid }}.service +After=zfs.target user@{{ atlas_admin_uid }}.service atlas-nextcloud-backup-recovery.service +{% if atlas_manage_monitoring | bool %} +OnFailure=atlas-monitor-failure@%n.service +{% endif %} + +[Service] +Type=oneshot +User=root +UMask=0077 +StateDirectory=atlas-nextcloud-backup +StateDirectoryMode=0700 +ExecStart=/usr/local/sbin/atlas-nextcloud-backup +ExecStopPost=/usr/local/sbin/atlas-nextcloud-backup --recover +TimeoutStartSec=3h +TimeoutStopSec=5min +Nice=10 +IOSchedulingClass=best-effort +IOSchedulingPriority=7 diff --git a/ansible/roles/profile_atlas/templates/atlas-nextcloud-backup.sh.j2 b/ansible/roles/profile_atlas/templates/atlas-nextcloud-backup.sh.j2 new file mode 100644 index 0000000..4758f5f --- /dev/null +++ b/ansible/roles/profile_atlas/templates/atlas-nextcloud-backup.sh.j2 @@ -0,0 +1,133 @@ +#!/usr/bin/env bash +set -Eeuo pipefail +export PATH=/usr/sbin:/usr/bin:/sbin:/bin +umask 077 +readonly dataset={{ atlas_nextcloud_dataset | quote }} +readonly source_root={{ atlas_nextcloud_root | quote }} +readonly backup_root={{ atlas_nextcloud_backup_root | quote }} +readonly state=/var/lib/atlas-nextcloud-backup +readonly keep={{ atlas_nextcloud_backup_keep | int }} +readonly owner={{ atlas_admin_username | quote }} +readonly uid={{ atlas_admin_uid | int }} + +user_run() { + runuser -u "$owner" -- env XDG_RUNTIME_DIR="/run/user/$uid" \ + DBUS_SESSION_BUS_ADDRESS="unix:path=/run/user/$uid/bus" "$@" +} +occ() { user_run podman exec --user 33 atlas-nextcloud php occ "$@"; } +exec 8>/run/lock/atlas-nextcloud-backup.lock +flock 8 +exec 9>/run/lock/atlas-zfs-snapshot.lock +mkdir -p "$state" +chmod 0700 "$state" + +resume() { + [[ -e "$state/paused" ]] || return 0 + user_run systemctl --user start atlas-nextcloud.service atlas-onlyoffice.service + local ready=false + for _ in {1..60}; do + if occ maintenance:mode --off >/dev/null 2>&1; then ready=true; break; fi + sleep 2 + done + [[ "$ready" == true ]] || { echo 'Nextcloud resume failed; recovery marker retained' >&2; return 1; } + user_run systemctl --user start atlas-nextcloud-cron.timer + # Persist maintenance-off before clearing durable interruption ownership. + sync -f "$source_root/app" + rm "$state/paused" + sync -f "$state" + echo 'Nextcloud/Office resumed and cron timer restored' +} + +recover() { + resume || return 1 + [[ -e "$state/stamp" ]] || return 0 + local stamp snapshot mount source + stamp=$(cat "$state/stamp") + [[ "$stamp" =~ ^[0-9]{8}T[0-9]{6}Z-[0-9]+$ ]] || return 65 + snapshot="nc-backup-$stamp" + flock 9 + for component in files app; do + mount="$source_root/$component/.zfs/snapshot/$snapshot" + source=$(findmnt -rn -M "$mount" -o SOURCE || true) + if [[ -n "$source" ]]; then + [[ "$source" == "$dataset/$component@$snapshot" ]] || return 65 + umount "$mount" || return 1 + fi + done + if zfs list -H -t snapshot "$dataset@$snapshot" >/dev/null 2>&1; then + zfs destroy -r "$dataset@$snapshot" || return 1 + fi + flock -u 9 + # Only this job's private, unpublished staging directory can be removed. + rm -rf -- "$backup_root/.partial-$stamp" + rm "$state/stamp" +} +if [[ "${1:-}" == --recover ]]; then recover; exit; fi +recover +[[ "$(zfs get -H -o value mounted "$dataset")" == yes ]] +[[ "$(zfs get -H -o value mountpoint "$dataset")" == "$source_root" ]] +[[ "$(zfs get -H -o value mounted {{ (atlas_zfs_pool ~ '/backup') | quote }})" == yes ]] +[[ "$(zfs get -H -o value mountpoint {{ (atlas_zfs_pool ~ '/backup') | quote }})" == {{ (atlas_mount_root ~ '/backup') | quote }} ]] +for component in app files; do + [[ "$(zfs get -H -o value mounted "$dataset/$component")" == yes ]] + [[ "$(zfs get -H -o value mountpoint "$dataset/$component")" == "$source_root/$component" ]] +done +for unit in atlas-nextcloud.service atlas-onlyoffice.service atlas-nextcloud-cron.timer; do + user_run systemctl --user is-active --quiet "$unit" +done +occ status --output=json | python3 -c 'import json,sys; s=json.load(sys.stdin); assert s["installed"] and not s["maintenance"] and not s["needsDbUpgrade"]' +mkdir -p "$backup_root/versions" +chmod 0700 "$backup_root" "$backup_root/versions" +stamp="$(date -u +%Y%m%dT%H%M%SZ)-$$" +snapshot="nc-backup-$stamp" +stage="$backup_root/.partial-$stamp" +mkdir "$stage" +printf '%s\n' "$stamp" > "$state/stamp" +sync -f "$state" +cleanup() { + local rc=$? + trap - EXIT + if ! recover; then rc=1; fi + exit "$rc" +} +trap cleanup EXIT +trap 'exit 143' HUP INT TERM +# Wait for snapshot serialization before interrupting application availability. +flock 9 +touch "$state/paused" +sync -f "$state" +user_run systemctl --user stop atlas-nextcloud-cron.timer atlas-nextcloud-cron.service +occ maintenance:mode --on +user_run systemctl --user stop atlas-onlyoffice.service atlas-nextcloud.service +user_run podman exec atlas-nextcloud-db pg_dumpall -U nextcloud --globals-only > "$stage/postgres-globals.sql" +user_run podman exec atlas-nextcloud-db pg_dump -U nextcloud -d nextcloud --format=custom > "$stage/database.dump" +zfs snapshot -r "$dataset@$snapshot" +flock -u 9 +resume +# Copy immutable snapshot views; hashing and transfer never extend the outage. +previous=$(readlink -f "$backup_root/latest" 2>/dev/null || true) +for component in app files; do + args=(-aHAX) + if [[ "$previous" == "$backup_root/versions/"* && -d "$previous/$component" ]]; then + args+=("--link-dest=$previous/$component") + fi + if [[ "$component" == app ]]; then args+=(--exclude=/data); fi + rsync "${args[@]}" "$source_root/$component/.zfs/snapshot/$snapshot/" "$stage/$component/" +done +user_run podman exec -i atlas-nextcloud-db pg_restore --list < "$stage/database.dump" > "$stage/database-toc.txt" +user_run podman inspect --format '{% raw %}{{.ImageName}}{% endraw %}' atlas-nextcloud atlas-nextcloud-db atlas-nextcloud-redis atlas-onlyoffice > "$stage/images.txt" +(cd "$stage"; find app files -type f -exec sha256sum '{}' +; sha256sum database.dump postgres-globals.sql images.txt) > "$stage/SHA256SUMS" +(cd "$stage"; sha256sum --quiet --check SHA256SUMS) +printf 'snapshot=%s@%s\ncreated_utc=%s\n' "$dataset" "$snapshot" "$stamp" > "$stage/manifest.txt" +mv "$stage" "$backup_root/versions/$stamp" +ln -s "versions/$stamp" "$backup_root/.latest-$stamp" +mv -Tf "$backup_root/.latest-$stamp" "$backup_root/latest" +sync -f "$backup_root" +# Prune only timestamped job-owned versions after verified atomic publication. +mapfile -t versions < <(find "$backup_root/versions" -mindepth 1 -maxdepth 1 -type d -printf '%f\n' | grep -E '^[0-9]{8}T[0-9]{6}Z-[0-9]+$' | sort -r) +{% raw %} +for ((index=keep; index<${#versions[@]}; index++)); do +{% endraw %} + rm -rf -- "$backup_root/versions/${versions[$index]}" +done +echo "Published verified consistent Nextcloud bundle $stamp; local retention=$keep" diff --git a/ansible/roles/profile_atlas/templates/atlas-nextcloud-external-scan.service.j2 b/ansible/roles/profile_atlas/templates/atlas-nextcloud-external-scan.service.j2 new file mode 100644 index 0000000..4816ea5 --- /dev/null +++ b/ansible/roles/profile_atlas/templates/atlas-nextcloud-external-scan.service.j2 @@ -0,0 +1,12 @@ +[Unit] +Description=Discover existing Archive documents and iCloud photos in Nextcloud +Requires=atlas-nextcloud.service +After=atlas-nextcloud.service + +[Service] +Type=oneshot +{% for mount in atlas_nextcloud_external_mounts %} +ExecStart=/usr/bin/podman exec --user 33 atlas-nextcloud php occ files:scan "--path={{ mount.user }}/files/{{ mount.name }}" --quiet +{% endfor %} +TimeoutStartSec=90min +NoNewPrivileges=true diff --git a/ansible/roles/profile_atlas/templates/atlas-nextcloud-external-scan.timer.j2 b/ansible/roles/profile_atlas/templates/atlas-nextcloud-external-scan.timer.j2 new file mode 100644 index 0000000..5824797 --- /dev/null +++ b/ansible/roles/profile_atlas/templates/atlas-nextcloud-external-scan.timer.j2 @@ -0,0 +1,10 @@ +[Unit] +Description=Periodic discovery of Archive changes made outside Nextcloud + +[Timer] +OnBootSec=15min +OnUnitInactiveSec=1h +Unit=atlas-nextcloud-external-scan.service + +[Install] +WantedBy=timers.target diff --git a/ansible/roles/profile_atlas/templates/atlas-nextcloud.container.j2 b/ansible/roles/profile_atlas/templates/atlas-nextcloud.container.j2 index af031ea..e724e95 100644 --- a/ansible/roles/profile_atlas/templates/atlas-nextcloud.container.j2 +++ b/ansible/roles/profile_atlas/templates/atlas-nextcloud.container.j2 @@ -2,7 +2,7 @@ Description=Atlas Nextcloud Requires=atlas-nextcloud-db.service atlas-nextcloud-redis.service After=atlas-nextcloud-db.service atlas-nextcloud-redis.service -RequiresMountsFor={{ atlas_nextcloud_root }}/app {{ atlas_nextcloud_root }}/files +RequiresMountsFor={{ atlas_nextcloud_root }}/app {{ atlas_nextcloud_root }}/files{% for mount in atlas_nextcloud_external_mounts %} {{ mount.source }}{% endfor %} [Container] ContainerName=atlas-nextcloud @@ -26,6 +26,10 @@ Environment=PHP_UPLOAD_LIMIT=2G Volume={{ atlas_nextcloud_root }}/app:/var/www/html:Z Volume={{ atlas_nextcloud_root }}/files:/var/www/html/data:Z Volume={{ atlas_nextcloud_app_cache }}:/mnt/atlas-apps:ro,z +{% for mount in atlas_nextcloud_external_mounts %} +# Shared Archive label, not the private :Z label used for internal state. +Volume={{ mount.source }}:{{ mount.target }}:{{ 'ro' if mount.readonly else 'rw' }},z +{% endfor %} Volume={{ atlas_nextcloud_private_dir }}/postgres-password:/run/secrets/postgres-password:ro,z Volume={{ atlas_nextcloud_private_dir }}/admin-password:/run/secrets/admin-password:ro,z Volume={{ atlas_nextcloud_private_dir }}/redis-password:/run/secrets/redis-password:ro,z diff --git a/docs/atlas-nextcloud-design.md b/docs/atlas-nextcloud-design.md index 3f68eca..0edef8e 100644 --- a/docs/atlas-nextcloud-design.md +++ b/docs/atlas-nextcloud-design.md @@ -2,8 +2,10 @@ Status: the empty stack was deployed on 2026-10-03, explicitly before the first scrub. The operator configured DNS/NPM and authorized public cutover; public TLS, -DAV and cross-user file checks passed. Client editing/sync acceptance and consistent -backup/restore validation remain open before family data. iCloud import remains a +DAV and cross-user file checks passed. The first scrub and a manual consistent +backup/isolated restore passed on 2026-10-04. Client editing/sync acceptance and +USB recovery validation remain open before family data. Recurring backup integration +and recovery from a new Borg archive passed. iCloud import remains a separate operation. See `docs/atlas-nextcloud.md` for observed runtime state. ## Confirmed requirements @@ -95,7 +97,8 @@ acceptance tests; the app is not treated as proof of server-side compatibility. 1. Validate desktop Office editing/saving, calendar/contact synchronization and mobile ONLYOFFICE app integration; public empty-stack cutover is verified. -2. Complete protection gates and application-consistent backup/recovery tests. +2. Complete recovery from a new offline USB version; recurring preparation and + encrypted Borg recovery have passed. 3. Plan the deferred iCloud migration when explicitly requested. ## Primary references diff --git a/docs/atlas-nextcloud-recovery-test.md b/docs/atlas-nextcloud-recovery-test.md new file mode 100644 index 0000000..7e8d206 --- /dev/null +++ b/docs/atlas-nextcloud-recovery-test.md @@ -0,0 +1,126 @@ +# Nextcloud manual backup/restore rehearsal — 2026-10-04 + +## Observed outcome + +The operator authorized testing consistent database/files backup and recovery. +This was executed directly, not added as a one-time playbook task or feature flag. +No production database was replaced and no iCloud import was performed. + +The existing ZFS scrub independently passed: completed at 05:02 CEST after +2:02:04, 0 B repaired, zero errors, successful service exit and healthy pool. + +## Backup artifact + +Retained on Atlas: +`/zpool/backup/nextcloud-rehearsal-20261004T092635Z` (676 MiB observed). +Its host parent is restricted to admin, mode 0700; dump and manifest files were +created with umask 077. It contains sensitive application configuration and +database contents, not just test data. No plaintext secret was saved to Git. + +- Complete application tree, including configuration, custom apps and themes; + the overlaid data directory was copied separately. +- Complete dedicated files tree. +- PostgreSQL custom-format database dump, role definitions, image references, + SHA-256 manifest and canary description. + +Cron was stopped, maintenance enabled, and Nextcloud/ONLYOFFICE stopped while +the database dump and application/files copies were taken. Rsync checksum and +metadata comparisons passed while writers were stopped. Live services resumed +with maintenance off, and the cron timer resumed. Role definitions were captured +read-only immediately afterward when the isolated restore exposed the separate +`oc_admin` database role. Its saved password was subsequently verified against +the copied application configuration using SCRAM authentication. The recurring +procedure should capture both database and role dumps during the same pause. + +## Isolated restoration + +- Fresh rootless PostgreSQL using the exact production image digest, not the + live database volume. Restored roles first, then the database with owners/ACLs. +- Copied application and files directories, using the matching Nextcloud image. +- Fresh, empty isolated Redis; cache contents are not a recovery requirement. +- One pod with `network=none`, no published ports and no live data bind mounts. + Components communicate only through their shared loopback interface. +- Only the restored configuration was adjusted for loopback database/cache, + localhost URLs and disabled mail. Production configuration was unchanged. +- No test cron or Office service was run. External connectivity and callbacks + were impossible from this pod. + +Checks passed: + +1. Backup SHA-256 verification before restoration and again after cleanup. +2. PostgreSQL role/database restore with failure-on-error enabled. +3. Nextcloud 33.0.9 installed, maintenance off, no pending database upgrade. +4. Restored admin/fabio/chiara accounts and Famiglia permission mask 15. +5. The uniquely named Fabio canary existed both in files and the database index. +6. Authenticated HTTP WebDAV retrieved that canary from the restored instance; + its SHA-256 matched the original uploaded contents. +7. The live canary remained unchanged and was then deleted through WebDAV. +8. Live Nextcloud, ONLYOFFICE and cron timer active; Office connection check + passed. Cloud/Git/Music/Syncthing public HTTPS returned 200 with valid TLS. + +Initial fixture failures established two prerequisites: wait for PostgreSQL's +final TCP listener, not the temporary initialization socket, and restore global +roles in addition to the database dump. Persistent Redis settings also require +a working isolated cache. Failed fixture pods were removed before retries. + +After success, the final test pod and temporary restore directory were removed. +No test network, live rollback or pool snapshot destruction was needed. + +## Limits and remaining work + +This validates manual recovery from the current local application/database/files +copy, not a production-size recovery, RPO/RTO compliance, client resynchronization, +Office editing-session recovery, or extraction from Borg/offline USB. ONLYOFFICE's +own persistent service state was not part of this Nextcloud artifact. + +The artifact has no dedicated automatic retention policy; do not call it the +recurring Nextcloud backup solution. Integrate a coordinated dump/copy or snapshot +procedure into recurring backups with locking, failure recovery, monitoring and +retention. Independently validate new offsite and offline versions before import. +Keep local/Vault recovery access independent of Nextcloud availability. + +The procedure follows the required configuration/apps/files/themes/database scope +and maintenance pause described in the [Nextcloud backup guide](https://docs.nextcloud.com/server/33/admin_manual/maintenance/backup.html) +and tests restoration into a separate environment rather than applying the +[restore procedure](https://docs.nextcloud.com/server/33/admin_manual/maintenance/restore.html) +to production. + +## Recurring integration and offsite recovery, later on 2026-10-04 + +The recurring preparation helper, system service, boot recovery and ordered +Borg/USB dependency drop-ins were deployed from Ansible. The existing snapshot +and backup jobs retain their lock, ownership, namespace, encryption and retention +policies. Two local verified bundles are the declared staging retention; this +supersedes the missing retention warning above for the managed bundle path only. +The earlier manually named rehearsal artifact remains separate and untouched. + +Preparation published `20261004T094945Z-3294831`, then starting the actual Borg job +required a fresh preparation and published `20261004T095158Z-3303565`. Application +availability resumed before immutable copy/hash processing completed. The Borg +service successfully published `atlas-20261004T095255Z`, completed pruning and +compaction, and removed its source snapshot after exit. + +The latter consistent bundle was extracted from that encrypted Hetzner archive, +not copied from the current local bundle. All SHA-256 checks passed. The extracted +application/files and PostgreSQL role/database dumps were recovered into a fresh +network-none pod with separate database/cache and matching image digests. It +reported installed Nextcloud 33.0.9 without pending upgrade. The three accounts, +Famiglia permissions and authenticated DAV PROPFIND for each account passed; +PostgreSQL used the saved role password with SCRAM on its isolated TCP listener. +The pod, extracted tree and its independent temporary Borg cache were removed. +Production services and the pool remained healthy. + +Failure validation used sandboxed helper mocks for maintenance/dump errors: the +original failure code propagated, services/cron resumed, and state/partial files +were removed. Separate transient systemd fixtures verified that a failed ordered +requirement prevents its consumer from executing. These are fault-injection tests, +not production failures or proof of a full host-crash recovery. A real boot with +an interrupted preparation remains untested. + +A third preparation published `20261004T100212Z-3366172`; exactly two managed +versions remained, with the oldest version pruned only after publication. Source +snapshots and persistent interruption markers were absent after success. + +A new UUID-bound USB version and recovery of its consistent Nextcloud bundle +remain to be verified after the operator connects/unlocks the configured disk. +Do not mark USB recovery complete merely because the dependency was installed. diff --git a/docs/atlas-nextcloud.md b/docs/atlas-nextcloud.md index 6f2ece4..6aee790 100644 --- a/docs/atlas-nextcloud.md +++ b/docs/atlas-nextcloud.md @@ -17,7 +17,8 @@ were added. An actual repeat run returned `changed=0`, with no failures. 10.2.1 and Team Folders 21.0.9 archives are pinned by version and SHA-256. - Dedicated ZFS namespace: `zpool/services/data/nextcloud`, with separate `app`, `files`, `database`, `cache` and `office` datasets. No writable SMB/Syncthing - access to the Nextcloud-managed file namespace is provided. + access to the Nextcloud-managed file namespace is provided. Existing Archive + directories are exposed separately through local external storage (see below). - The `admin` Nextcloud account is an application administrator, distinct from the host account. `fabio` and `chiara` are standard users in `famiglia`, each with no initial quota. Team folder `Famiglia` has unlimited quota and group @@ -101,20 +102,32 @@ Dry-run skips initial downloads, image pulls and runtime account/app commands; it is not proof of an installed or healthy stack. The deployed repeat run is the current idempotence evidence. +## Manual recovery evidence, 2026-10-04 + +The first monthly scrub completed successfully and was verified from both the +service result and pool scan (zero errors, 0 B repaired). A manual consistent +application/files copy and PostgreSQL dump were restored into a network-isolated +Nextcloud/PostgreSQL/Redis test pod. Account recovery, Famiglia permissions and +authenticated DAV retrieval of a checksum-matched canary passed. Live services +resumed normally; the test pod and restore workspace were removed. +See `atlas-nextcloud-recovery-test.md` for scope, retained artifact and limitations. + ## Gates before family data and full client acceptance -- Verify the first actual scrub and the outstanding protection checks. +- The first actual scrub passed on 2026-10-04; preserve the existing protection checks. - Public TLS, redirects, web login and WebDAV passed. Complete calendar/contact synchronization and Office editing/saving from a desktop. - Test opening, editing and saving from the iPhone/iPad ONLYOFFICE app; mobile browser editing is not a requirement. No such client test is claimed yet. - Private-space isolation and cross-user shared writes/deletes passed the public smoke test above; complete normal client acceptance as well. -- Integrate and test application-consistent database/files backups before import. +- The manual rehearsal and recurring integration passed, including recovery from + a new encrypted Borg archive. Complete recovery from a new USB version before import. The new datasets fall beneath existing recursive snapshot/backup scope, but - that alone does not verify a new Borg/USB version or a consistent Nextcloud restore. + that alone does not verify a new Borg/USB version or recovery through those versions. - For a consistent backup, coordinate pending Office saves, pause cron and writes, - take a verified PostgreSQL dump and matching application/files snapshot, and + take verified PostgreSQL database and role dumps plus a matching application/files + snapshot or quiesced copy, and resume services promptly even on failure. Extend recurring backup procedures, not the steady-state playbook with one-time migration tasks. Restore into an isolated environment using matching image/app versions, config, files and DB. @@ -125,3 +138,92 @@ the current idempotence evidence. against an upgraded database; use matching tested backups for recovery. - Future Uranus migration and iCloud import are separate, explicitly authorized operations. No source data deletion or automatic cross-system cutover is provided. + +## Recurring consistent bundles + +`atlas-nextcloud-backup.service` is now an ordered requirement of both +`atlas-borg-backup.service` and the operator-started `atlas-usb-backup.service`. +No additional backup timer is needed: the existing Borg schedule prepares a fresh +bundle before its pool snapshot, and a manual USB run does the same. Preparation +failure blocks the dependent job rather than silently using an old dump. + +The helper checks the mounted datasets and healthy active application state, +serializes preparations and briefly pauses cron, Nextcloud and ONLYOFFICE. It +captures database plus global roles and a recursive Nextcloud-only ZFS snapshot. +Services resume before the longer immutable-file copy and checksum verification. +Active editing sessions are interrupted; only committed Nextcloud state is covered. +This does not claim preservation of unsaved ONLYOFFICE editing sessions. + +Private bundles are published atomically under `/zpool/backup/nextcloud/versions`, +with a relative `latest` link. `atlas_nextcloud_backup_keep: 2` retains two local +verified versions, with hard links for unchanged files. Long-term Borg, USB and +ZFS policies are unchanged. A trap and `ExecStopPost` restore availability and +clean only this helper's named source snapshot/partial directory; root-private +persistent state permits boot recovery through the enabled recovery unit. Both +units have failure alerts and are included in the Atlas monitored failure units. +No automatic rollback, import, or repair of user application data is performed. + +Validation: + +```bash +ansible-playbook ansible/site.yml --limit atlas --tags nextcloud_backup,monitoring --check --diff +sudo systemctl start atlas-nextcloud-backup.service +sudo systemctl show atlas-nextcloud-backup.service -p Result -p ExecMainExitTimestamp +``` + +The second command briefly interrupts the applications and is an explicit manual +run of the recurring job, not a normal deployment side effect. The preparation +unit is not enabled as a boot backup; only interrupted-job recovery is enabled. +Do not stop a Borg/USB job, break its lock or unmount its source snapshot to run a test. + + +## Existing Archive storage (no import or duplicate originals) + +The Atlas declaration exposes only these existing directories to the rootless +Nextcloud container, using shared SELinux `:z` labels: + +| Nextcloud folder | Host directory | Access | +| --- | --- | --- | +| `Documenti` | `/zpool/archive/Documents` | Read/write | +| `Foto iCloud` | `/zpool/archive/Pictures/iCloudPD` | Read-only | + +Both system mounts are restricted to the Nextcloud user `fabio` only; `chiara` +and the `famiglia` group have no access through these mounts. +External re-sharing is disabled. The photos bind is also read-only at container +level, independently of Nextcloud's mount option. iCloudPD remains the photo +writer. Neither directory is copied into the internal data dataset, and the +existing `Famiglia` team folder remains separate and untouched. + +The Archive dataset enables persistent `acltype=posix` support; this does not +change pool features or vdev layout. Scoped ACLs grant the actual rootless-mapped web UID access to existing files and +inheritance on new directories/files. Document defaults retain host administrator +access to files created through Nextcloud; ownership is not changed recursively. +Symlinks are not followed when applying ACLs. Do not change Archive ownership or +apply private `:Z` relabeling to these shared paths. + +The user `atlas-nextcloud-external-scan.timer` discovers external changes for only +these mounts and their explicitly allowed user: first after boot at 15 minutes, then one hour +after the previous scan finishes. Nextcloud also checks for external changes on +access. Indexing and previews are not duplicate originals; document versions, +trash, ZFS snapshots and backups may retain additional data intentionally. +Avoid simultaneously editing the same document through SMB and Nextcloud. + +Archive originals retain their existing recursive ZFS/Borg/offline USB coverage. +The Nextcloud-only recovery bundle does **not** include these external originals: +a recovery must restore the corresponding Archive data as well as the application +and database. This change does not migrate iCloud Drive or remove anything there. + + +### Runtime validation, 2026-10-04 + +The mounts were applied on Atlas without importing originals. A disposable +application-level document create/read test reached the original bind directory; +the probe was deleted, including its trash entry. After the operator narrowed +access to Fabio only, fresh Nextcloud application checks confirmed Fabio can read +both mounts and create documents, while photo create/update/delete are denied. +Chiara cannot access either mount; the separate `Famiglia` team folder remains +available to both users. Container inspection independently confirmed the photo +bind is read-only. Public Nextcloud HTTPS returned 200 and the pool was healthy. +The targeted second Ansible run for mount applicability, options and discovery +unit returned `changed=0`, with no failures. This is focused idempotency evidence, +not a claim about a full Atlas playbook run.