Integrate consistent Nextcloud backups and recovery

This commit is contained in:
Fabio Scotto di Santolo
2026-10-04 17:00:48 +02:00
parent def3dbf313
commit 9f95e68190
20 changed files with 772 additions and 18 deletions

View File

@@ -68,6 +68,8 @@ Ansible-driven personal infrastructure repo for Fedora and Void desktops, Fedora
`ansible-playbook ansible/site.yml --limit atlas --tags gitea_public_domain --check --diff` `ansible-playbook ansible/site.yml --limit atlas --tags gitea_public_domain --check --diff`
- Atlas Nextcloud/ONLYOFFICE steady state: - Atlas Nextcloud/ONLYOFFICE steady state:
`ansible-playbook ansible/site.yml --limit atlas --tags nextcloud --check --diff` `ansible-playbook ansible/site.yml --limit atlas --tags nextcloud --check --diff`
- Atlas recurring consistent Nextcloud backup preparation:
`ansible-playbook ansible/site.yml --limit atlas --tags nextcloud_backup,monitoring --check --diff`
- Atlas iCloudPD storage and boot-started Quadlet: - Atlas iCloudPD storage and boot-started Quadlet:
`ansible-playbook ansible/site.yml --limit atlas --tags icloudpd --check --diff` `ansible-playbook ansible/site.yml --limit atlas --tags icloudpd --check --diff`
- Ongoing Gitea proxy configuration: - Ongoing Gitea proxy configuration:
@@ -210,15 +212,17 @@ and TCP reachability to Atlas were verified. Temporary Navidrome and Syncthing a
manual NPM Proxy Hosts; Syncthing uses `/data/Org` backed by the SMB-shared Archive dataset. Aegis has also manual NPM Proxy Hosts; Syncthing uses `/data/Org` backed by the SMB-shared Archive dataset. Aegis has also
validated NFSv4.2 read, write, delete, and `all_squash` mapping to UID/GID `1100` end-to-end. The ZFS validated NFSv4.2 read, write, delete, and `all_squash` mapping to UID/GID `1100` end-to-end. The ZFS
snapshot timers are active; a recursive hourly snapshot and scheduled retention prune completed snapshot timers are active; a recursive hourly snapshot and scheduled retention prune completed
successfully. The first monthly scrub remains a runtime check. successfully. The first monthly scrub completed successfully on 2026-10-04;
the actual service result and pool scan were independently verified.
### Priority 1 - Data protection ### Priority 1 - Data protection
- [x] Deploy Ansible-managed recursive ZFS snapshots with 24 hourly, 30 daily, 8 weekly, and 12 monthly - [x] Deploy Ansible-managed recursive ZFS snapshots with 24 hourly, 30 daily, 8 weekly, and 12 monthly
generations, plus a monthly scrub on the first Sunday at 03:00. The timers and first hourly snapshot were generations, plus a monthly scrub on the first Sunday at 03:00. The timers and first hourly snapshot were
verified on Atlas. Cockpit Scheduler is for visibility or manual operations only, and snapshot verified on Atlas. Cockpit Scheduler is for visibility or manual operations only, and snapshot
rollback is never automated. rollback is never automated.
- [ ] Verify the first monthly ZFS scrub from its actual service result. Scheduled retention pruning - [x] Verify the first monthly ZFS scrub from its actual service result. On 2026-10-04
was observed on 2026-09-30; timer activation alone does not establish a successful scrub. it completed at 05:02 CEST after 2:02:04, repairing 0 B with zero errors;
the service exited successfully and the pool reported no known data errors.
- [x] Activate and validate the encrypted offsite Borg backup to the Hetzner Storage Box. Atlas uses the - [x] Activate and validate the encrypted offsite Borg backup to the Hetzner Storage Box. Atlas uses the
dedicated SSH identity, pinned ED25519 host key, Vault-backed `repokey` encryption, and a locked dedicated SSH identity, pinned ED25519 host key, Vault-backed `repokey` encryption, and a locked
non-login `borg` account with no sudo or supplementary groups. The initial snapshot-consistent backup, non-login `borg` account with no sudo or supplementary groups. The initial snapshot-consistent backup,
@@ -376,9 +380,20 @@ successfully. The first monthly scrub remains a runtime check.
web login, WebDAV, private-file isolation, Famiglia cross-user create/read/update/delete and web login, WebDAV, private-file isolation, Famiglia cross-user create/read/update/delete and
CalDAV/CardDAV discovery passed. The Office connector and public health/API asset passed. CalDAV/CardDAV discovery passed. The Office connector and public health/API asset passed.
Temporary test files were removed; no iCloud data was imported. Temporary test files were removed; no iCloud data was imported.
- [ ] Complete Nextcloud desktop/mobile editing and synchronization acceptance, and - [x] Test a manual consistent Nextcloud backup and isolated restore on 2026-10-04.
application-consistent backup/restore validation. Close the first actual scrub and Paused application writers and cron, copied app/config/custom apps/themes and files,
protection checks before importing family data. dumped PostgreSQL, restored database roles and verified authenticated DAV contents,
account recovery and Famiglia permissions. Test containers had no external network,
published ports or live data mounts; they and the temporary restore copy were removed.
See `docs/atlas-nextcloud-recovery-test.md`; this is not recurring Borg/USB recovery evidence.
- [x] Integrate consistent Nextcloud bundles with recurring Borg and operator-started USB
backups, two-version local retention, interruption recovery and failure monitoring.
On 2026-10-04 Borg archive `atlas-20261004T095255Z` succeeded; its new bundle was
extracted from Hetzner and restored in isolation with checksums, accounts, Famiglia
permissions and authenticated DAV verified. No production database was replaced.
- [ ] Validate a new USB version and restore its consistent Nextcloud bundle after
operator connection/unlock. Dependency installation alone is not restore evidence.
- [ ] Complete Nextcloud desktop/mobile editing and synchronization acceptance before family import.
iCloud migration and future Uranus transfer remain separate operations, not playbook flags. iCloud migration and future Uranus transfer remain separate operations, not playbook flags.
- [x] Move Gitea canonical HTTPS and SSH hostname to `git.fscotto.co` on - [x] Move Gitea canonical HTTPS and SSH hostname to `git.fscotto.co` on
2026-10-03 through Ansible. Only Gitea restarted; second run changed nothing. 2026-10-03 through Ansible. Only Gitea restarted; second run changed nothing.

View File

@@ -50,6 +50,8 @@ atlas_zfs_dataset_photobook: media/photobook
atlas_mount_root: /zpool atlas_mount_root: /zpool
atlas_manage_storage: true atlas_manage_storage: true
atlas_manage_nextcloud: true atlas_manage_nextcloud: true
# Two local consistent bundles; long-term history stays in Borg/USB and ZFS.
atlas_nextcloud_backup_keep: 2
atlas_nextcloud_domain: cloud.fscotto.co atlas_nextcloud_domain: cloud.fscotto.co
atlas_onlyoffice_domain: office.fscotto.co atlas_onlyoffice_domain: office.fscotto.co
# Resolved official amd64 images on 2026-10-03; updates are deliberate. # Resolved official amd64 images on 2026-10-03; updates are deliberate.
@@ -64,6 +66,17 @@ atlas_nextcloud_users:
- username: chiara - username: chiara
display_name: Chiara display_name: Chiara
password: "{{ vault_nextcloud_chiara_password }}" password: "{{ vault_nextcloud_chiara_password }}"
atlas_nextcloud_external_mounts:
- name: Documenti
user: fabio
source: /zpool/archive/Documents
target: /mnt/archive-documents
readonly: false
- name: Foto iCloud
user: fabio
source: /zpool/archive/Pictures/iCloudPD
target: /mnt/archive-icloud
readonly: true
atlas_nextcloud_apps: atlas_nextcloud_apps:
- id: groupfolders - id: groupfolders
version: 21.0.9 version: 21.0.9

View File

@@ -16,9 +16,13 @@ atlas_nextcloud_image: ""
atlas_nextcloud_postgres_image: "" atlas_nextcloud_postgres_image: ""
atlas_nextcloud_redis_image: "" atlas_nextcloud_redis_image: ""
atlas_onlyoffice_image: "" atlas_onlyoffice_image: ""
atlas_nextcloud_backup_root: "{{ atlas_mount_root }}/backup/nextcloud"
atlas_nextcloud_backup_keep: 2
atlas_nextcloud_admin: admin atlas_nextcloud_admin: admin
atlas_nextcloud_users: [] atlas_nextcloud_users: []
atlas_nextcloud_apps: [] atlas_nextcloud_apps: []
# Existing Archive directories; never import into the internal data namespace.
atlas_nextcloud_external_mounts: []
atlas_nextcloud_services: atlas_nextcloud_services:
- atlas-nextcloud-db.service - atlas-nextcloud-db.service
- atlas-nextcloud-redis.service - atlas-nextcloud-redis.service
@@ -147,7 +151,9 @@ atlas_monitor_effective_timers: >-
atlas_monitor_effective_failure_units: >- atlas_monitor_effective_failure_units: >-
{{ atlas_monitor_failure_units {{ atlas_monitor_failure_units
+ (['atlas-prometheus-pull.service'] + (['atlas-prometheus-pull.service']
if atlas_manage_prometheus_backup_pull | bool else []) }} if atlas_manage_prometheus_backup_pull | bool else [])
+ (['atlas-nextcloud-backup.service', 'atlas-nextcloud-backup-recovery.service']
if atlas_manage_nextcloud | bool else []) }}
atlas_monitor_remote_capacity: {} atlas_monitor_remote_capacity: {}
atlas_monitor_pool_warning_percent: 80 atlas_monitor_pool_warning_percent: 80
atlas_monitor_pool_critical_percent: 90 atlas_monitor_pool_critical_percent: 90

View File

@@ -35,6 +35,9 @@
- name: Import Atlas offline USB backup tasks - name: Import Atlas offline USB backup tasks
ansible.builtin.import_tasks: usb_backup.yml ansible.builtin.import_tasks: usb_backup.yml
- name: Import recurring Nextcloud backup preparation
ansible.builtin.import_tasks: nextcloud_backup.yml
- name: Import Atlas Prometheus backup pull identity tasks - name: Import Atlas Prometheus backup pull identity tasks
ansible.builtin.import_tasks: prometheus_pull_identity.yml ansible.builtin.import_tasks: prometheus_pull_identity.yml

View File

@@ -39,6 +39,10 @@
(atlas_nextcloud_users | map(attribute='password') | list) }} (atlas_nextcloud_users | map(attribute='password') | list) }}
no_log: true no_log: true
- name: Prepare access to declared existing Archive directories
ansible.builtin.include_tasks: nextcloud_external_access.yml
when: atlas_nextcloud_external_mounts | length > 0
- name: Verify the existing application-data parent is mounted - name: Verify the existing application-data parent is mounted
community.general.zfs_facts: community.general.zfs_facts:
name: "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_app_data }}" name: "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_app_data }}"
@@ -197,7 +201,8 @@
owner: "{{ atlas_admin_username }}" owner: "{{ atlas_admin_username }}"
group: "{{ atlas_admin_group }}" group: "{{ atlas_admin_group }}"
mode: "0644" mode: "0644"
loop: [atlas-nextcloud-cron.service, atlas-nextcloud-cron.timer] loop: [atlas-nextcloud-cron.service, atlas-nextcloud-cron.timer,
atlas-nextcloud-external-scan.service, atlas-nextcloud-external-scan.timer]
register: atlas_nextcloud_cron_units register: atlas_nextcloud_cron_units
- name: Manage and verify rootless Nextcloud services - name: Manage and verify rootless Nextcloud services
@@ -227,7 +232,13 @@
scope: user scope: user
name: "{{ item }}" name: "{{ item }}"
state: >- state: >-
{{ 'restarted' if (atlas_nextcloud_quadlets is changed or {{ 'restarted' if (
atlas_nextcloud_quadlets.results |
selectattr('item', 'equalto', item | replace('.service', '.container')) |
selectattr('changed') | list | length > 0 or
atlas_nextcloud_quadlets.results |
selectattr('item', 'equalto', 'atlas-nextcloud.network') |
selectattr('changed') | list | length > 0 or
atlas_nextcloud_private_configuration is changed or atlas_nextcloud_private_configuration is changed or
atlas_nextcloud_secret_files is changed) else 'started' }} atlas_nextcloud_secret_files is changed) else 'started' }}
loop: "{{ atlas_nextcloud_services }}" loop: "{{ atlas_nextcloud_services }}"
@@ -299,6 +310,14 @@
state: "{{ 'restarted' if atlas_nextcloud_cron_units is changed else 'started' }}" state: "{{ 'restarted' if atlas_nextcloud_cron_units is changed else 'started' }}"
enabled: true enabled: true
- name: Enable periodic targeted Archive discovery
ansible.builtin.systemd:
scope: user
name: atlas-nextcloud-external-scan.timer
state: "{{ 'restarted' if atlas_nextcloud_cron_units is changed else 'started' }}"
enabled: true
when: atlas_nextcloud_external_mounts | length > 0
- name: Verify ONLYOFFICE local health without publishing the domain - name: Verify ONLYOFFICE local health without publishing the domain
ansible.builtin.uri: ansible.builtin.uri:
url: "http://127.0.0.1:{{ atlas_onlyoffice_http_port }}/healthcheck" url: "http://127.0.0.1:{{ atlas_onlyoffice_http_port }}/healthcheck"

View File

@@ -169,3 +169,18 @@
argv: [podman, exec, --user, '33', atlas-nextcloud, php, occ, background:cron] argv: [podman, exec, --user, '33', atlas-nextcloud, php, occ, background:cron]
when: atlas_nextcloud_background_mode.stdout | trim != 'cron' when: atlas_nextcloud_background_mode.stdout | trim != 'cron'
changed_when: true changed_when: true
- name: Enable shipped external storage support when required
ansible.builtin.command:
argv: [podman, exec, --user, '33', atlas-nextcloud, php, occ, app:enable, files_external]
when:
- atlas_nextcloud_external_mounts | length > 0
- "'files_external' not in (atlas_nextcloud_current_apps.stdout | from_json).enabled"
changed_when: true
- name: Maintain only the declared Archive mounts
ansible.builtin.include_tasks: nextcloud_external_mount.yml
loop: "{{ atlas_nextcloud_external_mounts }}"
loop_control:
loop_var: atlas_nextcloud_mount
label: "{{ atlas_nextcloud_mount.name }}"

View File

@@ -0,0 +1,55 @@
---
- name: Manage recurring consistent Nextcloud backup preparation
tags: [atlas, nextcloud_backup]
when: atlas_manage_nextcloud | bool
block:
- name: Validate private backup scope and local bundle retention
ansible.builtin.assert:
that:
- atlas_nextcloud_backup_root == atlas_mount_root ~ '/backup/nextcloud'
- atlas_nextcloud_backup_keep | int >= 2
- atlas_manage_borg_backup | bool
- atlas_manage_usb_backup | bool
- name: Install recurring backup helper with shell syntax validation
ansible.builtin.template:
src: atlas-nextcloud-backup.sh.j2
dest: /usr/local/sbin/atlas-nextcloud-backup
owner: root
group: root
mode: "0750"
validate: /bin/bash -n %s
- name: Install Nextcloud backup preparation and boot recovery units
ansible.builtin.template:
src: "{{ item }}.j2"
dest: "/etc/systemd/system/{{ item }}"
owner: root
group: root
mode: "0644"
loop: [atlas-nextcloud-backup.service, atlas-nextcloud-backup-recovery.service]
- name: Create backup dependency drop-in directories
ansible.builtin.file:
path: "/etc/systemd/system/{{ item }}.d"
state: directory
owner: root
group: root
mode: "0755"
loop: [atlas-borg-backup.service, atlas-usb-backup.service]
- name: Require a fresh consistent bundle before offsite and manual USB backups
ansible.builtin.template:
src: atlas-nextcloud-backup-dependency.conf.j2
dest: "/etc/systemd/system/{{ item }}.d/nextcloud.conf"
owner: root
group: root
mode: "0644"
loop: [atlas-borg-backup.service, atlas-usb-backup.service]
- name: Reload systemd and enable interruption recovery without running a backup
ansible.builtin.systemd:
daemon_reload: true
name: atlas-nextcloud-backup-recovery.service
enabled: true
when: not ansible_check_mode

View File

@@ -0,0 +1,100 @@
---
- name: Restrict external storage to explicit Archive directories
ansible.builtin.assert:
that:
- item.source in [atlas_archive_mountpoint ~ '/Documents', atlas_icloudpd_photos_dir]
- item.target is match('^/mnt/archive-[a-z]+$')
- item.name is match('^[A-Za-z][A-Za-z ]+$')
- item.readonly is boolean
- item.user in (atlas_nextcloud_users | map(attribute='username') | list)
- item.source != atlas_icloudpd_photos_dir or item.readonly
loop: "{{ atlas_nextcloud_external_mounts }}"
- name: Inspect existing sources without creating or moving data
ansible.builtin.stat:
path: "{{ item.source }}"
follow: false
loop: "{{ atlas_nextcloud_external_mounts }}"
register: atlas_nextcloud_external_sources
- name: Refuse missing sources and symlinks
ansible.builtin.assert:
that:
- item.stat.isdir | default(false)
- not (item.stat.islnk | default(false))
loop: "{{ atlas_nextcloud_external_sources.results }}"
loop_control:
label: "{{ item.item.source }}"
- name: Verify the Archive dataset before modifying its ACL capability
community.general.zfs_facts:
name: "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_archive }}"
properties: name,mounted,mountpoint
register: atlas_nextcloud_external_dataset
- name: Refuse an absent or unmounted Archive dataset
ansible.builtin.assert:
that:
- atlas_nextcloud_external_dataset.ansible_facts.ansible_zfs_datasets | length == 1
- atlas_nextcloud_external_dataset.ansible_facts.ansible_zfs_datasets[0].mounted == 'yes'
- atlas_nextcloud_external_dataset.ansible_facts.ansible_zfs_datasets[0].mountpoint == atlas_archive_mountpoint
- name: Enable persistent POSIX ACL support on the verified Archive dataset
community.general.zfs:
name: "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_archive }}"
state: present
extra_zfs_properties:
acltype: posix
- name: Derive actual rootless web UID for narrowly scoped Archive ACLs
become_user: "{{ atlas_admin_username }}"
environment:
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
ansible.builtin.command:
argv:
- podman
- unshare
- python3
- -c
- >-
print(next(int(b)+33-int(a) for a,b,n in
(l.split() for l in open('/proc/self/uid_map')) if int(a)<=33<int(a)+int(n)))
register: atlas_nextcloud_external_uid
changed_when: false
check_mode: false
- name: Grant web user access only inside the declared sources
ansible.posix.acl:
path: "{{ item.source }}"
entity: "{{ atlas_nextcloud_external_uid.stdout | trim }}"
etype: user
permissions: "{{ 'rX' if item.readonly else 'rwX' }}"
recursive: true
follow: false
state: present
loop: "{{ atlas_nextcloud_external_mounts }}"
- name: Inherit web access on new files and directories
ansible.posix.acl:
path: "{{ item.source }}"
entity: "{{ atlas_nextcloud_external_uid.stdout | trim }}"
etype: user
permissions: "{{ 'rX' if item.readonly else 'rwX' }}"
default: true
recursive: true
follow: false
state: present
loop: "{{ atlas_nextcloud_external_mounts }}"
- name: Preserve administrator access to documents created through Nextcloud
ansible.posix.acl:
path: "{{ item.source }}"
entity: "{{ atlas_admin_uid }}"
etype: user
permissions: rwX
default: true
recursive: true
follow: false
state: present
loop: "{{ atlas_nextcloud_external_mounts }}"
when: not item.readonly

View File

@@ -0,0 +1,94 @@
---
- name: Inspect current system mounts without exposing credentials
ansible.builtin.command:
argv: [podman, exec, --user, '33', atlas-nextcloud, php, occ, files_external:list, --output=json]
register: atlas_nextcloud_mount_list
changed_when: false
no_log: true
- name: Select only the matching mount name
ansible.builtin.set_fact:
atlas_nextcloud_matching_mounts: >-
{{ atlas_nextcloud_mount_list.stdout | from_json |
selectattr('mount_point', 'equalto', '/' ~ atlas_nextcloud_mount.name) | list }}
no_log: true
- name: Refuse duplicates or repurposing of existing unrelated storage
ansible.builtin.assert:
that:
- atlas_nextcloud_matching_mounts | length <= 1
- >-
atlas_nextcloud_matching_mounts | length == 0 or
(atlas_nextcloud_matching_mounts[0].configuration.datadir | default('') == atlas_nextcloud_mount.target
and atlas_nextcloud_matching_mounts[0].storage == '\\OC\\Files\\Storage\\Local')
fail_msg: Existing storage conflicts with the declared Archive mount; refusing an implicit replacement.
- name: Create an absent local mount restricted to its declared user
ansible.builtin.command:
argv:
- podman
- exec
- --user
- '33'
- atlas-nextcloud
- php
- occ
- files_external:create
- "{{ atlas_nextcloud_mount.name }}"
- local
- null::null
- --config
- "datadir={{ atlas_nextcloud_mount.target }}"
- --applicable-user
- "{{ atlas_nextcloud_mount.user }}"
- --output=json
when: atlas_nextcloud_matching_mounts | length == 0
register: atlas_nextcloud_mount_created
changed_when: true
- name: Record the managed mount ID and options
ansible.builtin.set_fact:
atlas_nextcloud_mount_id: >-
{{ atlas_nextcloud_mount_created.stdout | trim if atlas_nextcloud_matching_mounts | length == 0
else atlas_nextcloud_matching_mounts[0].mount_id }}
atlas_nextcloud_mount_options: >-
{{ {} if atlas_nextcloud_matching_mounts | length == 0 else atlas_nextcloud_matching_mounts[0].options }}
- name: Restrict the managed mount to exactly its declared user
ansible.builtin.command:
argv: >-
{{ ['podman', 'exec', '--user', '33', 'atlas-nextcloud', 'php', 'occ',
'files_external:applicable', atlas_nextcloud_mount_id | string,
'--add-user=' ~ atlas_nextcloud_mount.user] +
(atlas_nextcloud_matching_mounts[0].applicable_groups |
map('regex_replace', '^', '--remove-group=') | list) +
(atlas_nextcloud_matching_mounts[0].applicable_users |
reject('equalto', atlas_nextcloud_mount.user) |
map('regex_replace', '^', '--remove-user=') | list) }}
when:
- atlas_nextcloud_matching_mounts | length > 0
- >-
atlas_nextcloud_matching_mounts[0].applicable_groups | length > 0 or
atlas_nextcloud_matching_mounts[0].applicable_users != [atlas_nextcloud_mount.user]
changed_when: true
- name: Maintain read-only photos and external change detection
ansible.builtin.command:
argv: [podman, exec, --user, '33', atlas-nextcloud, php, occ, files_external:option,
"{{ atlas_nextcloud_mount_id }}", "{{ item.key }}", "{{ item.value | to_json }}"]
loop:
- {key: readonly, value: "{{ atlas_nextcloud_mount.readonly }}"}
- {key: filesystem_check_changes, value: 1}
- {key: enable_sharing, value: false}
# Nextcloud persists option values as strings ("1" / "" for booleans).
when: >-
item.key not in atlas_nextcloud_mount_options or
atlas_nextcloud_mount_options[item.key] | string !=
(('1' if item.value else '') if item.value is boolean else item.value | string)
changed_when: true
- name: Verify the managed local storage is accessible
ansible.builtin.command:
argv: [podman, exec, --user, '33', atlas-nextcloud, php, occ, files_external:verify,
"{{ atlas_nextcloud_mount_id }}"]
changed_when: false

View File

@@ -10,6 +10,7 @@
properties: properties:
compression: zstd compression: zstd
mountpoint: "{{ atlas_archive_mountpoint }}" mountpoint: "{{ atlas_archive_mountpoint }}"
acltype: posix
- name: "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_services }}" - name: "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_services }}"
mountpoint: "{{ atlas_services_mountpoint }}" mountpoint: "{{ atlas_services_mountpoint }}"
owner: "{{ atlas_admin_username }}" owner: "{{ atlas_admin_username }}"

View File

@@ -0,0 +1,3 @@
[Unit]
Requires=atlas-nextcloud-backup.service
After=atlas-nextcloud-backup.service

View File

@@ -0,0 +1,19 @@
[Unit]
Description=Recover interrupted Nextcloud backup preparation after boot
Requires=zfs.target user@{{ atlas_admin_uid }}.service
After=zfs.target user@{{ atlas_admin_uid }}.service
{% if atlas_manage_monitoring | bool %}
OnFailure=atlas-monitor-failure@%n.service
{% endif %}
[Service]
Type=oneshot
User=root
UMask=0077
StateDirectory=atlas-nextcloud-backup
StateDirectoryMode=0700
ExecStart=/usr/local/sbin/atlas-nextcloud-backup --recover
TimeoutStartSec=5min
[Install]
WantedBy=multi-user.target

View File

@@ -0,0 +1,21 @@
[Unit]
Description=Prepare a consistent Nextcloud bundle before Atlas backups
Requires=zfs.target user@{{ atlas_admin_uid }}.service
After=zfs.target user@{{ atlas_admin_uid }}.service atlas-nextcloud-backup-recovery.service
{% if atlas_manage_monitoring | bool %}
OnFailure=atlas-monitor-failure@%n.service
{% endif %}
[Service]
Type=oneshot
User=root
UMask=0077
StateDirectory=atlas-nextcloud-backup
StateDirectoryMode=0700
ExecStart=/usr/local/sbin/atlas-nextcloud-backup
ExecStopPost=/usr/local/sbin/atlas-nextcloud-backup --recover
TimeoutStartSec=3h
TimeoutStopSec=5min
Nice=10
IOSchedulingClass=best-effort
IOSchedulingPriority=7

View File

@@ -0,0 +1,133 @@
#!/usr/bin/env bash
set -Eeuo pipefail
export PATH=/usr/sbin:/usr/bin:/sbin:/bin
umask 077
readonly dataset={{ atlas_nextcloud_dataset | quote }}
readonly source_root={{ atlas_nextcloud_root | quote }}
readonly backup_root={{ atlas_nextcloud_backup_root | quote }}
readonly state=/var/lib/atlas-nextcloud-backup
readonly keep={{ atlas_nextcloud_backup_keep | int }}
readonly owner={{ atlas_admin_username | quote }}
readonly uid={{ atlas_admin_uid | int }}
user_run() {
runuser -u "$owner" -- env XDG_RUNTIME_DIR="/run/user/$uid" \
DBUS_SESSION_BUS_ADDRESS="unix:path=/run/user/$uid/bus" "$@"
}
occ() { user_run podman exec --user 33 atlas-nextcloud php occ "$@"; }
exec 8>/run/lock/atlas-nextcloud-backup.lock
flock 8
exec 9>/run/lock/atlas-zfs-snapshot.lock
mkdir -p "$state"
chmod 0700 "$state"
resume() {
[[ -e "$state/paused" ]] || return 0
user_run systemctl --user start atlas-nextcloud.service atlas-onlyoffice.service
local ready=false
for _ in {1..60}; do
if occ maintenance:mode --off >/dev/null 2>&1; then ready=true; break; fi
sleep 2
done
[[ "$ready" == true ]] || { echo 'Nextcloud resume failed; recovery marker retained' >&2; return 1; }
user_run systemctl --user start atlas-nextcloud-cron.timer
# Persist maintenance-off before clearing durable interruption ownership.
sync -f "$source_root/app"
rm "$state/paused"
sync -f "$state"
echo 'Nextcloud/Office resumed and cron timer restored'
}
recover() {
resume || return 1
[[ -e "$state/stamp" ]] || return 0
local stamp snapshot mount source
stamp=$(cat "$state/stamp")
[[ "$stamp" =~ ^[0-9]{8}T[0-9]{6}Z-[0-9]+$ ]] || return 65
snapshot="nc-backup-$stamp"
flock 9
for component in files app; do
mount="$source_root/$component/.zfs/snapshot/$snapshot"
source=$(findmnt -rn -M "$mount" -o SOURCE || true)
if [[ -n "$source" ]]; then
[[ "$source" == "$dataset/$component@$snapshot" ]] || return 65
umount "$mount" || return 1
fi
done
if zfs list -H -t snapshot "$dataset@$snapshot" >/dev/null 2>&1; then
zfs destroy -r "$dataset@$snapshot" || return 1
fi
flock -u 9
# Only this job's private, unpublished staging directory can be removed.
rm -rf -- "$backup_root/.partial-$stamp"
rm "$state/stamp"
}
if [[ "${1:-}" == --recover ]]; then recover; exit; fi
recover
[[ "$(zfs get -H -o value mounted "$dataset")" == yes ]]
[[ "$(zfs get -H -o value mountpoint "$dataset")" == "$source_root" ]]
[[ "$(zfs get -H -o value mounted {{ (atlas_zfs_pool ~ '/backup') | quote }})" == yes ]]
[[ "$(zfs get -H -o value mountpoint {{ (atlas_zfs_pool ~ '/backup') | quote }})" == {{ (atlas_mount_root ~ '/backup') | quote }} ]]
for component in app files; do
[[ "$(zfs get -H -o value mounted "$dataset/$component")" == yes ]]
[[ "$(zfs get -H -o value mountpoint "$dataset/$component")" == "$source_root/$component" ]]
done
for unit in atlas-nextcloud.service atlas-onlyoffice.service atlas-nextcloud-cron.timer; do
user_run systemctl --user is-active --quiet "$unit"
done
occ status --output=json | python3 -c 'import json,sys; s=json.load(sys.stdin); assert s["installed"] and not s["maintenance"] and not s["needsDbUpgrade"]'
mkdir -p "$backup_root/versions"
chmod 0700 "$backup_root" "$backup_root/versions"
stamp="$(date -u +%Y%m%dT%H%M%SZ)-$$"
snapshot="nc-backup-$stamp"
stage="$backup_root/.partial-$stamp"
mkdir "$stage"
printf '%s\n' "$stamp" > "$state/stamp"
sync -f "$state"
cleanup() {
local rc=$?
trap - EXIT
if ! recover; then rc=1; fi
exit "$rc"
}
trap cleanup EXIT
trap 'exit 143' HUP INT TERM
# Wait for snapshot serialization before interrupting application availability.
flock 9
touch "$state/paused"
sync -f "$state"
user_run systemctl --user stop atlas-nextcloud-cron.timer atlas-nextcloud-cron.service
occ maintenance:mode --on
user_run systemctl --user stop atlas-onlyoffice.service atlas-nextcloud.service
user_run podman exec atlas-nextcloud-db pg_dumpall -U nextcloud --globals-only > "$stage/postgres-globals.sql"
user_run podman exec atlas-nextcloud-db pg_dump -U nextcloud -d nextcloud --format=custom > "$stage/database.dump"
zfs snapshot -r "$dataset@$snapshot"
flock -u 9
resume
# Copy immutable snapshot views; hashing and transfer never extend the outage.
previous=$(readlink -f "$backup_root/latest" 2>/dev/null || true)
for component in app files; do
args=(-aHAX)
if [[ "$previous" == "$backup_root/versions/"* && -d "$previous/$component" ]]; then
args+=("--link-dest=$previous/$component")
fi
if [[ "$component" == app ]]; then args+=(--exclude=/data); fi
rsync "${args[@]}" "$source_root/$component/.zfs/snapshot/$snapshot/" "$stage/$component/"
done
user_run podman exec -i atlas-nextcloud-db pg_restore --list < "$stage/database.dump" > "$stage/database-toc.txt"
user_run podman inspect --format '{% raw %}{{.ImageName}}{% endraw %}' atlas-nextcloud atlas-nextcloud-db atlas-nextcloud-redis atlas-onlyoffice > "$stage/images.txt"
(cd "$stage"; find app files -type f -exec sha256sum '{}' +; sha256sum database.dump postgres-globals.sql images.txt) > "$stage/SHA256SUMS"
(cd "$stage"; sha256sum --quiet --check SHA256SUMS)
printf 'snapshot=%s@%s\ncreated_utc=%s\n' "$dataset" "$snapshot" "$stamp" > "$stage/manifest.txt"
mv "$stage" "$backup_root/versions/$stamp"
ln -s "versions/$stamp" "$backup_root/.latest-$stamp"
mv -Tf "$backup_root/.latest-$stamp" "$backup_root/latest"
sync -f "$backup_root"
# Prune only timestamped job-owned versions after verified atomic publication.
mapfile -t versions < <(find "$backup_root/versions" -mindepth 1 -maxdepth 1 -type d -printf '%f\n' | grep -E '^[0-9]{8}T[0-9]{6}Z-[0-9]+$' | sort -r)
{% raw %}
for ((index=keep; index<${#versions[@]}; index++)); do
{% endraw %}
rm -rf -- "$backup_root/versions/${versions[$index]}"
done
echo "Published verified consistent Nextcloud bundle $stamp; local retention=$keep"

View File

@@ -0,0 +1,12 @@
[Unit]
Description=Discover existing Archive documents and iCloud photos in Nextcloud
Requires=atlas-nextcloud.service
After=atlas-nextcloud.service
[Service]
Type=oneshot
{% for mount in atlas_nextcloud_external_mounts %}
ExecStart=/usr/bin/podman exec --user 33 atlas-nextcloud php occ files:scan "--path={{ mount.user }}/files/{{ mount.name }}" --quiet
{% endfor %}
TimeoutStartSec=90min
NoNewPrivileges=true

View File

@@ -0,0 +1,10 @@
[Unit]
Description=Periodic discovery of Archive changes made outside Nextcloud
[Timer]
OnBootSec=15min
OnUnitInactiveSec=1h
Unit=atlas-nextcloud-external-scan.service
[Install]
WantedBy=timers.target

View File

@@ -2,7 +2,7 @@
Description=Atlas Nextcloud Description=Atlas Nextcloud
Requires=atlas-nextcloud-db.service atlas-nextcloud-redis.service Requires=atlas-nextcloud-db.service atlas-nextcloud-redis.service
After=atlas-nextcloud-db.service atlas-nextcloud-redis.service After=atlas-nextcloud-db.service atlas-nextcloud-redis.service
RequiresMountsFor={{ atlas_nextcloud_root }}/app {{ atlas_nextcloud_root }}/files RequiresMountsFor={{ atlas_nextcloud_root }}/app {{ atlas_nextcloud_root }}/files{% for mount in atlas_nextcloud_external_mounts %} {{ mount.source }}{% endfor %}
[Container] [Container]
ContainerName=atlas-nextcloud ContainerName=atlas-nextcloud
@@ -26,6 +26,10 @@ Environment=PHP_UPLOAD_LIMIT=2G
Volume={{ atlas_nextcloud_root }}/app:/var/www/html:Z Volume={{ atlas_nextcloud_root }}/app:/var/www/html:Z
Volume={{ atlas_nextcloud_root }}/files:/var/www/html/data:Z Volume={{ atlas_nextcloud_root }}/files:/var/www/html/data:Z
Volume={{ atlas_nextcloud_app_cache }}:/mnt/atlas-apps:ro,z Volume={{ atlas_nextcloud_app_cache }}:/mnt/atlas-apps:ro,z
{% for mount in atlas_nextcloud_external_mounts %}
# Shared Archive label, not the private :Z label used for internal state.
Volume={{ mount.source }}:{{ mount.target }}:{{ 'ro' if mount.readonly else 'rw' }},z
{% endfor %}
Volume={{ atlas_nextcloud_private_dir }}/postgres-password:/run/secrets/postgres-password:ro,z Volume={{ atlas_nextcloud_private_dir }}/postgres-password:/run/secrets/postgres-password:ro,z
Volume={{ atlas_nextcloud_private_dir }}/admin-password:/run/secrets/admin-password:ro,z Volume={{ atlas_nextcloud_private_dir }}/admin-password:/run/secrets/admin-password:ro,z
Volume={{ atlas_nextcloud_private_dir }}/redis-password:/run/secrets/redis-password:ro,z Volume={{ atlas_nextcloud_private_dir }}/redis-password:/run/secrets/redis-password:ro,z

View File

@@ -2,8 +2,10 @@
Status: the empty stack was deployed on 2026-10-03, explicitly before the first Status: the empty stack was deployed on 2026-10-03, explicitly before the first
scrub. The operator configured DNS/NPM and authorized public cutover; public TLS, scrub. The operator configured DNS/NPM and authorized public cutover; public TLS,
DAV and cross-user file checks passed. Client editing/sync acceptance and consistent DAV and cross-user file checks passed. The first scrub and a manual consistent
backup/restore validation remain open before family data. iCloud import remains a backup/isolated restore passed on 2026-10-04. Client editing/sync acceptance and
USB recovery validation remain open before family data. Recurring backup integration
and recovery from a new Borg archive passed. iCloud import remains a
separate operation. See `docs/atlas-nextcloud.md` for observed runtime state. separate operation. See `docs/atlas-nextcloud.md` for observed runtime state.
## Confirmed requirements ## Confirmed requirements
@@ -95,7 +97,8 @@ acceptance tests; the app is not treated as proof of server-side compatibility.
1. Validate desktop Office editing/saving, calendar/contact synchronization and 1. Validate desktop Office editing/saving, calendar/contact synchronization and
mobile ONLYOFFICE app integration; public empty-stack cutover is verified. mobile ONLYOFFICE app integration; public empty-stack cutover is verified.
2. Complete protection gates and application-consistent backup/recovery tests. 2. Complete recovery from a new offline USB version; recurring preparation and
encrypted Borg recovery have passed.
3. Plan the deferred iCloud migration when explicitly requested. 3. Plan the deferred iCloud migration when explicitly requested.
## Primary references ## Primary references

View File

@@ -0,0 +1,126 @@
# Nextcloud manual backup/restore rehearsal — 2026-10-04
## Observed outcome
The operator authorized testing consistent database/files backup and recovery.
This was executed directly, not added as a one-time playbook task or feature flag.
No production database was replaced and no iCloud import was performed.
The existing ZFS scrub independently passed: completed at 05:02 CEST after
2:02:04, 0 B repaired, zero errors, successful service exit and healthy pool.
## Backup artifact
Retained on Atlas:
`/zpool/backup/nextcloud-rehearsal-20261004T092635Z` (676 MiB observed).
Its host parent is restricted to admin, mode 0700; dump and manifest files were
created with umask 077. It contains sensitive application configuration and
database contents, not just test data. No plaintext secret was saved to Git.
- Complete application tree, including configuration, custom apps and themes;
the overlaid data directory was copied separately.
- Complete dedicated files tree.
- PostgreSQL custom-format database dump, role definitions, image references,
SHA-256 manifest and canary description.
Cron was stopped, maintenance enabled, and Nextcloud/ONLYOFFICE stopped while
the database dump and application/files copies were taken. Rsync checksum and
metadata comparisons passed while writers were stopped. Live services resumed
with maintenance off, and the cron timer resumed. Role definitions were captured
read-only immediately afterward when the isolated restore exposed the separate
`oc_admin` database role. Its saved password was subsequently verified against
the copied application configuration using SCRAM authentication. The recurring
procedure should capture both database and role dumps during the same pause.
## Isolated restoration
- Fresh rootless PostgreSQL using the exact production image digest, not the
live database volume. Restored roles first, then the database with owners/ACLs.
- Copied application and files directories, using the matching Nextcloud image.
- Fresh, empty isolated Redis; cache contents are not a recovery requirement.
- One pod with `network=none`, no published ports and no live data bind mounts.
Components communicate only through their shared loopback interface.
- Only the restored configuration was adjusted for loopback database/cache,
localhost URLs and disabled mail. Production configuration was unchanged.
- No test cron or Office service was run. External connectivity and callbacks
were impossible from this pod.
Checks passed:
1. Backup SHA-256 verification before restoration and again after cleanup.
2. PostgreSQL role/database restore with failure-on-error enabled.
3. Nextcloud 33.0.9 installed, maintenance off, no pending database upgrade.
4. Restored admin/fabio/chiara accounts and Famiglia permission mask 15.
5. The uniquely named Fabio canary existed both in files and the database index.
6. Authenticated HTTP WebDAV retrieved that canary from the restored instance;
its SHA-256 matched the original uploaded contents.
7. The live canary remained unchanged and was then deleted through WebDAV.
8. Live Nextcloud, ONLYOFFICE and cron timer active; Office connection check
passed. Cloud/Git/Music/Syncthing public HTTPS returned 200 with valid TLS.
Initial fixture failures established two prerequisites: wait for PostgreSQL's
final TCP listener, not the temporary initialization socket, and restore global
roles in addition to the database dump. Persistent Redis settings also require
a working isolated cache. Failed fixture pods were removed before retries.
After success, the final test pod and temporary restore directory were removed.
No test network, live rollback or pool snapshot destruction was needed.
## Limits and remaining work
This validates manual recovery from the current local application/database/files
copy, not a production-size recovery, RPO/RTO compliance, client resynchronization,
Office editing-session recovery, or extraction from Borg/offline USB. ONLYOFFICE's
own persistent service state was not part of this Nextcloud artifact.
The artifact has no dedicated automatic retention policy; do not call it the
recurring Nextcloud backup solution. Integrate a coordinated dump/copy or snapshot
procedure into recurring backups with locking, failure recovery, monitoring and
retention. Independently validate new offsite and offline versions before import.
Keep local/Vault recovery access independent of Nextcloud availability.
The procedure follows the required configuration/apps/files/themes/database scope
and maintenance pause described in the [Nextcloud backup guide](https://docs.nextcloud.com/server/33/admin_manual/maintenance/backup.html)
and tests restoration into a separate environment rather than applying the
[restore procedure](https://docs.nextcloud.com/server/33/admin_manual/maintenance/restore.html)
to production.
## Recurring integration and offsite recovery, later on 2026-10-04
The recurring preparation helper, system service, boot recovery and ordered
Borg/USB dependency drop-ins were deployed from Ansible. The existing snapshot
and backup jobs retain their lock, ownership, namespace, encryption and retention
policies. Two local verified bundles are the declared staging retention; this
supersedes the missing retention warning above for the managed bundle path only.
The earlier manually named rehearsal artifact remains separate and untouched.
Preparation published `20261004T094945Z-3294831`, then starting the actual Borg job
required a fresh preparation and published `20261004T095158Z-3303565`. Application
availability resumed before immutable copy/hash processing completed. The Borg
service successfully published `atlas-20261004T095255Z`, completed pruning and
compaction, and removed its source snapshot after exit.
The latter consistent bundle was extracted from that encrypted Hetzner archive,
not copied from the current local bundle. All SHA-256 checks passed. The extracted
application/files and PostgreSQL role/database dumps were recovered into a fresh
network-none pod with separate database/cache and matching image digests. It
reported installed Nextcloud 33.0.9 without pending upgrade. The three accounts,
Famiglia permissions and authenticated DAV PROPFIND for each account passed;
PostgreSQL used the saved role password with SCRAM on its isolated TCP listener.
The pod, extracted tree and its independent temporary Borg cache were removed.
Production services and the pool remained healthy.
Failure validation used sandboxed helper mocks for maintenance/dump errors: the
original failure code propagated, services/cron resumed, and state/partial files
were removed. Separate transient systemd fixtures verified that a failed ordered
requirement prevents its consumer from executing. These are fault-injection tests,
not production failures or proof of a full host-crash recovery. A real boot with
an interrupted preparation remains untested.
A third preparation published `20261004T100212Z-3366172`; exactly two managed
versions remained, with the oldest version pruned only after publication. Source
snapshots and persistent interruption markers were absent after success.
A new UUID-bound USB version and recovery of its consistent Nextcloud bundle
remain to be verified after the operator connects/unlocks the configured disk.
Do not mark USB recovery complete merely because the dependency was installed.

View File

@@ -17,7 +17,8 @@ were added. An actual repeat run returned `changed=0`, with no failures.
10.2.1 and Team Folders 21.0.9 archives are pinned by version and SHA-256. 10.2.1 and Team Folders 21.0.9 archives are pinned by version and SHA-256.
- Dedicated ZFS namespace: `zpool/services/data/nextcloud`, with separate `app`, - Dedicated ZFS namespace: `zpool/services/data/nextcloud`, with separate `app`,
`files`, `database`, `cache` and `office` datasets. No writable SMB/Syncthing `files`, `database`, `cache` and `office` datasets. No writable SMB/Syncthing
access to the Nextcloud-managed file namespace is provided. access to the Nextcloud-managed file namespace is provided. Existing Archive
directories are exposed separately through local external storage (see below).
- The `admin` Nextcloud account is an application administrator, distinct from - The `admin` Nextcloud account is an application administrator, distinct from
the host account. `fabio` and `chiara` are standard users in `famiglia`, each the host account. `fabio` and `chiara` are standard users in `famiglia`, each
with no initial quota. Team folder `Famiglia` has unlimited quota and group with no initial quota. Team folder `Famiglia` has unlimited quota and group
@@ -101,20 +102,32 @@ Dry-run skips initial downloads, image pulls and runtime account/app commands;
it is not proof of an installed or healthy stack. The deployed repeat run is it is not proof of an installed or healthy stack. The deployed repeat run is
the current idempotence evidence. the current idempotence evidence.
## Manual recovery evidence, 2026-10-04
The first monthly scrub completed successfully and was verified from both the
service result and pool scan (zero errors, 0 B repaired). A manual consistent
application/files copy and PostgreSQL dump were restored into a network-isolated
Nextcloud/PostgreSQL/Redis test pod. Account recovery, Famiglia permissions and
authenticated DAV retrieval of a checksum-matched canary passed. Live services
resumed normally; the test pod and restore workspace were removed.
See `atlas-nextcloud-recovery-test.md` for scope, retained artifact and limitations.
## Gates before family data and full client acceptance ## Gates before family data and full client acceptance
- Verify the first actual scrub and the outstanding protection checks. - The first actual scrub passed on 2026-10-04; preserve the existing protection checks.
- Public TLS, redirects, web login and WebDAV passed. Complete calendar/contact - Public TLS, redirects, web login and WebDAV passed. Complete calendar/contact
synchronization and Office editing/saving from a desktop. synchronization and Office editing/saving from a desktop.
- Test opening, editing and saving from the iPhone/iPad ONLYOFFICE app; mobile - Test opening, editing and saving from the iPhone/iPad ONLYOFFICE app; mobile
browser editing is not a requirement. No such client test is claimed yet. browser editing is not a requirement. No such client test is claimed yet.
- Private-space isolation and cross-user shared writes/deletes passed the public - Private-space isolation and cross-user shared writes/deletes passed the public
smoke test above; complete normal client acceptance as well. smoke test above; complete normal client acceptance as well.
- Integrate and test application-consistent database/files backups before import. - The manual rehearsal and recurring integration passed, including recovery from
a new encrypted Borg archive. Complete recovery from a new USB version before import.
The new datasets fall beneath existing recursive snapshot/backup scope, but The new datasets fall beneath existing recursive snapshot/backup scope, but
that alone does not verify a new Borg/USB version or a consistent Nextcloud restore. that alone does not verify a new Borg/USB version or recovery through those versions.
- For a consistent backup, coordinate pending Office saves, pause cron and writes, - For a consistent backup, coordinate pending Office saves, pause cron and writes,
take a verified PostgreSQL dump and matching application/files snapshot, and take verified PostgreSQL database and role dumps plus a matching application/files
snapshot or quiesced copy, and
resume services promptly even on failure. Extend recurring backup procedures, resume services promptly even on failure. Extend recurring backup procedures,
not the steady-state playbook with one-time migration tasks. Restore into an not the steady-state playbook with one-time migration tasks. Restore into an
isolated environment using matching image/app versions, config, files and DB. isolated environment using matching image/app versions, config, files and DB.
@@ -125,3 +138,92 @@ the current idempotence evidence.
against an upgraded database; use matching tested backups for recovery. against an upgraded database; use matching tested backups for recovery.
- Future Uranus migration and iCloud import are separate, explicitly authorized - Future Uranus migration and iCloud import are separate, explicitly authorized
operations. No source data deletion or automatic cross-system cutover is provided. operations. No source data deletion or automatic cross-system cutover is provided.
## Recurring consistent bundles
`atlas-nextcloud-backup.service` is now an ordered requirement of both
`atlas-borg-backup.service` and the operator-started `atlas-usb-backup.service`.
No additional backup timer is needed: the existing Borg schedule prepares a fresh
bundle before its pool snapshot, and a manual USB run does the same. Preparation
failure blocks the dependent job rather than silently using an old dump.
The helper checks the mounted datasets and healthy active application state,
serializes preparations and briefly pauses cron, Nextcloud and ONLYOFFICE. It
captures database plus global roles and a recursive Nextcloud-only ZFS snapshot.
Services resume before the longer immutable-file copy and checksum verification.
Active editing sessions are interrupted; only committed Nextcloud state is covered.
This does not claim preservation of unsaved ONLYOFFICE editing sessions.
Private bundles are published atomically under `/zpool/backup/nextcloud/versions`,
with a relative `latest` link. `atlas_nextcloud_backup_keep: 2` retains two local
verified versions, with hard links for unchanged files. Long-term Borg, USB and
ZFS policies are unchanged. A trap and `ExecStopPost` restore availability and
clean only this helper's named source snapshot/partial directory; root-private
persistent state permits boot recovery through the enabled recovery unit. Both
units have failure alerts and are included in the Atlas monitored failure units.
No automatic rollback, import, or repair of user application data is performed.
Validation:
```bash
ansible-playbook ansible/site.yml --limit atlas --tags nextcloud_backup,monitoring --check --diff
sudo systemctl start atlas-nextcloud-backup.service
sudo systemctl show atlas-nextcloud-backup.service -p Result -p ExecMainExitTimestamp
```
The second command briefly interrupts the applications and is an explicit manual
run of the recurring job, not a normal deployment side effect. The preparation
unit is not enabled as a boot backup; only interrupted-job recovery is enabled.
Do not stop a Borg/USB job, break its lock or unmount its source snapshot to run a test.
## Existing Archive storage (no import or duplicate originals)
The Atlas declaration exposes only these existing directories to the rootless
Nextcloud container, using shared SELinux `:z` labels:
| Nextcloud folder | Host directory | Access |
| --- | --- | --- |
| `Documenti` | `/zpool/archive/Documents` | Read/write |
| `Foto iCloud` | `/zpool/archive/Pictures/iCloudPD` | Read-only |
Both system mounts are restricted to the Nextcloud user `fabio` only; `chiara`
and the `famiglia` group have no access through these mounts.
External re-sharing is disabled. The photos bind is also read-only at container
level, independently of Nextcloud's mount option. iCloudPD remains the photo
writer. Neither directory is copied into the internal data dataset, and the
existing `Famiglia` team folder remains separate and untouched.
The Archive dataset enables persistent `acltype=posix` support; this does not
change pool features or vdev layout. Scoped ACLs grant the actual rootless-mapped web UID access to existing files and
inheritance on new directories/files. Document defaults retain host administrator
access to files created through Nextcloud; ownership is not changed recursively.
Symlinks are not followed when applying ACLs. Do not change Archive ownership or
apply private `:Z` relabeling to these shared paths.
The user `atlas-nextcloud-external-scan.timer` discovers external changes for only
these mounts and their explicitly allowed user: first after boot at 15 minutes, then one hour
after the previous scan finishes. Nextcloud also checks for external changes on
access. Indexing and previews are not duplicate originals; document versions,
trash, ZFS snapshots and backups may retain additional data intentionally.
Avoid simultaneously editing the same document through SMB and Nextcloud.
Archive originals retain their existing recursive ZFS/Borg/offline USB coverage.
The Nextcloud-only recovery bundle does **not** include these external originals:
a recovery must restore the corresponding Archive data as well as the application
and database. This change does not migrate iCloud Drive or remove anything there.
### Runtime validation, 2026-10-04
The mounts were applied on Atlas without importing originals. A disposable
application-level document create/read test reached the original bind directory;
the probe was deleted, including its trash entry. After the operator narrowed
access to Fabio only, fresh Nextcloud application checks confirmed Fabio can read
both mounts and create documents, while photo create/update/delete are denied.
Chiara cannot access either mount; the separate `Famiglia` team folder remains
available to both users. Container inspection independently confirmed the photo
bind is read-only. Public Nextcloud HTTPS returned 200 and the pool was healthy.
The targeted second Ansible run for mount applicability, options and discovery
unit returned `changed=0`, with no failures. This is focused idempotency evidence,
not a claim about a full Atlas playbook run.