mirror of
https://github.com/fscotto/infra.git
synced 2026-10-04 22:09:50 +00:00
Integrate consistent Nextcloud backups and recovery
This commit is contained in:
27
AGENTS.md
27
AGENTS.md
@@ -68,6 +68,8 @@ Ansible-driven personal infrastructure repo for Fedora and Void desktops, Fedora
|
||||
`ansible-playbook ansible/site.yml --limit atlas --tags gitea_public_domain --check --diff`
|
||||
- Atlas Nextcloud/ONLYOFFICE steady state:
|
||||
`ansible-playbook ansible/site.yml --limit atlas --tags nextcloud --check --diff`
|
||||
- Atlas recurring consistent Nextcloud backup preparation:
|
||||
`ansible-playbook ansible/site.yml --limit atlas --tags nextcloud_backup,monitoring --check --diff`
|
||||
- Atlas iCloudPD storage and boot-started Quadlet:
|
||||
`ansible-playbook ansible/site.yml --limit atlas --tags icloudpd --check --diff`
|
||||
- Ongoing Gitea proxy configuration:
|
||||
@@ -210,15 +212,17 @@ and TCP reachability to Atlas were verified. Temporary Navidrome and Syncthing a
|
||||
manual NPM Proxy Hosts; Syncthing uses `/data/Org` backed by the SMB-shared Archive dataset. Aegis has also
|
||||
validated NFSv4.2 read, write, delete, and `all_squash` mapping to UID/GID `1100` end-to-end. The ZFS
|
||||
snapshot timers are active; a recursive hourly snapshot and scheduled retention prune completed
|
||||
successfully. The first monthly scrub remains a runtime check.
|
||||
successfully. The first monthly scrub completed successfully on 2026-10-04;
|
||||
the actual service result and pool scan were independently verified.
|
||||
|
||||
### Priority 1 - Data protection
|
||||
- [x] Deploy Ansible-managed recursive ZFS snapshots with 24 hourly, 30 daily, 8 weekly, and 12 monthly
|
||||
generations, plus a monthly scrub on the first Sunday at 03:00. The timers and first hourly snapshot were
|
||||
verified on Atlas. Cockpit Scheduler is for visibility or manual operations only, and snapshot
|
||||
rollback is never automated.
|
||||
- [ ] Verify the first monthly ZFS scrub from its actual service result. Scheduled retention pruning
|
||||
was observed on 2026-09-30; timer activation alone does not establish a successful scrub.
|
||||
- [x] Verify the first monthly ZFS scrub from its actual service result. On 2026-10-04
|
||||
it completed at 05:02 CEST after 2:02:04, repairing 0 B with zero errors;
|
||||
the service exited successfully and the pool reported no known data errors.
|
||||
- [x] Activate and validate the encrypted offsite Borg backup to the Hetzner Storage Box. Atlas uses the
|
||||
dedicated SSH identity, pinned ED25519 host key, Vault-backed `repokey` encryption, and a locked
|
||||
non-login `borg` account with no sudo or supplementary groups. The initial snapshot-consistent backup,
|
||||
@@ -376,9 +380,20 @@ successfully. The first monthly scrub remains a runtime check.
|
||||
web login, WebDAV, private-file isolation, Famiglia cross-user create/read/update/delete and
|
||||
CalDAV/CardDAV discovery passed. The Office connector and public health/API asset passed.
|
||||
Temporary test files were removed; no iCloud data was imported.
|
||||
- [ ] Complete Nextcloud desktop/mobile editing and synchronization acceptance, and
|
||||
application-consistent backup/restore validation. Close the first actual scrub and
|
||||
protection checks before importing family data.
|
||||
- [x] Test a manual consistent Nextcloud backup and isolated restore on 2026-10-04.
|
||||
Paused application writers and cron, copied app/config/custom apps/themes and files,
|
||||
dumped PostgreSQL, restored database roles and verified authenticated DAV contents,
|
||||
account recovery and Famiglia permissions. Test containers had no external network,
|
||||
published ports or live data mounts; they and the temporary restore copy were removed.
|
||||
See `docs/atlas-nextcloud-recovery-test.md`; this is not recurring Borg/USB recovery evidence.
|
||||
- [x] Integrate consistent Nextcloud bundles with recurring Borg and operator-started USB
|
||||
backups, two-version local retention, interruption recovery and failure monitoring.
|
||||
On 2026-10-04 Borg archive `atlas-20261004T095255Z` succeeded; its new bundle was
|
||||
extracted from Hetzner and restored in isolation with checksums, accounts, Famiglia
|
||||
permissions and authenticated DAV verified. No production database was replaced.
|
||||
- [ ] Validate a new USB version and restore its consistent Nextcloud bundle after
|
||||
operator connection/unlock. Dependency installation alone is not restore evidence.
|
||||
- [ ] Complete Nextcloud desktop/mobile editing and synchronization acceptance before family import.
|
||||
iCloud migration and future Uranus transfer remain separate operations, not playbook flags.
|
||||
- [x] Move Gitea canonical HTTPS and SSH hostname to `git.fscotto.co` on
|
||||
2026-10-03 through Ansible. Only Gitea restarted; second run changed nothing.
|
||||
|
||||
@@ -50,6 +50,8 @@ atlas_zfs_dataset_photobook: media/photobook
|
||||
atlas_mount_root: /zpool
|
||||
atlas_manage_storage: true
|
||||
atlas_manage_nextcloud: true
|
||||
# Two local consistent bundles; long-term history stays in Borg/USB and ZFS.
|
||||
atlas_nextcloud_backup_keep: 2
|
||||
atlas_nextcloud_domain: cloud.fscotto.co
|
||||
atlas_onlyoffice_domain: office.fscotto.co
|
||||
# Resolved official amd64 images on 2026-10-03; updates are deliberate.
|
||||
@@ -64,6 +66,17 @@ atlas_nextcloud_users:
|
||||
- username: chiara
|
||||
display_name: Chiara
|
||||
password: "{{ vault_nextcloud_chiara_password }}"
|
||||
atlas_nextcloud_external_mounts:
|
||||
- name: Documenti
|
||||
user: fabio
|
||||
source: /zpool/archive/Documents
|
||||
target: /mnt/archive-documents
|
||||
readonly: false
|
||||
- name: Foto iCloud
|
||||
user: fabio
|
||||
source: /zpool/archive/Pictures/iCloudPD
|
||||
target: /mnt/archive-icloud
|
||||
readonly: true
|
||||
atlas_nextcloud_apps:
|
||||
- id: groupfolders
|
||||
version: 21.0.9
|
||||
|
||||
@@ -16,9 +16,13 @@ atlas_nextcloud_image: ""
|
||||
atlas_nextcloud_postgres_image: ""
|
||||
atlas_nextcloud_redis_image: ""
|
||||
atlas_onlyoffice_image: ""
|
||||
atlas_nextcloud_backup_root: "{{ atlas_mount_root }}/backup/nextcloud"
|
||||
atlas_nextcloud_backup_keep: 2
|
||||
atlas_nextcloud_admin: admin
|
||||
atlas_nextcloud_users: []
|
||||
atlas_nextcloud_apps: []
|
||||
# Existing Archive directories; never import into the internal data namespace.
|
||||
atlas_nextcloud_external_mounts: []
|
||||
atlas_nextcloud_services:
|
||||
- atlas-nextcloud-db.service
|
||||
- atlas-nextcloud-redis.service
|
||||
@@ -147,7 +151,9 @@ atlas_monitor_effective_timers: >-
|
||||
atlas_monitor_effective_failure_units: >-
|
||||
{{ atlas_monitor_failure_units
|
||||
+ (['atlas-prometheus-pull.service']
|
||||
if atlas_manage_prometheus_backup_pull | bool else []) }}
|
||||
if atlas_manage_prometheus_backup_pull | bool else [])
|
||||
+ (['atlas-nextcloud-backup.service', 'atlas-nextcloud-backup-recovery.service']
|
||||
if atlas_manage_nextcloud | bool else []) }}
|
||||
atlas_monitor_remote_capacity: {}
|
||||
atlas_monitor_pool_warning_percent: 80
|
||||
atlas_monitor_pool_critical_percent: 90
|
||||
|
||||
@@ -35,6 +35,9 @@
|
||||
- name: Import Atlas offline USB backup tasks
|
||||
ansible.builtin.import_tasks: usb_backup.yml
|
||||
|
||||
- name: Import recurring Nextcloud backup preparation
|
||||
ansible.builtin.import_tasks: nextcloud_backup.yml
|
||||
|
||||
- name: Import Atlas Prometheus backup pull identity tasks
|
||||
ansible.builtin.import_tasks: prometheus_pull_identity.yml
|
||||
|
||||
|
||||
@@ -39,6 +39,10 @@
|
||||
(atlas_nextcloud_users | map(attribute='password') | list) }}
|
||||
no_log: true
|
||||
|
||||
- name: Prepare access to declared existing Archive directories
|
||||
ansible.builtin.include_tasks: nextcloud_external_access.yml
|
||||
when: atlas_nextcloud_external_mounts | length > 0
|
||||
|
||||
- name: Verify the existing application-data parent is mounted
|
||||
community.general.zfs_facts:
|
||||
name: "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_app_data }}"
|
||||
@@ -197,7 +201,8 @@
|
||||
owner: "{{ atlas_admin_username }}"
|
||||
group: "{{ atlas_admin_group }}"
|
||||
mode: "0644"
|
||||
loop: [atlas-nextcloud-cron.service, atlas-nextcloud-cron.timer]
|
||||
loop: [atlas-nextcloud-cron.service, atlas-nextcloud-cron.timer,
|
||||
atlas-nextcloud-external-scan.service, atlas-nextcloud-external-scan.timer]
|
||||
register: atlas_nextcloud_cron_units
|
||||
|
||||
- name: Manage and verify rootless Nextcloud services
|
||||
@@ -227,7 +232,13 @@
|
||||
scope: user
|
||||
name: "{{ item }}"
|
||||
state: >-
|
||||
{{ 'restarted' if (atlas_nextcloud_quadlets is changed or
|
||||
{{ 'restarted' if (
|
||||
atlas_nextcloud_quadlets.results |
|
||||
selectattr('item', 'equalto', item | replace('.service', '.container')) |
|
||||
selectattr('changed') | list | length > 0 or
|
||||
atlas_nextcloud_quadlets.results |
|
||||
selectattr('item', 'equalto', 'atlas-nextcloud.network') |
|
||||
selectattr('changed') | list | length > 0 or
|
||||
atlas_nextcloud_private_configuration is changed or
|
||||
atlas_nextcloud_secret_files is changed) else 'started' }}
|
||||
loop: "{{ atlas_nextcloud_services }}"
|
||||
@@ -299,6 +310,14 @@
|
||||
state: "{{ 'restarted' if atlas_nextcloud_cron_units is changed else 'started' }}"
|
||||
enabled: true
|
||||
|
||||
- name: Enable periodic targeted Archive discovery
|
||||
ansible.builtin.systemd:
|
||||
scope: user
|
||||
name: atlas-nextcloud-external-scan.timer
|
||||
state: "{{ 'restarted' if atlas_nextcloud_cron_units is changed else 'started' }}"
|
||||
enabled: true
|
||||
when: atlas_nextcloud_external_mounts | length > 0
|
||||
|
||||
- name: Verify ONLYOFFICE local health without publishing the domain
|
||||
ansible.builtin.uri:
|
||||
url: "http://127.0.0.1:{{ atlas_onlyoffice_http_port }}/healthcheck"
|
||||
|
||||
@@ -169,3 +169,18 @@
|
||||
argv: [podman, exec, --user, '33', atlas-nextcloud, php, occ, background:cron]
|
||||
when: atlas_nextcloud_background_mode.stdout | trim != 'cron'
|
||||
changed_when: true
|
||||
|
||||
- name: Enable shipped external storage support when required
|
||||
ansible.builtin.command:
|
||||
argv: [podman, exec, --user, '33', atlas-nextcloud, php, occ, app:enable, files_external]
|
||||
when:
|
||||
- atlas_nextcloud_external_mounts | length > 0
|
||||
- "'files_external' not in (atlas_nextcloud_current_apps.stdout | from_json).enabled"
|
||||
changed_when: true
|
||||
|
||||
- name: Maintain only the declared Archive mounts
|
||||
ansible.builtin.include_tasks: nextcloud_external_mount.yml
|
||||
loop: "{{ atlas_nextcloud_external_mounts }}"
|
||||
loop_control:
|
||||
loop_var: atlas_nextcloud_mount
|
||||
label: "{{ atlas_nextcloud_mount.name }}"
|
||||
|
||||
55
ansible/roles/profile_atlas/tasks/nextcloud_backup.yml
Normal file
55
ansible/roles/profile_atlas/tasks/nextcloud_backup.yml
Normal file
@@ -0,0 +1,55 @@
|
||||
---
|
||||
- name: Manage recurring consistent Nextcloud backup preparation
|
||||
tags: [atlas, nextcloud_backup]
|
||||
when: atlas_manage_nextcloud | bool
|
||||
block:
|
||||
- name: Validate private backup scope and local bundle retention
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- atlas_nextcloud_backup_root == atlas_mount_root ~ '/backup/nextcloud'
|
||||
- atlas_nextcloud_backup_keep | int >= 2
|
||||
- atlas_manage_borg_backup | bool
|
||||
- atlas_manage_usb_backup | bool
|
||||
|
||||
- name: Install recurring backup helper with shell syntax validation
|
||||
ansible.builtin.template:
|
||||
src: atlas-nextcloud-backup.sh.j2
|
||||
dest: /usr/local/sbin/atlas-nextcloud-backup
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0750"
|
||||
validate: /bin/bash -n %s
|
||||
|
||||
- name: Install Nextcloud backup preparation and boot recovery units
|
||||
ansible.builtin.template:
|
||||
src: "{{ item }}.j2"
|
||||
dest: "/etc/systemd/system/{{ item }}"
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0644"
|
||||
loop: [atlas-nextcloud-backup.service, atlas-nextcloud-backup-recovery.service]
|
||||
|
||||
- name: Create backup dependency drop-in directories
|
||||
ansible.builtin.file:
|
||||
path: "/etc/systemd/system/{{ item }}.d"
|
||||
state: directory
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0755"
|
||||
loop: [atlas-borg-backup.service, atlas-usb-backup.service]
|
||||
|
||||
- name: Require a fresh consistent bundle before offsite and manual USB backups
|
||||
ansible.builtin.template:
|
||||
src: atlas-nextcloud-backup-dependency.conf.j2
|
||||
dest: "/etc/systemd/system/{{ item }}.d/nextcloud.conf"
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0644"
|
||||
loop: [atlas-borg-backup.service, atlas-usb-backup.service]
|
||||
|
||||
- name: Reload systemd and enable interruption recovery without running a backup
|
||||
ansible.builtin.systemd:
|
||||
daemon_reload: true
|
||||
name: atlas-nextcloud-backup-recovery.service
|
||||
enabled: true
|
||||
when: not ansible_check_mode
|
||||
100
ansible/roles/profile_atlas/tasks/nextcloud_external_access.yml
Normal file
100
ansible/roles/profile_atlas/tasks/nextcloud_external_access.yml
Normal file
@@ -0,0 +1,100 @@
|
||||
---
|
||||
- name: Restrict external storage to explicit Archive directories
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- item.source in [atlas_archive_mountpoint ~ '/Documents', atlas_icloudpd_photos_dir]
|
||||
- item.target is match('^/mnt/archive-[a-z]+$')
|
||||
- item.name is match('^[A-Za-z][A-Za-z ]+$')
|
||||
- item.readonly is boolean
|
||||
- item.user in (atlas_nextcloud_users | map(attribute='username') | list)
|
||||
- item.source != atlas_icloudpd_photos_dir or item.readonly
|
||||
loop: "{{ atlas_nextcloud_external_mounts }}"
|
||||
|
||||
- name: Inspect existing sources without creating or moving data
|
||||
ansible.builtin.stat:
|
||||
path: "{{ item.source }}"
|
||||
follow: false
|
||||
loop: "{{ atlas_nextcloud_external_mounts }}"
|
||||
register: atlas_nextcloud_external_sources
|
||||
|
||||
- name: Refuse missing sources and symlinks
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- item.stat.isdir | default(false)
|
||||
- not (item.stat.islnk | default(false))
|
||||
loop: "{{ atlas_nextcloud_external_sources.results }}"
|
||||
loop_control:
|
||||
label: "{{ item.item.source }}"
|
||||
|
||||
- name: Verify the Archive dataset before modifying its ACL capability
|
||||
community.general.zfs_facts:
|
||||
name: "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_archive }}"
|
||||
properties: name,mounted,mountpoint
|
||||
register: atlas_nextcloud_external_dataset
|
||||
|
||||
- name: Refuse an absent or unmounted Archive dataset
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- atlas_nextcloud_external_dataset.ansible_facts.ansible_zfs_datasets | length == 1
|
||||
- atlas_nextcloud_external_dataset.ansible_facts.ansible_zfs_datasets[0].mounted == 'yes'
|
||||
- atlas_nextcloud_external_dataset.ansible_facts.ansible_zfs_datasets[0].mountpoint == atlas_archive_mountpoint
|
||||
|
||||
- name: Enable persistent POSIX ACL support on the verified Archive dataset
|
||||
community.general.zfs:
|
||||
name: "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_archive }}"
|
||||
state: present
|
||||
extra_zfs_properties:
|
||||
acltype: posix
|
||||
|
||||
- name: Derive actual rootless web UID for narrowly scoped Archive ACLs
|
||||
become_user: "{{ atlas_admin_username }}"
|
||||
environment:
|
||||
XDG_RUNTIME_DIR: "/run/user/{{ atlas_admin_uid }}"
|
||||
ansible.builtin.command:
|
||||
argv:
|
||||
- podman
|
||||
- unshare
|
||||
- python3
|
||||
- -c
|
||||
- >-
|
||||
print(next(int(b)+33-int(a) for a,b,n in
|
||||
(l.split() for l in open('/proc/self/uid_map')) if int(a)<=33<int(a)+int(n)))
|
||||
register: atlas_nextcloud_external_uid
|
||||
changed_when: false
|
||||
check_mode: false
|
||||
|
||||
- name: Grant web user access only inside the declared sources
|
||||
ansible.posix.acl:
|
||||
path: "{{ item.source }}"
|
||||
entity: "{{ atlas_nextcloud_external_uid.stdout | trim }}"
|
||||
etype: user
|
||||
permissions: "{{ 'rX' if item.readonly else 'rwX' }}"
|
||||
recursive: true
|
||||
follow: false
|
||||
state: present
|
||||
loop: "{{ atlas_nextcloud_external_mounts }}"
|
||||
|
||||
- name: Inherit web access on new files and directories
|
||||
ansible.posix.acl:
|
||||
path: "{{ item.source }}"
|
||||
entity: "{{ atlas_nextcloud_external_uid.stdout | trim }}"
|
||||
etype: user
|
||||
permissions: "{{ 'rX' if item.readonly else 'rwX' }}"
|
||||
default: true
|
||||
recursive: true
|
||||
follow: false
|
||||
state: present
|
||||
loop: "{{ atlas_nextcloud_external_mounts }}"
|
||||
|
||||
- name: Preserve administrator access to documents created through Nextcloud
|
||||
ansible.posix.acl:
|
||||
path: "{{ item.source }}"
|
||||
entity: "{{ atlas_admin_uid }}"
|
||||
etype: user
|
||||
permissions: rwX
|
||||
default: true
|
||||
recursive: true
|
||||
follow: false
|
||||
state: present
|
||||
loop: "{{ atlas_nextcloud_external_mounts }}"
|
||||
when: not item.readonly
|
||||
@@ -0,0 +1,94 @@
|
||||
---
|
||||
- name: Inspect current system mounts without exposing credentials
|
||||
ansible.builtin.command:
|
||||
argv: [podman, exec, --user, '33', atlas-nextcloud, php, occ, files_external:list, --output=json]
|
||||
register: atlas_nextcloud_mount_list
|
||||
changed_when: false
|
||||
no_log: true
|
||||
|
||||
- name: Select only the matching mount name
|
||||
ansible.builtin.set_fact:
|
||||
atlas_nextcloud_matching_mounts: >-
|
||||
{{ atlas_nextcloud_mount_list.stdout | from_json |
|
||||
selectattr('mount_point', 'equalto', '/' ~ atlas_nextcloud_mount.name) | list }}
|
||||
no_log: true
|
||||
|
||||
- name: Refuse duplicates or repurposing of existing unrelated storage
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- atlas_nextcloud_matching_mounts | length <= 1
|
||||
- >-
|
||||
atlas_nextcloud_matching_mounts | length == 0 or
|
||||
(atlas_nextcloud_matching_mounts[0].configuration.datadir | default('') == atlas_nextcloud_mount.target
|
||||
and atlas_nextcloud_matching_mounts[0].storage == '\\OC\\Files\\Storage\\Local')
|
||||
fail_msg: Existing storage conflicts with the declared Archive mount; refusing an implicit replacement.
|
||||
|
||||
- name: Create an absent local mount restricted to its declared user
|
||||
ansible.builtin.command:
|
||||
argv:
|
||||
- podman
|
||||
- exec
|
||||
- --user
|
||||
- '33'
|
||||
- atlas-nextcloud
|
||||
- php
|
||||
- occ
|
||||
- files_external:create
|
||||
- "{{ atlas_nextcloud_mount.name }}"
|
||||
- local
|
||||
- null::null
|
||||
- --config
|
||||
- "datadir={{ atlas_nextcloud_mount.target }}"
|
||||
- --applicable-user
|
||||
- "{{ atlas_nextcloud_mount.user }}"
|
||||
- --output=json
|
||||
when: atlas_nextcloud_matching_mounts | length == 0
|
||||
register: atlas_nextcloud_mount_created
|
||||
changed_when: true
|
||||
|
||||
- name: Record the managed mount ID and options
|
||||
ansible.builtin.set_fact:
|
||||
atlas_nextcloud_mount_id: >-
|
||||
{{ atlas_nextcloud_mount_created.stdout | trim if atlas_nextcloud_matching_mounts | length == 0
|
||||
else atlas_nextcloud_matching_mounts[0].mount_id }}
|
||||
atlas_nextcloud_mount_options: >-
|
||||
{{ {} if atlas_nextcloud_matching_mounts | length == 0 else atlas_nextcloud_matching_mounts[0].options }}
|
||||
|
||||
- name: Restrict the managed mount to exactly its declared user
|
||||
ansible.builtin.command:
|
||||
argv: >-
|
||||
{{ ['podman', 'exec', '--user', '33', 'atlas-nextcloud', 'php', 'occ',
|
||||
'files_external:applicable', atlas_nextcloud_mount_id | string,
|
||||
'--add-user=' ~ atlas_nextcloud_mount.user] +
|
||||
(atlas_nextcloud_matching_mounts[0].applicable_groups |
|
||||
map('regex_replace', '^', '--remove-group=') | list) +
|
||||
(atlas_nextcloud_matching_mounts[0].applicable_users |
|
||||
reject('equalto', atlas_nextcloud_mount.user) |
|
||||
map('regex_replace', '^', '--remove-user=') | list) }}
|
||||
when:
|
||||
- atlas_nextcloud_matching_mounts | length > 0
|
||||
- >-
|
||||
atlas_nextcloud_matching_mounts[0].applicable_groups | length > 0 or
|
||||
atlas_nextcloud_matching_mounts[0].applicable_users != [atlas_nextcloud_mount.user]
|
||||
changed_when: true
|
||||
|
||||
- name: Maintain read-only photos and external change detection
|
||||
ansible.builtin.command:
|
||||
argv: [podman, exec, --user, '33', atlas-nextcloud, php, occ, files_external:option,
|
||||
"{{ atlas_nextcloud_mount_id }}", "{{ item.key }}", "{{ item.value | to_json }}"]
|
||||
loop:
|
||||
- {key: readonly, value: "{{ atlas_nextcloud_mount.readonly }}"}
|
||||
- {key: filesystem_check_changes, value: 1}
|
||||
- {key: enable_sharing, value: false}
|
||||
# Nextcloud persists option values as strings ("1" / "" for booleans).
|
||||
when: >-
|
||||
item.key not in atlas_nextcloud_mount_options or
|
||||
atlas_nextcloud_mount_options[item.key] | string !=
|
||||
(('1' if item.value else '') if item.value is boolean else item.value | string)
|
||||
changed_when: true
|
||||
|
||||
- name: Verify the managed local storage is accessible
|
||||
ansible.builtin.command:
|
||||
argv: [podman, exec, --user, '33', atlas-nextcloud, php, occ, files_external:verify,
|
||||
"{{ atlas_nextcloud_mount_id }}"]
|
||||
changed_when: false
|
||||
@@ -10,6 +10,7 @@
|
||||
properties:
|
||||
compression: zstd
|
||||
mountpoint: "{{ atlas_archive_mountpoint }}"
|
||||
acltype: posix
|
||||
- name: "{{ atlas_zfs_pool }}/{{ atlas_zfs_dataset_services }}"
|
||||
mountpoint: "{{ atlas_services_mountpoint }}"
|
||||
owner: "{{ atlas_admin_username }}"
|
||||
|
||||
@@ -0,0 +1,3 @@
|
||||
[Unit]
|
||||
Requires=atlas-nextcloud-backup.service
|
||||
After=atlas-nextcloud-backup.service
|
||||
@@ -0,0 +1,19 @@
|
||||
[Unit]
|
||||
Description=Recover interrupted Nextcloud backup preparation after boot
|
||||
Requires=zfs.target user@{{ atlas_admin_uid }}.service
|
||||
After=zfs.target user@{{ atlas_admin_uid }}.service
|
||||
{% if atlas_manage_monitoring | bool %}
|
||||
OnFailure=atlas-monitor-failure@%n.service
|
||||
{% endif %}
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
User=root
|
||||
UMask=0077
|
||||
StateDirectory=atlas-nextcloud-backup
|
||||
StateDirectoryMode=0700
|
||||
ExecStart=/usr/local/sbin/atlas-nextcloud-backup --recover
|
||||
TimeoutStartSec=5min
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
@@ -0,0 +1,21 @@
|
||||
[Unit]
|
||||
Description=Prepare a consistent Nextcloud bundle before Atlas backups
|
||||
Requires=zfs.target user@{{ atlas_admin_uid }}.service
|
||||
After=zfs.target user@{{ atlas_admin_uid }}.service atlas-nextcloud-backup-recovery.service
|
||||
{% if atlas_manage_monitoring | bool %}
|
||||
OnFailure=atlas-monitor-failure@%n.service
|
||||
{% endif %}
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
User=root
|
||||
UMask=0077
|
||||
StateDirectory=atlas-nextcloud-backup
|
||||
StateDirectoryMode=0700
|
||||
ExecStart=/usr/local/sbin/atlas-nextcloud-backup
|
||||
ExecStopPost=/usr/local/sbin/atlas-nextcloud-backup --recover
|
||||
TimeoutStartSec=3h
|
||||
TimeoutStopSec=5min
|
||||
Nice=10
|
||||
IOSchedulingClass=best-effort
|
||||
IOSchedulingPriority=7
|
||||
@@ -0,0 +1,133 @@
|
||||
#!/usr/bin/env bash
|
||||
set -Eeuo pipefail
|
||||
export PATH=/usr/sbin:/usr/bin:/sbin:/bin
|
||||
umask 077
|
||||
readonly dataset={{ atlas_nextcloud_dataset | quote }}
|
||||
readonly source_root={{ atlas_nextcloud_root | quote }}
|
||||
readonly backup_root={{ atlas_nextcloud_backup_root | quote }}
|
||||
readonly state=/var/lib/atlas-nextcloud-backup
|
||||
readonly keep={{ atlas_nextcloud_backup_keep | int }}
|
||||
readonly owner={{ atlas_admin_username | quote }}
|
||||
readonly uid={{ atlas_admin_uid | int }}
|
||||
|
||||
user_run() {
|
||||
runuser -u "$owner" -- env XDG_RUNTIME_DIR="/run/user/$uid" \
|
||||
DBUS_SESSION_BUS_ADDRESS="unix:path=/run/user/$uid/bus" "$@"
|
||||
}
|
||||
occ() { user_run podman exec --user 33 atlas-nextcloud php occ "$@"; }
|
||||
exec 8>/run/lock/atlas-nextcloud-backup.lock
|
||||
flock 8
|
||||
exec 9>/run/lock/atlas-zfs-snapshot.lock
|
||||
mkdir -p "$state"
|
||||
chmod 0700 "$state"
|
||||
|
||||
resume() {
|
||||
[[ -e "$state/paused" ]] || return 0
|
||||
user_run systemctl --user start atlas-nextcloud.service atlas-onlyoffice.service
|
||||
local ready=false
|
||||
for _ in {1..60}; do
|
||||
if occ maintenance:mode --off >/dev/null 2>&1; then ready=true; break; fi
|
||||
sleep 2
|
||||
done
|
||||
[[ "$ready" == true ]] || { echo 'Nextcloud resume failed; recovery marker retained' >&2; return 1; }
|
||||
user_run systemctl --user start atlas-nextcloud-cron.timer
|
||||
# Persist maintenance-off before clearing durable interruption ownership.
|
||||
sync -f "$source_root/app"
|
||||
rm "$state/paused"
|
||||
sync -f "$state"
|
||||
echo 'Nextcloud/Office resumed and cron timer restored'
|
||||
}
|
||||
|
||||
recover() {
|
||||
resume || return 1
|
||||
[[ -e "$state/stamp" ]] || return 0
|
||||
local stamp snapshot mount source
|
||||
stamp=$(cat "$state/stamp")
|
||||
[[ "$stamp" =~ ^[0-9]{8}T[0-9]{6}Z-[0-9]+$ ]] || return 65
|
||||
snapshot="nc-backup-$stamp"
|
||||
flock 9
|
||||
for component in files app; do
|
||||
mount="$source_root/$component/.zfs/snapshot/$snapshot"
|
||||
source=$(findmnt -rn -M "$mount" -o SOURCE || true)
|
||||
if [[ -n "$source" ]]; then
|
||||
[[ "$source" == "$dataset/$component@$snapshot" ]] || return 65
|
||||
umount "$mount" || return 1
|
||||
fi
|
||||
done
|
||||
if zfs list -H -t snapshot "$dataset@$snapshot" >/dev/null 2>&1; then
|
||||
zfs destroy -r "$dataset@$snapshot" || return 1
|
||||
fi
|
||||
flock -u 9
|
||||
# Only this job's private, unpublished staging directory can be removed.
|
||||
rm -rf -- "$backup_root/.partial-$stamp"
|
||||
rm "$state/stamp"
|
||||
}
|
||||
if [[ "${1:-}" == --recover ]]; then recover; exit; fi
|
||||
recover
|
||||
[[ "$(zfs get -H -o value mounted "$dataset")" == yes ]]
|
||||
[[ "$(zfs get -H -o value mountpoint "$dataset")" == "$source_root" ]]
|
||||
[[ "$(zfs get -H -o value mounted {{ (atlas_zfs_pool ~ '/backup') | quote }})" == yes ]]
|
||||
[[ "$(zfs get -H -o value mountpoint {{ (atlas_zfs_pool ~ '/backup') | quote }})" == {{ (atlas_mount_root ~ '/backup') | quote }} ]]
|
||||
for component in app files; do
|
||||
[[ "$(zfs get -H -o value mounted "$dataset/$component")" == yes ]]
|
||||
[[ "$(zfs get -H -o value mountpoint "$dataset/$component")" == "$source_root/$component" ]]
|
||||
done
|
||||
for unit in atlas-nextcloud.service atlas-onlyoffice.service atlas-nextcloud-cron.timer; do
|
||||
user_run systemctl --user is-active --quiet "$unit"
|
||||
done
|
||||
occ status --output=json | python3 -c 'import json,sys; s=json.load(sys.stdin); assert s["installed"] and not s["maintenance"] and not s["needsDbUpgrade"]'
|
||||
mkdir -p "$backup_root/versions"
|
||||
chmod 0700 "$backup_root" "$backup_root/versions"
|
||||
stamp="$(date -u +%Y%m%dT%H%M%SZ)-$$"
|
||||
snapshot="nc-backup-$stamp"
|
||||
stage="$backup_root/.partial-$stamp"
|
||||
mkdir "$stage"
|
||||
printf '%s\n' "$stamp" > "$state/stamp"
|
||||
sync -f "$state"
|
||||
cleanup() {
|
||||
local rc=$?
|
||||
trap - EXIT
|
||||
if ! recover; then rc=1; fi
|
||||
exit "$rc"
|
||||
}
|
||||
trap cleanup EXIT
|
||||
trap 'exit 143' HUP INT TERM
|
||||
# Wait for snapshot serialization before interrupting application availability.
|
||||
flock 9
|
||||
touch "$state/paused"
|
||||
sync -f "$state"
|
||||
user_run systemctl --user stop atlas-nextcloud-cron.timer atlas-nextcloud-cron.service
|
||||
occ maintenance:mode --on
|
||||
user_run systemctl --user stop atlas-onlyoffice.service atlas-nextcloud.service
|
||||
user_run podman exec atlas-nextcloud-db pg_dumpall -U nextcloud --globals-only > "$stage/postgres-globals.sql"
|
||||
user_run podman exec atlas-nextcloud-db pg_dump -U nextcloud -d nextcloud --format=custom > "$stage/database.dump"
|
||||
zfs snapshot -r "$dataset@$snapshot"
|
||||
flock -u 9
|
||||
resume
|
||||
# Copy immutable snapshot views; hashing and transfer never extend the outage.
|
||||
previous=$(readlink -f "$backup_root/latest" 2>/dev/null || true)
|
||||
for component in app files; do
|
||||
args=(-aHAX)
|
||||
if [[ "$previous" == "$backup_root/versions/"* && -d "$previous/$component" ]]; then
|
||||
args+=("--link-dest=$previous/$component")
|
||||
fi
|
||||
if [[ "$component" == app ]]; then args+=(--exclude=/data); fi
|
||||
rsync "${args[@]}" "$source_root/$component/.zfs/snapshot/$snapshot/" "$stage/$component/"
|
||||
done
|
||||
user_run podman exec -i atlas-nextcloud-db pg_restore --list < "$stage/database.dump" > "$stage/database-toc.txt"
|
||||
user_run podman inspect --format '{% raw %}{{.ImageName}}{% endraw %}' atlas-nextcloud atlas-nextcloud-db atlas-nextcloud-redis atlas-onlyoffice > "$stage/images.txt"
|
||||
(cd "$stage"; find app files -type f -exec sha256sum '{}' +; sha256sum database.dump postgres-globals.sql images.txt) > "$stage/SHA256SUMS"
|
||||
(cd "$stage"; sha256sum --quiet --check SHA256SUMS)
|
||||
printf 'snapshot=%s@%s\ncreated_utc=%s\n' "$dataset" "$snapshot" "$stamp" > "$stage/manifest.txt"
|
||||
mv "$stage" "$backup_root/versions/$stamp"
|
||||
ln -s "versions/$stamp" "$backup_root/.latest-$stamp"
|
||||
mv -Tf "$backup_root/.latest-$stamp" "$backup_root/latest"
|
||||
sync -f "$backup_root"
|
||||
# Prune only timestamped job-owned versions after verified atomic publication.
|
||||
mapfile -t versions < <(find "$backup_root/versions" -mindepth 1 -maxdepth 1 -type d -printf '%f\n' | grep -E '^[0-9]{8}T[0-9]{6}Z-[0-9]+$' | sort -r)
|
||||
{% raw %}
|
||||
for ((index=keep; index<${#versions[@]}; index++)); do
|
||||
{% endraw %}
|
||||
rm -rf -- "$backup_root/versions/${versions[$index]}"
|
||||
done
|
||||
echo "Published verified consistent Nextcloud bundle $stamp; local retention=$keep"
|
||||
@@ -0,0 +1,12 @@
|
||||
[Unit]
|
||||
Description=Discover existing Archive documents and iCloud photos in Nextcloud
|
||||
Requires=atlas-nextcloud.service
|
||||
After=atlas-nextcloud.service
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
{% for mount in atlas_nextcloud_external_mounts %}
|
||||
ExecStart=/usr/bin/podman exec --user 33 atlas-nextcloud php occ files:scan "--path={{ mount.user }}/files/{{ mount.name }}" --quiet
|
||||
{% endfor %}
|
||||
TimeoutStartSec=90min
|
||||
NoNewPrivileges=true
|
||||
@@ -0,0 +1,10 @@
|
||||
[Unit]
|
||||
Description=Periodic discovery of Archive changes made outside Nextcloud
|
||||
|
||||
[Timer]
|
||||
OnBootSec=15min
|
||||
OnUnitInactiveSec=1h
|
||||
Unit=atlas-nextcloud-external-scan.service
|
||||
|
||||
[Install]
|
||||
WantedBy=timers.target
|
||||
@@ -2,7 +2,7 @@
|
||||
Description=Atlas Nextcloud
|
||||
Requires=atlas-nextcloud-db.service atlas-nextcloud-redis.service
|
||||
After=atlas-nextcloud-db.service atlas-nextcloud-redis.service
|
||||
RequiresMountsFor={{ atlas_nextcloud_root }}/app {{ atlas_nextcloud_root }}/files
|
||||
RequiresMountsFor={{ atlas_nextcloud_root }}/app {{ atlas_nextcloud_root }}/files{% for mount in atlas_nextcloud_external_mounts %} {{ mount.source }}{% endfor %}
|
||||
|
||||
[Container]
|
||||
ContainerName=atlas-nextcloud
|
||||
@@ -26,6 +26,10 @@ Environment=PHP_UPLOAD_LIMIT=2G
|
||||
Volume={{ atlas_nextcloud_root }}/app:/var/www/html:Z
|
||||
Volume={{ atlas_nextcloud_root }}/files:/var/www/html/data:Z
|
||||
Volume={{ atlas_nextcloud_app_cache }}:/mnt/atlas-apps:ro,z
|
||||
{% for mount in atlas_nextcloud_external_mounts %}
|
||||
# Shared Archive label, not the private :Z label used for internal state.
|
||||
Volume={{ mount.source }}:{{ mount.target }}:{{ 'ro' if mount.readonly else 'rw' }},z
|
||||
{% endfor %}
|
||||
Volume={{ atlas_nextcloud_private_dir }}/postgres-password:/run/secrets/postgres-password:ro,z
|
||||
Volume={{ atlas_nextcloud_private_dir }}/admin-password:/run/secrets/admin-password:ro,z
|
||||
Volume={{ atlas_nextcloud_private_dir }}/redis-password:/run/secrets/redis-password:ro,z
|
||||
|
||||
@@ -2,8 +2,10 @@
|
||||
|
||||
Status: the empty stack was deployed on 2026-10-03, explicitly before the first
|
||||
scrub. The operator configured DNS/NPM and authorized public cutover; public TLS,
|
||||
DAV and cross-user file checks passed. Client editing/sync acceptance and consistent
|
||||
backup/restore validation remain open before family data. iCloud import remains a
|
||||
DAV and cross-user file checks passed. The first scrub and a manual consistent
|
||||
backup/isolated restore passed on 2026-10-04. Client editing/sync acceptance and
|
||||
USB recovery validation remain open before family data. Recurring backup integration
|
||||
and recovery from a new Borg archive passed. iCloud import remains a
|
||||
separate operation. See `docs/atlas-nextcloud.md` for observed runtime state.
|
||||
|
||||
## Confirmed requirements
|
||||
@@ -95,7 +97,8 @@ acceptance tests; the app is not treated as proof of server-side compatibility.
|
||||
|
||||
1. Validate desktop Office editing/saving, calendar/contact synchronization and
|
||||
mobile ONLYOFFICE app integration; public empty-stack cutover is verified.
|
||||
2. Complete protection gates and application-consistent backup/recovery tests.
|
||||
2. Complete recovery from a new offline USB version; recurring preparation and
|
||||
encrypted Borg recovery have passed.
|
||||
3. Plan the deferred iCloud migration when explicitly requested.
|
||||
|
||||
## Primary references
|
||||
|
||||
126
docs/atlas-nextcloud-recovery-test.md
Normal file
126
docs/atlas-nextcloud-recovery-test.md
Normal file
@@ -0,0 +1,126 @@
|
||||
# Nextcloud manual backup/restore rehearsal — 2026-10-04
|
||||
|
||||
## Observed outcome
|
||||
|
||||
The operator authorized testing consistent database/files backup and recovery.
|
||||
This was executed directly, not added as a one-time playbook task or feature flag.
|
||||
No production database was replaced and no iCloud import was performed.
|
||||
|
||||
The existing ZFS scrub independently passed: completed at 05:02 CEST after
|
||||
2:02:04, 0 B repaired, zero errors, successful service exit and healthy pool.
|
||||
|
||||
## Backup artifact
|
||||
|
||||
Retained on Atlas:
|
||||
`/zpool/backup/nextcloud-rehearsal-20261004T092635Z` (676 MiB observed).
|
||||
Its host parent is restricted to admin, mode 0700; dump and manifest files were
|
||||
created with umask 077. It contains sensitive application configuration and
|
||||
database contents, not just test data. No plaintext secret was saved to Git.
|
||||
|
||||
- Complete application tree, including configuration, custom apps and themes;
|
||||
the overlaid data directory was copied separately.
|
||||
- Complete dedicated files tree.
|
||||
- PostgreSQL custom-format database dump, role definitions, image references,
|
||||
SHA-256 manifest and canary description.
|
||||
|
||||
Cron was stopped, maintenance enabled, and Nextcloud/ONLYOFFICE stopped while
|
||||
the database dump and application/files copies were taken. Rsync checksum and
|
||||
metadata comparisons passed while writers were stopped. Live services resumed
|
||||
with maintenance off, and the cron timer resumed. Role definitions were captured
|
||||
read-only immediately afterward when the isolated restore exposed the separate
|
||||
`oc_admin` database role. Its saved password was subsequently verified against
|
||||
the copied application configuration using SCRAM authentication. The recurring
|
||||
procedure should capture both database and role dumps during the same pause.
|
||||
|
||||
## Isolated restoration
|
||||
|
||||
- Fresh rootless PostgreSQL using the exact production image digest, not the
|
||||
live database volume. Restored roles first, then the database with owners/ACLs.
|
||||
- Copied application and files directories, using the matching Nextcloud image.
|
||||
- Fresh, empty isolated Redis; cache contents are not a recovery requirement.
|
||||
- One pod with `network=none`, no published ports and no live data bind mounts.
|
||||
Components communicate only through their shared loopback interface.
|
||||
- Only the restored configuration was adjusted for loopback database/cache,
|
||||
localhost URLs and disabled mail. Production configuration was unchanged.
|
||||
- No test cron or Office service was run. External connectivity and callbacks
|
||||
were impossible from this pod.
|
||||
|
||||
Checks passed:
|
||||
|
||||
1. Backup SHA-256 verification before restoration and again after cleanup.
|
||||
2. PostgreSQL role/database restore with failure-on-error enabled.
|
||||
3. Nextcloud 33.0.9 installed, maintenance off, no pending database upgrade.
|
||||
4. Restored admin/fabio/chiara accounts and Famiglia permission mask 15.
|
||||
5. The uniquely named Fabio canary existed both in files and the database index.
|
||||
6. Authenticated HTTP WebDAV retrieved that canary from the restored instance;
|
||||
its SHA-256 matched the original uploaded contents.
|
||||
7. The live canary remained unchanged and was then deleted through WebDAV.
|
||||
8. Live Nextcloud, ONLYOFFICE and cron timer active; Office connection check
|
||||
passed. Cloud/Git/Music/Syncthing public HTTPS returned 200 with valid TLS.
|
||||
|
||||
Initial fixture failures established two prerequisites: wait for PostgreSQL's
|
||||
final TCP listener, not the temporary initialization socket, and restore global
|
||||
roles in addition to the database dump. Persistent Redis settings also require
|
||||
a working isolated cache. Failed fixture pods were removed before retries.
|
||||
|
||||
After success, the final test pod and temporary restore directory were removed.
|
||||
No test network, live rollback or pool snapshot destruction was needed.
|
||||
|
||||
## Limits and remaining work
|
||||
|
||||
This validates manual recovery from the current local application/database/files
|
||||
copy, not a production-size recovery, RPO/RTO compliance, client resynchronization,
|
||||
Office editing-session recovery, or extraction from Borg/offline USB. ONLYOFFICE's
|
||||
own persistent service state was not part of this Nextcloud artifact.
|
||||
|
||||
The artifact has no dedicated automatic retention policy; do not call it the
|
||||
recurring Nextcloud backup solution. Integrate a coordinated dump/copy or snapshot
|
||||
procedure into recurring backups with locking, failure recovery, monitoring and
|
||||
retention. Independently validate new offsite and offline versions before import.
|
||||
Keep local/Vault recovery access independent of Nextcloud availability.
|
||||
|
||||
The procedure follows the required configuration/apps/files/themes/database scope
|
||||
and maintenance pause described in the [Nextcloud backup guide](https://docs.nextcloud.com/server/33/admin_manual/maintenance/backup.html)
|
||||
and tests restoration into a separate environment rather than applying the
|
||||
[restore procedure](https://docs.nextcloud.com/server/33/admin_manual/maintenance/restore.html)
|
||||
to production.
|
||||
|
||||
## Recurring integration and offsite recovery, later on 2026-10-04
|
||||
|
||||
The recurring preparation helper, system service, boot recovery and ordered
|
||||
Borg/USB dependency drop-ins were deployed from Ansible. The existing snapshot
|
||||
and backup jobs retain their lock, ownership, namespace, encryption and retention
|
||||
policies. Two local verified bundles are the declared staging retention; this
|
||||
supersedes the missing retention warning above for the managed bundle path only.
|
||||
The earlier manually named rehearsal artifact remains separate and untouched.
|
||||
|
||||
Preparation published `20261004T094945Z-3294831`, then starting the actual Borg job
|
||||
required a fresh preparation and published `20261004T095158Z-3303565`. Application
|
||||
availability resumed before immutable copy/hash processing completed. The Borg
|
||||
service successfully published `atlas-20261004T095255Z`, completed pruning and
|
||||
compaction, and removed its source snapshot after exit.
|
||||
|
||||
The latter consistent bundle was extracted from that encrypted Hetzner archive,
|
||||
not copied from the current local bundle. All SHA-256 checks passed. The extracted
|
||||
application/files and PostgreSQL role/database dumps were recovered into a fresh
|
||||
network-none pod with separate database/cache and matching image digests. It
|
||||
reported installed Nextcloud 33.0.9 without pending upgrade. The three accounts,
|
||||
Famiglia permissions and authenticated DAV PROPFIND for each account passed;
|
||||
PostgreSQL used the saved role password with SCRAM on its isolated TCP listener.
|
||||
The pod, extracted tree and its independent temporary Borg cache were removed.
|
||||
Production services and the pool remained healthy.
|
||||
|
||||
Failure validation used sandboxed helper mocks for maintenance/dump errors: the
|
||||
original failure code propagated, services/cron resumed, and state/partial files
|
||||
were removed. Separate transient systemd fixtures verified that a failed ordered
|
||||
requirement prevents its consumer from executing. These are fault-injection tests,
|
||||
not production failures or proof of a full host-crash recovery. A real boot with
|
||||
an interrupted preparation remains untested.
|
||||
|
||||
A third preparation published `20261004T100212Z-3366172`; exactly two managed
|
||||
versions remained, with the oldest version pruned only after publication. Source
|
||||
snapshots and persistent interruption markers were absent after success.
|
||||
|
||||
A new UUID-bound USB version and recovery of its consistent Nextcloud bundle
|
||||
remain to be verified after the operator connects/unlocks the configured disk.
|
||||
Do not mark USB recovery complete merely because the dependency was installed.
|
||||
@@ -17,7 +17,8 @@ were added. An actual repeat run returned `changed=0`, with no failures.
|
||||
10.2.1 and Team Folders 21.0.9 archives are pinned by version and SHA-256.
|
||||
- Dedicated ZFS namespace: `zpool/services/data/nextcloud`, with separate `app`,
|
||||
`files`, `database`, `cache` and `office` datasets. No writable SMB/Syncthing
|
||||
access to the Nextcloud-managed file namespace is provided.
|
||||
access to the Nextcloud-managed file namespace is provided. Existing Archive
|
||||
directories are exposed separately through local external storage (see below).
|
||||
- The `admin` Nextcloud account is an application administrator, distinct from
|
||||
the host account. `fabio` and `chiara` are standard users in `famiglia`, each
|
||||
with no initial quota. Team folder `Famiglia` has unlimited quota and group
|
||||
@@ -101,20 +102,32 @@ Dry-run skips initial downloads, image pulls and runtime account/app commands;
|
||||
it is not proof of an installed or healthy stack. The deployed repeat run is
|
||||
the current idempotence evidence.
|
||||
|
||||
## Manual recovery evidence, 2026-10-04
|
||||
|
||||
The first monthly scrub completed successfully and was verified from both the
|
||||
service result and pool scan (zero errors, 0 B repaired). A manual consistent
|
||||
application/files copy and PostgreSQL dump were restored into a network-isolated
|
||||
Nextcloud/PostgreSQL/Redis test pod. Account recovery, Famiglia permissions and
|
||||
authenticated DAV retrieval of a checksum-matched canary passed. Live services
|
||||
resumed normally; the test pod and restore workspace were removed.
|
||||
See `atlas-nextcloud-recovery-test.md` for scope, retained artifact and limitations.
|
||||
|
||||
## Gates before family data and full client acceptance
|
||||
|
||||
- Verify the first actual scrub and the outstanding protection checks.
|
||||
- The first actual scrub passed on 2026-10-04; preserve the existing protection checks.
|
||||
- Public TLS, redirects, web login and WebDAV passed. Complete calendar/contact
|
||||
synchronization and Office editing/saving from a desktop.
|
||||
- Test opening, editing and saving from the iPhone/iPad ONLYOFFICE app; mobile
|
||||
browser editing is not a requirement. No such client test is claimed yet.
|
||||
- Private-space isolation and cross-user shared writes/deletes passed the public
|
||||
smoke test above; complete normal client acceptance as well.
|
||||
- Integrate and test application-consistent database/files backups before import.
|
||||
- The manual rehearsal and recurring integration passed, including recovery from
|
||||
a new encrypted Borg archive. Complete recovery from a new USB version before import.
|
||||
The new datasets fall beneath existing recursive snapshot/backup scope, but
|
||||
that alone does not verify a new Borg/USB version or a consistent Nextcloud restore.
|
||||
that alone does not verify a new Borg/USB version or recovery through those versions.
|
||||
- For a consistent backup, coordinate pending Office saves, pause cron and writes,
|
||||
take a verified PostgreSQL dump and matching application/files snapshot, and
|
||||
take verified PostgreSQL database and role dumps plus a matching application/files
|
||||
snapshot or quiesced copy, and
|
||||
resume services promptly even on failure. Extend recurring backup procedures,
|
||||
not the steady-state playbook with one-time migration tasks. Restore into an
|
||||
isolated environment using matching image/app versions, config, files and DB.
|
||||
@@ -125,3 +138,92 @@ the current idempotence evidence.
|
||||
against an upgraded database; use matching tested backups for recovery.
|
||||
- Future Uranus migration and iCloud import are separate, explicitly authorized
|
||||
operations. No source data deletion or automatic cross-system cutover is provided.
|
||||
|
||||
## Recurring consistent bundles
|
||||
|
||||
`atlas-nextcloud-backup.service` is now an ordered requirement of both
|
||||
`atlas-borg-backup.service` and the operator-started `atlas-usb-backup.service`.
|
||||
No additional backup timer is needed: the existing Borg schedule prepares a fresh
|
||||
bundle before its pool snapshot, and a manual USB run does the same. Preparation
|
||||
failure blocks the dependent job rather than silently using an old dump.
|
||||
|
||||
The helper checks the mounted datasets and healthy active application state,
|
||||
serializes preparations and briefly pauses cron, Nextcloud and ONLYOFFICE. It
|
||||
captures database plus global roles and a recursive Nextcloud-only ZFS snapshot.
|
||||
Services resume before the longer immutable-file copy and checksum verification.
|
||||
Active editing sessions are interrupted; only committed Nextcloud state is covered.
|
||||
This does not claim preservation of unsaved ONLYOFFICE editing sessions.
|
||||
|
||||
Private bundles are published atomically under `/zpool/backup/nextcloud/versions`,
|
||||
with a relative `latest` link. `atlas_nextcloud_backup_keep: 2` retains two local
|
||||
verified versions, with hard links for unchanged files. Long-term Borg, USB and
|
||||
ZFS policies are unchanged. A trap and `ExecStopPost` restore availability and
|
||||
clean only this helper's named source snapshot/partial directory; root-private
|
||||
persistent state permits boot recovery through the enabled recovery unit. Both
|
||||
units have failure alerts and are included in the Atlas monitored failure units.
|
||||
No automatic rollback, import, or repair of user application data is performed.
|
||||
|
||||
Validation:
|
||||
|
||||
```bash
|
||||
ansible-playbook ansible/site.yml --limit atlas --tags nextcloud_backup,monitoring --check --diff
|
||||
sudo systemctl start atlas-nextcloud-backup.service
|
||||
sudo systemctl show atlas-nextcloud-backup.service -p Result -p ExecMainExitTimestamp
|
||||
```
|
||||
|
||||
The second command briefly interrupts the applications and is an explicit manual
|
||||
run of the recurring job, not a normal deployment side effect. The preparation
|
||||
unit is not enabled as a boot backup; only interrupted-job recovery is enabled.
|
||||
Do not stop a Borg/USB job, break its lock or unmount its source snapshot to run a test.
|
||||
|
||||
|
||||
## Existing Archive storage (no import or duplicate originals)
|
||||
|
||||
The Atlas declaration exposes only these existing directories to the rootless
|
||||
Nextcloud container, using shared SELinux `:z` labels:
|
||||
|
||||
| Nextcloud folder | Host directory | Access |
|
||||
| --- | --- | --- |
|
||||
| `Documenti` | `/zpool/archive/Documents` | Read/write |
|
||||
| `Foto iCloud` | `/zpool/archive/Pictures/iCloudPD` | Read-only |
|
||||
|
||||
Both system mounts are restricted to the Nextcloud user `fabio` only; `chiara`
|
||||
and the `famiglia` group have no access through these mounts.
|
||||
External re-sharing is disabled. The photos bind is also read-only at container
|
||||
level, independently of Nextcloud's mount option. iCloudPD remains the photo
|
||||
writer. Neither directory is copied into the internal data dataset, and the
|
||||
existing `Famiglia` team folder remains separate and untouched.
|
||||
|
||||
The Archive dataset enables persistent `acltype=posix` support; this does not
|
||||
change pool features or vdev layout. Scoped ACLs grant the actual rootless-mapped web UID access to existing files and
|
||||
inheritance on new directories/files. Document defaults retain host administrator
|
||||
access to files created through Nextcloud; ownership is not changed recursively.
|
||||
Symlinks are not followed when applying ACLs. Do not change Archive ownership or
|
||||
apply private `:Z` relabeling to these shared paths.
|
||||
|
||||
The user `atlas-nextcloud-external-scan.timer` discovers external changes for only
|
||||
these mounts and their explicitly allowed user: first after boot at 15 minutes, then one hour
|
||||
after the previous scan finishes. Nextcloud also checks for external changes on
|
||||
access. Indexing and previews are not duplicate originals; document versions,
|
||||
trash, ZFS snapshots and backups may retain additional data intentionally.
|
||||
Avoid simultaneously editing the same document through SMB and Nextcloud.
|
||||
|
||||
Archive originals retain their existing recursive ZFS/Borg/offline USB coverage.
|
||||
The Nextcloud-only recovery bundle does **not** include these external originals:
|
||||
a recovery must restore the corresponding Archive data as well as the application
|
||||
and database. This change does not migrate iCloud Drive or remove anything there.
|
||||
|
||||
|
||||
### Runtime validation, 2026-10-04
|
||||
|
||||
The mounts were applied on Atlas without importing originals. A disposable
|
||||
application-level document create/read test reached the original bind directory;
|
||||
the probe was deleted, including its trash entry. After the operator narrowed
|
||||
access to Fabio only, fresh Nextcloud application checks confirmed Fabio can read
|
||||
both mounts and create documents, while photo create/update/delete are denied.
|
||||
Chiara cannot access either mount; the separate `Famiglia` team folder remains
|
||||
available to both users. Container inspection independently confirmed the photo
|
||||
bind is read-only. Public Nextcloud HTTPS returned 200 and the pool was healthy.
|
||||
The targeted second Ansible run for mount applicability, options and discovery
|
||||
unit returned `changed=0`, with no failures. This is focused idempotency evidence,
|
||||
not a claim about a full Atlas playbook run.
|
||||
|
||||
Reference in New Issue
Block a user