mirror of
https://github.com/fscotto/infra.git
synced 2026-10-04 22:09:50 +00:00
Integrate consistent Nextcloud backups and recovery
This commit is contained in:
@@ -2,8 +2,10 @@
|
||||
|
||||
Status: the empty stack was deployed on 2026-10-03, explicitly before the first
|
||||
scrub. The operator configured DNS/NPM and authorized public cutover; public TLS,
|
||||
DAV and cross-user file checks passed. Client editing/sync acceptance and consistent
|
||||
backup/restore validation remain open before family data. iCloud import remains a
|
||||
DAV and cross-user file checks passed. The first scrub and a manual consistent
|
||||
backup/isolated restore passed on 2026-10-04. Client editing/sync acceptance and
|
||||
USB recovery validation remain open before family data. Recurring backup integration
|
||||
and recovery from a new Borg archive passed. iCloud import remains a
|
||||
separate operation. See `docs/atlas-nextcloud.md` for observed runtime state.
|
||||
|
||||
## Confirmed requirements
|
||||
@@ -95,7 +97,8 @@ acceptance tests; the app is not treated as proof of server-side compatibility.
|
||||
|
||||
1. Validate desktop Office editing/saving, calendar/contact synchronization and
|
||||
mobile ONLYOFFICE app integration; public empty-stack cutover is verified.
|
||||
2. Complete protection gates and application-consistent backup/recovery tests.
|
||||
2. Complete recovery from a new offline USB version; recurring preparation and
|
||||
encrypted Borg recovery have passed.
|
||||
3. Plan the deferred iCloud migration when explicitly requested.
|
||||
|
||||
## Primary references
|
||||
|
||||
126
docs/atlas-nextcloud-recovery-test.md
Normal file
126
docs/atlas-nextcloud-recovery-test.md
Normal file
@@ -0,0 +1,126 @@
|
||||
# Nextcloud manual backup/restore rehearsal — 2026-10-04
|
||||
|
||||
## Observed outcome
|
||||
|
||||
The operator authorized testing consistent database/files backup and recovery.
|
||||
This was executed directly, not added as a one-time playbook task or feature flag.
|
||||
No production database was replaced and no iCloud import was performed.
|
||||
|
||||
The existing ZFS scrub independently passed: completed at 05:02 CEST after
|
||||
2:02:04, 0 B repaired, zero errors, successful service exit and healthy pool.
|
||||
|
||||
## Backup artifact
|
||||
|
||||
Retained on Atlas:
|
||||
`/zpool/backup/nextcloud-rehearsal-20261004T092635Z` (676 MiB observed).
|
||||
Its host parent is restricted to admin, mode 0700; dump and manifest files were
|
||||
created with umask 077. It contains sensitive application configuration and
|
||||
database contents, not just test data. No plaintext secret was saved to Git.
|
||||
|
||||
- Complete application tree, including configuration, custom apps and themes;
|
||||
the overlaid data directory was copied separately.
|
||||
- Complete dedicated files tree.
|
||||
- PostgreSQL custom-format database dump, role definitions, image references,
|
||||
SHA-256 manifest and canary description.
|
||||
|
||||
Cron was stopped, maintenance enabled, and Nextcloud/ONLYOFFICE stopped while
|
||||
the database dump and application/files copies were taken. Rsync checksum and
|
||||
metadata comparisons passed while writers were stopped. Live services resumed
|
||||
with maintenance off, and the cron timer resumed. Role definitions were captured
|
||||
read-only immediately afterward when the isolated restore exposed the separate
|
||||
`oc_admin` database role. Its saved password was subsequently verified against
|
||||
the copied application configuration using SCRAM authentication. The recurring
|
||||
procedure should capture both database and role dumps during the same pause.
|
||||
|
||||
## Isolated restoration
|
||||
|
||||
- Fresh rootless PostgreSQL using the exact production image digest, not the
|
||||
live database volume. Restored roles first, then the database with owners/ACLs.
|
||||
- Copied application and files directories, using the matching Nextcloud image.
|
||||
- Fresh, empty isolated Redis; cache contents are not a recovery requirement.
|
||||
- One pod with `network=none`, no published ports and no live data bind mounts.
|
||||
Components communicate only through their shared loopback interface.
|
||||
- Only the restored configuration was adjusted for loopback database/cache,
|
||||
localhost URLs and disabled mail. Production configuration was unchanged.
|
||||
- No test cron or Office service was run. External connectivity and callbacks
|
||||
were impossible from this pod.
|
||||
|
||||
Checks passed:
|
||||
|
||||
1. Backup SHA-256 verification before restoration and again after cleanup.
|
||||
2. PostgreSQL role/database restore with failure-on-error enabled.
|
||||
3. Nextcloud 33.0.9 installed, maintenance off, no pending database upgrade.
|
||||
4. Restored admin/fabio/chiara accounts and Famiglia permission mask 15.
|
||||
5. The uniquely named Fabio canary existed both in files and the database index.
|
||||
6. Authenticated HTTP WebDAV retrieved that canary from the restored instance;
|
||||
its SHA-256 matched the original uploaded contents.
|
||||
7. The live canary remained unchanged and was then deleted through WebDAV.
|
||||
8. Live Nextcloud, ONLYOFFICE and cron timer active; Office connection check
|
||||
passed. Cloud/Git/Music/Syncthing public HTTPS returned 200 with valid TLS.
|
||||
|
||||
Initial fixture failures established two prerequisites: wait for PostgreSQL's
|
||||
final TCP listener, not the temporary initialization socket, and restore global
|
||||
roles in addition to the database dump. Persistent Redis settings also require
|
||||
a working isolated cache. Failed fixture pods were removed before retries.
|
||||
|
||||
After success, the final test pod and temporary restore directory were removed.
|
||||
No test network, live rollback or pool snapshot destruction was needed.
|
||||
|
||||
## Limits and remaining work
|
||||
|
||||
This validates manual recovery from the current local application/database/files
|
||||
copy, not a production-size recovery, RPO/RTO compliance, client resynchronization,
|
||||
Office editing-session recovery, or extraction from Borg/offline USB. ONLYOFFICE's
|
||||
own persistent service state was not part of this Nextcloud artifact.
|
||||
|
||||
The artifact has no dedicated automatic retention policy; do not call it the
|
||||
recurring Nextcloud backup solution. Integrate a coordinated dump/copy or snapshot
|
||||
procedure into recurring backups with locking, failure recovery, monitoring and
|
||||
retention. Independently validate new offsite and offline versions before import.
|
||||
Keep local/Vault recovery access independent of Nextcloud availability.
|
||||
|
||||
The procedure follows the required configuration/apps/files/themes/database scope
|
||||
and maintenance pause described in the [Nextcloud backup guide](https://docs.nextcloud.com/server/33/admin_manual/maintenance/backup.html)
|
||||
and tests restoration into a separate environment rather than applying the
|
||||
[restore procedure](https://docs.nextcloud.com/server/33/admin_manual/maintenance/restore.html)
|
||||
to production.
|
||||
|
||||
## Recurring integration and offsite recovery, later on 2026-10-04
|
||||
|
||||
The recurring preparation helper, system service, boot recovery and ordered
|
||||
Borg/USB dependency drop-ins were deployed from Ansible. The existing snapshot
|
||||
and backup jobs retain their lock, ownership, namespace, encryption and retention
|
||||
policies. Two local verified bundles are the declared staging retention; this
|
||||
supersedes the missing retention warning above for the managed bundle path only.
|
||||
The earlier manually named rehearsal artifact remains separate and untouched.
|
||||
|
||||
Preparation published `20261004T094945Z-3294831`, then starting the actual Borg job
|
||||
required a fresh preparation and published `20261004T095158Z-3303565`. Application
|
||||
availability resumed before immutable copy/hash processing completed. The Borg
|
||||
service successfully published `atlas-20261004T095255Z`, completed pruning and
|
||||
compaction, and removed its source snapshot after exit.
|
||||
|
||||
The latter consistent bundle was extracted from that encrypted Hetzner archive,
|
||||
not copied from the current local bundle. All SHA-256 checks passed. The extracted
|
||||
application/files and PostgreSQL role/database dumps were recovered into a fresh
|
||||
network-none pod with separate database/cache and matching image digests. It
|
||||
reported installed Nextcloud 33.0.9 without pending upgrade. The three accounts,
|
||||
Famiglia permissions and authenticated DAV PROPFIND for each account passed;
|
||||
PostgreSQL used the saved role password with SCRAM on its isolated TCP listener.
|
||||
The pod, extracted tree and its independent temporary Borg cache were removed.
|
||||
Production services and the pool remained healthy.
|
||||
|
||||
Failure validation used sandboxed helper mocks for maintenance/dump errors: the
|
||||
original failure code propagated, services/cron resumed, and state/partial files
|
||||
were removed. Separate transient systemd fixtures verified that a failed ordered
|
||||
requirement prevents its consumer from executing. These are fault-injection tests,
|
||||
not production failures or proof of a full host-crash recovery. A real boot with
|
||||
an interrupted preparation remains untested.
|
||||
|
||||
A third preparation published `20261004T100212Z-3366172`; exactly two managed
|
||||
versions remained, with the oldest version pruned only after publication. Source
|
||||
snapshots and persistent interruption markers were absent after success.
|
||||
|
||||
A new UUID-bound USB version and recovery of its consistent Nextcloud bundle
|
||||
remain to be verified after the operator connects/unlocks the configured disk.
|
||||
Do not mark USB recovery complete merely because the dependency was installed.
|
||||
@@ -17,7 +17,8 @@ were added. An actual repeat run returned `changed=0`, with no failures.
|
||||
10.2.1 and Team Folders 21.0.9 archives are pinned by version and SHA-256.
|
||||
- Dedicated ZFS namespace: `zpool/services/data/nextcloud`, with separate `app`,
|
||||
`files`, `database`, `cache` and `office` datasets. No writable SMB/Syncthing
|
||||
access to the Nextcloud-managed file namespace is provided.
|
||||
access to the Nextcloud-managed file namespace is provided. Existing Archive
|
||||
directories are exposed separately through local external storage (see below).
|
||||
- The `admin` Nextcloud account is an application administrator, distinct from
|
||||
the host account. `fabio` and `chiara` are standard users in `famiglia`, each
|
||||
with no initial quota. Team folder `Famiglia` has unlimited quota and group
|
||||
@@ -101,20 +102,32 @@ Dry-run skips initial downloads, image pulls and runtime account/app commands;
|
||||
it is not proof of an installed or healthy stack. The deployed repeat run is
|
||||
the current idempotence evidence.
|
||||
|
||||
## Manual recovery evidence, 2026-10-04
|
||||
|
||||
The first monthly scrub completed successfully and was verified from both the
|
||||
service result and pool scan (zero errors, 0 B repaired). A manual consistent
|
||||
application/files copy and PostgreSQL dump were restored into a network-isolated
|
||||
Nextcloud/PostgreSQL/Redis test pod. Account recovery, Famiglia permissions and
|
||||
authenticated DAV retrieval of a checksum-matched canary passed. Live services
|
||||
resumed normally; the test pod and restore workspace were removed.
|
||||
See `atlas-nextcloud-recovery-test.md` for scope, retained artifact and limitations.
|
||||
|
||||
## Gates before family data and full client acceptance
|
||||
|
||||
- Verify the first actual scrub and the outstanding protection checks.
|
||||
- The first actual scrub passed on 2026-10-04; preserve the existing protection checks.
|
||||
- Public TLS, redirects, web login and WebDAV passed. Complete calendar/contact
|
||||
synchronization and Office editing/saving from a desktop.
|
||||
- Test opening, editing and saving from the iPhone/iPad ONLYOFFICE app; mobile
|
||||
browser editing is not a requirement. No such client test is claimed yet.
|
||||
- Private-space isolation and cross-user shared writes/deletes passed the public
|
||||
smoke test above; complete normal client acceptance as well.
|
||||
- Integrate and test application-consistent database/files backups before import.
|
||||
- The manual rehearsal and recurring integration passed, including recovery from
|
||||
a new encrypted Borg archive. Complete recovery from a new USB version before import.
|
||||
The new datasets fall beneath existing recursive snapshot/backup scope, but
|
||||
that alone does not verify a new Borg/USB version or a consistent Nextcloud restore.
|
||||
that alone does not verify a new Borg/USB version or recovery through those versions.
|
||||
- For a consistent backup, coordinate pending Office saves, pause cron and writes,
|
||||
take a verified PostgreSQL dump and matching application/files snapshot, and
|
||||
take verified PostgreSQL database and role dumps plus a matching application/files
|
||||
snapshot or quiesced copy, and
|
||||
resume services promptly even on failure. Extend recurring backup procedures,
|
||||
not the steady-state playbook with one-time migration tasks. Restore into an
|
||||
isolated environment using matching image/app versions, config, files and DB.
|
||||
@@ -125,3 +138,92 @@ the current idempotence evidence.
|
||||
against an upgraded database; use matching tested backups for recovery.
|
||||
- Future Uranus migration and iCloud import are separate, explicitly authorized
|
||||
operations. No source data deletion or automatic cross-system cutover is provided.
|
||||
|
||||
## Recurring consistent bundles
|
||||
|
||||
`atlas-nextcloud-backup.service` is now an ordered requirement of both
|
||||
`atlas-borg-backup.service` and the operator-started `atlas-usb-backup.service`.
|
||||
No additional backup timer is needed: the existing Borg schedule prepares a fresh
|
||||
bundle before its pool snapshot, and a manual USB run does the same. Preparation
|
||||
failure blocks the dependent job rather than silently using an old dump.
|
||||
|
||||
The helper checks the mounted datasets and healthy active application state,
|
||||
serializes preparations and briefly pauses cron, Nextcloud and ONLYOFFICE. It
|
||||
captures database plus global roles and a recursive Nextcloud-only ZFS snapshot.
|
||||
Services resume before the longer immutable-file copy and checksum verification.
|
||||
Active editing sessions are interrupted; only committed Nextcloud state is covered.
|
||||
This does not claim preservation of unsaved ONLYOFFICE editing sessions.
|
||||
|
||||
Private bundles are published atomically under `/zpool/backup/nextcloud/versions`,
|
||||
with a relative `latest` link. `atlas_nextcloud_backup_keep: 2` retains two local
|
||||
verified versions, with hard links for unchanged files. Long-term Borg, USB and
|
||||
ZFS policies are unchanged. A trap and `ExecStopPost` restore availability and
|
||||
clean only this helper's named source snapshot/partial directory; root-private
|
||||
persistent state permits boot recovery through the enabled recovery unit. Both
|
||||
units have failure alerts and are included in the Atlas monitored failure units.
|
||||
No automatic rollback, import, or repair of user application data is performed.
|
||||
|
||||
Validation:
|
||||
|
||||
```bash
|
||||
ansible-playbook ansible/site.yml --limit atlas --tags nextcloud_backup,monitoring --check --diff
|
||||
sudo systemctl start atlas-nextcloud-backup.service
|
||||
sudo systemctl show atlas-nextcloud-backup.service -p Result -p ExecMainExitTimestamp
|
||||
```
|
||||
|
||||
The second command briefly interrupts the applications and is an explicit manual
|
||||
run of the recurring job, not a normal deployment side effect. The preparation
|
||||
unit is not enabled as a boot backup; only interrupted-job recovery is enabled.
|
||||
Do not stop a Borg/USB job, break its lock or unmount its source snapshot to run a test.
|
||||
|
||||
|
||||
## Existing Archive storage (no import or duplicate originals)
|
||||
|
||||
The Atlas declaration exposes only these existing directories to the rootless
|
||||
Nextcloud container, using shared SELinux `:z` labels:
|
||||
|
||||
| Nextcloud folder | Host directory | Access |
|
||||
| --- | --- | --- |
|
||||
| `Documenti` | `/zpool/archive/Documents` | Read/write |
|
||||
| `Foto iCloud` | `/zpool/archive/Pictures/iCloudPD` | Read-only |
|
||||
|
||||
Both system mounts are restricted to the Nextcloud user `fabio` only; `chiara`
|
||||
and the `famiglia` group have no access through these mounts.
|
||||
External re-sharing is disabled. The photos bind is also read-only at container
|
||||
level, independently of Nextcloud's mount option. iCloudPD remains the photo
|
||||
writer. Neither directory is copied into the internal data dataset, and the
|
||||
existing `Famiglia` team folder remains separate and untouched.
|
||||
|
||||
The Archive dataset enables persistent `acltype=posix` support; this does not
|
||||
change pool features or vdev layout. Scoped ACLs grant the actual rootless-mapped web UID access to existing files and
|
||||
inheritance on new directories/files. Document defaults retain host administrator
|
||||
access to files created through Nextcloud; ownership is not changed recursively.
|
||||
Symlinks are not followed when applying ACLs. Do not change Archive ownership or
|
||||
apply private `:Z` relabeling to these shared paths.
|
||||
|
||||
The user `atlas-nextcloud-external-scan.timer` discovers external changes for only
|
||||
these mounts and their explicitly allowed user: first after boot at 15 minutes, then one hour
|
||||
after the previous scan finishes. Nextcloud also checks for external changes on
|
||||
access. Indexing and previews are not duplicate originals; document versions,
|
||||
trash, ZFS snapshots and backups may retain additional data intentionally.
|
||||
Avoid simultaneously editing the same document through SMB and Nextcloud.
|
||||
|
||||
Archive originals retain their existing recursive ZFS/Borg/offline USB coverage.
|
||||
The Nextcloud-only recovery bundle does **not** include these external originals:
|
||||
a recovery must restore the corresponding Archive data as well as the application
|
||||
and database. This change does not migrate iCloud Drive or remove anything there.
|
||||
|
||||
|
||||
### Runtime validation, 2026-10-04
|
||||
|
||||
The mounts were applied on Atlas without importing originals. A disposable
|
||||
application-level document create/read test reached the original bind directory;
|
||||
the probe was deleted, including its trash entry. After the operator narrowed
|
||||
access to Fabio only, fresh Nextcloud application checks confirmed Fabio can read
|
||||
both mounts and create documents, while photo create/update/delete are denied.
|
||||
Chiara cannot access either mount; the separate `Famiglia` team folder remains
|
||||
available to both users. Container inspection independently confirmed the photo
|
||||
bind is read-only. Public Nextcloud HTTPS returned 200 and the pool was healthy.
|
||||
The targeted second Ansible run for mount applicability, options and discovery
|
||||
unit returned `changed=0`, with no failures. This is focused idempotency evidence,
|
||||
not a claim about a full Atlas playbook run.
|
||||
|
||||
Reference in New Issue
Block a user