Files
infra/docs/prometheus-backup.md
2026-09-30 21:21:48 +02:00

107 lines
6.3 KiB
Markdown

# Prometheus to Atlas backup pull
The playbook and both hosts have the dedicated identity, restricted SSH
access, helpers, and systemd units. A manual export, pull, and temporary
restore passed on 2026-09-30. Both timers are enabled; their first scheduled
runs are pending, so daily operation is not yet verified.
## Declared design
- Prometheus prepares a tar archive of Nginx Proxy Manager and Gitea data,
their managed Compose configuration, SSH/firewalld/WireGuard configuration,
and the Gitea SSH path. NPM access logs and regenerable Gitea logs, sessions,
temporary files, and indexers are excluded. The archive contains credentials,
certificates, and the WireGuard private key: protect both copies accordingly.
- The approved consistency mode stops the managed Compose stack for the local
tar creation at 02:00 Europe/Rome, then restarts it even if archiving fails.
A manual test outside that window requires separate approval.
- Prometheus publishes the archive with its checksum as a versioned, read-only
source under `/var/lib/prometheus-backup-export`. A locked service account
has no sudo or supplementary groups. Its only authorized SSH key is forced
through Rocky's `rrsync -ro`; root owns the key file and export directories,
so the account cannot add an unrestricted key or change prepared data.
- Atlas generates and retains the private Ed25519 identity under
`/etc/atlas-prometheus-pull`. Its pinned Prometheus host key came through
the controller's already strict SSH trust; the observed fingerprint was
`SHA256:rfedk7DHI9mLB3UHk/4F3HHlSIiswtCAFsAXvfh6iXk` on 2026-09-30.
Atlas pulls only the prepared `current/` version, verifies SHA-256, tar
readability, metadata, and source freshness, then publishes atomically
below `/zpool/backup/hosts/prometheus/snapshots`. Long-term retention runs
only after publication. A local `rrsync` fixture verified the in-tree
`current` symlink. A live Atlas-to-Prometheus SSH test verified that the
account could list only the prepared versions directory,
cannot obtain a shell, and cannot write to the export. The key is restricted
to `/var/lib/prometheus-backup-export/versions`, not the account's `.ssh`.
- Approved source preparation is 02:00 Europe/Rome, pull 03:00, three source
versions, and 30 daily/8 weekly/12 monthly Atlas versions. The source
timer is non-persistent to avoid an unexpected outage after a missed run.
Atlas rejects a prepared source older than 24 hours.
- The Atlas pull joins the existing health monitor's timer/failure checks
only when enabled. Its failure hook uses 45Drives Alerts; email delivery
is not claimed. A failed source preparation should produce a stale-source
pull failure, not a silently successful reuse of an old archive.
## Activation and verification
1. The user confirmed downtime/consistency mode, schedule, retention, and
targeted configuration scope. Review the tar path list and exclusions
against the actual containers.
2. The identity and units are deployed. Re-run the targeted
check, confirm the Atlas public key remains only the restricted Prometheus
account's key, and verify `sshd -T -C user=prometheus-backup,...` plus
read-only SSH denial tests after any SSH configuration change.
3. During an agreed window, start the Prometheus export service manually.
Confirm Compose is healthy afterward, inspect the archive without exposing
file contents, and verify the checksum/metadata.
4. Start the Atlas pull service manually. Confirm the SSH host pin, source
freshness, checksum, tar listing, published `latest`, retention behavior,
clean temporary directories, and healthy pool.
5. Independently restore the selected archive to an empty staging directory
(never `/`) and compare the SQLite databases, Git repositories, NPM data,
Compose file, permissions, and representative files. Test application
startup only in an isolated environment or an approved restore window.
6. The two timers were enabled after the manual test. Verify their calendars
and the next actual run. A successful manual test is not proof of scheduled
operation.
Narrow static validation:
```bash
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
ansible-playbook ansible/site.yml --syntax-check
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
ansible-playbook ansible/site.yml --limit prometheus,atlas \
--tags prometheus_backup --check --diff
```
Do not run the export service as part of a routine playbook deployment. The
service restart and any restore/cutover require separate operator decisions.
On 2026-09-30 the initial targeted `--check --diff` run ended `changed=0`
with gates false. After enabling **implementation only**, a targeted real run
installed the identities and units; both timers were confirmed `disabled` and
`inactive`, the Compose stack stayed active, and the new account was locked
with no supplementary groups. No application was stopped.
The rendered shell helpers passed `bash -n` and ShellCheck; the retention
helper passed an isolated 400-version fixture. These static/isolated checks
were followed by live SSH, export, pull, and temporary restore checks.
Read-only preflight on 2026-09-30 found the Compose service active, all
declared source paths present, both timers inactive, and no prepared versions.
The source filesystem had about 6.0 GB free. Of the 2.1 GB NPM data tree,
2.1 GB was excluded access logs, so the expected archive is much smaller than
the raw tree size; capacity still needs verification after actual exports.
The manual export produced a 285,777,920-byte tar (273 MiB allocated at the
source), and Prometheus retained about 5.8 GB free. NPM and Gitea restarted;
both containers were running and their local HTTP endpoints returned 200.
Atlas pulled the same version, verified SHA-256, published `latest`, and kept
the pool healthy. A full extract to `/var/tmp` yielded 4,747 files; both
SQLite databases passed `PRAGMA integrity_check`, and one restored Gitea Git
repository passed `git fsck`. The temporary restore directory was removed.
This did not test application startup on an isolated host.
After these checks, Ansible enabled the Prometheus 02:00 Europe/Rome export
timer and Atlas 03:00 Europe/Rome pull timer. The next scheduled occurrences
were displayed for 2026-10-01. Atlas' health monitor now includes the pull
timer. Check both actual service results after the first scheduled run before
claiming unattended operation.