mirror of
https://github.com/fscotto/infra.git
synced 2026-10-03 21:39:50 +00:00
107 lines
6.3 KiB
Markdown
107 lines
6.3 KiB
Markdown
# Prometheus to Atlas backup pull
|
|
|
|
The playbook and both hosts have the dedicated identity, restricted SSH
|
|
access, helpers, and systemd units. A manual export, pull, and temporary
|
|
restore passed on 2026-09-30. Both timers are enabled; their first scheduled
|
|
runs are pending, so daily operation is not yet verified.
|
|
|
|
## Declared design
|
|
|
|
- Prometheus prepares a tar archive of Nginx Proxy Manager and Gitea data,
|
|
their managed Compose configuration, SSH/firewalld/WireGuard configuration,
|
|
and the Gitea SSH path. NPM access logs and regenerable Gitea logs, sessions,
|
|
temporary files, and indexers are excluded. The archive contains credentials,
|
|
certificates, and the WireGuard private key: protect both copies accordingly.
|
|
- The approved consistency mode stops the managed Compose stack for the local
|
|
tar creation at 02:00 Europe/Rome, then restarts it even if archiving fails.
|
|
A manual test outside that window requires separate approval.
|
|
- Prometheus publishes the archive with its checksum as a versioned, read-only
|
|
source under `/var/lib/prometheus-backup-export`. A locked service account
|
|
has no sudo or supplementary groups. Its only authorized SSH key is forced
|
|
through Rocky's `rrsync -ro`; root owns the key file and export directories,
|
|
so the account cannot add an unrestricted key or change prepared data.
|
|
- Atlas generates and retains the private Ed25519 identity under
|
|
`/etc/atlas-prometheus-pull`. Its pinned Prometheus host key came through
|
|
the controller's already strict SSH trust; the observed fingerprint was
|
|
`SHA256:rfedk7DHI9mLB3UHk/4F3HHlSIiswtCAFsAXvfh6iXk` on 2026-09-30.
|
|
Atlas pulls only the prepared `current/` version, verifies SHA-256, tar
|
|
readability, metadata, and source freshness, then publishes atomically
|
|
below `/zpool/backup/hosts/prometheus/snapshots`. Long-term retention runs
|
|
only after publication. A local `rrsync` fixture verified the in-tree
|
|
`current` symlink. A live Atlas-to-Prometheus SSH test verified that the
|
|
account could list only the prepared versions directory,
|
|
cannot obtain a shell, and cannot write to the export. The key is restricted
|
|
to `/var/lib/prometheus-backup-export/versions`, not the account's `.ssh`.
|
|
- Approved source preparation is 02:00 Europe/Rome, pull 03:00, three source
|
|
versions, and 30 daily/8 weekly/12 monthly Atlas versions. The source
|
|
timer is non-persistent to avoid an unexpected outage after a missed run.
|
|
Atlas rejects a prepared source older than 24 hours.
|
|
- The Atlas pull joins the existing health monitor's timer/failure checks
|
|
only when enabled. Its failure hook uses 45Drives Alerts; email delivery
|
|
is not claimed. A failed source preparation should produce a stale-source
|
|
pull failure, not a silently successful reuse of an old archive.
|
|
|
|
## Activation and verification
|
|
|
|
1. The user confirmed downtime/consistency mode, schedule, retention, and
|
|
targeted configuration scope. Review the tar path list and exclusions
|
|
against the actual containers.
|
|
2. The identity and units are deployed. Re-run the targeted
|
|
check, confirm the Atlas public key remains only the restricted Prometheus
|
|
account's key, and verify `sshd -T -C user=prometheus-backup,...` plus
|
|
read-only SSH denial tests after any SSH configuration change.
|
|
3. During an agreed window, start the Prometheus export service manually.
|
|
Confirm Compose is healthy afterward, inspect the archive without exposing
|
|
file contents, and verify the checksum/metadata.
|
|
4. Start the Atlas pull service manually. Confirm the SSH host pin, source
|
|
freshness, checksum, tar listing, published `latest`, retention behavior,
|
|
clean temporary directories, and healthy pool.
|
|
5. Independently restore the selected archive to an empty staging directory
|
|
(never `/`) and compare the SQLite databases, Git repositories, NPM data,
|
|
Compose file, permissions, and representative files. Test application
|
|
startup only in an isolated environment or an approved restore window.
|
|
6. The two timers were enabled after the manual test. Verify their calendars
|
|
and the next actual run. A successful manual test is not proof of scheduled
|
|
operation.
|
|
|
|
Narrow static validation:
|
|
|
|
```bash
|
|
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
|
|
ansible-playbook ansible/site.yml --syntax-check
|
|
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
|
|
ansible-playbook ansible/site.yml --limit prometheus,atlas \
|
|
--tags prometheus_backup --check --diff
|
|
```
|
|
|
|
Do not run the export service as part of a routine playbook deployment. The
|
|
service restart and any restore/cutover require separate operator decisions.
|
|
|
|
On 2026-09-30 the initial targeted `--check --diff` run ended `changed=0`
|
|
with gates false. After enabling **implementation only**, a targeted real run
|
|
installed the identities and units; both timers were confirmed `disabled` and
|
|
`inactive`, the Compose stack stayed active, and the new account was locked
|
|
with no supplementary groups. No application was stopped.
|
|
The rendered shell helpers passed `bash -n` and ShellCheck; the retention
|
|
helper passed an isolated 400-version fixture. These static/isolated checks
|
|
were followed by live SSH, export, pull, and temporary restore checks.
|
|
Read-only preflight on 2026-09-30 found the Compose service active, all
|
|
declared source paths present, both timers inactive, and no prepared versions.
|
|
The source filesystem had about 6.0 GB free. Of the 2.1 GB NPM data tree,
|
|
2.1 GB was excluded access logs, so the expected archive is much smaller than
|
|
the raw tree size; capacity still needs verification after actual exports.
|
|
The manual export produced a 285,777,920-byte tar (273 MiB allocated at the
|
|
source), and Prometheus retained about 5.8 GB free. NPM and Gitea restarted;
|
|
both containers were running and their local HTTP endpoints returned 200.
|
|
Atlas pulled the same version, verified SHA-256, published `latest`, and kept
|
|
the pool healthy. A full extract to `/var/tmp` yielded 4,747 files; both
|
|
SQLite databases passed `PRAGMA integrity_check`, and one restored Gitea Git
|
|
repository passed `git fsck`. The temporary restore directory was removed.
|
|
This did not test application startup on an isolated host.
|
|
|
|
After these checks, Ansible enabled the Prometheus 02:00 Europe/Rome export
|
|
timer and Atlas 03:00 Europe/Rome pull timer. The next scheduled occurrences
|
|
were displayed for 2026-10-01. Atlas' health monitor now includes the pull
|
|
timer. Check both actual service results after the first scheduled run before
|
|
claiming unattended operation.
|