Skip to content

fix(backup): drop local copies that are already off-site, and say why one stays - #209

Merged
nechodom merged 4 commits into
mainfrom
claude/backups-server-retention-7a2828
Oct 5, 2026
Merged

nechodom merged 4 commits into
mainfrom
claude/backups-server-retention-7a2828

Conversation

@nechodom

@nechodom nechodom commented Oct 5, 2026 •

Copy link
Copy Markdown
Owner

What

Backups marked off-site ✓ stayed on the server, and the panel never said why. The drop-local step now works on a site's existing copies too. Each row on the Backups tab now says why its local copy is still there. The change also fixes the problems an adversarial review found around the drop (S3-only nodes, restore races, unreachable targets, old nodes).

Why

An operator asked why backups that "say they're on FTP" were still on disk. Production check on hos-1:

  • agent.toml has no [backup] section, so "Delete the local copy once it is verified off-site" is off (the default).
  • Even when on, it only ever acted on the backup that had just finished, so turning it on never freed anything already on disk.
  • The Settings hint pointed at "Copy existing backups off-site". That sweep skipped every backup already marked off-site, so it freed nothing either.
  • The hosting.backup.drop_local ok audit rows came from manual Off-site & drop. That action re-uploaded ~267 MiB per backup before deleting anything.

Operator-facing change?

  • Operator-facing change:
    • Settings → Backups & trash → Local backup copies:
      • New Keep the newest local copies anyway ([backup] keep_local_latest, per site, default 0).
      • With dropping on, each backup now also drops the site's older local copies that are on record off-site, past the newest N. Turning the option on frees a site's existing copies at its next archive backup.
    • Backups tab:
      • Each off-site row gets a small line saying why the local copy is still there: option off, one of the newest N, not confirmed off-site (reason stored on the row), or drops at the next backup.
      • A card note appears when the owning node has the option off.
      • S3-only backups now show uploaded instead of "—".
    • Off-site & drop (per row and bulk): an already-off-site backup is re-checked where it is and dropped without uploading it again. It falls back to a push only if the re-check fails.
    • Copy off-site & delete local (Settings, this node's sites):
      • It now also covers copies that were already off-site, and honours the keep count.
      • Its job has a heartbeat, so a long sweep is no longer reaped at 1 h.
    • Both jobs now finish red when a copy was kept, instead of green next to a backup still on disk.
    • off-site ✓ (FTP verified) now also requires the database dump to size-check. That is what the drop step always demanded.
    • Rollback snapshots (pre-change/pre-push) are no longer pushed to FTP. S3 and the drop already skipped them.
    • Error text said [remote_backup]; the section is [backup_remote].

Data-safety details (reviewers: please read)

The drop rules are unchanged: a local copy goes only after every file is independently re-verified off-site. The DB paths are cleared first, then the files are deleted. New in this PR:

  • s3_ok_targets (migration 076) records the S3 targets a push of the run reached in full. Each entry is the target name plus a fingerprint of endpoint, bucket and age recipient. Editing a target in place (for example a new age recipient after the old identity was lost) therefore makes older evidence inert.
    • The S3 push never wrote any off-site state. So S3-only nodes could never drop an older copy, and keep_local_latest > 0 disabled dropping there entirely.
    • An age-encrypted object is still trusted by existence only when a complete push is on record.
    • A failed re-push removes that target from the list, since the key may now hold a truncated object.
  • Busy-archive registry: an archive that a restore, a backup or a push is using is never dropped or pruned.
    • Right before deleting, the drop atomically reserves the archive. The reservation is refused while anyone but the caller holds the archive, and while it stands no restore or push can start on that file. The restore then points to the Off-site panel.
    • The drop also re-reads the row and refuses if its paths or its S3 evidence changed since the check.
    • Restore also checks for the .sql before extracting. If it vanishes mid-way, the restore now fails. Before, it reported success with old files over the current database.
  • Unreachable target: the re-check is three-state (confirmed / not confirmed / unreachable), and the sweep stops on unreachable. Before, a blackholed FTP cost 60 s × files × older rows per backup in the serial scheduled sweep. verify_remote now has connect-timeout = 15.
  • Files listed but not on disk: nothing is deleted and the row keeps its paths, since an unmounted backup volume looks exactly the same. The row gets a note ("local files not on disk") and is counted as missing, not kept.
  • FTP "no such file/dir" (curl 9/19/21/78) counts as an answer about that file, so the sweep moves on to the next run instead of treating the server as down.
  • Unreadable keep_local_latest (a quoted string, a float, a negative number) disables dropping with a warning, including in the estate sweep. It no longer silently means 0, the most aggressive value.
  • list_needing_offsite excludes runs with a completed S3 push, so the backfill stops re-uploading them.
  • A run that is on record off-site but whose local archive is gone is no longer stamped "off-site FAILED".
  • A stored "could not confirm" note is cleared once the copy is kept on purpose.
  • The audit row and WARN for a kept copy are written only when its reason changes, not after every backup.
  • The .manifest.json is deleted along with its archive on drop and on prune. It is write-only (nothing reads it), and it used to be orphaned.
  • Mixed-version clusters: an older node rejects the whole [backup] save over the unknown key. The panel resends without keep_local_latest only when that changes nothing on that node (keep 0, or dropping switched off). Otherwise it reports the node as needing an upgrade.

Two rounds of multi-lens adversarial review ran on this (data safety, behaviour, ops/perf, web/IA, regressions):

  • Round 1: 25 confirmed findings, fixed in the first commit.
  • Round 2, on those fixes: 11 confirmed findings, fixed in f2fdccd.

The residuals are listed under Anti-scope.

Test plan

  • cargo test --workspace passes, including after merging current main.
  • cargo clippy --workspace --all-targets passes with no warnings.
  • cargo fmt --all -- --check passes.
  • Manually tested on a real Debian node. Suggested check on hos-1:
    1. Tick the option (optionally keep 1).
    2. Run Back up now on www.centrumsrdicko.cz.
    3. The older off-site copies drop: the rows lose their local path, and hosting.backup.drop_local ok audit rows appear.
  • Added / updated tests:
    • Selection (local_copies_to_drop, including S3 evidence and keep N).
    • Per-row reasons, the FTP verdict, the manifest path.
    • Policy parsing, including a Settings write followed by a read-back, and the bad values.
    • The busy-archive registry.
    • Service-level: kept with an unreachable FTP, reported once; busy / missing / dump-only.
    • The on-demand path; backup_list reasons.
    • State: write-then-read-back for local_note and s3_ok_targets, and the sweep query.
    • Web: the old-node field stripping.

Anti-scope

  • No FTP-side retention: FTP copies still accumulate. S3 has its own retention_daily.
  • The estate sweep still covers only the node it runs on. The text now says so.
  • Rollback snapshots taken before migration 072 cannot be told apart from normal backups (no_offsite = 0). If they were FTP-pushed, they can be dropped once they are verified off-site.
  • A chunked Download of an older archive can still race a drop.

🤖 Generated with Claude Code

mkn and others added 4 commits October 5, 2026 14:09
… one stays

An operator saw backups marked "off-site ✓" still sitting on the server and
read it as a broken delete. It was not: "Delete the local copy once it is
verified off-site" is off by default, it only ever acted on the backup that
had just finished, and nothing on the Backups tab said either.

- After each backup, with the option on, the site's OLDER local copies that
  are on record off-site are re-verified and dropped too, past a new
  `[backup] keep_local_latest` (keep the N newest locally, default 0). Turning
  the option on now frees existing copies at the site's next archive backup.
- "Off-site & drop" (per row and bulk) re-checks an already-off-site backup
  where it is and drops it without uploading it again; it only falls back to
  a push when the re-check fails. The estate-wide "Copy off-site & delete
  local" now also covers copies that were already off-site (it used to skip
  them, while the Settings hint pointed operators at it), honours the keep
  count, and heartbeats its job.
- The Backups tab says, per row, why a local copy is still there (option off,
  one of the newest N kept, not confirmed off-site, or dropping at the next
  backup), from the owning node's settings; a refused drop stores its reason
  on the row (`local_note`, migration 076). Jobs go red when a copy was kept.
- S3 pushes are recorded on the run (`s3_ok_targets`, migration 076): the S3
  push never wrote any off-site state, so S3-only nodes could never drop an
  older copy, and a keep count disabled dropping there entirely. A failed
  re-push removes the target again (age objects are trusted by existence only
  with a complete push on record).
- Safety: an archive a restore, backup or push is using is never dropped or
  pruned; a restore whose dump vanishes mid-way now fails instead of
  reporting success with the files of one moment and the current database;
  an unreachable target stops the sweep instead of waiting out every timeout
  (curl connect-timeout 15s); an unreadable keep count disables dropping
  rather than meaning "keep 0"; an archive already gone from disk is reported
  missing and its row set straight, not counted as kept forever.
- "off-site ✓" (FTP verified) now requires the database dump to size-check
  too, as the drop step always did. Rollback snapshots are no longer pushed
  to FTP (S3 and the drop already skipped them). The manifest is removed with
  its archive on drop and prune instead of being orphaned.
- An older node that does not know `keep_local_latest` still gets the on/off
  switch when leaving the count out changes nothing there.
- Error text named `[remote_backup]`; the section is `[backup_remote]`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Second adversarial review of the local-copy drop:

- "Off-site & drop" pushed and then deleted through the gate without asking
  the in-use registry, so it could delete an archive a restore was reading.
  The gate now reserves the archive atomically right before deleting (refused
  while anyone but the caller holds it), and a restore or push cannot start
  on an archive being deleted (the restore says to use the Off-site panel).
  The early busy check alone missed restores that started during the
  minutes-long off-site re-check.
- Before deleting, the gate re-reads the row and refuses if its paths or S3
  evidence moved since it was checked; a sweep reads each run fresh instead
  of from a list that can be minutes old (a failed re-push may have taken an
  age target off the run's evidence meanwhile).
- S3 evidence is recorded as the target name plus a fingerprint of endpoint,
  bucket and age recipient. Editing a target in place to a new recipient
  (old identity lost) no longer lets the old, now undecryptable objects
  "confirm" a copy by existence.
- Files listed on a row but not on disk are no longer cleared from the row:
  an unmounted backup volume looks exactly like that. The row keeps its
  paths, gets a note, and shows "local files not on disk".
- FTP "no such file/dir" answers (curl 9/19/21/78) count as an answer about
  that file, so the sweep moves on instead of treating the server as down.
- An unreadable keep count also stops the estate "Copy off-site & delete
  local" sweep (it read it as 0) and the job says why.
- push_run_offsite no longer stamps "off-site FAILED" on a run whose local
  archive is gone but which is on record off-site.
- Backfill no longer re-uploads runs with a completed S3 push.
- A stored "could not confirm" note is cleared once the copy is kept on
  purpose (inside the keep window or without an off-site record).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@nechodom
nechodom merged commit e8ae558 into main Oct 5, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant