Repository navigation
fix(backup): drop local copies that are already off-site, and say why one stays - #209
Merged
Merged
Conversation
… one stays An operator saw backups marked "off-site ✓" still sitting on the server and read it as a broken delete. It was not: "Delete the local copy once it is verified off-site" is off by default, it only ever acted on the backup that had just finished, and nothing on the Backups tab said either. - After each backup, with the option on, the site's OLDER local copies that are on record off-site are re-verified and dropped too, past a new `[backup] keep_local_latest` (keep the N newest locally, default 0). Turning the option on now frees existing copies at the site's next archive backup. - "Off-site & drop" (per row and bulk) re-checks an already-off-site backup where it is and drops it without uploading it again; it only falls back to a push when the re-check fails. The estate-wide "Copy off-site & delete local" now also covers copies that were already off-site (it used to skip them, while the Settings hint pointed operators at it), honours the keep count, and heartbeats its job. - The Backups tab says, per row, why a local copy is still there (option off, one of the newest N kept, not confirmed off-site, or dropping at the next backup), from the owning node's settings; a refused drop stores its reason on the row (`local_note`, migration 076). Jobs go red when a copy was kept. - S3 pushes are recorded on the run (`s3_ok_targets`, migration 076): the S3 push never wrote any off-site state, so S3-only nodes could never drop an older copy, and a keep count disabled dropping there entirely. A failed re-push removes the target again (age objects are trusted by existence only with a complete push on record). - Safety: an archive a restore, backup or push is using is never dropped or pruned; a restore whose dump vanishes mid-way now fails instead of reporting success with the files of one moment and the current database; an unreachable target stops the sweep instead of waiting out every timeout (curl connect-timeout 15s); an unreadable keep count disables dropping rather than meaning "keep 0"; an archive already gone from disk is reported missing and its row set straight, not counted as kept forever. - "off-site ✓" (FTP verified) now requires the database dump to size-check too, as the drop step always did. Rollback snapshots are no longer pushed to FTP (S3 and the drop already skipped them). The manifest is removed with its archive on drop and prune instead of being orphaned. - An older node that does not know `keep_local_latest` still gets the on/off switch when leaving the count out changes nothing there. - Error text named `[remote_backup]`; the section is `[backup_remote]`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-retention-7a2828
Second adversarial review of the local-copy drop: - "Off-site & drop" pushed and then deleted through the gate without asking the in-use registry, so it could delete an archive a restore was reading. The gate now reserves the archive atomically right before deleting (refused while anyone but the caller holds it), and a restore or push cannot start on an archive being deleted (the restore says to use the Off-site panel). The early busy check alone missed restores that started during the minutes-long off-site re-check. - Before deleting, the gate re-reads the row and refuses if its paths or S3 evidence moved since it was checked; a sweep reads each run fresh instead of from a list that can be minutes old (a failed re-push may have taken an age target off the run's evidence meanwhile). - S3 evidence is recorded as the target name plus a fingerprint of endpoint, bucket and age recipient. Editing a target in place to a new recipient (old identity lost) no longer lets the old, now undecryptable objects "confirm" a copy by existence. - Files listed on a row but not on disk are no longer cleared from the row: an unmounted backup volume looks exactly like that. The row keeps its paths, gets a note, and shows "local files not on disk". - FTP "no such file/dir" answers (curl 9/19/21/78) count as an answer about that file, so the sweep moves on instead of treating the server as down. - An unreadable keep count also stops the estate "Copy off-site & delete local" sweep (it read it as 0) and the job says why. - push_run_offsite no longer stamps "off-site FAILED" on a run whose local archive is gone but which is on record off-site. - Backfill no longer re-uploads runs with a completed S3 push. - A stored "could not confirm" note is cleared once the copy is kept on purpose (inside the keep window or without an off-site record). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-retention-7a2828
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Backups marked off-site ✓ stayed on the server, and the panel never said why. The drop-local step now works on a site's existing copies too. Each row on the Backups tab now says why its local copy is still there. The change also fixes the problems an adversarial review found around the drop (S3-only nodes, restore races, unreachable targets, old nodes).
Why
An operator asked why backups that "say they're on FTP" were still on disk. Production check on hos-1:
agent.tomlhas no[backup]section, so "Delete the local copy once it is verified off-site" is off (the default).hosting.backup.drop_local okaudit rows came from manual Off-site & drop. That action re-uploaded ~267 MiB per backup before deleting anything.Operator-facing change?
[backup] keep_local_latest, per site, default 0).[remote_backup]; the section is[backup_remote].Data-safety details (reviewers: please read)
The drop rules are unchanged: a local copy goes only after every file is independently re-verified off-site. The DB paths are cleared first, then the files are deleted. New in this PR:
s3_ok_targets(migration 076) records the S3 targets a push of the run reached in full. Each entry is the target name plus a fingerprint of endpoint, bucket and age recipient. Editing a target in place (for example a new age recipient after the old identity was lost) therefore makes older evidence inert.keep_local_latest > 0disabled dropping there entirely..sqlbefore extracting. If it vanishes mid-way, the restore now fails. Before, it reported success with old files over the current database.verify_remotenow hasconnect-timeout = 15.keep_local_latest(a quoted string, a float, a negative number) disables dropping with a warning, including in the estate sweep. It no longer silently means 0, the most aggressive value.list_needing_offsiteexcludes runs with a completed S3 push, so the backfill stops re-uploading them..manifest.jsonis deleted along with its archive on drop and on prune. It is write-only (nothing reads it), and it used to be orphaned.[backup]save over the unknown key. The panel resends withoutkeep_local_latestonly when that changes nothing on that node (keep 0, or dropping switched off). Otherwise it reports the node as needing an upgrade.Two rounds of multi-lens adversarial review ran on this (data safety, behaviour, ops/perf, web/IA, regressions):
f2fdccd.The residuals are listed under Anti-scope.
Test plan
cargo test --workspacepasses, including after merging currentmain.cargo clippy --workspace --all-targetspasses with no warnings.cargo fmt --all -- --checkpasses.hosting.backup.drop_local okaudit rows appear.local_copies_to_drop, including S3 evidence and keep N).backup_listreasons.local_noteands3_ok_targets, and the sweep query.Anti-scope
retention_daily.no_offsite = 0). If they were FTP-pushed, they can be dropped once they are verified off-site.🤖 Generated with Claude Code