fix: preserve concurrent external-update slots with re-apply retry - #192
Merged
Merged
Conversation
Concurrent external updates from multiple upstream artifacts to one primary repo lost a slot to a non-fast-forward push: the losing run failed with "! [rejected] main -> main (fetch first)" and dropped its external slot. The generated external-update workflow's concurrency group only serializes runs; a queued run still holds a stale-parent checkout and is rejected on push. The plain push had no recovery, and a git-level rebase cannot resolve the textual conflict two updates produce in the same re-marshaled YAML region. Make cascade external update self-heal at the data-structure level: on a rejected push, fetch the remote tip, reset onto it, and re-apply this update's external slot before retrying. Different artifacts write different slot keys, so re-applying merges cleanly and no slot is lost. Signed-off-by: Joshua Temple <joshua.temple@stablekernel.com>
joshua-temple
force-pushed
the
fix/external-update-push-retry
branch
from
June 16, 2026 21:30
3554db4 to
0ddd5c9
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Concurrent
cascade external updateruns (multiple upstream artifacts notifying the same primary repo) lost an external slot to a non-fast-forward push. The losing run failed with! [rejected] main -> main (fetch first)and dropped its slot from the manifest, a silent state loss.The generated external-update workflow's
concurrency:group only serializes runs; a queued run still holds a stale-parent checkout and is rejected on push. The push had no recovery.Root cause and why a simple retry is not enough
The verb pushed via the plain
git.CommitAndPush(no fetch/rebase/retry). Swapping to the existinggit.CommitAndPushWithRetrywas the obvious fix, but verification proved it insufficient: its retry doesgit pull --rebase, and two external updates both mutate the same region of the re-marshaled YAML (state.<env>.external), so the rebase hits a textual conflict (UU cascade.yaml), wedges the repo mid-rebase, and after the retries returns an error with no state written.The merge has to happen at the data-structure level, not the text level: different artifacts write different keys under
state.<env>.external, so re-applying the mutation onto the fresh remote manifest merges cleanly.Fix
internal/external/command.go:runUpdatenow validates once up front, then drives the manifest mutation throughcommitWithApplicationRetry. On a rejected push it fetches the remote tip,reset --hard origin/<branch>, re-reads the manifest fresh, re-applies this update's slot, recommits, and retries (up to 5, short backoff). The mutation closure re-parses the file each attempt so it merges onto rebased remote content. Anothing to commitoutcome short-circuits to success (idempotent re-runs).internal/git/git.go: addedCurrentBranch()helper.CommitAndPushandCommitAndPushWithRetryare left untouched for their other callers (promote/hotfix).internal/generate/external.go: updated thewriteConcurrencydoc comment to reflect that serialization alone is not sufficient and the verb self-heals via rebase-and-retry.Tests
internal/git/retry_test.go: provesCommitAndPushWithRetryrecovers from a plain non-fast-forward (non-conflicting case) with both slots preserved.internal/external/retry_test.go: drives the realrunUpdatethrough the concurrent race (competitor landscdkfrom a stale checkout, this run landslambda) and asserts both slots survive in origin HEAD. Failing-first verified: on the pre-fix code the push is rejected non-fast-forward. Also covers the sequential stale-checkout case and idempotent re-apply.e2e/harness/multi_repo_scenario_test.go:TestMultiRepoRunner_ConcurrentExternalUpdatesPreserveBothSlotsdrives two real external-update dispatches against one primary under act/gitea and asserts both external slots land in the committed manifest.E2E limitation: gitea has no live
workflow_dispatchAPI, so two truly simultaneous pushes cannot be staged deterministically. The scenario uses sequential dispatches with two distinct external artifacts, which exercises the same fetch+reset+re-apply path a stale checkout hits. The unit/integration tests cover the conflicting concurrent write directly.Verification
No generated-workflow drift (the only generator change is a doc comment; generate golden tests stay green).
Residual note
A same-deploy-key concurrent write (two notifications for the same deploy) resolves last-writer-wins after the reset and re-read, which is correct: a duplicate notification should land the latest SHA/version in that one slot.