Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
35 commits
Select commit Hold shift + click to select a range
aaf64c3
docs(cargo-anvil): add implementation plan 0003 for benchmark regress…
martin-kolinek Aug 3, 2026
3ac9a23
docs(cargo-anvil): design the benchmark regression detection subsystem
martin-kolinek Aug 3, 2026
222e3c9
docs(cargo-anvil): specify benchmark history persistence in the backe…
martin-kolinek Aug 3, 2026
36a74ab
docs(cargo-anvil): describe desired state, not design evolution, in b…
martin-kolinek Aug 3, 2026
b5f208b
docs(cargo-anvil): reflect the scheduled-benchmarks group in diagrams…
martin-kolinek Aug 3, 2026
96e7135
docs(cargo-anvil): confine design to the design docs; make 0003 a pur…
martin-kolinek Aug 4, 2026
eb0dd16
docs(cargo-anvil): make artifacts the sole store; add local behavior;…
martin-kolinek Aug 4, 2026
6871c23
docs(cargo-anvil): add scheduled-benchmarks to the local recipe surface
martin-kolinek Aug 4, 2026
f18f4e7
feat(cargo-anvil): detect benchmark regressions on the scheduled tier
martin-kolinek Aug 4, 2026
efedf1e
merge: integrate origin/main into PR branch (auto-resolved conflicts)
martin-kolinek Aug 4, 2026
2d73e3a
fix(cargo-anvil): regenerate the README for the new catalog row
martin-kolinek Aug 4, 2026
f61b558
merge: integrate origin/main into PR branch (auto-resolved conflicts)
martin-kolinek Aug 5, 2026
2a883f4
refactor(cargo-anvil): defer benchmark failure reporting to the sched…
martin-kolinek Aug 7, 2026
d870aca
merge: integrate origin/main into PR branch (auto-resolved conflicts)
martin-kolinek Aug 7, 2026
2237b07
merge: integrate origin/main into PR branch (auto-resolved conflicts)
martin-kolinek Aug 20, 2026
b7322c0
fix(cargo-anvil): address review of the benchmark regression capability
martin-kolinek Aug 21, 2026
e43fd78
fix(cargo-anvil): keep the ADO publish inside the wrapper seam; cover…
martin-kolinek Aug 21, 2026
418dac9
merge: integrate origin/main and wire benchmarks into the failure pub…
martin-kolinek Aug 21, 2026
751d1f7
fix(cargo-anvil): sync the dogfooded scheduled workflows with the cat…
martin-kolinek Aug 21, 2026
edd5fae
refactor(cargo-anvil): keep per-group steps out of the tier templates
martin-kolinek Aug 21, 2026
5225d8f
Address review round 3: fail-closed restore, template fragments, per-…
martin-kolinek Aug 24, 2026
884152e
Merge remote-tracking branch 'origin/main' into anvil-benchmarks
martin-kolinek Aug 25, 2026
d9b9519
Merge remote-tracking branch 'origin/main' into anvil-benchmarks
martin-kolinek Aug 25, 2026
e397a46
fix(cargo-anvil): scope ADO restore temp paths and fail closed on sto…
martin-kolinek Aug 25, 2026
aa856be
Merge remote-tracking branch 'origin/main' into anvil-benchmarks
martin-kolinek Aug 27, 2026
a3b4973
docs(cargo-anvil): record the ADO artifact-retention asymmetry
martin-kolinek Aug 27, 2026
4837202
refactor(cargo-anvil): make the history round-trip an input of the sh…
martin-kolinek Aug 31, 2026
5deed42
Merge remote-tracking branch 'origin/main' into anvil-benchmarks
martin-kolinek Aug 31, 2026
1e97e48
Merge remote-tracking branch 'origin/main' into anvil-benchmarks
martin-kolinek Aug 31, 2026
7870c57
chore(cargo-anvil): regenerate after merging main
martin-kolinek Aug 31, 2026
14921d9
Merge remote-tracking branch 'origin/main' into anvil-benchmarks
martin-kolinek Aug 31, 2026
75974da
fix(cargo-anvil): guard the empty-artifact restore copy and read CI m…
martin-kolinek Aug 31, 2026
afe9a9a
fix(cargo-anvil): treat a blank CI marker as unset
martin-kolinek Aug 31, 2026
0c8e47c
test(cargo-anvil): refresh snapshots for the blank-marker gate
martin-kolinek Aug 31, 2026
cbd9b9f
fix(cargo-anvil): tolerate a nested artifact layout when restoring hi…
martin-kolinek Aug 31, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 16 additions & 8 deletions .anvil.lock
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
version = 1
tool = "anvil"
tool_version = "0.5.0"
catalog_checksum = "sha256:1388eee0dd1074adb96e6a944c65f2542ec4d6da90d17ad401ee453270359d3b"
catalog_checksum = "sha256:1a7aa05ccee55e5bad46ea953d460a3c4cad4e2737ccec514bfdfa4736558dc9"

[[file]]
path = ".anvil/container/Dockerfile.dockerignore"
Expand All @@ -17,7 +17,7 @@ checksum = "sha256:9940d1947482150ac08fcb9b4150da99f5ae60642f4caeea137577ce0e709

[[file]]
path = ".github/actions/anvil-run-group/action.yml"
checksum = "sha256:ff8def6c0786b6e146c4b633dfe38cb9b8ede398345516cca32bbcd5586087af"
checksum = "sha256:23a323e87f5ac0bb72a4e64e33692a71acf827956356e74698ac1b97a5e0da0b"

[[file]]
path = ".github/actions/anvil-setup/action.yml"
Expand All @@ -37,11 +37,11 @@ checksum = "sha256:0c2530d9a38e6a74e0a7fd4f999b4a1790f97de30b58b68c6c2344600da19

[[file]]
path = ".github/workflows/anvil-scheduled-impl.yml"
checksum = "sha256:ac70061acf594c8c212c45ed97c3b653e7b8de68f4e1dcc9695a2628e4e2596d"
checksum = "sha256:1c5250ead07c7915c978b52e425cc950e2ebe151a30cfc9a1398f354d1209ee9"

[[file]]
path = ".github/workflows/anvil-scheduled.yml"
checksum = "sha256:d4d3bd645a5586e9a1cc3a5fc27e38c93e59b1e29f83ede0ccd610273eec00f4"
checksum = "sha256:bc8120b035051db9e3db875495d6570238824339b9694234e388e759c0a84ff0"

[[file]]
path = "justfiles/anvil/checks/aprz.just"
Expand All @@ -51,6 +51,10 @@ checksum = "sha256:0f9f3dd3c8a2f034ebf4f2ac0344cba3db79612ff2fc2ea037bfabbaacb1b
path = "justfiles/anvil/checks/audit.just"
checksum = "sha256:54abf96a320bb4b35a3c0ddf2f30b0f4a30e0673e482ca3a71242fa383536415"

[[file]]
path = "justfiles/anvil/checks/bench-history.just"
checksum = "sha256:d24d92356347c3215b9db87eaddac821e689d71708761624ad61e698e19d15dd"

[[file]]
path = "justfiles/anvil/checks/bench.just"
checksum = "sha256:50f04b4ea6c99df8ad7434d320f34dccdb6db6a77b83e3090e35de0ca3a15f83"
Expand Down Expand Up @@ -191,6 +195,10 @@ checksum = "sha256:ea96d29e261b454a585c0ba3dc7954a35d0c726e9e94f6bb7c82f15261531
path = "justfiles/anvil/groups/scheduled-advisories.just"
checksum = "sha256:4f9940bb54fd7cd1d622f3207f7c43f38f95232a370f7537b15546271d88805d"

[[file]]
path = "justfiles/anvil/groups/scheduled-benchmarks.just"
checksum = "sha256:38aa0554499568b97be95dd8c73227f5a097e9e4a1c0fa144af8544c22d3be12"

[[file]]
path = "justfiles/anvil/groups/scheduled-exhaustive.just"
checksum = "sha256:0b9023c614ae400c30f7131fc939b318bf6ee2ba086b18d45491fa30691694e7"
Expand All @@ -213,19 +221,19 @@ checksum = "sha256:5714a138135154b91b234f6bad06ec3ade0e5780f761d631c1eca92b6010e

[[file]]
path = "justfiles/anvil/mod.just"
checksum = "sha256:2f7f0187f8c45716a1bed85f0c45ffbc71cdf592ed6414c9bc4cd72cf716a3a5"
checksum = "sha256:3cace5e08fd88200b1feac2345ec697d63993819dc46ee2dfdde0dad29cfa055"

[[file]]
path = "justfiles/anvil/tiers.just"
checksum = "sha256:00453a12cbb34811ee6a2c083dade5f6198575e3b0610f49e4743366326cdd18"
checksum = "sha256:a63a2cd4b5a63353c08a7d98e8b23683c87ef7ad60a23c33e4dd789204b0480d"

[[file]]
path = "justfiles/anvil/tools.just"
checksum = "sha256:a1e44ca16f172b487afa3997f102512733d3b65a4418cf894cbd749a3abc17dc"
checksum = "sha256:d7dc849363748f40df2768f13b364a3237767eda84525a1c901d4caddbacec3c"

[[file]]
path = "justfiles/anvil/versions.just"
checksum = "sha256:983e6732348188becb8f22a4af5237a13ac2579d566a9c4ff4bab6e724336077"
checksum = "sha256:69d843dd7fdfee808d74d46485ab3b290a839418db390233eeb89ef8893e9f5d"

[[region]]
host = ".anvil/container/Dockerfile"
Expand Down
155 changes: 155 additions & 0 deletions .github/actions/anvil-run-group/action.yml
Original file line number Diff line number Diff line change
Expand Up @@ -34,6 +34,29 @@ inputs:
Clean runs only supersede prior failures.
default: "false"
required: false
bench_history:
description: >-
Round-trip a cargo-bench-history store around the group run, so the
regression analysis has cross-run history to compare against. The
caller must grant actions: read and check out full history
(fetch-depth: 0), neither of which an action can request for itself.
default: "false"
required: false
bench_artifact:
description: >-
Artifact name carrying this leg's history. Must be unique per matrix
leg: the history is partitioned per machine, and merging two runners'
samples into one series destroys the comparison. Supplied by the
caller, which is the only place the matrix value is in scope.
default: ""
required: false
bench_machine_key:
description: >-
Overrides cargo-bench-history's hardware fingerprint with a stable
pool label, for runner pools heterogeneous enough to fragment a series
into partitions too sparse to analyze.
default: ""
required: false
runs:
using: composite
steps:
Expand All @@ -43,6 +66,110 @@ runs:
group: ${{ inputs.group }}
free-disk-space: ${{ inputs.free-disk-space }}

- name: Restore benchmark history
if: inputs.bench_history == 'true'
shell: bash
env:
GH_TOKEN: ${{ github.token }}
ARTIFACT: ${{ inputs.bench_artifact }}
DEFAULT_BRANCH: ${{ github.event.repository.default_branch }}
# The caller's workflow *name*, not a hardcoded filename: the root
# scheduled workflow is an owned, renameable file, and a rename must
# not silently reset the series.
WORKFLOW: ${{ github.workflow }}
REPO: ${{ github.repository }}
WINDOW: "30"
run: |
set -euo pipefail

if [ -z "$ARTIFACT" ]; then
echo "::error::bench_history is enabled but bench_artifact is empty;" \
"refusing to continue, since an unnamed store cannot be restored" \
"or published and every run would report a false clean."
exit 1
fi

# Staged first. The store path is created only once the restore has
# reached a known state, so an operational failure leaves no store
# at all and a publisher that runs unconditionally has nothing to
# upload over the accumulated chain.
staging="$(mktemp -d)"

complete_restore() {
mkdir -p target/anvil/bench-history
# `gh run download` with a single `--name` extracts into --dir
# directly, but with several it nests under one directory per
# artifact. Lift a nested layout if we ever see one: getting this
# wrong loses the history silently and every run reports a clean
# cold start, which is the one failure this feature must not have.
src="$staging"
if [ -d "$staging/$ARTIFACT" ]; then
src="$staging/$ARTIFACT"
fi
if [ -n "$(ls -A "$src" 2>/dev/null)" ]; then
cp -R "$src/." target/anvil/bench-history/
fi
echo "ANVIL_BENCH_RESTORE=$1" >> "$GITHUB_ENV"
# The recipe refuses to write anywhere but here, so an override
# cannot silently detach it from the store that is published.
echo "ANVIL_BENCH_WIRED_STORE=target/anvil/bench-history" >> "$GITHUB_ENV"
}

# Assigned rather than consumed directly by `for`: `set -e` ignores
# the exit status of a command substitution used as a word list, so
# a failed listing would yield an empty list and fall through to a
# cold start -- publishing a truncated store over the chain and
# reporting green for want of the history needed to report red.
if ! run_ids="$(gh run list --workflow "$WORKFLOW" \
--branch "$DEFAULT_BRANCH" --limit "$WINDOW" \
--json databaseId --jq '.[].databaseId')"; then
echo "::error::could not list $WORKFLOW runs on $DEFAULT_BRANCH;" \
"refusing to continue, since treating this as a cold start" \
"would publish a truncated store over the existing chain."
exit 1
fi

# Walk back from the newest run and take the first that carries this
# leg's artifact. Restoring from the latest *successful* run would
# drop every sample collected while the pipeline was red from a
# regression -- precisely the window that matters.
#
# Absence and failure are kept distinct. A run is only a candidate
# once the artifacts API confirms the artifact exists and has not
# expired; a download that then fails is an operational error
# (token, API, corrupt payload) and fails the job rather than being
# silently downgraded to a cold start.
for run_id in $run_ids; do
artifact_id=$(gh api --paginate \
"repos/$REPO/actions/runs/$run_id/artifacts" \
--jq ".artifacts[] | select(.name == \"$ARTIFACT\" and .expired == false) | .id" \
| head -n1)
[ -n "$artifact_id" ] || continue

if ! gh run download "$run_id" --name "$ARTIFACT" --dir "$staging"; then
echo "::error::found $ARTIFACT in run $run_id but could not download it;" \
"failing rather than continuing with an empty history, which would" \
"publish a truncated store over the existing chain."
exit 1
fi
echo "restored benchmark history from run $run_id"
complete_restore restored
exit 0
done

# No run in the window carried the artifact: a genuine cold start
# (first run, or the chain lapsed), which is a valid empty store.
# Surfaced on the summary rather than only in this log -- "history
# quietly restarted" must not look like "no regressions".
complete_restore cold-start
echo "no $ARTIFACT artifact in the last $WINDOW scheduled runs; starting a new history"
{
printf '### Benchmark history: cold start\n\n'
printf 'No `%s` artifact was found in the last %s `%s` runs on `%s`, ' \
"$ARTIFACT" "$WINDOW" "$WORKFLOW" "$DEFAULT_BRANCH"
printf 'so this run starts a new series. Trend detection needs several runs of history.\n'
} >> "$GITHUB_STEP_SUMMARY"

- name: Run Anvil group
id: run
if: steps.setup.outcome == 'success'
Expand All @@ -53,6 +180,7 @@ runs:
# target/anvil/impact cache (read via `_anvil-impact-include`), not
# threaded --package strings. This action only fixes the mode.
ANVIL_IMPACT: ${{ inputs.impact_mode }}
ANVIL_BENCH_MACHINE_KEY: ${{ inputs.bench_machine_key }}
# Some checks (e.g. cargo-aprz) call GitHub's API. The built-in token
# gives them the authenticated quota without adding group knowledge
# to this action.
Expand All @@ -75,6 +203,33 @@ runs:

# Reporting is supplemental: run after success or failure, but never let
# an API outage determine the authoritative workflow-job result.
- name: Save benchmark history
# always(): the run's own samples belong in the history even when the
# analysis flagged a regression and failed the group.
#
# Guarded on the restore having reached a known state: if the restore
# failed operationally the store is not a continuation of the chain,
# and publishing it would overwrite good history with a truncated
# snapshot.
if: always() && inputs.bench_history == 'true' && env.ANVIL_BENCH_RESTORE != ''
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: ${{ inputs.bench_artifact }}
path: target/anvil/bench-history
# Comfortably longer than the scheduled cadence, so a paused or
# infrequent schedule does not break the chain.
retention-days: 90
Comment on lines +214 to +221

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Copilot speaking]

Align the durable-store claim with implemented support

The pull request description says the CI-artifact rolling window has an opt-in durable Azure Blob option for longer history. The reviewed GitHub and Azure DevOps workflows only restore and republish target/anvil/bench-history through CI-native artifacts, while the benchmark design says durable cargo-bench-history backends such as Azure Blob are outside cargo-anvil's supported scope.

Reproducible reasoning: The GitHub action uploads the local directory as an Actions artifact. Azure DevOps restores and publishes the same directory as a pipeline artifact. No reviewed surface selects an Azure Blob backend or supplies its endpoint, credentials, restore, or publication behavior. That implementation matches the design's explicit exclusion but not the pull request description's substantive capability claim.

Consequence: Reviewers and adopters can plan around a supported long-retention option that the generated workflows cannot configure, then lose history when CI artifact retention expires.

Recommended action: Reconcile the pull request description, benchmark design, and generated workflows without assuming which side is authoritative. Either describe CI-native artifacts as the only supported store in this change, or add and validate the claimed durable backend with explicit configuration, credential handling, documentation, and equivalent behavior on both workflow backends.

References:

  • Pull request 68 description
  • crates/cargo-anvil/docs/design/benchmarks.md, sections 4 and 8
  • crates/cargo-anvil/docs/design/ado.md, Benchmark regression detection

Impacted locations:

  • .github/actions/anvil-run-group/action.yml:214-221
  • crates/cargo-anvil/docs/design/benchmarks.md:88-111
  • crates/cargo-anvil/docs/design/benchmarks.md:188-193
  • crates/cargo-anvil/docs/design/ado.md:940-963
  • crates/cargo-anvil/templates/justfiles/anvil/checks/bench-history.just
  • crates/cargo-anvil/tests/snapshots/snapshots__ado_backend.snap:617-630
  • crates/cargo-anvil/tests/snapshots/snapshots__ado_backend.snap:732-909
  • crates/cargo-anvil/templates/ado/scheduled-stages.yml:127-140

if-no-files-found: ignore

- name: Publish benchmark findings
if: always() && inputs.bench_history == 'true'
shell: bash
run: |
set -euo pipefail
if [ -f target/anvil/bench/findings.md ]; then
cat target/anvil/bench/findings.md >> "$GITHUB_STEP_SUMMARY"
fi

- name: Publish supplemental Anvil commit status
if: always() && inputs.publish_commit_statuses == 'true' && github.event_name == 'pull_request' && github.event.pull_request.head.repo.full_name == github.repository
continue-on-error: true
Expand Down
38 changes: 38 additions & 0 deletions .github/workflows/anvil-scheduled-impl.yml
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,14 @@ on:
description: Runner label for aarch64 Windows jobs.
type: string
default: windows-11-arm
bench_machine_key:
description: |
Machine key the benchmark history is partitioned by. Leave empty to
use cargo-bench-history's hardware fingerprint. Set a stable pool
label when the runner pool is heterogeneous enough to fragment a
series into partitions too sparse to analyze.
type: string
default: ""
Comment thread
martin-kolinek marked this conversation as resolved.
secrets:
CODECOV_TOKEN:
description: |
Expand Down Expand Up @@ -141,13 +149,43 @@ jobs:
with:
group: scheduled-exhaustive

scheduled-benchmarks:
strategy:
fail-fast: false
matrix:
os: [linux, windows]
runs-on: ${{ matrix.os == 'linux' && inputs.linux_runner || inputs.windows_runner }}
permissions:
contents: read
# Restoring the history walks the Actions runs/artifacts API. An action
# cannot request permissions, so this has to be granted here.
actions: read
Comment on lines +159 to +162

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Copilot speaking]

Align the permission summary with benchmark restoration

The reusable workflow says that after resetting ordinary jobs to contents: read, only the publisher's issue permission is restored. The new scheduled-benchmarks job also restores actions: read so the shared action can list and download prior benchmark-history artifacts.

Reproducible reasoning: The file-level comment describes the permission model for the whole reusable workflow. The benchmark job now adds actions: read, while publish-failure adds issues: write; the nearby benchmark comment explains why an action cannot request the former for itself. The high-level statement is therefore incomplete even though the job-level grants are explicit.

Consequence: A security or maintenance review that relies on the summary can overlook an Actions API capability that must be preserved for benchmark restoration.

Recommended action: Make the high-level summary and the workflow permission model agree. If Actions API restoration remains, describe the reset as followed by capability-specific job grants, including actions: read for benchmark restoration and issues: write for failure publication. If the publisher-only summary is the intended contract, redesign restoration so the benchmark job no longer needs that grant.

References:

  • crates/cargo-anvil/docs/design/github.md, Benchmark regression detection

Impacted locations:

  • .github/workflows/anvil-scheduled-impl.yml:42-43
  • .github/workflows/anvil-scheduled-impl.yml:159-162
  • crates/cargo-anvil/templates/github/scheduled-impl-workflow.yml:42-43
  • crates/cargo-anvil/templates/github/scheduled-impl-workflow.yml:159-162

steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
# The analysis orders each series by first-parent commit topology
# and locates the merge-base, so it needs the whole commit graph.
# The checkout has already happened by the time an action runs, so
# this too has to be set here.
fetch-depth: 0
lfs: true
- uses: ./.github/actions/anvil-run-group
with:
group: scheduled-benchmarks
bench_history: true
# Per-leg identity: the matrix value is in scope here and nowhere
# inside the action.
bench_artifact: bench-history-${{ matrix.os }}
bench_machine_key: ${{ inputs.bench_machine_key }}

publish-failure:
name: Publish scheduled failure
needs:
- scheduled-test
- scheduled-advisories
- scheduled-runtime-analysis
- scheduled-exhaustive
- scheduled-benchmarks
if: ${{ always() && vars.ANVIL_PUBLISH_FAILURE_ISSUE != 'false'
&& contains(needs.*.result, 'failure') }}
runs-on: ${{ inputs.linux_runner }}
Expand Down
3 changes: 3 additions & 0 deletions .github/workflows/anvil-scheduled.yml
Original file line number Diff line number Diff line change
Expand Up @@ -21,5 +21,8 @@ jobs:
# and restores that scope only on publish-failure.
permissions:
contents: read
# The scheduled-benchmarks job restores its history artifact, which
# reads the Actions runs/artifacts API. Narrowed to that job inside.
actions: read
issues: write
secrets: inherit
Loading
Loading