Repository navigation
⚡ Remeasure test weights at 752e8c28 and recalibrate the runtime shards - #837
Merged
Merged
Conversation
The committed weights dated from a7f60c0 and the corpus has grown past them: 81 Deno, 43 Node and 43 Bun applicable files had no recorded weight and were charged the corpus maximum. A fallback that large does not balance anything — priced with a fresh measurement, `test-deno (6/13)` was carrying 471s of work while `test-deno (1/13)` carried 183s, and `test-node (5/7)` carried 521s. Every runtime had been running over the 300s ceiling for four consecutive runs on main. `test-weights.json` is byte-for-byte the artifact of measure-test-weights run 35671233627, dispatched at 752e8c2 on ubuntu-latest. Every applicable file in all three runtimes is measured; the fallback set is empty. The serial totals are Deno 3805s over 416 files, Node 2012s over 322, Bun 1092s over 321, which puts the measured floors at 13/7/4 — where the counts already stood, and where the partition leaves a shard nothing under the ceiling but its own rounding. So the counts move to 15/10/5, the smallest ones whose predicted worst shard lands near 235s under the per-runtime ratio the last four main runs measured between an isolated weight and the step that ran it. The balanced split predicts 254s, 201s and 218s of work per shard, spread within a second across every shard of a runtime. The five-run sequence the spec calls for runs on this branch.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Every runtime's shards have been running over the 300 s ceiling
specs/testing-spec.mdsets, on every recent run ofmain:3efedc3385b5d2efc77a4b88c00fa822The cause is not runner overhead — it is that the committed weights had stopped describing the corpus. They dated from
a7f60c02; since then 81 Deno, 43 Node and 43 Bun applicable files had arrived with no recorded weight, and each was charged the corpus maximum (115 s / 103 s / 70 s). A fallback set that large balances nothing: the partition was placing files by a number that had no relation to what they cost.Priced with a fresh measurement, the split CI was actually running looked like this:
1/13— 183 s6/13— 471 s2/7— 164 s5/7— 521 s3/4— 218 s2/4— 312 stest-deno (6/13)carried two and a half times shard 1's work, andtest-node (5/7)three times shard 2's. The observed 398 s and 629 s steps are those shards, not a slow runner.What changes
Before:
test-weights.jsonmeasured ata7f60c02, with 167 of 1059 applicable (runtime, file) pairs unmeasured; counts 13/7/4; heaviest true shard 471 s (Deno), 521 s (Node), 312 s (Bun).After:
test-weights.jsonis byte-for-byte the artifact of Measure test weights run 35671233627, dispatched at752e8c28onubuntu-latest(deno 2.9.5 / node v22.23.2 / bun 1.4.0 — provenance in the file'ssourceblock, sha25667667d9e…ff3f). Every applicable file under all three runtimes is measured and the fallback set is empty. Counts move to 15/10/5, and the partition spreads each runtime's corpus to within a second across its shards:How it works
Why 15/10/5 and not the floors
The measured floors are 13/7/4 — exactly the counts already installed. That is the whole problem with reading the floor as an answer here:
floor(sum / 300000) + 1is the count at which a perfectly balanced shard's isolated work reaches the ceiling, so at 13/7/4 the balanced split predicts 294 s, 288 s and 273 s per shard and leaves a step nothing but rounding before it goes over.Sizing above the floor needs the ratio between an isolated per-file weight and the step that actually ran those files. Re-pricing run 35662675197's shards with the fresh weights gives one per runtime:
Node and Bun steps run above their isolated weights and Deno's run below, so 13/7/4 projects 274 s, 335 s and 294 s — Node misses outright and Bun sits on the line. 15/10/5 is the smallest triple whose projected worst step lands near 235 s in all three runtimes (Deno 237 s, Node 234 s, Bun 235 s), which leaves the 19–49 s the window clause has historically added over the worst step.
A count cannot fix the window clause past a point, and this is near it. #721 recorded that the account runs about 20 jobs at once, and that a run holding 33 lost a third of its shards to queueing — a 409 s window with every step under 300 s. This change takes the shard jobs from 24 to 30, so the window will be measuring the queue as much as the partition; a window miss with every step comfortably under is that, and is not answered by another increment.
Calibration
Per the spec's rule: five consecutive runs on one fixed head keep every shard
Teststep and each runtime's execution window under 300 s, and a miss increments only that runtime and starts a fresh five. This branch's own runs are the sequence.What must stay true
partitionTestsand checked byscripts/tests/ci-workflow.test.ts"the actual partition the workflow installs", which reads the committed corpus, exclusions, weights and the counts inci.yml.ci.ymlstays the single place a count is written — the job name, the matrix and theTestargument agree, checked by the same file's "shows and invokes the same index out of the same count".test-weights.jsonis only ever written bydeno task weights:measure, and its provenance names the run that produced it.How to verify it
ok | 27 passed (118 steps) | 0 failed, exit 0 — the count-sensitive battery at this head:ci-workflow,test-shards,test-weights,runtime-tests,shard-execution,measure-test-weights,runtime-exclusions,test-file-discovery. A count inci.ymlthat no longer matches the matrix, an overlapping split, or an empty shard fails there rather than in CI.752e8c28and the measurement is already behind the corpus.mainruns they have to beat.Scope
Included
test-weights.json, byte-for-byte from the measurement artifact.ci.ymlshard counts for all three runtimes.Intentionally unchanged
AGENTS.md'stest-deno (3/6)example, which illustrates the check-name shape rather than the installed count.Generated or mechanical changes
test-weights.jsoncomes from.github/workflows/measure-test-weights.yml; no millisecond was edited by hand.Risks and limitations
Scope confirmation