Skip to content

⚡ Remeasure test weights at 752e8c28 and recalibrate the runtime shards - #837

Merged
taras merged 1 commit into
mainfrom
agent/remeasure-test-weights
Sep 22, 2026
Merged

taras merged 1 commit into
mainfrom
agent/remeasure-test-weights

Conversation

@taras

@taras taras commented Sep 22, 2026

Copy link
Copy Markdown
Owner

Why

Every runtime's shards have been running over the 300 s ceiling specs/testing-spec.md sets, on every recent run of main:

run head deno worst / window node worst / window bun worst / window
35662675197 3efedc33 398 s / 438 s 629 s / 678 s 273 s / 292 s
35656847269 85b5d2ef 404 s / 447 s 578 s / 621 s 303 s / 322 s
35643769997 c77a4b88 457 s / 499 s 547 s / 600 s 290 s / 308 s
35240007908 c00fa822 406 s / 446 s 411 s / 455 s 333 s / 356 s

The cause is not runner overhead — it is that the committed weights had stopped describing the corpus. They dated from a7f60c02; since then 81 Deno, 43 Node and 43 Bun applicable files had arrived with no recorded weight, and each was charged the corpus maximum (115 s / 103 s / 70 s). A fallback set that large balances nothing: the partition was placing files by a number that had no relation to what they cost.

Priced with a fresh measurement, the split CI was actually running looked like this:

lightest shard heaviest shard
deno (13) 1/13 — 183 s 6/13 — 471 s
node (7) 2/7 — 164 s 5/7 — 521 s
bun (4) 3/4 — 218 s 2/4 — 312 s

test-deno (6/13) carried two and a half times shard 1's work, and test-node (5/7) three times shard 2's. The observed 398 s and 629 s steps are those shards, not a slow runner.

What changes

Before: test-weights.json measured at a7f60c02, with 167 of 1059 applicable (runtime, file) pairs unmeasured; counts 13/7/4; heaviest true shard 471 s (Deno), 521 s (Node), 312 s (Bun).

After: test-weights.json is byte-for-byte the artifact of Measure test weights run 35671233627, dispatched at 752e8c28 on ubuntu-latest (deno 2.9.5 / node v22.23.2 / bun 1.4.0 — provenance in the file's source block, sha256 67667d9e…ff3f). Every applicable file under all three runtimes is measured and the fallback set is empty. Counts move to 15/10/5, and the partition spreads each runtime's corpus to within a second across its shards:

files serial total floor installed predicted per shard
deno 416 3805 s 13 15 253.3–254.0 s
node 322 2012 s 7 10 201.158–201.170 s
bun 321 1092 s 4 5 218.416–218.446 s

How it works

measure-test-weights at 752e8c28 → test-weights.json → partitionTests(applicable, weights, count) → ci.yml matrix

Why 15/10/5 and not the floors

The measured floors are 13/7/4 — exactly the counts already installed. That is the whole problem with reading the floor as an answer here: floor(sum / 300000) + 1 is the count at which a perfectly balanced shard's isolated work reaches the ceiling, so at 13/7/4 the balanced split predicts 294 s, 288 s and 273 s per shard and leaves a step nothing but rounding before it goes over.

Sizing above the floor needs the ratio between an isolated per-file weight and the step that actually ran those files. Re-pricing run 35662675197's shards with the fresh weights gives one per runtime:

ratio min median max
deno 0.83 0.93 1.07
node 0.68 1.16 1.28
bun 0.80 1.08 1.09

Node and Bun steps run above their isolated weights and Deno's run below, so 13/7/4 projects 274 s, 335 s and 294 s — Node misses outright and Bun sits on the line. 15/10/5 is the smallest triple whose projected worst step lands near 235 s in all three runtimes (Deno 237 s, Node 234 s, Bun 235 s), which leaves the 19–49 s the window clause has historically added over the worst step.

A count cannot fix the window clause past a point, and this is near it. #721 recorded that the account runs about 20 jobs at once, and that a run holding 33 lost a third of its shards to queueing — a 409 s window with every step under 300 s. This change takes the shard jobs from 24 to 30, so the window will be measuring the queue as much as the partition; a window miss with every step comfortably under is that, and is not answered by another increment.

Calibration

Per the spec's rule: five consecutive runs on one fixed head keep every shard Test step and each runtime's execution window under 300 s, and a miss increments only that runtime and starts a fresh five. This branch's own runs are the sequence.

head counts run/attempt result
(this PR) 15/10/5 pending

What must stay true

  • Each runtime's applicable corpus is covered exactly once, with no empty shard — enforced by partitionTests and checked by scripts/tests/ci-workflow.test.ts "the actual partition the workflow installs", which reads the committed corpus, exclusions, weights and the counts in ci.yml.
  • ci.yml stays the single place a count is written — the job name, the matrix and the Test argument agree, checked by the same file's "shows and invokes the same index out of the same count".
  • No weight is hand-edited. test-weights.json is only ever written by deno task weights:measure, and its provenance names the run that produced it.

How to verify it

  • ok | 27 passed (118 steps) | 0 failed, exit 0 — the count-sensitive battery at this head: ci-workflow, test-shards, test-weights, runtime-tests, shard-execution, measure-test-weights, runtime-exclusions, test-file-discovery. A count in ci.yml that no longer matches the matrix, an overlapping split, or an empty shard fails there rather than in CI.
  • Every shard logs its runtime, its selection, its predicted total and its unmeasured files before it runs. Every shard on this PR should log 0 unmeasured — a non-zero count means a test file landed after 752e8c28 and the measurement is already behind the corpus.
  • The five-run sequence above is the evidence for the counts themselves; the numbers in "Why" are the four main runs they have to beat.

Scope

Included

  • test-weights.json, byte-for-byte from the measurement artifact.
  • The ci.yml shard counts for all three runtimes.

Intentionally unchanged

  • Runner commands, the corpus, the exclusions, and in-shard concurrency — a ceiling miss is answered with a count, never with these.
  • AGENTS.md's test-deno (3/6) example, which illustrates the check-name shape rather than the installed count.

Generated or mechanical changes

  • test-weights.json comes from .github/workflows/measure-test-weights.yml; no millisecond was edited by hand.

Risks and limitations

  • The counts are a projection until the five-run sequence fills the table above. A Node or Bun miss increments that runtime; a Deno step near 300 s would mean the 0.93 ratio does not hold at 15 shards.
  • Thirty shard jobs against an account that runs about 20 at once means queue skew shows up in the window, not in a step. Recovery: revert this commit to return to 13/7/4 with the stale weights, or keep the weights and lower a count on its own.

Scope confirmation

  • Every changed file supports the purpose described above.
  • Unrelated cleanup and formatting changes are excluded.
  • Generated or mechanical changes are clearly identified.
  • The description matches the final diff and test results.

The committed weights dated from a7f60c0 and the corpus has grown past them:
81 Deno, 43 Node and 43 Bun applicable files had no recorded weight and were
charged the corpus maximum. A fallback that large does not balance anything —
priced with a fresh measurement, `test-deno (6/13)` was carrying 471s of work
while `test-deno (1/13)` carried 183s, and `test-node (5/7)` carried 521s. Every
runtime had been running over the 300s ceiling for four consecutive runs on main.

`test-weights.json` is byte-for-byte the artifact of measure-test-weights run
35671233627, dispatched at 752e8c2 on ubuntu-latest. Every applicable file in
all three runtimes is measured; the fallback set is empty. The serial totals are
Deno 3805s over 416 files, Node 2012s over 322, Bun 1092s over 321, which puts
the measured floors at 13/7/4 — where the counts already stood, and where the
partition leaves a shard nothing under the ceiling but its own rounding.

So the counts move to 15/10/5, the smallest ones whose predicted worst shard
lands near 235s under the per-runtime ratio the last four main runs measured
between an isolated weight and the step that ran it. The balanced split predicts
254s, 201s and 218s of work per shard, spread within a second across every shard
of a runtime. The five-run sequence the spec calls for runs on this branch.
@taras
taras merged commit cd0d27b into main Sep 22, 2026
43 of 44 checks passed
@taras
taras deleted the agent/remeasure-test-weights branch September 22, 2026 02:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant