test(v4): measure and guard mounting at page scale - #829
Conversation
Code ReviewRisk: Low — No concrete defects were found in the changed code; the performance guards and CI comparison workflow are safe to merge. Adds browser-based, page-scale v4 mounting benchmarks, shared fixtures, blocking regression guards, and a reusable action that compares interleaved base and head measurements. The workflow also distinguishes absent baseline suites from broken benchmark runs and reports timing differences without making noisy wall-clock thresholds block CI. Review usage: 25,481 in (3,056 cached) / 664 out tokens — $0.0162 (openrouter/openai/gpt-5.6-luna, thinking: low) Reviewed by @weareikko/code-review v0.9.5 for commit 4f2857e. Previous review runsPrevious run archived 2026-08-16T14:31:50ZCode ReviewRisk: Low — No concrete defects were found in the reviewed changes; the MR is safe to merge aside from any environment-specific benchmark instability. Adds page-scale Chromium benchmarks and blocking mounting guards for v4, sharing fixtures between throughput measurements and regression tests. It also adds a composite action and workflow that alternates base/head benchmark runs, aggregates reports, and posts pull-request comparisons while failing broken measurements. Review usage: 105,498 in (80,405 cached) / 1,206 out tokens — $0.0233 (openrouter/openai/gpt-5.6-luna, thinking: low) Reviewed by @weareikko/code-review v0.9.5 for commit ebb3d48. Previous run archived 2026-08-16T14:21:47ZCode ReviewRisk: Medium — issues that should be addressed before merge. Adds a reusable composite action for alternating base/head benchmark runs, aggregates Vitest JSON output, and updates a sticky pull-request comment. It also adds Chromium at-scale mounting benchmarks and blocking page-scale guards for v4, while excluding fixture files from builds and packages. 1 issue found:
Still open from earlier reviews (2 findings):
Review usage: 39,742 in (3,056 cached) / 980 out tokens — $0.0260 (openrouter/openai/gpt-5.6-luna, thinking: low) Reviewed by @weareikko/code-review v0.9.5 for commit 2ee80d6. Previous run archived 2026-08-16T14:17:58ZCode ReviewRisk: Medium — issues that should be addressed before merge. Adds a reusable composite action for alternating base/head Vitest benchmark runs, report aggregation, and sticky pull-request comments. It also adds Chromium mounting benchmarks, scale fixtures, regression guards, and a v4 benchmark workflow. 1 issue found:
Still open from earlier reviews (2 findings):
Review usage: 73,028 in (38,914 cached) / 1,885 out tokens — $0.0289 (openrouter/openai/gpt-5.6-luna, thinking: low) Reviewed by @weareikko/code-review v0.9.5 for commit e144747. Previous run archived 2026-08-16T13:56:33ZCode ReviewRisk: Medium — issues that should be addressed before merge. The MR introduces shared scale fixtures, Vitest browser benchmarks, page-scale mounting assertions, and a composite GitHub Action that compares interleaved base/head benchmark rounds in a sticky pull-request comment. It also excludes fixture files from the v4 build and package checks. 1 issue found:
Still open from earlier reviews (1 finding):
Review usage: 185,625 in (150,668 cached) / 2,484 out tokens — $0.0368 (openrouter/openai/gpt-5.6-luna, thinking: low) Reviewed by @weareikko/code-review v0.9.5 for commit 93010d7. Previous run archived 2026-08-16T13:47:25ZCode ReviewRisk: Low — The change adds v4 browser-scale benchmarks, blocking scale guards, and a non-blocking base-versus-head benchmark report in CI; no concrete defects were found in the reviewed diff. The new fixtures are shared by the benchmark and guard suites, covering mounting, nesting, batching, responsive options, visibility controllers, and teardown at larger page sizes. The composite action measures both commits on one runner, aggregates round medians, and updates a pull-request comment without making timing regressions block CI. Still open from earlier reviews (1 finding):
Review usage: 120,647 in (102,765 cached) / 1,918 out tokens — $0.0221 (openrouter/openai/gpt-5.6-luna, thinking: low) Reviewed by @weareikko/code-review v0.9.5 for commit 78351b6. Previous run archived 2026-08-16T13:45:08ZCode ReviewRisk: Low — no blocking issues; safe to merge aside from nits. This MR adds a reusable benchmark-diff composite action, a v4 Chromium benchmark workflow, at-scale mounting fixtures, and blocking performance specs. It also excludes fixture files from package builds and adds a local v4 reporting command. Still open from earlier reviews (1 finding):
Review usage: 112,786 in (79,733 cached) / 2,034 out tokens — $0.0307 (openrouter/openai/gpt-5.6-luna, thinking: low) Reviewed by @weareikko/code-review v0.9.5 for commit 322d876. Previous run archived 2026-08-16T13:38:39ZCode ReviewRisk: Low — no blocking issues; safe to merge aside from nits. Adds Chromium benchmarks and scale-oriented mounting guards for v4, plus a reusable GitHub Action that alternates base and head measurements and comments on timing differences. It also excludes fixture files from the package build and adds a local v4 benchmark reporting command. Still open from earlier reviews (1 finding):
Review usage: 33,687 in (3,056 cached) / 1,062 out tokens — $0.0225 (openrouter/openai/gpt-5.6-luna, thinking: low) Reviewed by @weareikko/code-review v0.9.5 for commit 2ebc9a4. Previous run archived 2026-08-16T13:34:47ZCode ReviewRisk: Medium — issues that should be addressed before merge. Adds page-scale Chromium benchmarks and blocking mounting guards for v4, plus a composite action that compares base and head benchmark results in pull requests. The intended CI coverage is currently ineffective because each benchmark invocation fails to resolve its output path and failures are explicitly ignored. 1 issue found:
Review usage: 32,964 in / 1,019 out tokens — $0.0237 (openrouter/openai/gpt-5.6-luna, thinking: low) Reviewed by @weareikko/code-review v0.9.5 for commit 8d3824f. |
Export sizeBundled per export with peer dependencies left external, dynamic imports excluded and the output minified; sizes are gzipped. ✅ No export size changes. Unchanged (388)@studiometa/js-toolkit
@studiometa/js-toolkit-v4
|
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #829 +/- ##
=======================================
Coverage 97.16% 97.16%
=======================================
Files 170 170
Lines 4133 4133
Branches 1151 1151
=======================================
Hits 4016 4016
Misses 106 106
Partials 11 11
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Merging this PR will not alter performance
Comparing Footnotes
|
`mount.bench.ts` measures one instance. Nothing measured a page of them, and v4 had no performance coverage in CI at all, so the demo page that stalled while mounting had no number to point at. Two layers. `mount-at-scale.bench.ts` reports throughput in a real Chromium for flat, four-deep, realistic, batched, destroyed, `in-view` and responsive-option pages at 100, 1 000 and 5 000 components, against a control of declared-but-unregistered elements. `mount-at-scale.spec.ts` turns the findings into blocking guards. The measurement corrects the premise. Eager mounting already chunks: `applyMountStrategy()` posts one background task per element, drained inside the scheduler's 5 ms budget. What does not chunk is the observer batch — `processMutations()` scans an inserted subtree and tears down a removed one in one synchronous pass — so blocking cost tracks DOM nodes, not component weight. The long-task cliff sits near 12 000 inserted nodes and 13 000 removed ones, and 4 000 realistic components settle in 363 ms while blocking for only 59 ms of it. Guards therefore stay far under that cliff wherever the threshold would depend on machine speed, and are tight only where it does not: five times the components must stay within 2.5x per component, and a four-deep tree within 1.8x of a flat one. Both sides of each ratio are measured in the tens of milliseconds, because taking the best of several repeats favours the shorter side — a 3 ms run dodges the interruption a 70 ms run absorbs — and a ratio inherits that bias. All four guards survive four repeats quiet and three with every core saturated. `V4_BENCH_SIZES` narrows the matrix so CI can compare a subset, and the config's reference to a `packages/tests` CodSpeed suite — which does not exist in this repository — is replaced by what actually guards these. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011nNdFD3aQhzfdm3EsCSbS9
78351b6 to
93010d7
Compare
v4 had no performance coverage in CI. It cannot join the v3 CodSpeed job: simulation mode forces `pool: "forks"` and instruments the Node process, while a browser-mode benchmark body runs in Chromium over CDP, so it would measure the driver. So this borrows the shape `weareikko/export-size` already proved in this repository — measure head, check the base sha into a subdirectory of the same runner, measure it with the same action-provided script, diff, upsert a sticky comment — and adds what a stopwatch needs that a byte counter does not. Sides alternate within a round and the order flips between rounds, so drift lands on both. Every value is a median of round medians. A benchmark under 5 ms is reported but never flagged, because a sub-millisecond measurement cannot resolve a 25 % move. The threshold is measured, not chosen. Running one commit against itself with three interleaved rounds per side moves benchmarks by 4.5 % median and 12.5 % p90; every case above 13 % is a benchmark under 2 ms, which the floor already excludes. 25 % leaves room for a shared runner. No cached baseline, deliberately. Timed here, `npm ci` is 13.7 s and no build is needed at all — vitest runs the browser suite from TypeScript — against ~30 s per benchmark run, six of them. Caching would save under a tenth of the job and would reintroduce the cross-machine noise the whole design exists to remove, since a stored number cannot be interleaved. The action knows nothing about environments: it takes a directory, a command and an output path, and reads the JSON vitest emits, which is the same schema in Node, in happy-dom and in a browser. Whose benchmarks these are is an input — `id`, `title`, `unit` — so a second suite is a step with its own sticky comment rather than a branch inside the action. It comments and never blocks: a wall-clock gate on a shared runner is a gate that gets deleted. `npm run bench:v4` prints the same table locally. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011nNdFD3aQhzfdm3EsCSbS9
93010d7 to
e144747
Compare
v4 mount benchmarksBase and head measured on this runner, alternating over 3 rounds each; every value is the median of the round medians. Running both sides on one machine is what removes cross-machine noise — a cached baseline from another runner would put it back. A move under 25%, or on a benchmark under 5 ms, is not reported as a change: it is inside the measured noise of a shared runner. The base commit has no benchmark suite, so every benchmark below is new. A base that failed to build or measure would have failed this job rather than appearing here.
|
`mountCost()` took the minimum of every signal, which is right for a cost and backwards for a defect: one lucky repeat out of four made `expect(longestTask).toBe(0)` pass on a page that blocked the main thread in the other three. The maximum is wrong the other way. A `longtask` entry attributes any 50 ms task in the frame, mounting's or not, so a single GC pause would fail the build — and a guard that flakes gets deleted. Measured across the cliff, four repeats each, the two failure modes are visible side by side: | flat components | repeats (ms) | min | second worst | max | | --------------- | ------------- | --- | ------------ | --- | | 2 000 (guarded) | 0, 0, 0, 0 | 0 | 0 | 0 | | 9 000 | 0, 0, 0, 67 | 0 | 0 | 67 | | 11 000 | 0, 0, 51, 58 | 0 | 51 | 58 | | 12 000 | 50, 53, 60, 61| 50 | 60 | 61 | At 11 000 the minimum passes a page that blocks half the time. At 9 000 the maximum fails on one sample out of four. The second worst is the only reduction that separates them, so that is what a long task uses; wall time keeps the minimum, where the least interrupted run really is the best estimate of the work. Noise is absorbed in how many repeats blocked, never by allowing some blocking, so the assertion still reads "no long task". At the guarded sizes it costs nothing: 240 repeats across ten runs, quiet and with all eight cores saturated, every one of them clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011nNdFD3aQhzfdm3EsCSbS9
`${{ inputs.install }} || true` and the same on `prepare` came from
`weareikko/export-size`, where degrading is harmless: a missing byte count
is obviously missing. A missing *baseline* is not obviously missing. It
becomes "everything is new", which reads as a successful comparison that
never happened.
Absent and broken are now different states, decided by counting output
rather than by the action knowing anything about the command it was given:
| side | every round measured | none measured | some measured |
| ---- | -------------------- | ---------------------------- | ------------- |
| head | compare | fail | fail |
| base | compare | empty base, warn, comment says so | fail |
A base with no benchmark suite yet is real — it is the state of every
pull request that adds one, including this one — so it still degrades,
and the comment now says that a base which failed to build or measure
would have failed the job rather than appearing there. A base that
measured some rounds and not others is a broken suite, not an absent one,
and fails. Base `install` and `prepare` are no longer suppressed at all.
All six combinations exercised against the step's own shell.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011nNdFD3aQhzfdm3EsCSbS9
`bench-comment.mjs` caught every read and parse error and returned `[]`. That fallback existed for the absent base, but an absent base is already expressed as a report which *is* `[]`, written deliberately by the aggregation step. So the catch only ever hid the other case: a report that could not be read or parsed became an empty base, and the comment announced a comparison that never happened — the same silent degradation just removed from the base install. Let it throw. A malformed or missing report now fails the step, while a base that legitimately measured nothing still comes through as `[]` and still comments. The shell flags are spelled out too. `shell: bash` already runs with `-e -o pipefail` — confirmed in the runner log, not just the docs — so a failing `bench-report.mjs` did abort the step, but a script that says only `set -u` invites the reader to conclude otherwise. It now says `set -euo pipefail`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011nNdFD3aQhzfdm3EsCSbS9
The nesting guard read 1.91 on a CI runner against a 1.8 ceiling, while the benchmark suite put the same pair at 1.01 on the same runner class. The difference is the harness, not the code: the spec measured one side to completion and then the other, so any drift between the two blocks landed entirely in the ratio. Measure them alternately, exactly as `bench-diff` alternates base and head, and compare 5 000 against 5 000 rather than 2 000 so one scheduling gap is a small fraction of each side. The nesting ratio is now 0.898 to 1.024 across ten runs, quiet and with all eight cores saturated, against 0.82 to 1.15 before. The linearity guard keeps a bias interleaving cannot remove, because its two sides must differ in size: the best of several repeats favours the shorter one, since a 14 ms run slips between interruptions an 80 ms run absorbs. Measured 1.11 to 1.20 per component quiet, up to 1.57 contended. Its threshold therefore moves from 2.5 to 3, placed between the worst noise and the signal it exists to catch — five times the components read about 1.6 per component when linear and would read 5 if quadratic — rather than just above the noise. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011nNdFD3aQhzfdm3EsCSbS9
v4 had no performance coverage in CI at all —
.github/workflows/benchmarks.ymlruns the v3 suite only, and the three v4 benchmark files ran nowhere but a developer's machine. This adds the missing at-scale dimension, turns the findings into blocking assertions, and gives CI a way to notice a regression.This PR only measures. It changes nothing in the mounting path.
The numbers
Chromium via Playwright, median of nine rounds on a developer machine. µs per component, so a non-linear curve would be legible as a number.
Mounting one insertion of N components
data-mount="in-view"— a controller per elementPer-component cost does not grow. Flat mounting goes 21 → 18 → 17 µs from 100 to 5 000; 50× the components cost 50× the time. In absolute terms, 5 000 flat components settle in 85 ms and 5 000 realistic ones in 540 ms.
Nesting is free. Four-deep costs 1.06× flat at 5 000 and 0.91× at 1 000 — inside the noise. Nothing orchestrates a child's mount from its parent, and this is now the benchmark that keeps it that way.
Weight multiplier. A realistic component — five refs, three declared options, four handlers, a
mounted()reading both — costs 6.3× a bare one. The declared surface, not the count, is what makes a page expensive.in-viewcosts 2.9× an eager mount and never mounts anything: it is theIntersectionObservercontroller per element. It is also the noisiest row, becausewhenDOMSettled()deliberately does not await visibility conditions, so intersection delivery lands partly outside the measured window — that is why 100 elements read cheaper than 1 000.A responsive option costs 1.5× a plain mount (26.5 vs 17.1 µs).
Batching and teardown
Observer batching does not change per-component cost above 1 000. Ten separate morphs cost what one insertion costs.
Destroy throughput, flat: 4.0 / 4.2 / 4.9 µs per component at 100 / 1 000 / 5 000 — roughly a quarter of what mounting costs.
What the measurement corrected
The brief's premise was that mounting does not chunk. Half of it does.
applyMountStrategy()posts one background task per element for theeagerstrategy, andScheduler.#drainBackground()runs them under a 5 ms budget. Mounting is already time-sliced. What is not sliced is the observer batch itself:processMutations()scans an inserted subtree and tears down a removed one in one synchronous pass.So the blocking cost of a page tracks its DOM node count, not its component weight — and it is a much smaller share of the total than the settle times above suggest. Long-task probe, best of three:
The cliff sits near 12 000 inserted DOM nodes and 13 000 removed ones. Mounting 4 000 realistic components takes 363 ms in total but blocks for only 59 ms of it. Control: a deliberate 120 ms busy loop reports 122 ms through the same observer, so the absence of entries is real.
Whether to chunk
processMutationsis still your call — but the numbers say the un-chunked pass is a scan, not the mounting, and it takes ~12 000 nodes to become a problem.The guards
src/mount-at-scale.spec.ts, in the normal browser test suite, blocking:How the repeats reduce, per signal. Wall time takes the minimum — it is a cost, and the least interrupted run is the best estimate of the work. A long task takes the second worst, because it is a defect, not a cost. Review caught that taking the minimum of both made the blocking guards unable to fail; the maximum would have been the opposite mistake. Measured across the cliff, four repeats each:
At 11 000 the minimum passes a page that blocks half the time; at 9 000 the maximum fails on one spurious sample. The second worst is the only reduction that separates them. Noise is absorbed in how many repeats blocked, never by allowing some blocking, so the assertion still reads
toBe(0). At the guarded sizes it costs nothing: 240 repeats across ten runs, quiet and fully contended, every one clean.Each threshold is a ceiling that tolerates one bad repeat, so a failure needs a consistent signal. All four guards pass four consecutive local runs, three more with all eight cores saturated by busy loops, and CI.
The
ubuntu-latestrunner turned out to be faster than the machine these thresholds were set on — 5 000 flat components settle in 58 ms there against 85 ms here, 5 000 realistic in 336 ms against 540 ms — so the wall-clock headroom is wider in CI than the table assumes.Two guards flaked, CI caught both, and neither was fixed by moving a number. Both were measurement bugs in the same family — a ratio of two things that were not measured under the same conditions.
The linearity check compared 500 against 4 000 and read 2.22 on the runner against 1.0 here: taking the best of several repeats systematically favours the shorter side, because a 14 ms run slips between interruptions that an 80 ms run absorbs. Comparable magnitudes shrink it; it cannot be removed entirely, since differing sizes are what the guard is for, so its threshold is now placed between the measured noise ceiling (1.57 fully contended) and the signal (5 if quadratic).
The nesting check read 1.91 on the runner while the benchmark suite put the same pair at 1.01 on the same runner class — the tell that it was the harness, not the code. The spec measured one side to completion and then the other, so drift between the blocks landed in the ratio. It now alternates the two sides, exactly as
bench-diffalternates base and head, and reads 0.898–1.024 across ten runs quiet and fully contended, against 0.82–1.15 before.Deliberately not asserted: that the cliff exists at 12 000. A faster machine would never reach it and a slower one would reach it early, so it would fail on runner speed rather than on code. The figure lives in a comment in the file instead; the ratio guard is the machine-independent half of the same protection.
CI: same-runner base-vs-head, no service
.github/actions/bench-diff— a local composite action mirroringweareikko/export-size, which this repo already uses. Measure head; check the base sha into a subdirectory of the same runner; measure it with the same action-provided script; diff; upsert a sticky comment..github/workflows/benchmarks-v4.ymlwires it up with apaths:filter.benchmarks.ymlis untouched — a workflow-levelpaths:filter is the only way to keep this off unrelated PRs, and v3's CodSpeed job has to keep running.What a byte counter does not need and a stopwatch does:
Alternate, do not run sequentially. Export-size measures head then base; for timings that puts all thermal drift on one side. Sides alternate within a round and the order flips between rounds.
Median of round medians, never a mean.
A resolution floor. Chromium clamps
performance.now()to 100 µs, so a benchmark under 5 ms is reported but never flagged.A measured threshold, below.
Comment on regressions, fail on broken measurements. A timing threshold that fails a build on a shared runner gets deleted, so regressions are reported. A measurement that did not happen is the opposite — a green report of a comparison that never ran — so absent and broken are separated by counting output:
A base with no benchmark suite is real — it is the state of this very PR — so that one case degrades, and the comment says which case it is. Base
install/prepareare never suppressed, and a report that cannot be parsed fails rather than becoming an empty baseline. The hard performance gate remains the guard specs.The noise floor
Six identical runs of the same commit, aggregated as three interleaved rounds per side and compared as if they were base and head:
Every case above 13 % is a benchmark under 2 ms, which the 5 ms floor already excludes. At
threshold: 25andfloor: 5, running that same-commit data through the real comment script reports zero changes — which is the point of the exercise.It runs, on this PR
The job is green and has posted its comment above. Because these benchmarks do not exist on
mainyet, the base side legitimately produces nothing and the run degrades to "everything is new" — the same failure modeexport-sizedesigns for, exercised for real on the first try. Three head rounds took 66 s — 22 s each on the runner — and the whole job 2 min 9 s. Oncemaincarries these benchmarks it is six runs rather than three, so roughly 4 min steady state.The first attempt is worth recording, because it is the failure this design is prone to.
BENCH_JSON="$out" <command>is a prefix assignment, so the"$BENCH_JSON"written inside the command was expanded by the calling shell before the assignment existed: sixunbound variableerrors, and a green check, because each side is allowed to fail. Fixed by exporting it as its own command, and by making an empty head an error rather than an empty report. A base that cannot build still degrades gracefully; an action that measured nothing now says so.What was rejected, and why
npm run bench:v4vitest bench, nothing trackedperformancehyperfinearound a headless runIncidentally, the
CodSpeed Performance Analysischeck on this very pull request has now reported fail → fail → pass → fail across four pushes, on a branch whosepackages/js-toolkitcontent never changed at all; the first failure came with CodSpeed's own warning that "different runtime environments" were compared. That is the noise a same-runner comparison exists to remove, demonstrated for free.CodSpeed simulation is disqualified outright:
@codspeed/vitest-pluginforcespool: "forks"and instruments the Node process, while a browser-mode benchmark body runs in Chromium over CDP, so it would measure the driver. Walltime does consume vitest's own timings and would work — but it compares across machines and runs, which is precisely the noise a same-runner diff removes for free, and it wants an account and a token in a repo whose other performance tooling deliberately has neither. Tachometer has the best statistics of anything here and is the option worth revisiting if this proves too coarse; the cost today is a second benchmark harness beside the one we already run.Plain
vitest benchwith no tracking was a genuinely acceptable outcome. It lost by about 250 lines of plain node.Wall clock, and why nothing is cached
Timed on this machine:
npm ci, coldnode_modules, warm npm cachenpx playwright install chromium, cached~/.cache/ms-playwrightnpm run build, all packagesThe measurement dominates by roughly 6×, and the build is not on the path at all. The job is 6 × 30 s of benchmarking plus two installs: about 3.5 min of work here, ~4 min on the runner, and it samples 1 000 and 5 000 rather than the full matrix. Extending it to the other three v4 bench files would add ~55 s per run — 5.5 min across six runs — for micro-benchmarks a wall-clock diff can barely resolve, so it does not.
On caching the base side, three variants, priced:
mainrun came from a different runner under different contention, which re-imports the 10–20 % cross-machine noise this design exists to remove. That is larger than most regressions worth catching: we would keep "3× slower" and lose "15 % slower". Interleaving also becomes impossible by construction, since a stored number cannot alternate with anything.Recommended follow-up, not in this PR: a trend tier. A per-PR diff cannot see slow drift across many merges, by construction. Storing each
mainresult —actions/cache, or an orphan branch the waybenchmark-action/github-action-benchmarkdoes — and comparing with a deliberately wide threshold is where a cache belongs, because imprecision does not matter there.Built to take v3 later, without doing it here
The action knows nothing about environments. It takes a directory, a command and an output path, and reads the JSON vitest emits — the same schema whether bodies run in Node, in happy-dom or in Chromium over CDP. Nothing hardcodes a browser flag, a config path or
packages/v4. Whose benchmarks these are is an input, so a second suite is a step, not a branch:idnamespaces the sticky comment and the temp files, so the two suites keep two comments rather than overwriting each other. That parameterisation is in this PR precisely because retrofitting it after a second caller exists is expensive; the second caller is not..github/workflows/benchmarks.ymland the CodSpeed job are untouched. The migration is its own change, to be judged on this one working first.What migrating v3 would cost. CodSpeed's simulation mode measures instruction counts under CPU simulation rather than wall time. It is near-deterministic and resolves single-digit percentages; a wall-clock diff on a shared runner resolves ~25 % and cannot be made to do better by design. For micro-benchmarks — one
$emit, one ref read — that is a real loss, and it is the honest argument for keeping CodSpeed on v3.Whether v3's benchmarks still earn it, measured. Of the last 12 merged pull requests, not one changed a single file under
packages/js-toolkit/. CodSpeed nonetheless failed on five of them — #823, #824, #826, #827, #828 — and all five were merged anyway. This PR reproduced it live, four times over. Its v3 content never changed by one byte, and the CodSpeed check read fail (-14.94 %, with its own "different runtime environments" warning) → fail → pass → fail across four pushes. Precision you cannot act on is not precision. A check that is always dismissed is worse than no check, because it teaches everyone to dismiss the next one — and that, more than the noise, is the case for the migration.My reading: v4 supersedes that package, and the v3 benchmarks defend code that is no longer being changed. Retiring the job is defensible on its own; if any of it is worth keeping, port the two or three benchmarks that still describe a live decision and accept the coarser resolution.
Extraction
Yes, once it has run on a few PRs. The action is already generic and the workflow passes everything in. Moving it to
weareikko/bench-diffbesideexport-sizewould be a move, not a rewrite.Also
vitest.bench.config.jsreferred to a CodSpeed suite inpackages/tests, which does not exist in this repository. Replaced with what actually guards these.V4_BENCH_SIZESnarrows the at-scale matrix so CI can sample it.npm run bench:v4prints the µs-per-component table locally.*.fixtures.tsjoins specs, benchmarks and test utilities as source-only, enforced bycheck:package.Verified:
npm run lint,npm run lint:types,npm run test:v4(1 066 tests),npm run check:package.🤖 Generated with Claude Code
https://claude.ai/code/session_011nNdFD3aQhzfdm3EsCSbS9