Skip to content

docs(bench): refresh three-engine matrix for 889f7c1 (P0–P4) - #31

Merged
huacnlee merged 2 commits into
mainfrom
docs/roadmap-matrix-889f7c1
Sep 30, 2026
Merged

huacnlee merged 2 commits into
mainfrom
docs/roadmap-matrix-889f7c1

Conversation

@huacnlee

Copy link
Copy Markdown
Member

Summary

Full 35-scenario paired benchmark of main (889f7c1, PR #30) against the previous JIT (82d3808 runtime, same v3 harness), same-version QuickJS (JIT detached) and Bun 1.4.0 (default flags). README's JIT_MATRIX is replaced; the 6165383 v2 matrix moves to a labeled historical block (its Bun column was inflated by the bun -e stall).

  • Protocol shared-js-multibatch-v3, 16 timed batches/process, 5 discarded + 30 retained processes per configuration, CPU 8 pinned, i7-13700KF, 6,125 process records.
  • vs previous JIT: 28 faster, 2 slower (quickjs-fibonacci 0.967x, calls-recursion-closures 0.992x), 5 tied.
  • vs QuickJS: 25 faster, 5 slower, 5 tied.
  • vs Bun: 0/35 faster; closest fibonacci-recursive 0.72x, widest gap methods-dynamic 0.013x.

Evidence: benchmarks/results/roadmap-889f7c1-{paired-x86_64.tar.gz,paired-x86_64.json,metadata-x86_64.json,manifest-x86_64.json}.

Caveat: the candidate and previous builds used their own untracked Cargo.lock files; both hashes are recorded in the manifest.

🤖 Generated with Claude Code

huacnlee and others added 2 commits September 30, 2026 10:31
…89f7c1

Full paired shared-js-multibatch-v3 collection (5 discarded + 30 retained
processes per configuration, CPU 8 pinned) comparing the P0-P4 roadmap
candidate with the 82d3808 runtime, same-version QuickJS, and Bun 1.4.0.
The 6165383 v2 matrix moves to a labeled historical block.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@huacnlee
huacnlee merged commit 0487818 into main Sep 30, 2026
13 of 14 checks passed
@huacnlee
huacnlee deleted the docs/roadmap-matrix-889f7c1 branch September 30, 2026 02:41
huacnlee added a commit that referenced this pull request Sep 30, 2026
## Summary

0.12.10 regressed real gpui-shell renders (acceptance sample, one core):
compute ran at 0.13x, mixed at 0.49x and panel at 0.82x the speed of
0.12.9. The 35-kernel jit-bench matrix could not show it: it polls the
JIT after every batch and warms 640 calls, while gpui-shell never calls
`Jit::poll` and is measured from the first frames.

Three root causes, one commit each:

- **`de40dca` Baseline demoted during the funded Tier2 trial.** Five
rejected Tier2 profitability retries returned every Baseline to the
interpreter until the trial installed. The retries only say the modeled
Tier2 saving has not paid off; the runtime never measures the
interpreter. Since P2 widened Tier 1 coverage, the trial waited behind
other first Baselines on the single worker, so the layout kernel ran at
interpreter speed for about 100 frames. Now only a Baseline whose
fastest complete run is under 10 µs (entry bookkeeping can dominate,
e.g. `return a + b`) is demoted. A longer one keeps running, and its
trial moves to the front of the compile queue. The worker channel never
holds more jobs than there are workers, so the backlog stays
reorderable.
- **`1680c2c` P2 opcodes slower than the interpreter in host code.** The
panel builder's short helpers (closures, `push_this`, `typeof`, …)
became native, with about 220 entries per render, at 0.8x the
interpreter's speed. Stable-IC trials bypass the retry demotion, so
nothing returned them. Automatic tiering now rejects functions
containing any of the 52 opcodes Tier 1 gained in P2, with the reason
0.12.9 reported. `BaselineOnly`, `Optimize` and forced test tiers still
compile them. (Chosen as the conservative option for 0.12.11; letting
profitability measure the interpreter is the follow-up.)
- **`8fa2a48` Retry backoff stretched without host polling.** Retries
are evaluated only by maintenance, which ran every 64 hot ticks (about
31 frames) once the kernel was native. Maintenance is now due at the
next retry, and exhausted retries queue the trial at the next
maintenance. The demotion rule uses the fastest complete (non-OSR)
Baseline run, so preemption or a partial OSR run cannot flip it.

A generic "Tier2 first" queue order was tried and dropped. It exposed
ordering hazards: a callee reaching Tier2 before its caller's Baseline
loses call-link feedback, and `tagged_call_link` caught it.

## Performance

**gpui-shell** (`emit_one_jit_acceptance_sample`, 10 retained process
pairs, mean render µs):

| Workload | Cores | 0.12.9 | 0.12.10 | this PR | vs 0.12.9 | vs 0.12.10
|
| --- | --- | ---: | ---: | ---: | ---: | ---: |
| compute | 1 | 15.1 | 127.1 | 10.3 | 1.5x (faster) | 12.3x (faster) |
| compute | 2 | 16.3 | 62.4 | 10.2 | 1.6x (faster) | 6.6x (faster) |
| mixed | 1 | 178.5 | 363.3 | 92.7 | 1.9x (faster) | 3.9x (faster) |
| mixed | 2 | 179.2 | 180.1 | 92.4 | 1.9x (faster) | 2.0x (faster) |
| panel | 1 | 812.0 | 981.4 | 803.7 | 1.03x (faster) | 1.3x (faster) |
| panel | 2 | 806.9 | 987.9 | 805.5 | 1.00x (tied) | 1.2x (faster) |

**jit-bench**, all 35 scenarios, `shared-js-multibatch-v3`, 30 retained
processes per configuration, CPU 8. Full tables are in the README.
- vs v0.12.10: 8 faster, 11 slower, 16 tied. The two large losses are
the deliberate P2 deferral: `exceptions-sync` 0.21x and `for-of-array`
0.36x, back to QuickJS speed. The other nine are 0.5–3.6% slower; for
three of them the unchanged interpreter moved by a similar amount.
- vs QuickJS: 23 faster, 5 slower, 7 tied.
- vs Bun: 0 of 35 faster.

## Tests

- Regression tests, each verified to fail without its fix:
`automatic_gpui_layout_kernel_keeps_baseline_until_tier2_trial`,
`automatic_layout_kernel_reaches_tier2_without_host_polling`,
`prioritized_request_is_taken_ahead_of_earlier_requests`,
`dispatch_keeps_the_backlog_in_the_coordinator_while_workers_are_busy`.
- Exception lowering tests moved to `BaselineOnly`. The settle tests
assert the automatic deferral. The background tests follow bounded
dispatch.
- `cargo test -p quickjs-jit-runtime --release --features
compiler,test-support`: 56 binaries, 987 tests pass, run twice. `cargo
clippy --tests`: no warnings.

## Sanitizer CI

The nightly sanitizer job had failed since the P0–P4 merge (#30),
including on docs-only #31 and #32. It stopped at the first slow MSan
test, so the ASan and TSan steps never ran. A local no-fail-fast run of
the CI commands, also pinned to two cores to mimic the runner, found
every remaining issue. No ASan memory error and no TSan data race was
reported.

- `2166391`: sanitizer builds default to a ten-minute compile budget.
One MSan-instrumented Tier 1 compile took over 30 s on the runner and
timed out on every retry. `object_fast_paths` now waits 300 s. A forced
test tier ignores timed Tier2 losses, which timing noise alone could
trigger.
- `7f88490`: every fixed `let deadline = now + N s` in the JIT tests
allows 5x under sanitizers. The `tier2_ownership` runner keeps calling
until Tier 2 when the test expects it. `fast_entry`'s intentionally
leaked state stays reachable for LeakSanitizer. The absolute-time
entry-dominated demotion test is ignored under sanitizers:
instrumentation slows the bookkeeping but not the generated code.
- `36ddc59`: `frame_inline` waits for the linked Tier2 caller before
counting generic calls.

CI: MSan, ASan and TSan all pass. Locally, on two cores, 62/62 JIT test
binaries pass under each sanitizer.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant