You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
[**Submit a Router**](https://github.com/CommonstackAI/TwinRouterBench/issues/new?template=leaderboard_submission.yml)
20
+
21
+
</div>
7
22
8
-
**Twin Router Bench** is a single benchmark suite for **per-step LLM routing**: a *router* chooses which pooled `model_id` to use on every agent step, under locked pricing and cache rules. The suite ships in one Python distribution (**`twinrouterbench`**) and one source tree (**`TwinRouterBench/`**).
23
+
<palign="center">
24
+
<imgsrc="leaderboard/assets/overview.svg"width="100%"alt="TwinRouterBench combines execution-verified static supervision with live SWE-bench routing evaluation." />
25
+
</p>
9
26
10
-
It contains **two tracks** inside the same product—same protocol and locked tables—not two separate benchmarks:
27
+
## Why TwinRouterBench?
11
28
12
-
| Track | Role | CLI entry |
13
-
|-------|------|-----------|
14
-
|**Static**| Fast validation on a fixed supervision bank (tier labels + nominal cost metrics). |`twinrouterbench static …`|
15
-
|**Dynamic**| End-to-end evaluation on **SWE-bench Verified** with real tool use—**mini-swe-agent** scaffold or **editor** scaffold. |`twinrouterbench dynamic …` / `twinrouterbench swe …`|
29
+
Most router benchmarks make one model choice for a complete prompt. Real coding,
30
+
research, and computer-use agents make many calls under changing tool state and
31
+
conversation history. TwinRouterBench evaluates the decision the router actually
32
+
faces: **which model should handle the next call given the full current prefix?**
33
+
34
+
| Track | What it measures | Public artifact |
35
+
|---|---|---|
36
+
|**Static**| Cheapest-sufficient tier at each routed step, verified through downgrade search and mixed-model execution. |**970 labels**, **520 trajectories**, **5 workloads**|
37
+
|**Dynamic**| End-to-end task success and realized API spend while routing every live agent call. |**100 held-out SWE-bench Verified cases**|
38
+
39
+
> **Headline result.** The trained router resolves **75/100** held-out cases for
40
+
> **$25.66**, compared with **74/100** for the **$54.73** all-Opus reference—a
41
+
> **53.1% cost reduction** with comparable task success.
42
+
43
+
## Quick start
44
+
45
+
```bash
46
+
pip install -e .
47
+
48
+
# Reproduce the complete offline data-construction flow without API keys.
49
+
twinrouterbench data generate \
50
+
--benchmark all \
51
+
--backend mock \
52
+
--output-dir runs/data-generation/mock-all
53
+
54
+
# Add and run a new benchmark using only data + JSON configuration.
|`twinrouterbench/data_generation/`| Static data construction, review, replay, provenance, and publish pipeline. |
437
+
|`leaderboard/`| Project page and interactive static/dynamic leaderboard published through GitHub Pages. |
438
+
|`twinrouterbench/`| Unified CLI dispatcher. |
376
439
|`.env.example`| Template for gateway credentials. |
377
440
378
441
---
@@ -394,15 +457,28 @@ Do **not** upload run logs to the Hugging Face dataset page.
394
457
395
458
## Citation
396
459
397
-
If you use Twin Router Bench in research, please cite the associated paper. **Bibliographic details are withheld for anonymous review** and will be added after publication (no preprint URL in this release).
460
+
If you use TwinRouterBench, please cite the paper and this repository. Machine-readable metadata is available in [`CITATION.cff`](CITATION.cff).
461
+
462
+
```bibtex
463
+
@misc{yang2026twinrouterbench,
464
+
title = {TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing},
465
+
author = {Pei Yang and Wanyi Chen and Tongyun Yang and Pengbin Feng and Jiarong Xing and Wentao Guo and Yuhang Yao and Yuhang Han and Hanchen Li and Xu Wang and Zeyu Wang and Jie Xiao and Anjie Yang and Liang Tian and Lynn Ai and Eric Yang and Tianyu Shi},
466
+
year = {2026},
467
+
eprint = {2605.18859},
468
+
archivePrefix = {arXiv},
469
+
primaryClass = {cs.LG},
470
+
doi = {10.48550/arXiv.2605.18859},
471
+
url = {https://arxiv.org/abs/2605.18859}
472
+
}
473
+
```
398
474
399
475
400
476
## Implementation note (CLI forwarding)
401
477
402
-
`twinrouterbench static|dynamic|swe` dispatches in-process to the existing CLIs. For debugging, you may still invoke `python -m miniswerouter.cli` or `python -m swerouter.cli` with `PYTHONPATH` set to `TwinRouterBench/`.
478
+
`twinrouterbench data|static|dynamic|swe` dispatches in-process to the corresponding CLIs. For debugging, you may still invoke `python -m twinrouterbench.data_generation.cli`, `python -m miniswerouter.cli`, or `python -m swerouter.cli` with `PYTHONPATH` set to `TwinRouterBench/`.
403
479
404
480
---
405
481
406
482
## Appendix: migration
407
483
408
-
This tree unifies the static and dynamic router benchmark tracks in one install. Use **Twin Router Bench**naming in new scripts; keep legacy consolescript names and pathnames only where required for backward compatibility.
484
+
This tree unifies the static and dynamic router benchmark tracks in one install. Use **TwinRouterBench** in new documentation and scripts; keep legacy console-script names and pathnames only where backward compatibility requires them.
{"id":"capital-france","question":"What is the capital of France?","answer":"Paris","target_tier":"low","hint":"low"}
2
+
{"id":"explain-cache","question":"Explain why a cache TTL is useful in one sentence.","answer":"A cache TTL limits how long stale data can be reused.","target_tier":"mid","hint":"low"}
Copy file name to clipboardExpand all lines: data/README.md
+4Lines changed: 4 additions & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -6,3 +6,7 @@
6
6
|**`dynamic/`**| Dynamic | Locked JSON for SWE runs: model pool, pricing, TTL policy, tier map, SR-KNN label mapping, etc. Also `dynamic_heldout100_ids.txt` — the paper’s fixed 100 SWE-bench Verified held-out instance IDs (see top-level README). |
0 commit comments