Skip to content

Commit 7cbb0de

Browse files
authored
Merge pull request #7 from CommonstackAI/feat/open-data-pipeline-site
Open data pipeline and project site
2 parents 31bcac5 + 57d16aa commit 7cbb0de

44 files changed

Lines changed: 5552 additions & 149 deletions

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎CHANGELOG.md‎

Lines changed: 20 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -2,6 +2,26 @@
22

33
All notable changes to **TwinRouterBench** are documented in this file.
44

5+
## [Unreleased]
6+
7+
### Added
8+
9+
- Benchmark-agnostic pipeline configuration and the stable
10+
`run_pipeline(config)` / `twinrouterbench data run --config` interfaces.
11+
- Configurable normalized and single-turn loaders; execution, backend-judge,
12+
exact-match, and contains evaluators; loader/evaluator/backend plugin contracts.
13+
- Generic OpenAI-compatible executor and a config-only custom QA example.
14+
- Pipeline/suite validation commands and rejection of inline credentials.
15+
16+
- Versioned static data-construction package and `twinrouterbench data` CLI.
17+
- Paper-aligned sequential-locking downgrade search, tier-pool cascade,
18+
mixed-prefix reconstruction, hardened open-ended judge, and manual audit flow.
19+
- SWE-bench, BFCL, mtRAG, QMSum, and PinchBench adapters with deterministic
20+
offline fixtures.
21+
- Mock, replay, and optional live-plugin backends; guarded release-candidate
22+
publication and source/license provenance validation.
23+
- Offline unit, end-to-end, CLI, replay-equivalence, and golden-hash tests.
24+
525
## [0.1.0] - 2026-04-30
626

727
### Changed

‎CITATION.cff‎

Lines changed: 88 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,88 @@
1+
cff-version: 1.2.0
2+
message: "If you use TwinRouterBench, please cite the paper and this software."
3+
title: "TwinRouterBench"
4+
type: software
5+
version: 0.1.0
6+
repository-code: "https://github.com/CommonstackAI/TwinRouterBench"
7+
url: "https://commonstackai.github.io/TwinRouterBench/"
8+
license: Apache-2.0
9+
authors:
10+
- family-names: Yang
11+
given-names: Pei
12+
- family-names: Chen
13+
given-names: Wanyi
14+
- family-names: Yang
15+
given-names: Tongyun
16+
- family-names: Feng
17+
given-names: Pengbin
18+
- family-names: Xing
19+
given-names: Jiarong
20+
- family-names: Guo
21+
given-names: Wentao
22+
- family-names: Yao
23+
given-names: Yuhang
24+
- family-names: Han
25+
given-names: Yuhang
26+
- family-names: Li
27+
given-names: Hanchen
28+
- family-names: Wang
29+
given-names: Xu
30+
- family-names: Wang
31+
given-names: Zeyu
32+
- family-names: Xiao
33+
given-names: Jie
34+
- family-names: Yang
35+
given-names: Anjie
36+
- family-names: Tian
37+
given-names: Liang
38+
- family-names: Ai
39+
given-names: Lynn
40+
- family-names: Yang
41+
given-names: Eric
42+
- family-names: Shi
43+
given-names: Tianyu
44+
identifiers:
45+
- type: doi
46+
value: "10.48550/arXiv.2605.18859"
47+
description: "arXiv-issued DOI"
48+
preferred-citation:
49+
type: article
50+
title: "TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing"
51+
year: 2026
52+
doi: "10.48550/arXiv.2605.18859"
53+
url: "https://arxiv.org/abs/2605.18859"
54+
authors:
55+
- family-names: Yang
56+
given-names: Pei
57+
- family-names: Chen
58+
given-names: Wanyi
59+
- family-names: Yang
60+
given-names: Tongyun
61+
- family-names: Feng
62+
given-names: Pengbin
63+
- family-names: Xing
64+
given-names: Jiarong
65+
- family-names: Guo
66+
given-names: Wentao
67+
- family-names: Yao
68+
given-names: Yuhang
69+
- family-names: Han
70+
given-names: Yuhang
71+
- family-names: Li
72+
given-names: Hanchen
73+
- family-names: Wang
74+
given-names: Xu
75+
- family-names: Wang
76+
given-names: Zeyu
77+
- family-names: Xiao
78+
given-names: Jie
79+
- family-names: Yang
80+
given-names: Anjie
81+
- family-names: Tian
82+
given-names: Liang
83+
- family-names: Ai
84+
given-names: Lynn
85+
- family-names: Yang
86+
given-names: Eric
87+
- family-names: Shi
88+
given-names: Tianyu

‎README.md‎

Lines changed: 90 additions & 14 deletions
Original file line numberDiff line numberDiff line change
@@ -1,18 +1,78 @@
1-
# Twin Router Bench
1+
<div align="center">
22

3-
[![Paper](https://img.shields.io/badge/Paper-arXiv-b31b1b)](https://arxiv.org/html/2605.18859)
4-
[![Code](https://img.shields.io/badge/Code-GitHub-24292f)](https://github.com/CommonstackAI/TwinRouterBench)
3+
# TwinRouterBench
4+
5+
### Fast static supervision. Live agentic routing evaluation.
6+
7+
**A two-track benchmark for choosing the cheapest sufficient LLM at every agent step.**
8+
9+
[![Paper](https://img.shields.io/badge/Paper-arXiv%3A2605.18859-b31b1b)](https://arxiv.org/abs/2605.18859)
510
[![Dataset](https://img.shields.io/badge/Dataset-Hugging%20Face-ffcc4d)](https://huggingface.co/datasets/Amorph/TwinRouterBench)
6-
[![Website](https://img.shields.io/badge/Website-Project-2ea44f)](https://commonstackai.github.io/TwinRouterBench/)
11+
[![Leaderboard](https://img.shields.io/badge/Leaderboard-Live-0f766e)](https://commonstackai.github.io/TwinRouterBench/)
12+
[![License](https://img.shields.io/badge/License-Apache--2.0-2563eb)](LICENSE)
13+
[![GitHub Stars](https://img.shields.io/github/stars/CommonstackAI/TwinRouterBench?style=flat&color=f97316)](https://github.com/CommonstackAI/TwinRouterBench/stargazers)
14+
15+
[**Project Page**](https://commonstackai.github.io/TwinRouterBench/) ·
16+
[**Paper**](https://arxiv.org/abs/2605.18859) ·
17+
[**Dataset**](https://huggingface.co/datasets/Amorph/TwinRouterBench) ·
18+
[**Data Pipeline**](docs/DATA_GENERATION.md) ·
19+
[**Submit a Router**](https://github.com/CommonstackAI/TwinRouterBench/issues/new?template=leaderboard_submission.yml)
20+
21+
</div>
722

8-
**Twin Router Bench** is a single benchmark suite for **per-step LLM routing**: a *router* chooses which pooled `model_id` to use on every agent step, under locked pricing and cache rules. The suite ships in one Python distribution (**`twinrouterbench`**) and one source tree (**`TwinRouterBench/`**).
23+
<p align="center">
24+
<img src="leaderboard/assets/overview.svg" width="100%" alt="TwinRouterBench combines execution-verified static supervision with live SWE-bench routing evaluation." />
25+
</p>
926

10-
It contains **two tracks** inside the same product—same protocol and locked tables—not two separate benchmarks:
27+
## Why TwinRouterBench?
1128

12-
| Track | Role | CLI entry |
13-
|-------|------|-----------|
14-
| **Static** | Fast validation on a fixed supervision bank (tier labels + nominal cost metrics). | `twinrouterbench static …` |
15-
| **Dynamic** | End-to-end evaluation on **SWE-bench Verified** with real tool use—**mini-swe-agent** scaffold or **editor** scaffold. | `twinrouterbench dynamic …` / `twinrouterbench swe …` |
29+
Most router benchmarks make one model choice for a complete prompt. Real coding,
30+
research, and computer-use agents make many calls under changing tool state and
31+
conversation history. TwinRouterBench evaluates the decision the router actually
32+
faces: **which model should handle the next call given the full current prefix?**
33+
34+
| Track | What it measures | Public artifact |
35+
|---|---|---|
36+
| **Static** | Cheapest-sufficient tier at each routed step, verified through downgrade search and mixed-model execution. | **970 labels**, **520 trajectories**, **5 workloads** |
37+
| **Dynamic** | End-to-end task success and realized API spend while routing every live agent call. | **100 held-out SWE-bench Verified cases** |
38+
39+
> **Headline result.** The trained router resolves **75/100** held-out cases for
40+
> **$25.66**, compared with **74/100** for the **$54.73** all-Opus reference—a
41+
> **53.1% cost reduction** with comparable task success.
42+
43+
## Quick start
44+
45+
```bash
46+
pip install -e .
47+
48+
# Reproduce the complete offline data-construction flow without API keys.
49+
twinrouterbench data generate \
50+
--benchmark all \
51+
--backend mock \
52+
--output-dir runs/data-generation/mock-all
53+
54+
# Add and run a new benchmark using only data + JSON configuration.
55+
twinrouterbench data run \
56+
--config configs/data_generation/custom_qa_pipeline.json \
57+
--output-dir runs/data-generation/custom-qa
58+
```
59+
60+
The package exposes four connected surfaces:
61+
62+
```bash
63+
twinrouterbench data ... # config-driven construction/review/publish pipeline
64+
twinrouterbench static ... # deterministic offline router scoring
65+
twinrouterbench dynamic ... # live mini-swe-agent evaluation
66+
twinrouterbench swe ... # editor-scaffold SWE-bench evaluation
67+
```
68+
69+
| Resource | Link |
70+
|---|---|
71+
| Interactive leaderboard and protocol | [Project page](https://commonstackai.github.io/TwinRouterBench/) |
72+
| Static question bank | [Hugging Face](https://huggingface.co/datasets/Amorph/TwinRouterBench) |
73+
| Data-generation methodology and CLI | [docs/DATA_GENERATION.md](docs/DATA_GENERATION.md) |
74+
| Leaderboard submission rules | [docs/SUBMISSION.md](docs/SUBMISSION.md) |
75+
| Citation metadata | [CITATION.cff](CITATION.cff) |
1676

1777
---
1878

@@ -100,6 +160,7 @@ Ensure a checkout exists at `semantic-router/` relative to the checkout parent d
100160
Primary entrypoint:
101161

102162
```bash
163+
twinrouterbench data <subcommand> [args] # construct/review/publish data
103164
twinrouterbench static <subcommand> [args] # static track
104165
twinrouterbench dynamic <subcommand> [args] # mini-swe-agent harness
105166
twinrouterbench swe <subcommand> [args] # editor-scaffold harness
@@ -372,7 +433,9 @@ Under `TwinRouterBench/scripts/examples/`:
372433
| `swerouter/` | Router protocol, pricing, cache simulation, harness, and leaderboard. |
373434
| `data/static/` | Static track JSONL: `question_bank.jsonl`, `manifest.json`. |
374435
| `data/dynamic/` | Dynamic track locked JSON: `model_pool.json`, `model_pricing.json`, `ttl_policy.json`, `tier_to_model.json`, `sr_knn_to_pool.json`, plus `dynamic_heldout100_ids.txt` (paper’s 100-case SWE-bench Verified held-out split). |
375-
| `twinrouterbench/` | Meta-CLI dispatcher. |
436+
| `twinrouterbench/data_generation/` | Static data construction, review, replay, provenance, and publish pipeline. |
437+
| `leaderboard/` | Project page and interactive static/dynamic leaderboard published through GitHub Pages. |
438+
| `twinrouterbench/` | Unified CLI dispatcher. |
376439
| `.env.example` | Template for gateway credentials. |
377440

378441
---
@@ -394,15 +457,28 @@ Do **not** upload run logs to the Hugging Face dataset page.
394457

395458
## Citation
396459

397-
If you use Twin Router Bench in research, please cite the associated paper. **Bibliographic details are withheld for anonymous review** and will be added after publication (no preprint URL in this release).
460+
If you use TwinRouterBench, please cite the paper and this repository. Machine-readable metadata is available in [`CITATION.cff`](CITATION.cff).
461+
462+
```bibtex
463+
@misc{yang2026twinrouterbench,
464+
title = {TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing},
465+
author = {Pei Yang and Wanyi Chen and Tongyun Yang and Pengbin Feng and Jiarong Xing and Wentao Guo and Yuhang Yao and Yuhang Han and Hanchen Li and Xu Wang and Zeyu Wang and Jie Xiao and Anjie Yang and Liang Tian and Lynn Ai and Eric Yang and Tianyu Shi},
466+
year = {2026},
467+
eprint = {2605.18859},
468+
archivePrefix = {arXiv},
469+
primaryClass = {cs.LG},
470+
doi = {10.48550/arXiv.2605.18859},
471+
url = {https://arxiv.org/abs/2605.18859}
472+
}
473+
```
398474

399475

400476
## Implementation note (CLI forwarding)
401477

402-
`twinrouterbench static|dynamic|swe` dispatches in-process to the existing CLIs. For debugging, you may still invoke `python -m miniswerouter.cli` or `python -m swerouter.cli` with `PYTHONPATH` set to `TwinRouterBench/`.
478+
`twinrouterbench data|static|dynamic|swe` dispatches in-process to the corresponding CLIs. For debugging, you may still invoke `python -m twinrouterbench.data_generation.cli`, `python -m miniswerouter.cli`, or `python -m swerouter.cli` with `PYTHONPATH` set to `TwinRouterBench/`.
403479

404480
---
405481

406482
## Appendix: migration
407483

408-
This tree unifies the static and dynamic router benchmark tracks in one install. Use **Twin Router Bench** naming in new scripts; keep legacy console script names and pathnames only where required for backward compatibility.
484+
This tree unifies the static and dynamic router benchmark tracks in one install. Use **TwinRouterBench** in new documentation and scripts; keep legacy console-script names and pathnames only where backward compatibility requires them.
Lines changed: 15 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,15 @@
1+
{
2+
"run_id": "commonstack-live-smoke-2026-08-18",
3+
"backend_name": "live",
4+
"seed": 20260818,
5+
"review_rate": 0,
6+
"cascade_size": 1,
7+
"max_cases_per_benchmark": 1,
8+
"collector": "twinrouterbench-commonstack-live-smoke",
9+
"collected_at": "2026-08-18",
10+
"model_pool_path": "configs/data_generation/commonstack_smoke_models.json",
11+
"generation_parameters": {
12+
"temperature": 0,
13+
"top_p": 1
14+
}
15+
}
Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,9 @@
1+
{
2+
"version": "commonstack-live-smoke-2026-08-18",
3+
"tiers": {
4+
"low": ["openai/gpt-4o-mini"],
5+
"mid": ["deepseek/deepseek-v3.2"],
6+
"mid_high": ["google/gemini-3-flash-preview"],
7+
"high": ["openai/gpt-5.4-nano-2026-03-17"]
8+
}
9+
}
Lines changed: 48 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,48 @@
1+
{
2+
"schema": "twinrouterbench.data_pipeline.v1",
3+
"suite_version": "custom-qa-example-v1",
4+
"generation": {
5+
"run_id": "custom-qa-config-only",
6+
"backend_name": "mock",
7+
"review_rate": 0,
8+
"cascade_size": 1,
9+
"collector": "config-only-example",
10+
"collected_at": "2026-08-18",
11+
"model_pool_path": "commonstack_smoke_models.json"
12+
},
13+
"backend": "mock",
14+
"benchmarks": {
15+
"custom_qa": {
16+
"display_name": "Custom QA",
17+
"scenario": "single_turn_qa",
18+
"multi_step": false,
19+
"manual_review": false,
20+
"source": {
21+
"uri": "local://custom_qa_tasks.jsonl",
22+
"license": "CC0-1.0",
23+
"version": "example-v1"
24+
},
25+
"loader": {
26+
"type": "single_turn",
27+
"options": {
28+
"system_prompt": "Answer the question concisely.",
29+
"field_map": {
30+
"instance_id": "id",
31+
"prompt": "question",
32+
"reference": "answer",
33+
"target_tier": "target_tier",
34+
"hint": "hint"
35+
}
36+
}
37+
},
38+
"evaluation": {
39+
"trial": "execution",
40+
"final": "execution"
41+
}
42+
}
43+
},
44+
"run": {
45+
"benchmarks": "all",
46+
"output_dir": "../../runs/data-generation/custom-qa"
47+
}
48+
}
Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,2 @@
1+
{"id":"capital-france","question":"What is the capital of France?","answer":"Paris","target_tier":"low","hint":"low"}
2+
{"id":"explain-cache","question":"Explain why a cache TTL is useful in one sentence.","answer":"A cache TTL limits how long stale data can be reused.","target_tier":"mid","hint":"low"}
Lines changed: 13 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,13 @@
1+
{
2+
"run_id": "mock-all-v1",
3+
"backend_name": "mock",
4+
"seed": 20260508,
5+
"review_rate": 0.1,
6+
"cascade_size": 3,
7+
"collector": "twinrouterbench-reference-pipeline",
8+
"collected_at": "2026-05-08",
9+
"generation_parameters": {
10+
"temperature": 0,
11+
"top_p": 1
12+
}
13+
}

‎data/README.md‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -6,3 +6,7 @@
66
| **`dynamic/`** | Dynamic | Locked JSON for SWE runs: model pool, pricing, TTL policy, tier map, SR-KNN label mapping, etc. Also `dynamic_heldout100_ids.txt` — the paper’s fixed 100 SWE-bench Verified held-out instance IDs (see top-level README). |
77

88
Code defaults: `main.dataset.STATIC_DATA_DIR` → `TwinRouterBench/data/static/`; harness defaults for pool/pricing → `TwinRouterBench/data/dynamic/`.
9+
10+
The construction code that produces static-track candidates lives under
11+
`twinrouterbench/data_generation/`. It writes isolated run directories and does
12+
not overwrite this directory. See [`../docs/DATA_GENERATION.md`](../docs/DATA_GENERATION.md).

0 commit comments

Comments
 (0)