Skip to content
 
 

Latest commit

 

History

91 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

OpenRouter Benchmark Harness

OpenRouter's internal benchmarking harness, externalized for transparency. We port benchmarks here so we can run them scalably on our infrastructure and iterate quickly.

bun install
OPENROUTER_API_KEY=... bun run bench -- --benchmark gpqa_diamond --model openai/gpt-4o-mini --limit 5

See CONTRIBUTING.md before proposing changes. Report security issues privately as described in SECURITY.md.

Preparing SWE-bench Verified task folders

SWE-bench Verified tasks are generated by Harbor in a separate checkout and then synchronized into this repository so they are present on the controller Droplet. Harbor's Python adapter is only a preparation tool; the managed benchmark runtime remains TypeScript.

Keep the repositories next to each other:

github.com/
├── harbor/
└── benchmark-harness/

Generate and validate the complete 500-task bundle:

cd ../harbor/adapters/swebench
uv run swebench
cd ../..
git rev-parse HEAD > datasets/swebench-verified/.harbor-commit

python3 - <<'PY'
from pathlib import Path

root = Path("datasets/swebench-verified")
tasks = [path for path in root.iterdir() if path.is_dir()]
assert len(tasks) == 500, f"expected 500 tasks, found {len(tasks)}"
required = (
    "task.toml",
    "instruction.md",
    "environment/Dockerfile",
    "tests/test.sh",
    "tests/config.json",
    "solution/solve.sh",
)
for task in tasks:
    missing = [name for name in required if not (task / name).is_file()]
    assert not missing, f"{task.name}: missing {missing}"
print("Validated 500 SWE-bench Verified tasks")
PY

Synchronize the generated folders into this checkout:

cd ../harbor
mkdir -p ../benchmark-harness/datasets/swebench-verified
rsync -a --delete \
  datasets/swebench-verified/ \
  ../benchmark-harness/datasets/swebench-verified/

On the controller, the TypeScript runtime reads:

export BENCH_SWE_BENCH_TASKS_DIR="$PWD/datasets/swebench-verified"

After the DigitalOcean sandbox variables below are configured, run a one-task smoke test with:

bun run bench -- \
  --benchmark swe_bench_verified \
  --model kimi-k3 \
  --limit 1 \
  --epochs 1 \
  --concurrency 1

When the Harbor adapter or source dataset is intentionally updated, regenerate with uv run swebench --overwrite, update .harbor-commit, rerun validation, and synchronize again. Review the resulting task-bundle diff separately from application code.

Each generated solution/solve.sh contains the reference solution. The controller may retain it for provenance and oracle validation, but the runtime must never mount or upload solution/ into the candidate container. It must capture the candidate patch first and upload only tests/ for verification.

Running GPQA against DigitalOcean's inference-proxy (fix/do-inference-proxy-support)

This branch/fork fixes two things needed to point the harness at an OpenAI-compatible endpoint other than openrouter.ai -- see PR #1 for details. To run it yourself:

git clone git@github.com:jdigitalocean/benchmark-harness.git
cd benchmark-harness
git checkout fix/do-inference-proxy-support
bun install

You'll need a DigitalOcean MODEL_ACCESS_KEY for the inference-proxy (ask your team lead if you don't have one -- it's not a real openrouter.ai key, despite the env var name below).

OPENROUTER_API_KEY=<your DO MODEL_ACCESS_KEY> \
OPENROUTER_BASE_URL=https://inference.do-ai.run/v1 \
bun run bench -- --benchmark gpqa_diamond --model kimi-k3 --epochs 3

# For high concurency
bun run bench -- --benchmark gpqa_diamond --model kimi-k3 --epochs 3 --concurrency 16
  • --model is the raw model id as DO's inference-proxy expects it (e.g. kimi-k3) -- no openrouter/ prefix.
  • Swap inference.do-ai.run for inference.do-ai-test.run to run against the test environment instead of prod.
  • Add --limit N to cap the number of questions for a quick smoke test before committing to a full run.
  • Results are written to bench-results/ as parquet, and the run summary (accuracy, token usage, per-sample scores) prints to stdout as JSON.

Managed benchmark API

See API.md for the complete consumer-facing HTTP API specification, including run discovery, request and response contracts, artifact downloads, reports, and retry workflows.

The API accepts GPQA Diamond, TAU Bench Verified Airline, Deep SWE, SWE-bench Verified, Terminal-Bench 2.1, and all three SWE Atlas tracks (QA, Test Writing, and Refactoring). Omit benchmark to default to GPQA Diamond. Configure the server with an API token, DigitalOcean Managed MySQL, and a private DigitalOcean Spaces bucket:

export BENCH_API_TOKEN='secret'
export BENCH_RUN_TRIGGER_SECRET='replace-me'
export BENCH_API_MAX_RUNS='8'
# DigitalOcean personal access token with the genai:read scope. Used only by
# the server to populate the production/test model selector; never sent to browsers.
export DO_MODEL_CATALOG_TOKEN='replace-me'
# OpenRouter API key used only to list OpenRouter models.
# Its successful catalog response is cached by the server for two hours.
export OPENROUTER_MODEL_CATALOG_TOKEN='replace-me'
export MYSQL_HOST='replace-me.db.ondigitalocean.com'
export MYSQL_PORT='25060'
export MYSQL_USER='doadmin'
export MYSQL_PASSWORD='replace-me'
export MYSQL_DATABASE='defaultdb'
export MYSQL_SSL_MODE='required'
# For certificate verification, use MYSQL_SSL_MODE='verify-ca' and set
# MYSQL_CA_CERT_PATH to the downloaded DigitalOcean CA certificate.

export SPACES_ENDPOINT='https://nyc3.digitaloceanspaces.com'
export SPACES_REGION='nyc3'
export SPACES_BUCKET='model-benchmarks-do-not-delete'
export SPACES_ACCESS_KEY_ID='replace-me'
export SPACES_SECRET_ACCESS_KEY='replace-me'
export SPACES_PREFIX='benchmark-runs'
# Set SPACES_FORCE_PATH_STYLE=1 only for local S3-compatible test servers.

# Required when Deep SWE, SWE-bench Verified, Terminal-Bench, or SWE Atlas
# uses disposable DigitalOcean Droplets.
export BENCH_HARBOR_SANDBOX='digitalocean'
export DO_SANDBOX_TOKEN='replace-me'
export DO_SANDBOX_SSH_KEY_ID='replace-me'
export DO_SANDBOX_SSH_PRIVATE_KEY_PATH='/root/.ssh/benchmark-harness-do'
export DO_SANDBOX_REGION='nyc3'
export DO_SANDBOX_SIZE='s-4vcpu-8gb'
# SWE-bench task bundles synced from Harbor. The optional size override takes
# precedence over DO_SANDBOX_SIZE only for swe_bench_verified.
export BENCH_SWE_BENCH_TASKS_DIR="$PWD/datasets/swebench-verified"
export SWE_BENCH_DO_SANDBOX_SIZE='s-4vcpu-8gb'
# Optional Atlas-only override; choose a size available in the region with at
# least 16 vCPUs and 16 GiB RAM.
export SWE_ATLAS_DO_SANDBOX_SIZE='c-16'
export DO_SANDBOX_IMAGE='ubuntu-24-04-x64'

# Required for SWE Atlas's independently configured verifier judge.
export SWE_ATLAS_JUDGE_API_KEY='replace-me'

bun run serve

MySQL stores queryable run metadata and is required for API startup and lifecycle updates. Failed runs persist a human-readable failure_reason; pending SQL migrations, including this column, are applied automatically during API startup. BENCH_RUN_TRIGGER_SECRET independently protects run creation and is required in the X-Bench-Run-Secret header. At most eight runs can be active; BENCH_API_MAX_RUNS may lower but cannot raise that cap. Every run requires a triggeredByEmail ending in @digitalocean.com, which is persisted with its metadata. MYSQL_SSL_MODE accepts required, verify-ca, or disabled; use disabled only for local development. Inference base URLs are currently unrestricted. The inference API key is accepted in the run payload, passed only to the benchmark child process, and never returned, logged, or persisted.

Open http://<server>:8080/gpqa-benchmarks for the runs dashboard. Enter BENCH_API_TOKEN in the browser to start or cancel supported runs and view model, base URL, status, live CLI completion percentage and completed/skipped evaluation counts, Parquet accuracy as the quality score, and full run artifacts. Evaluation totals include all epochs, so 198 GPQA questions over three epochs display as 594 evaluations. The token is kept in tab-scoped session storage and is not placed in URLs.

Logs, run state, and inference request events can be viewed or downloaded while a run is active. A download taken during execution is a snapshot of the file currently stored on the Droplet; download it again to include newer entries.

The inference requests viewer keeps both pending and completed attempts collapsed by default and refreshes every two seconds while any request is pending. Expanding an attempt shows the sanitized request body, partial response text, provider-exposed reasoning, tool-call deltas, bytes received, and time to first model output as streamed events arrive. The page returns to a one-minute refresh interval after all attempts complete. Request bodies larger than 256,000 characters are stored as bounded previews so request observability cannot create unbounded log files. Models and providers that do not expose reasoning cannot display hidden chain-of-thought. After a GPQA report becomes available, this same page derives latency-versus-correctness buckets, epoch consistency, response-quality diagnostics, response/reasoning length correlations, hardest questions, and epoch trends entirely from the existing Parquet report.

Start a run:

curl -X POST http://127.0.0.1:8080/runs \
  -H "Authorization: Bearer ${BENCH_API_TOKEN}" \
  -H "X-Bench-Run-Secret: ${BENCH_RUN_TRIGGER_SECRET}" \
  -H 'Content-Type: application/json' \
  -d '{
    "benchmark": "gpqa_diamond",
    "triggeredByEmail": "user@digitalocean.com",
    "inference": {
      "baseUrl": "https://inference.do-ai.run/v1",
      "apiKey": "replace-me",
      "model": "kimi-k3",
      "temperature": 1,
      "maxTokens": 8192,
      "reasoningEffort": "high",
      "timeoutMs": 120000,
      "completionTimeoutMs": 3600000
    },
    "execution": {
      "epochs": 3,
      "concurrency": 3,
      "unordered": true,
      "limit": 198,
      "maxRetries": 6
    }
  }'

The API and dashboard default GPQA and TAU to three epochs and concurrency three. Deep SWE, SWE-bench Verified, Terminal-Bench, and SWE Atlas default to one epoch and concurrency one because each evaluation provisions disposable sandbox workers. All benchmarks default to reasoning effort high, a one-hour full-response timeout per attempt, six retries for retryable inference errors, and retry-on-error behavior. GPQA defaults to temperature 1; TAU, Deep SWE, SWE-bench Verified, Terminal-Bench, and SWE Atlas default to temperature 0. Request timeouts (408), rate limits (429), transport failures, and server errors (5xx) are retryable; maxRetries: 0 disables inference retries. Every value can be overridden. Set "benchmark": "tau_bench_verified_airline" to run the 50-task TAU suite, "benchmark": "deep_swe" for the 113-task Deep SWE suite, "benchmark": "swe_bench_verified" for the 500-task SWE-bench Verified suite, "benchmark": "terminal_bench" for the 89-task Terminal-Bench 2.1 suite, or "benchmark": "swe_atlas_qa" for the 124-task SWE Atlas QA suite. SWE Atlas runs accept a top-level "judgeModel" selected independently from inference.model; the judge always uses https://inference.do-ai.run/v1 and the server-side SWE_ATLAS_JUDGE_API_KEY. The dashboard shows this field only for SWE Atlas. SWE-bench uses Harbor's deterministic verifier and does not use a judge model. Each SWE-bench task uses one worker for both candidate work and verification; hidden tests are uploaded only after the candidate patch is captured. The patch and verifier report are saved in sample metadata in the result Parquet and therefore included in the normal Spaces run bundle. Terminal-Bench and SWE Atlas also use one worker per concurrent evaluation, with agent and verifier sharing the worker. SWE Atlas task metadata requests 16 CPUs and 16 GiB RAM, so configure its dedicated size override before increasing concurrency. TAU Airline always uses openai-gpt-5.4-mini for its simulated customer through DigitalOcean inference. A production or test DigitalOcean candidate reuses its inference.baseUrl and inference.apiKey; an OpenRouter or custom candidate requires a separate top-level simulatorApiKey, which is used only with https://inference.do-ai.run/v1. The dashboard requests this token only when required, and neither secret is persisted. Optional inference fields are temperature, endpointId, costTier, sort, providerOnly, allowFallbacks, cloudflareVersion, costQualityTradeoff, pinModel, and completionTimeoutMs. completionTimeoutMs limits the complete inference attempt, including connection and streamed-body processing. The existing timeoutMs remains the connection/response-header timeout. When OpenRouter is selected, the dashboard exposes provider routing only inside Advanced configuration. Selecting DigitalOcean sends providerOnly: ["digitalocean"] and disables provider fallback by default, so OpenRouter must use DigitalOcean or fail the request. Execution can use start plus either end or limit. Set execution.unordered to true for rolling concurrency without input-order head-of-line blocking; omitted or false preserves ordered result emission.

Use "swe_atlas_qa" for the 124-task Codebase Q&A track, "swe_atlas_tw" for the 90-task Test Writing track, or "swe_atlas_rf" for the 65-task Refactoring track. All three use the same independently selected judge model and default to one epoch and concurrency one in managed runs.

SWE_ATLAS_DO_SANDBOX_SIZE overrides DO_SANDBOX_SIZE only for SWE Atlas. If it is omitted, Atlas uses the shared size. DigitalOcean containers receive the task's declared CPU and memory limits without additional CPU scaling.

SWE_BENCH_DO_SANDBOX_SIZE similarly overrides the shared size only for SWE-bench Verified. Its generated tasks request 1 CPU, 4 GiB RAM, and 10 GiB working storage; s-4vcpu-8gb is the recommended Droplet size. The current Droplet backend does not enforce Harbor's storage quota separately, so this is local Droplet/container disk, not a persistent volume.

Managed Deep SWE, SWE-bench Verified, Terminal-Bench, and SWE Atlas runs are capped at concurrency six because every concurrent evaluation provisions a worker. The dashboard warns for these benchmarks when more than one epoch is requested because such runs can take a long time and are not recommended.

Start a one-task Deep SWE smoke run with a per-request DigitalOcean inference key:

curl -X POST http://127.0.0.1:8080/runs \
  -H "Authorization: Bearer ${BENCH_API_TOKEN}" \
  -H "X-Bench-Run-Secret: ${BENCH_RUN_TRIGGER_SECRET}" \
  -H 'Content-Type: application/json' \
  -d '{
    "benchmark": "deep_swe",
    "triggeredByEmail": "user@digitalocean.com",
    "inference": {
      "baseUrl": "https://inference.do-ai.run/v1",
      "apiKey": "replace-with-model-access-key",
      "model": "replace-with-responses-compatible-model"
    },
    "execution": {
      "epochs": 1,
      "concurrency": 1,
      "limit": 1
    }
  }'

Authenticated endpoints:

GPQA and TAU report rows include a diagnostic Retry action. The retry panel asks for the inference API key and run-trigger password, reuses every other inference setting from the original run, and displays the result without modifying the original run, Parquet, score, or Spaces artifacts. GPQA retries one question; TAU reruns the complete stateful scenario. Up to three diagnostic retries can run concurrently, and their in-memory status is not recovered after an API restart.

  • GET /runs/:id/gpqa-retry-comparisons — list durable GPQA retry comparisons and the available k/N source failure bands
  • POST /runs/:id/gpqa-retry-comparisons — launch an original-configuration arm and optional alternate-provider/model arm; requires X-Bench-Run-Secret
  • GET /runs/:id/gpqa-retry-comparisons/:comparisonId — poll comparison and child-run progress
  • POST /runs/:id/gpqa-retry-comparisons/:comparisonId/cancel — cancel all running comparison arms; requires X-Bench-Run-Secret
  • GET /runs/:id/gpqa-retry-comparisons/:comparisonId/report — view or download the persisted side-by-side comparison report
  • GET /runs and GET /runs/:id — MySQL-backed run and upload metadata
  • GET /runs?view=page&page=1&pageSize=50 — paginated run metadata; dashboard filters are sent as benchmark, model, modelExact, durationGt, triggeredBy, status, qualityLt, hideCanary, canaryOnly, fullSuiteOnly, and showDisabled
  • GET /model-catalog?baseUrl=<supported inference URL> — all model IDs returned by the selected provider's catalog for the dashboard selectors
  • GET /runs/:id/logs — combined process output
  • GET /runs/:id/state — local restart-recovery state
  • GET /runs/:id/request-records — raw inference request-event JSONL
  • GET /runs/:id/parquet — Parquet result file
  • GET /runs/:id/summary — detailed Parquet-derived run summary
  • GET /runs/:id/gpqa-report — GPQA question, model response, correct answer, reasoning, per-evaluation model latency, and result details; add ?download=1 to download JSON
  • GET /runs/:id/tau-airline-report — TAU Airline scenario, expected criteria, conversation, tool calls, reasoning, per-evaluation agent latency, and reward details; add ?download=1 to download JSON
  • POST /runs/:id/diagnostic-retries — asynchronously rerun one report item using the original run configuration plus a newly supplied inference API key; requires X-Bench-Run-Secret
  • GET /runs/:id/diagnostic-retries/:retryId — poll a diagnostic retry and retrieve its result

The GPQA report has a Retry failed questions workflow for durable batch comparisons. For an N-epoch source run, questions are grouped into selectable k/N bands by counting every non-correct source outcome, including skipped evaluations. A retry comparison always reruns the selected questions with the source configuration and can also run an alternate endpoint, model, or OpenRouter provider configuration. Each arm has an independent call count (implemented as epochs), so the comparison report can show average accuracy, recovery, consistency, answer distribution, latency, full responses, and provider-exposed reasoning across repeated calls. Comparison arm runs use the normal MySQL, restart-recovery, request-log, Parquet, and Spaces lifecycle, count toward the active-run limit, stay hidden from the main run list, and never modify the source run's Parquet or score. Inference API keys are launch-only and are not persisted.

curl -X POST "http://127.0.0.1:8080/runs/<SOURCE_RUN_ID>/gpqa-retry-comparisons" \
  -H "Authorization: Bearer $BENCH_API_TOKEN" \
  -H "X-Bench-Run-Secret: $BENCH_RUN_TRIGGER_SECRET" \
  -H "Content-Type: application/json" \
  -d '{
    "selectedFailureCounts": [3, 2],
    "triggeredByEmail": "engineer@digitalocean.com",
    "original": {
      "apiKey": "<ORIGINAL_INFERENCE_API_KEY>",
      "repetitions": 3,
      "concurrency": 3,
      "unordered": true
    },
    "comparison": {
      "apiKey": "<ALTERNATE_INFERENCE_API_KEY>",
      "repetitions": 3,
      "concurrency": 3,
      "unordered": true,
      "maxRetries": 6,
      "inference": {
        "baseUrl": "https://openrouter.ai/api/v1",
        "model": "deepseek/deepseek-v4-flash-0731",
        "reasoningEffort": "high",
        "providerOnly": ["digitalocean"],
        "allowFallbacks": false
      }
    }
  }'
  • GET /runs/:id/results — Parquet metadata
  • POST /runs/:id/cancel — request cancellation
  • POST /runs/:id/disable — hide a run from the dashboard by default; send {"disabled": false} to re-enable it
  • GET /results and GET /summary — results across API runs

Each active terminal run is written locally under logs/api/<run-id>/. After it finishes, its artifacts are uploaded privately to:

s3://<bucket>/<prefix>/<benchmark-folder>/YYYY/MM/DD/<run-id>/
  run.json
  manifest.json
  logs/run.log
  requests/requests.jsonl
  results/*.parquet

<benchmark-folder> is gpqa for GPQA Diamond and tau_bench_verified_airline for TAU Bench Verified Airline.

The request JSONL writes a started event before each network attempt and a completed event when the attempt terminates. The dashboard correlates them so interrupted attempts remain visible as pending. Records include request summary, start and completion timestamps, duration, HTTP status, and error details. Successful response bodies and full prompt payloads are not stored; failed non-200/201 response bodies are retained for diagnosis. The run log combines child-process output with structured API lifecycle entries for run acceptance, process start/exit, 10% progress milestones, the post-100% result-persistence stage, cancellation, restart recovery, MySQL failures, final failure reasons, and artifact uploads. Abrupt exits receive an actionable fallback reason that calls out possible OOM/SIGKILL or service restart. Large artifacts use multipart uploads; transient upload failures retry with a fresh file stream and record the HTTP status, Spaces request IDs, and nested error details in the API log and uploadError. Benchmark status and artifact-upload status are independent: a valid Parquet containing every expected non-skipped evaluation is successful even when the Spaces upload fails. The dashboard shows failed-run reasons inline and a compact warning icon for upload failures. manifest.json is uploaded last and acts as the completion marker. Completed run views and summaries read from Spaces, and dashboard downloads use five-minute signed Spaces URLs so file data does not pass through the Droplet. After MySQL confirms the completed upload metadata, the server removes the local logs, requests, Parquet, progress, and manifest files while retaining only the small restart-recovery run.json. Failed or in-progress uploads remain local and continue to use local live snapshots. MySQL contains run configuration, model, lifecycle timestamps, status, failure reason, and Spaces pointers; it never contains inference credentials or full artifacts.

  • To convert a parquet result file to a readable markdown report:
bun src/cli/parquet-to-md.ts bench-results/<filename>.parquet > bench-results/<filename>.md

Capturing transient errors (5xx, decode failures)

Transient errors (HTTP 503s, decode errors, retries) are logged to stderr during the run but not persisted in the parquet results. To capture them, redirect stderr to a log file:

bun run bench -- --benchmark gpqa_diamond --model kimi-k3 --epochs 3 --concurrency 16 2> bench-results/run.log

To get a quick summary of 5xx error rates from the log:

# Total retry count
grep -c "Retrying after transient error" bench-results/run.log

# 5xx errors specifically
grep -c "error_status: 503" bench-results/run.log

# Decode errors (200 OK but unreadable body)
grep -c "Decode error" bench-results/run.log

About

OpenRouter TypeScript harness for reproducible LLM benchmarks and evaluations.

Resources

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages