Skip to content

feat(harness): run a single epoch of a run - #132

Closed
abhinav-pola wants to merge 1 commit into
mainfrom
feat/run-single-epoch
Closed

abhinav-pola wants to merge 1 commit into
mainfrom
feat/run-single-epoch

Conversation

@abhinav-pola

@abhinav-pola abhinav-pola commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Adds an optional epoch to RunConfig and runBenchmarkById. When it is set, the run evaluates only that epoch index of epochs, and sample scores, request session IDs and response-cache keys carry that index. When it is unset, behavior is unchanged.

openrouter-web runs each Kepler trial pod as one sample. Today that pod runs all of the sample's epochs, so one slow epoch keeps the pod past its deadline and the kill discards the epochs that had finished. With this change openrouter-web can schedule one pod per (sample, epoch): OpenRouterTeam/openrouter-web#48753, which picks this up through a subtree pull once it merges.

Testing

bun run format:check, bun run check, bun run typecheck, bun test (1617 pass) and bun run build. New test: runBenchmark({ epochs: 3, epoch: 2 }) scores each sample once, at epoch 2.

No score changes: a run without epoch evaluates exactly the same sample and epoch pairs as before.

🤖 Generated with Claude Code


Devin Review

An optional epoch on RunConfig and runBenchmarkById runs only that epoch index; unset runs every epoch as before. openrouter-web runs each Kepler trial pod as one sample's one epoch, so a slow epoch can no longer hold up or discard the others.
@abhinav-pola
abhinav-pola requested a review from a team as a code owner October 2, 2026 00:18

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 2 potential issues.

1 flag not posted on this PR by your GitHub settings — view it in Devin Review. (Configure)

Devin Review

Comment thread src/runner/run-by-id.ts
const model = modelFromConfig(input.benchmarkConfig);
const runConfig: RunConfig = definedValues({
epochs: input.epochs,
epoch: input.epoch,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Single-epoch result files overwrite each other

When separate epoch runs reuse a session ID, makeLocalResultStore writes them to the same filename. The later write discards the earlier epoch's scores.

Learn more

The local result store names files by benchmark, model, and session ID, not epoch. The new single-epoch mode permits distinct calls with the same session ID and different epoch indices. Both calls write valid results but target one file, so the last write replaces the first. This affects callers that use resultStore for multiple epoch runs under the same benchmark, model, and session ID; callers without a result store are unaffected.

Example: Run epoch 0 and then epoch 1 with epochs: 2, sessionId: "trial", and the same local store. Both runs return scores, but the resulting parquet file contains only epoch 1.

Recommended fix: Pass the selected epoch through the result-store write options and include it in the filename when present. Keep the existing filename for runs without a selected epoch, and test two writes sharing a session ID.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Comment thread src/harness/run.ts
Comment on lines +341 to +343
config.epoch === undefined
? Array.from({ length: config.epochs }, (_, epoch) => epoch)
: [config.epoch],

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Out-of-range epoch produces invalid scores

When epoch is negative or at least epochs, runBenchmark still evaluates that nonexistent epoch. Its scores and request IDs carry the invalid index.

Learn more

epochs defines the number of valid zero-based epoch indices. The selected epoch bypasses that count, so an invalid index still flows through scoring, logging, request session IDs, and cache salts. The public input accepts numbers without checking their range.

Example: runBenchmark({ epochs: 3, epoch: 3, maxConcurrency: 1 }) evaluates every sample as epoch 3 instead of rejecting the request; the valid indices are 0, 1, and 2.

Recommended fix: Validate the selected index as a safe integer in [0, config.epochs) before streaming. Surface an explicit input error to runBenchmarkById callers and add tests for negative, fractional, and out-of-range indices.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant