feat(harness): run a single epoch of a run - #132
abhinav-pola wants to merge 1 commit into
Conversation
An optional epoch on RunConfig and runBenchmarkById runs only that epoch index; unset runs every epoch as before. openrouter-web runs each Kepler trial pod as one sample's one epoch, so a slow epoch can no longer hold up or discard the others.
There was a problem hiding this comment.
Devin Review found 2 potential issues.
1 flag not posted on this PR by your GitHub settings — view it in Devin Review. (Configure)
| const model = modelFromConfig(input.benchmarkConfig); | ||
| const runConfig: RunConfig = definedValues({ | ||
| epochs: input.epochs, | ||
| epoch: input.epoch, |
There was a problem hiding this comment.
🔴 Single-epoch result files overwrite each other
When separate epoch runs reuse a session ID, makeLocalResultStore writes them to the same filename. The later write discards the earlier epoch's scores.
Learn more
The local result store names files by benchmark, model, and session ID, not epoch. The new single-epoch mode permits distinct calls with the same session ID and different epoch indices. Both calls write valid results but target one file, so the last write replaces the first. This affects callers that use resultStore for multiple epoch runs under the same benchmark, model, and session ID; callers without a result store are unaffected.
Example: Run epoch 0 and then epoch 1 with epochs: 2, sessionId: "trial", and the same local store. Both runs return scores, but the resulting parquet file contains only epoch 1.
Recommended fix: Pass the selected epoch through the result-store write options and include it in the filename when present. Keep the existing filename for runs without a selected epoch, and test two writes sharing a session ID.
Was this helpful? React with 👍 or 👎 to provide feedback.
| config.epoch === undefined | ||
| ? Array.from({ length: config.epochs }, (_, epoch) => epoch) | ||
| : [config.epoch], |
There was a problem hiding this comment.
🟡 Out-of-range epoch produces invalid scores
When epoch is negative or at least epochs, runBenchmark still evaluates that nonexistent epoch. Its scores and request IDs carry the invalid index.
Learn more
epochs defines the number of valid zero-based epoch indices. The selected epoch bypasses that count, so an invalid index still flows through scoring, logging, request session IDs, and cache salts. The public input accepts numbers without checking their range.
Example: runBenchmark({ epochs: 3, epoch: 3, maxConcurrency: 1 }) evaluates every sample as epoch 3 instead of rejecting the request; the valid indices are 0, 1, and 2.
Recommended fix: Validate the selected index as a safe integer in [0, config.epochs) before streaming. Surface an explicit input error to runBenchmarkById callers and add tests for negative, fractional, and out-of-range indices.
Was this helpful? React with 👍 or 👎 to provide feedback.
Summary
Adds an optional
epochtoRunConfigandrunBenchmarkById. When it is set, the run evaluates only that epoch index ofepochs, and sample scores, request session IDs and response-cache keys carry that index. When it is unset, behavior is unchanged.openrouter-web runs each Kepler trial pod as one sample. Today that pod runs all of the sample's epochs, so one slow epoch keeps the pod past its deadline and the kill discards the epochs that had finished. With this change openrouter-web can schedule one pod per (sample, epoch): OpenRouterTeam/openrouter-web#48753, which picks this up through a subtree pull once it merges.
Testing
bun run format:check,bun run check,bun run typecheck,bun test(1617 pass) andbun run build. New test:runBenchmark({ epochs: 3, epoch: 2 })scores each sample once, at epoch 2.No score changes: a run without
epochevaluates exactly the same sample and epoch pairs as before.🤖 Generated with Claude Code