Skip to content

Run-level integration: deterministic record/replay + statistical delta for benchmark runs #12

Description

@arkh-node

Following up on the discussion in #11 — here's a concrete spec before any implementation, so we can agree on the shape.

Problem

The harness runs identical multi-turn prompt sequences across many models. But the model calls in server/routes.ts go straight to the vendor SDKs with default sampling, and a run only persists responses[] (content + tokens + latency). It records which models ran and what they said, but not enough to reproduce a run or to tell a real behavioral change from sampling noise. That makes the evals hard to test, hard to re-score after a parser change, and hard to trust over time.

The key distinction (two layers, usually conflated)

The hard part isn't "add a seed" — most provider APIs don't expose one, and model snapshots drift server-side. So I'd split this into two independent layers:

Layer A — deterministic record/replay (code/scoring correctness). Capture every model call as a run artifact: the full request envelope + the raw response. Then parsing, scoring, and UI can be re-run offline, deterministically, without re-spending API budget. This is what makes the harness testable and CI-able, and lets you re-score an old run after changing the parser — with no model calls at all.

Layer B — statistical delta / rescoring (the genuinely non-deterministic measurement). Re-run the same (model, config) N times, capture the distribution, and compute the delta between runs (or model versions) so real signal is separated from sampling variance. This is the same shape bioinformatics-eval needs for "rescore to assess the delta," so it should live in a reusable form.

Proposed artifact (grounded in the Drizzle schema)

A sibling capture table to runs (or an extension), one row per model call:

run_artifacts: {
  id, runId, chatbotId, stepOrder,
  provider, model,
  modelVersion,            // resolved snapshot from the response, if the API returns it
  request: jsonb,          // { messages, params: { temperature, top_p, seed?, max_tokens, stop } }
  response_raw: jsonb,     // unparsed provider payload
  content, finishReason,
  promptTokens, completionTokens, totalTokens,
  latencyMs, createdAt,
  requestHash              // stable hash of {provider, model, request} -> replay key + dedup
}

requestHash is the replay key: in replay mode you look up the artifact instead of calling the API.

Hook point

Wrap the vendor clients in server/routes.ts behind a thin ModelClient with two modes:

  • live — call the API, persist a run_artifact;
  • replay — resolve by requestHash, return the recorded response, no network.

Everything downstream (parsing COOPERATE/DEFECT/..., scoring, UI) is unchanged.

Operations

  • replayRun(runId) → re-parse/re-score from artifacts, deterministic, zero API.
  • sampleRun(config, N) → run N times, store N artifacts → distribution per step.
  • compareRuns(a, b) → per-step agreement + distributional delta (Layer B).

Phasing

  1. Phase 1 (Layer A): capture + replay. Unblocks testing and re-scoring immediately.
  2. Phase 2 (Layer B): N-sample distribution + delta/compare. Carries straight over to bioinformatics-eval.

Happy to take Phase 1 as the first PR if the shape looks right — ModelClient wrapper + run_artifacts table + a replayRun path, behind a flag, no behavior change for existing runs.

Disclosure: human–AI pair — I work with Lor, my Claude Code partner, and we read, run, and own everything we submit.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions