Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
21 commits
Select commit Hold shift + click to select a range
071ffd2
feat: validate runtime repeatability
burtenshaw Sep 24, 2026
96f3fc6
Merge commit 'f1575e9856aa022b7c0c89fe79e41f2d6cf2906d' into ben/rfc0…
burtenshaw Sep 24, 2026
7e7c092
Merge commit '776dace7608385937c5273da3dc0d6a605317b5b' into ben/rfc0…
burtenshaw Sep 24, 2026
0d461e5
style: format provider capabilities
burtenshaw Sep 24, 2026
ddb3ddb
fix: distinguish incomplete replay evidence
burtenshaw Sep 24, 2026
0d44431
fix: isolate runtime fixture faults
burtenshaw Sep 24, 2026
2ba5adb
Merge branch 'ben/rfc008-l2-04-telemetry' into ben/rfc008-l2-05-repea…
burtenshaw Sep 24, 2026
9f9b780
fix: preserve grading after replay cleanup failure
burtenshaw Sep 24, 2026
26f6bb1
fix: conceal keys in divergence reports
burtenshaw Sep 25, 2026
7dc4d96
fix: require confirmed replay cleanup
burtenshaw Sep 25, 2026
a6a6f7d
fix: integrate validated runtime parent
burtenshaw Oct 1, 2026
b2ca1eb
fix: honor repeatability startup dependencies
burtenshaw Oct 1, 2026
cde9177
Merge commit '94494cc8e764837f813681693f205d693605053b' into HEAD
burtenshaw Oct 1, 2026
7d4b030
fix: accept incomplete replay skip reasons
burtenshaw Oct 1, 2026
9eccc69
fix: inherit bounded HTTP transport
burtenshaw Oct 1, 2026
dde52d0
test: inherit bounded cancellation check
burtenshaw Oct 1, 2026
274ebf1
fix: handle unscored replay steps
burtenshaw Oct 2, 2026
89f7b59
fix: sync replay contracts
burtenshaw Oct 2, 2026
1282eac
test: sync installed-wheel regression
burtenshaw Oct 2, 2026
c362304
test: isolate factory failures
burtenshaw Oct 2, 2026
a6c5f1b
docs: sync review clarifications
burtenshaw Oct 2, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
28 changes: 19 additions & 9 deletions docs/source/reference/cli.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,11 +45,15 @@ openenv validate ./my_env --level runtime --local --output report.json
Runtime validation requires Docker and a `validation.execution` declaration in
`openenv.yaml`. The declaration points to a bounded JSON replay plan containing
one reset and a sequence of actions. The validator builds an immutable image,
opens one WebSocket session, measures reward values, observation schemas and
state continuity, and removes the container even if a check fails.

This first runtime slice implements startup, reward, observation and state
checks. Other applicable Level 2 checks appear explicitly as `SKIP`; they make
collects the original episode, repeats the plan in a fresh WebSocket session and
with a different seed, then replays it in a fresh container when the provider
supports that capability. Docker validation uses two containers concurrently
during the fresh-container replay. Both use the same immutable image and are
removed even if a check fails.

The seven implemented runtime checks cover startup, rewards, observation schemas,
state continuity, seed control, determinism and emitted trajectory records. Other
applicable Level 2 checks appear as `SKIP`; they make
the result `WARN`, which exits zero and does not mean Level 2 is complete.
`FAIL` exits 1, unsupported package formats exit 2, and internal or policy errors
exit 3. `--level semantic`
Expand All @@ -58,9 +62,14 @@ includes the available lower-level checks but does not claim semantic execution.
The Docker provider currently supports CPU workloads and `public` network mode;
unsupported network or GPU requirements skip runtime before building.
`--local` explicitly selects this default package mode and rejects a remote URL.
The declared `episode_timeout_s` bounds collection; a reset or judged step may use
its remaining budget. Collection failures appear once under `runtime.startup`,
with dependent contract checks skipped and the completed trace retained.
The declared `episode_timeout_s` bounds each episode; a reset or judged step may use
its remaining budget. Original and replay collection share a 300-second deadline
and a 32 MiB evidence budget. `llm_judged` requests 20 identical-input episodes,
including the original, plus the different-seed episode. Validation repeats the
actions and any external service calls they make, so budget for the additional
runtime and cost. Missing replay samples cannot pass determinism. Primary episode
collection failures appear under `runtime.startup`, with dependent contract checks
skipped and the completed trace retained.

Validation does not inherit host credentials, and the CLI currently has no secret
injection mechanism. Environments requiring a judge API key can therefore fail at
Expand All @@ -72,7 +81,8 @@ reset returns a different observation type; it defaults to the step observation
class. Older servers without a reset schema use the step schema for both.
The Level 2 profile permits null rewards on reset and nonterminal steps. Terminal
steps require numeric rewards, and every numeric reward must be finite and within
the declared range; boolean rewards are invalid.
the declared range; boolean rewards are invalid. Judged reward variance needs scored steps: entirely unscored episodes produce
`SKIP` for that check.

Cleanup removes run-owned containers. Built images remain in Docker's local cache
for reuse; the report records their immutable image IDs. Remove an unwanted image
Expand Down
19 changes: 16 additions & 3 deletions rfcs/008-environment-auto-validation.md
Original file line number Diff line number Diff line change
Expand Up @@ -584,7 +584,7 @@ limitations. Alternate DNS, unmatched addresses and unsupported address families
must fail closed. This requires a separate reviewed enforcement implementation;
parsing these declarations does not claim they are enforced.

### Later runtime evidence contracts
### Runtime replay evidence (PR5)

Seed acceptance and empirical determinism are separate findings. A reset that
silently drops its seed does not establish seed control, while a deterministic
Expand All @@ -594,11 +594,24 @@ policy-owned volatile metadata can be excluded. Authors cannot exclude fields.
For `llm_judged`, the bounded variance path uses 20 completed identical-input fresh
replays and population reward variance in reward-squared units, compared to the
declared bound. The total run budget bounds all samples; fewer than 20 is incomplete.
An unscored non-terminal step contributes no reward variance. Its null position must
agree across replays; a null/numeric mismatch is a divergence. Variance is measured
separately at each numeric step position, and an entirely unscored episode is
incomplete rather than passing. Terminal null rewards remain invalid.
This procedure is a runtime check, not a statistical confidence claim.

The initial implementation compares the baseline against a new session and an
independently inspected new container, and separately requests a different seed.
The judged sample count includes the completed baseline. Container identity must
change while image identity remains fixed. There are no volatile-field exclusions
in this version. Collection shares a 300-second deadline and retains at most
32 MiB across baseline and replay evidence. Missing samples, missing provider
capabilities and failed cleanup cannot produce a passing determinism finding.
`replays.json` preserves completed traces, telemetry, identity and cleanup outcomes.

Session telemetry for seed handling, named rubric/configuration, child attribution
and subject-emitted record references is orchestrator-only. A future protocol
slice must authorize access with an opt-in, random per-run/per-session capability
and subject-emitted record references is orchestrator-only. The protocol
authorizes access with an opt-in, random per-run/per-session capability
attached to the **same** replay connection, reject unauthorized/cross-session
reads and never expose telemetry as agent MCP tools. A second WebSocket creates
another environment and cannot supply evidence for the measured instance.
Expand Down
Loading
Loading