Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .claude/docs/REPO_WALKTHROUGH.md
Original file line number Diff line number Diff line change
Expand Up @@ -78,6 +78,7 @@ src/
│ │ ├── report.py # Versioned validation report models
│ │ ├── policy.py # Severity-policy loading and application
│ │ ├── runner.py # Local validation orchestration
│ │ ├── runtime/ # Bounded plans, launch requests and raw evidence
│ │ ├── parsers/ # Manifest parser registry
│ │ ├── providers/ # Validation provider contracts
│ │ ├── graders/ # Grader registry and protocols
Expand Down
431 changes: 431 additions & 0 deletions docs/plans/rfc008-level2-implementation.md

Large diffs are not rendered by default.

157 changes: 157 additions & 0 deletions rfcs/008-environment-auto-validation.md
Original file line number Diff line number Diff line change
Expand Up @@ -459,6 +459,163 @@ Stacked PRs, each vertically testable:
The PostTrain `task.md` parser and the LLM-judged grader path (variance-mode determinism, rubric
deepening) are fully specified with contracts and fixtures in PR2 and implemented separately.

## Level 2 execution amendment

This amendment specifies the missing execution contracts for the Docker-local
runtime implementation. It preserves the existing parser → normalized manifest →
grader → policy → report architecture. The first delivery wave consists of three
stacked slices: versioned contracts/shared assets; reproducible Docker supervision;
and a runnable CLI with startup, reward, observation and state checks. Later slices
add the remaining runtime checks. A partial implementation must expose the missing
requested checks as named SKIPs, never imply full Level 2 coverage from a PASS on
the implemented subset. Existing WARN exit semantics remain unchanged.

### Versioned declarations and public probe inputs

`validation.execution` opts a package into normalized manifest schema **2**:

```yaml
validation:
execution:
kind: openenv_ws
probe_path: validation/runtime.json
dockerfile: Dockerfile
context: .
agent_boundary: api
```

The other manifest sections remain authoritative for capabilities, resources,
network policy, reward bounds and grader applicability. The initial execution
binding is `openenv_ws` with API-only agent access. Process/filesystem agent
identities require a later explicit declaration; privileged provider exec is not
evidence of agent access. Paths are package-relative and portable. Readers reject
parent traversal, absolute paths and resolved symlink escapes.

Packages without `validation.execution` continue to produce the unchanged schema-1
manifest and static report. `NormalizedManifestV2` and `ValidationReportV2` have
separate committed JSON schemas. Report 2 accepts manifest 1, manifest 2, or null
after a parse failure so a runtime request can report missing prerequisites without
dropping diagnostics. No field is silently added to schema 1.

The public sidecar is data, not imported Python or a second capability manifest:

```json
{
"plan_schema_version": "1",
"reset": {"episode_id": "validation-probe", "seed": 42, "options": {}},
"actions": [{"increment": 1}, {"increment": 1}]
}
```

The trusted loader bounds the file at 65,536 bytes, JSON nesting at 32 levels,
node count at 10,000 and actions at 1–100. It rejects duplicate keys, non-finite
numbers, non-JSON objects and reset options overriding `seed` or `episode_id`.
The seed is an explicit unsigned 32-bit integer. Missing or invalid probe inputs
are reported as an unmet prerequisite or an input failure; they never deselect
graders. The plan's content digest belongs in the reproduction bundle. Neither
privileged oracle inputs nor host callbacks may be substituted for public actions.

One collector owns reset, ordered actions until termination, and state reads on
**one** WebSocket session. It records immutable raw request/response strings before
client defaults or Pydantic coercion can hide malformed responses. Graders consume
that evidence and cannot mutate the measured episode. The advertised observation
schema is recorded alongside the transcript; reconstruct observation plus the
separate reward/done envelope before validating it. Reset reward may be null; every
step reward must be a finite JSON number, excluding booleans, within the declared
range. State must retain the requested episode identity, reset its step count to
zero and advance it coherently for successful steps.

### Startup, policy and provider supervision

Policy **v2** retains every v1 entry and bound and adds `runtime.startup` at level 2,
local lane, failure severity. Runtime defaults to v2 and rejects an explicitly
incompatible policy before executing a subject. Static v1 remains supported.
Build, start or readiness failure caused by the subject yields startup FAIL with
the failed phase; a validator/provider defect yields ERROR. A missing prerequisite
or unsupported enforcement capability yields a named SKIP. One successful build
does not establish `static.reproducible_build`.

The validation-specific `LaunchSpec` names an immutable image ID/digest, requested
resources and network policy, an explicit run identity, startup deadline and
explicit environment variables. It inherits no host environment or credentials.
The provider owns build, readiness, bounded exec/logs, inspected settings and
idempotent cleanup on success, failure, timeout and cancellation. Cleanup evidence
must establish that no run-owned subject remains. Core provider ABCs are unchanged.

Launches use no privilege, host namespaces, host-directory mounts, Docker socket
or forwarded credentials. They drop Linux capabilities, enable no-new-privileges,
use a read-only image and explicitly bounded writable roots, and expose only a
loopback control port. CPU is an allocation ceiling; memory includes an explicit
swap setting. `disk_mb` means aggregate writable subject storage, excluding image
layers. The initial provider splits this allowance between bounded writable `/tmp`
and `/dev/shm`; another writable root requires its own accounted budget. Episode
deadlines are externally enforced and
cover descendants; PID, log and build limits are additional supervisor budgets.
Unsupported enforcement is disclosed and must never trigger a weaker retry.

The initial provider may support only `public` networking and CPU subjects.
`no-network`, `allowlist` and GPU requests must then be refused before launch with
their missing capability named. Public networking permits egress; it does not
establish a host/private-address deny policy. A later no-network provider uses a
trusted helper sharing only the subject network namespace, with separate image and
filesystem, no external interface, Docker socket or outbound-proxy API. Merely
publishing a port on a no-network container is insufficient for control access.

Allowlist semantics are **destination-address enforcement**: exact DNS hostnames,
`*.example.com` subdomain patterns (not the bare apex), and IPv4/IPv6 CIDRs resolve
through a controlled resolver; matching DNS requests add their recorded addresses
to the run-owned allow set. Direct addresses are allowed only if included in a CIDR
or the recorded set. Entries allow all ports/protocols at those destinations; they
do not authenticate TLS/HTTP host identity and therefore disclose shared-IP
limitations. Alternate DNS, unmatched addresses and unsupported address families
must fail closed. This requires a separate reviewed enforcement implementation;
parsing these declarations does not claim they are enforced.

### Later runtime evidence contracts

Seed acceptance and empirical determinism are separate findings. A reset that
silently drops its seed does not establish seed control, while a deterministic
environment may legitimately behave identically under different seeds. Replay
comparison includes observation, reward, done, tool output and state; only
policy-owned volatile metadata can be excluded. Authors cannot exclude fields.
For `llm_judged`, the bounded variance path uses 20 completed identical-input fresh
replays and population reward variance in reward-squared units, compared to the
declared bound. The total run budget bounds all samples; fewer than 20 is incomplete.
This procedure is a runtime check, not a statistical confidence claim.

Session telemetry for seed handling, named rubric/configuration, child attribution
and subject-emitted record references is orchestrator-only. A future protocol
slice must authorize access with an opt-in, random per-run/per-session capability
attached to the **same** replay connection, reject unauthorized/cross-session
reads and never expose telemetry as agent MCP tools. A second WebSocket creates
another environment and cannot supply evidence for the measured instance.
A validator transcript alone cannot pass subject-emitted trajectory recording.
These authorization requirements do not add new public wire messages in this slice.

Applicability predicates must distinguish empty declared sets from absent
capabilities. Missing subject features, missing provider support and checks whose
implementation has not shipped are distinct outcomes. Independent graders must
not step a shared environment; invasive containment and resource probes use
disposable subjects. Episode-isolation checks require same-server reset evidence,
not just a new container, and oracle access checks use the declared agent boundary.

### Shared reproducibility and review assets

`tests/fixtures/validation/runtime/cases.json` is the versioned acceptance catalog.
Every runtime policy ID has positive/negative expectations, applicability, required
provider features and its implementation slice. Future cases are labelled planned;
they do not count as measured coverage. A small shared served subject supplies
deterministic public actions and test-only faults. Preserve the existing static
fixtures and run their schema-1 round trips alongside schema-2 tests.

The Docker/reference lane and local reproduction command use the same locked test
project, digest-pinned base, exact-revision wheel and case catalog. An evidence
bundle records wheel/source/plan/image identities, platform details, reports,
transcripts and cleanup findings. CI separately asserts inventory and expected
findings: a WARN exit 0 does not prove completeness. Product publish gating and
operator certification remain outside this feature; the test workflow only proves
the implementation's own acceptance cases.

## Explicitly out of scope

No hub/runner/queue/coordinator; no submission or auth APIs; no statistical-level implementations
Expand Down
8 changes: 6 additions & 2 deletions scripts/sync_validation_schemas.py
Original file line number Diff line number Diff line change
Expand Up @@ -25,12 +25,16 @@

def rendered_schemas():
"""Return {filename: rendered JSON text} for every exported schema."""
from openenv.validation.manifest import NormalizedManifest
from openenv.validation.report import ValidationReport
from openenv.validation.manifest import NormalizedManifest, NormalizedManifestV2
from openenv.validation.report import ValidationReport, ValidationReportV2
from openenv.validation.runtime.contracts import RuntimePlan

exports = {
"manifest.schema.json": NormalizedManifest,
"report.schema.json": ValidationReport,
"manifest-v2.schema.json": NormalizedManifestV2,
"report-v2.schema.json": ValidationReportV2,
"runtime-plan.schema.json": RuntimePlan,
}
return {
fname: json.dumps(model.model_json_schema(), indent=2, sort_keys=True) + "\n"
Expand Down
82 changes: 81 additions & 1 deletion src/openenv/validation/manifest.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,9 +5,11 @@
report provenance only — grader selection never reads it.
"""

import math
from pathlib import PurePosixPath
from typing import Annotated, Any, Literal

from pydantic import BaseModel, ConfigDict, Field, model_validator
from pydantic import BaseModel, ConfigDict, Field, field_validator, model_validator

from .types import SignatureKind

Expand Down Expand Up @@ -304,3 +306,81 @@ def _judge_pin_iff_llm_judged(self) -> "NormalizedManifest":
"a judge pin is declared but capabilities.llm_judged is false"
)
return self


class ExecutionDeclaration(BaseModel):
"""
Author-declared runtime binding, introduced by manifest schema 2.

All paths are portable package-relative paths. Resolving the probe or build
context must additionally reject symlink escapes before reading source files.
The data-only probe supplies actions; it cannot change declared capabilities.

Attributes:
kind (`str`):
The implemented transport binding, `"openenv_ws"`.
probe_path (`str`):
JSON file containing the versioned runtime plan.
dockerfile (`str`):
Dockerfile relative to the package root.
context (`str`):
Build context relative to the package root.
agent_boundary (`str`):
The access granted to an agent; currently only API access is supported.
"""

model_config = ConfigDict(extra="forbid", frozen=True)

kind: Literal["openenv_ws"] = "openenv_ws"
probe_path: str = "validation/runtime.json"
dockerfile: str = "Dockerfile"
context: str = "."
agent_boundary: Literal["api"] = "api"

@field_validator("probe_path", "dockerfile", "context")
@classmethod
def _package_relative(cls, value: str) -> str:
path = PurePosixPath(value)
if (
not value
or "\\" in value
or ":" in value
or "\x00" in value
or path.is_absolute()
or ".." in path.parts
):
raise ValueError("must be a portable package-relative path without '..'")
return value

@field_validator("probe_path", "dockerfile")
@classmethod
def _file_path(cls, value: str) -> str:
if PurePosixPath(value) == PurePosixPath("."):
raise ValueError("must name a file")
return value


class NormalizedManifestV2(NormalizedManifest):
"""
Version 2 adds the runtime execution declaration without changing schema 1.

Static packages without `validation.execution` continue to normalize to
[`~openenv.validation.manifest.NormalizedManifest`]. An explicit execution
declaration opts into this model and the corresponding version 2 report.
"""

manifest_schema_version: Literal["2"]
execution: ExecutionDeclaration

@model_validator(mode="after")
def _finite_declarations(self) -> "NormalizedManifestV2":
pending = [self.model_dump()]
while pending:
value = pending.pop()
if isinstance(value, float) and not math.isfinite(value):
raise ValueError("schema 2 numeric declarations must be finite")
if isinstance(value, dict):
pending.extend(value.values())
elif isinstance(value, (list, tuple)):
pending.extend(value)
return self
13 changes: 9 additions & 4 deletions src/openenv/validation/parsers/openenv_yaml.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@
import yaml
from pydantic import ValidationError

from ..manifest import ManifestError, NormalizedManifest
from ..manifest import ManifestError, NormalizedManifest, NormalizedManifestV2
from ..types import SignatureKind

VALIDATION_BLOCK_REMEDIATION = (
Expand Down Expand Up @@ -41,7 +41,7 @@ class OpenEnvYamlParser:

signature = SignatureKind.OPENENV_SERVED

def parse(self, package_root: Path) -> NormalizedManifest:
def parse(self, package_root: Path) -> NormalizedManifest | NormalizedManifestV2:
"""
Parse a served-environment package.

Expand Down Expand Up @@ -69,8 +69,9 @@ def parse(self, package_root: Path) -> NormalizedManifest:
if not isinstance(validation, dict):
raise ManifestError(["`validation:` must be a mapping"])

has_execution = "execution" in validation
data: dict = {
"manifest_schema_version": "1",
"manifest_schema_version": "2" if has_execution else "1",
"signature": SignatureKind.OPENENV_SERVED,
"version": raw.get("version"),
"judge": validation.get("judge"),
Expand All @@ -87,8 +88,12 @@ def parse(self, package_root: Path) -> NormalizedManifest:
if value is not None:
data[key] = value

if has_execution:
data["execution"] = validation["execution"]

try:
return NormalizedManifest.model_validate(data)
model = NormalizedManifestV2 if has_execution else NormalizedManifest
return model.model_validate(data)
except ValidationError as exc:
raise ManifestError(
_format_validation_error(exc),
Expand Down
Loading
Loading