From b35550217dc27d45120e4cc3d91e50c614ddcb90 Mon Sep 17 00:00:00 2001 From: muratkeremozcan Date: Tue, 22 Sep 2026 10:21:38 -0500 Subject: [PATCH 1/4] docs: redesign walkthrough pipeline map and clarify execution boundaries - Replace competing conceptual diagrams with ONE canonical numbered pipeline map matching sections 1 through 7. - Visibly distinguish executed CLI stages from read-only inspection sections. - Add explicit callouts for caller/harness execution boundaries and the two skipped execution gaps (preflight active probing and SUT/evaluator execution). - Rename Step 5 to 'Reduce the preflight observations' and clarify that the CLI command executes preflightFromObservations over prepared fixtures. - Rename Step 6 to 'Score the evaluation record' and make explicit that score evaluates prepared evidence from a prior run. - Update the artifact map to distinguish newly minted artifacts from replayed prepared fixtures. - Update summary flow in key takeaways to mirror the numbered pipeline. --- docs/how-to/author-behavioral-contracts.md | 213 ++++++++++----------- 1 file changed, 106 insertions(+), 107 deletions(-) diff --git a/docs/how-to/author-behavioral-contracts.md b/docs/how-to/author-behavioral-contracts.md index fcf9c18a..ccf819e7 100644 --- a/docs/how-to/author-behavioral-contracts.md +++ b/docs/how-to/author-behavioral-contracts.md @@ -64,95 +64,102 @@ Every digest inside it is computed from bytes this repository ships. ## How this walkthrough fits together -[How It Works](/explanation/behavioral-evaluation-contracts/) described the [four-layer model](/explanation/behavioral-evaluation-contracts/#four-things-and-the-boundaries-between-them) conceptually. -This walkthrough runs those stages over one real contract and lets you inspect the artifact produced at each point. +[How It Works](/explanation/behavioral-evaluation-contracts/) describes the conceptual model connecting the System Under Test (SUT), the evaluation harness, and `eval-quality`. +This hands-on walkthrough takes you through one complete scored evaluation arm, step by step, using the seven numbered sections below. -```text -Behavioral Evaluation Contract - ↓ - COMPILE - Is the design valid? - ↓ - SEAL - Create evaluator-safe brief - ↓ - PREFLIGHT - Is the environment measurable? - ↓ - EVALUATION - caller/evaluator produces - observations + findings - ↓ - SCORE - Does evidence support - the evaluator's claims? - ↓ - Evidence Artifact - verdict + outcomes + strength -``` - -### Artifact map - -This walkthrough generates four outputs on disk while replaying prepared caller evidence and configuration files. - -| Artifact | Producer in this exercise | Job | -| --- | --- | --- | -| `contract.json` | Committed author input | Defines behavior, interfaces, evidence relationships, and checks. | -| `eval-contract.json` | `compile` | Validated, canonical contract consumed by later stages. | -| `sealed-evaluator-brief.json` | `seal` | Evaluator-facing directions and permitted context. | -| `probes.json` + `observations.json` | Committed caller inputs | Inputs for this preflight reduction. | -| `preflight-verdict.json` | `preflight` | Records which measurability checks were satisfied. | -| `probe.json` + `sealed-run-record.json` | Committed probe and trial inputs | Known failure description plus what the evaluator observed and reported. | -| Policy, configuration, isolation manifest, corpus digest | Committed caller inputs and attested value | Thresholds, run context, and identity or integrity information. | -| `evidence-artifact.json` | `score` | Scored outcomes, verdict, reasons, trials, and strength. | - -The separately generated `piped-brief.json` is an equivalence verification rather than an additional stage. - -Be explicit about ownership: +### The walkthrough pipeline -```text -eval-quality: -compile -seal -preflight reduction -score - -caller / harness: -run the SUT -run the evaluator -collect observations -produce the sealed run record -``` - -The walkthrough uses committed caller-produced artifacts where appropriate. - -The full [twin-run model](/explanation/behavioral-evaluation-contracts/#the-twin-run) compares a clean system with a deliberately mutated system. -This walkthrough takes you through one scored arm in detail so you can first understand `compile`, `seal`, `preflight`, and `score`. -The later section shows how the same chain is repeated for both clean and mutated arms. +The numbered sections of this guide correspond to a single, continuous pipeline: ```text -Full twin run + Authored Behavioral Evaluation Contract + (examples/tutorials/walkthrough/contract.json) + │ + ▼ + 1. COMPILE THE CONTRACT + [eval-quality CLI] + │ + ├── 2. Inspect what the contract declares + │ [Read-only contract explanation] + │ + ├── 3. Inspect a compile rejection + │ [Read-only discipline explanation] + │ + ▼ + 4. SEAL THE BRIEF + [eval-quality CLI] + │ + ▼ + ┌──────────────────────────────────────────────┐ + │ Caller / Harness Boundary: Preflight Probing │ + │ (The SUT is not launched in this tutorial.) │ + │ Preflight requests were executed beforehand │ + │ to produce prepared probes + observations. │ + └──────────────────────────────────────────────┘ + │ + ▼ + 5. REDUCE PREFLIGHT OBSERVATIONS + [eval-quality CLI] + │ + ▼ + ┌──────────────────────────────────────────────┐ + │ Caller / Harness Boundary: SUT + Evaluator │ + │ (Neither SUT nor LLM runs in this tutorial.) │ + │ The defective SUT and evaluator were run in │ + │ a harness to produce a sealed run record. │ + └──────────────────────────────────────────────┘ + │ + ▼ + 6. SCORE THE EVALUATION RECORD + [eval-quality CLI] + │ + ▼ + Evidence Artifact + (/tmp/eval-quality-run/evidence-artifact.json) + │ + ▼ + 7. READ THE RUN YOU PRODUCED + [Read-only verdict & strength triage] +``` + +### Pipeline ownership and the two skipped execution gaps + +To follow the walkthrough without confusion, keep in mind who executes what: + +* **`eval-quality` owns four deterministic offline processing stages:** + `compile`, `seal`, `preflight` (evidence reduction), and `score`. None of these CLI commands launch background processes, spin up servers, query AI models, or issue network calls. +* **Your harness owns active environment and evaluator execution:** + Running the System Under Test (SUT), hosting services, dispatching active probe calls to verify measurability, running the evaluator model against the sealed brief, capturing interaction observations, and minting the sealed run record. + +Because this tutorial focuses on learning the `eval-quality` contract and verification tools, it uses **prepared fixture files** in place of dynamic harness execution. There are two explicit execution gaps: + +1. **Preflight probing gap (before step 5):** + In this tutorial, `eval-quality preflight` does not start the Notes API service or send HTTP requests. It executes `preflightFromObservations`: it reads the planned probe legs from the contract and checks them against the responses recorded in `observations.json`. In production, when using the TypeScript library, `runPreflight` can actively drive a caller-supplied `EnvironmentProbePort` to send planned legs directly to a running service. +2. **Evaluator and SUT execution gap (before step 6):** + In this tutorial, neither the defective Notes API nor an LLM evaluator runs live. In a live system, your harness executes the defective SUT, presents `sealed-evaluator-brief.json` to the evaluator, captures the resulting tool calls and responses, and seals them into `sealed-run-record.json`. Here, that work was executed in advance, and the resulting record is committed in `examples/tutorials/walkthrough/sealed-run-record.json` for replay. + +### How this maps to a full twin run + +The full [twin-run model](/explanation/behavioral-evaluation-contracts/#the-twin-run) stress-tests an evaluation by comparing two conditions: a clean system and a mutated system with a seeded defect. + +This walkthrough executes **one scored arm** in depth (the defective SUT arm) so you can understand the artifacts and decision rules firsthand. Once you master this sequence, repeating it across both arms is straightforward: you compile and seal the contract once, run preflight reduction on each arm, and score each arm against its own probe ([Next: run a real clean and mutated experiment](#next-run-a-real-clean-and-mutated-experiment)). - Clean SUT Mutated SUT - ↓ ↓ - Evaluation Evaluation - ↓ ↓ - SCORE SCORE - └──────────┬──────────────┘ - ↓ - Did the evaluation discriminate? +### Artifact map +This walkthrough generates four new artifacts on disk while consuming committed inputs and replaying prepared harness evidence: -This walkthrough +| Artifact | Role in this exercise | Job | +| --- | --- | --- | +| `contract.json` | Committed authored input | Defines behavior, interfaces, evidence relationships, and checks. | +| `eval-contract.json` | **Generated by step 1 (`compile`)** | Validated, canonical contract consumed by later stages. | +| `sealed-evaluator-brief.json` | **Generated by step 4 (`seal`)** | Evaluator-facing directions and permitted context (withholds answer key). | +| `probes.json` + `observations.json` | Committed prepared fixtures | Replayed inputs representing preflight probing interactions. | +| `preflight-verdict.json` | **Generated by step 5 (`preflight`)** | Records which environment measurability checks were satisfied. | +| `probe.json` + `sealed-run-record.json` | Committed prepared fixtures | Seeded defect description plus evaluator observations and findings. | +| Policy, configuration, isolation manifest, corpus digest | Committed inputs and attested digest | Thresholds, run controls, and cryptographic integrity tokens. | +| `evidence-artifact.json` | **Generated by step 6 (`score`)** | Scored outcomes, verdict, reasons, trials, and strength vector. | - Known defective SUT - ↓ - Evaluation - ↓ - SCORE - ↓ - Read the evidence -``` +The separately generated `piped-brief.json` in step 4 is an equivalence check verifying that streaming through standard I/O matches file-based execution byte for byte. ## 1. Compile the contract @@ -406,24 +413,23 @@ identical > `--in` left out reads stdin, which is what makes the pipe work. > `-` names stdin explicitly, and at most one input per command may be `-`. -## 5. Preflight the environment +## 5. Reduce the preflight observations > **Question:** Can this environment actually produce the evidence the contract depends on? -`preflight` answers one question: is the environment fit to be measured? +Conceptually, preflight determines whether the environment is measurable before running an expensive evaluation. + +In this walkthrough, the target Notes API service is not launched, and the CLI command issues zero network requests. The external probe interactions were recorded beforehand into a prepared fixture. The CLI `preflight` command executes `preflightFromObservations`: it plans the probe legs the contract implies, compares them against the supplied observations, and mints a `PreflightVerdict` for a named run. -It plans the probe legs the contract implies, reduces the observations you hand it, and mints a verdict for a named run. -The command issues no requests of its own. -The observations come from whatever system called the target. -The library's `runPreflight` can drive a caller-supplied `EnvironmentProbePort` instead. +When running programmatically via the library API, `runPreflight` can actively drive a caller-supplied `EnvironmentProbePort` (such as an HTTP or CLI adapter) to probe an active service directly. ### Preflight inputs -Preflight takes two files beyond the contract: +This reduction takes two prepared files beyond the contract: - **`probes.json`** is the probe list the plan builds from. This chain seeds no faults that preflight must watch fire, so the list is empty (`[]`). -- **`observations.json`** contains what the environment answered, one entry per planned leg. +- **`observations.json`** contains what the environment answered during prior probing, one entry per planned leg. In contrast, the later scoring step consumes `probe.json`, which declares the seeded defect `P-001`. Because `probes.json` is empty in this preflight invocation, preflight evaluates sensitivity and control legs from the contract without evaluating `seeded-fault-fired` or `seeded-faults-scoped` checks for the defect. @@ -559,11 +565,12 @@ Environment fit to score YES Preflight does not decide whether the evaluation is good. It establishes that the environment is fit enough for the resulting evidence to mean something. -## 6. Score it +## 6. Score the evaluation record > **Question:** Does the recorded evidence support what the evaluator claimed? `score` chains `ingest`, `score`, and `emit` over a trial set and mints an evidence artifact carrying the verdict. +In this walkthrough, `score` evaluates the prepared evidence from `sealed-run-record.json` rather than executing the evaluator or defective SUT live. This walkthrough supplies one record, so its result records one completed trial. Repeat `--record` with independently sealed records to meet a multi-trial policy minimum. Every record carries a distinct `trialIndex`; every record agrees on `contractDigest`, `evaluatorConfigurationDigest`, `mode`, `evaluatorRecommendation`, and `runId`. @@ -834,23 +841,15 @@ Reading only the strength vector would provide an incomplete picture of contract * Verdict and strength evaluate distinct dimensions of an evaluation. ```text -BEC - ↓ -COMPILE - ↓ -SEAL - ↓ -PREFLIGHT - ↓ -EVALUATE - ↓ -SEALED RUN RECORD - ↓ -SCORE - ↓ -EVIDENCE ARTIFACT - ↓ -verdict + strength +1. COMPILE (validate contract specification) + ├── 2. Inspect declared behavior, interfaces, and oracles + └── 3. Inspect schema and discipline rule rejections +4. SEAL (mint evaluator-safe brief) + └── [Harness records preflight interactions] +5. REDUCE PREFLIGHT (verify environment measurability) + └── [Harness executes SUT + evaluator, seals run record] +6. SCORE (evaluate empirical evidence against probe) +7. READ RESULT (triage verdict reasons and strength vector) ``` --- From 7b4729c755edd1222a31fbece41da44f4da59736 Mon Sep 17 00:00:00 2001 From: muratkeremozcan Date: Tue, 22 Sep 2026 10:27:31 -0500 Subject: [PATCH 2/4] docs: address walkthrough review feedback on step labeling and fixture provenance --- docs/how-to/author-behavioral-contracts.md | 17 ++++++++++------- 1 file changed, 10 insertions(+), 7 deletions(-) diff --git a/docs/how-to/author-behavioral-contracts.md b/docs/how-to/author-behavioral-contracts.md index ccf819e7..84d0d160 100644 --- a/docs/how-to/author-behavioral-contracts.md +++ b/docs/how-to/author-behavioral-contracts.md @@ -65,7 +65,7 @@ Every digest inside it is computed from bytes this repository ships. ## How this walkthrough fits together [How It Works](/explanation/behavioral-evaluation-contracts/) describes the conceptual model connecting the System Under Test (SUT), the evaluation harness, and `eval-quality`. -This hands-on walkthrough takes you through one complete scored evaluation arm, step by step, using the seven numbered sections below. +This walkthrough replays one complete scored evaluation arm through the `eval-quality` stages, step by step, using the seven numbered sections below. ### The walkthrough pipeline @@ -83,7 +83,7 @@ The numbered sections of this guide correspond to a single, continuous pipeline: │ [Read-only contract explanation] │ ├── 3. Inspect a compile rejection - │ [Read-only discipline explanation] + │ [compile CLI demonstration] │ ▼ 4. SEAL THE BRIEF @@ -93,8 +93,10 @@ The numbered sections of this guide correspond to a single, continuous pipeline: ┌──────────────────────────────────────────────┐ │ Caller / Harness Boundary: Preflight Probing │ │ (The SUT is not launched in this tutorial.) │ - │ Preflight requests were executed beforehand │ - │ to produce prepared probes + observations. │ + │ The preflight interactions were executed │ + │ beforehand. Their responses are committed as │ + │ observations.json. probes.json is the │ + │ committed probe-list input to planning. │ └──────────────────────────────────────────────┘ │ ▼ @@ -134,7 +136,7 @@ To follow the walkthrough without confusion, keep in mind who executes what: Because this tutorial focuses on learning the `eval-quality` contract and verification tools, it uses **prepared fixture files** in place of dynamic harness execution. There are two explicit execution gaps: 1. **Preflight probing gap (before step 5):** - In this tutorial, `eval-quality preflight` does not start the Notes API service or send HTTP requests. It executes `preflightFromObservations`: it reads the planned probe legs from the contract and checks them against the responses recorded in `observations.json`. In production, when using the TypeScript library, `runPreflight` can actively drive a caller-supplied `EnvironmentProbePort` to send planned legs directly to a running service. + In this tutorial, `eval-quality preflight` does not start the Notes API service or send HTTP requests. It executes `preflightFromObservations`: it plans the legs implied by the contract and supplied probe list, then reduces the prepared observations. In production, when using the TypeScript library, `runPreflight` can actively drive a caller-supplied `EnvironmentProbePort` to send planned legs directly to a running service. 2. **Evaluator and SUT execution gap (before step 6):** In this tutorial, neither the defective Notes API nor an LLM evaluator runs live. In a live system, your harness executes the defective SUT, presents `sealed-evaluator-brief.json` to the evaluator, captures the resulting tool calls and responses, and seals them into `sealed-run-record.json`. Here, that work was executed in advance, and the resulting record is committed in `examples/tutorials/walkthrough/sealed-run-record.json` for replay. @@ -146,7 +148,7 @@ This walkthrough executes **one scored arm** in depth (the defective SUT arm) so ### Artifact map -This walkthrough generates four new artifacts on disk while consuming committed inputs and replaying prepared harness evidence: +This walkthrough generates four primary pipeline artifacts on disk while consuming committed inputs and replaying prepared harness evidence: | Artifact | Role in this exercise | Job | | --- | --- | --- | @@ -843,7 +845,7 @@ Reading only the strength vector would provide an incomplete picture of contract ```text 1. COMPILE (validate contract specification) ├── 2. Inspect declared behavior, interfaces, and oracles - └── 3. Inspect schema and discipline rule rejections + └── 3. Inspect a compile rejection (CLI demonstration) 4. SEAL (mint evaluator-safe brief) └── [Harness records preflight interactions] 5. REDUCE PREFLIGHT (verify environment measurability) @@ -880,6 +882,7 @@ eval-quality seal --in contract.json --out run/sealed-evaluator-brief.json ``` Preflight each arm. +Before invoking the CLI, have the harness execute the planned preflight interactions for each arm and record their responses as `clean-observations.json` and `mutated-observations.json`. The commands below reduce those prepared observations into preflight verdicts. An arm that does not pass exits `3` and stops there, because a measurement over an unfit environment says nothing about the contract: ```text From 81bdfd9e243fdeddc7e1c782169e439425b271d6 Mon Sep 17 00:00:00 2001 From: muratkeremozcan Date: Tue, 22 Sep 2026 10:39:29 -0500 Subject: [PATCH 3/4] docs: clarify that walkthrough fixtures are prepared deterministic replay artifacts --- docs/how-to/author-behavioral-contracts.md | 24 +++++++++++----------- 1 file changed, 12 insertions(+), 12 deletions(-) diff --git a/docs/how-to/author-behavioral-contracts.md b/docs/how-to/author-behavioral-contracts.md index 84d0d160..c2c0c1c6 100644 --- a/docs/how-to/author-behavioral-contracts.md +++ b/docs/how-to/author-behavioral-contracts.md @@ -93,10 +93,9 @@ The numbered sections of this guide correspond to a single, continuous pipeline: ┌──────────────────────────────────────────────┐ │ Caller / Harness Boundary: Preflight Probing │ │ (The SUT is not launched in this tutorial.) │ - │ The preflight interactions were executed │ - │ beforehand. Their responses are committed as │ - │ observations.json. probes.json is the │ - │ committed probe-list input to planning. │ + │ The tutorial builder supplies prepared │ + │ observations representing harness responses. │ + │ probes.json is the committed planning input. │ └──────────────────────────────────────────────┘ │ ▼ @@ -107,12 +106,13 @@ The numbered sections of this guide correspond to a single, continuous pipeline: ┌──────────────────────────────────────────────┐ │ Caller / Harness Boundary: SUT + Evaluator │ │ (Neither SUT nor LLM runs in this tutorial.) │ - │ The defective SUT and evaluator were run in │ - │ a harness to produce a sealed run record. │ + │ The tutorial uses a prepared run record │ + │ representing a defective-SUT/evaluator run. │ + │ Neither is executed by this repository. │ └──────────────────────────────────────────────┘ │ ▼ - 6. SCORE THE EVALUATION RECORD + 6. SCORE THE EVALUATION RECORD [eval-quality CLI] │ ▼ @@ -136,9 +136,9 @@ To follow the walkthrough without confusion, keep in mind who executes what: Because this tutorial focuses on learning the `eval-quality` contract and verification tools, it uses **prepared fixture files** in place of dynamic harness execution. There are two explicit execution gaps: 1. **Preflight probing gap (before step 5):** - In this tutorial, `eval-quality preflight` does not start the Notes API service or send HTTP requests. It executes `preflightFromObservations`: it plans the legs implied by the contract and supplied probe list, then reduces the prepared observations. In production, when using the TypeScript library, `runPreflight` can actively drive a caller-supplied `EnvironmentProbePort` to send planned legs directly to a running service. + In this tutorial, `eval-quality preflight` does not start the Notes API service or send HTTP requests. The tutorial builder supplies prepared preflight observations representing the responses a harness would collect. It executes `preflightFromObservations`: it plans the legs implied by the contract and supplied probe list, then reduces the prepared observations. In production, when using the TypeScript library, `runPreflight` can actively drive a caller-supplied `EnvironmentProbePort` to send planned legs directly to a running service. 2. **Evaluator and SUT execution gap (before step 6):** - In this tutorial, neither the defective Notes API nor an LLM evaluator runs live. In a live system, your harness executes the defective SUT, presents `sealed-evaluator-brief.json` to the evaluator, captures the resulting tool calls and responses, and seals them into `sealed-run-record.json`. Here, that work was executed in advance, and the resulting record is committed in `examples/tutorials/walkthrough/sealed-run-record.json` for replay. + In this tutorial, neither the defective Notes API nor an LLM evaluator runs live. In a live system, your harness executes the defective SUT, presents `sealed-evaluator-brief.json` to the evaluator, captures the resulting tool calls and responses, and seals them into `sealed-run-record.json`. Here, the tutorial uses a prepared sealed run record representing a defective-SUT/evaluator run; the SUT and evaluator are not executed by this repository. The resulting record is committed in `examples/tutorials/walkthrough/sealed-run-record.json` for deterministic replay. ### How this maps to a full twin run @@ -421,7 +421,7 @@ identical Conceptually, preflight determines whether the environment is measurable before running an expensive evaluation. -In this walkthrough, the target Notes API service is not launched, and the CLI command issues zero network requests. The external probe interactions were recorded beforehand into a prepared fixture. The CLI `preflight` command executes `preflightFromObservations`: it plans the probe legs the contract implies, compares them against the supplied observations, and mints a `PreflightVerdict` for a named run. +In this walkthrough, the target Notes API service is not launched, and the CLI command issues zero network requests. The tutorial builder supplies prepared preflight observations representing the responses a harness would collect. The CLI `preflight` command executes `preflightFromObservations`: it plans the probe legs the contract implies, compares them against the supplied observations, and mints a `PreflightVerdict` for a named run. When running programmatically via the library API, `runPreflight` can actively drive a caller-supplied `EnvironmentProbePort` (such as an HTTP or CLI adapter) to probe an active service directly. @@ -431,7 +431,7 @@ This reduction takes two prepared files beyond the contract: - **`probes.json`** is the probe list the plan builds from. This chain seeds no faults that preflight must watch fire, so the list is empty (`[]`). -- **`observations.json`** contains what the environment answered during prior probing, one entry per planned leg. +- **`observations.json`** contains the prepared preflight responses, one entry per planned leg, representing what a harness would collect from the environment. In contrast, the later scoring step consumes `probe.json`, which declares the seeded defect `P-001`. Because `probes.json` is empty in this preflight invocation, preflight evaluates sensitivity and control legs from the contract without evaluating `seeded-fault-fired` or `seeded-faults-scoped` checks for the defect. @@ -613,7 +613,7 @@ Before scoring, inspect the inputs that originate outside the four-stage CLI pip - `findings`: Finding `F-001` reports that the note kept its old title, citing probe `P-001`, oracle `O-001`, and observation `obs-002`. - `oracleDispositions`: Evaluator judgments for each oracle (`violated` for O-001; `held` for O-002, O-003, and O-004). - `evaluatorRecommendation`: Records `FAIL` for the system under test. -- This file is a prepared record from an earlier evaluation run. +- This file is a prepared record representing a defective-SUT and evaluator run; the SUT and evaluator are not executed by this repository. Scoring replays this evidence and does not launch a fresh evaluator. ```text From a15b56dae37b834616ba6857bb5b3c4046d57dbf Mon Sep 17 00:00:00 2001 From: muratkeremozcan Date: Tue, 22 Sep 2026 10:54:13 -0500 Subject: [PATCH 4/4] docs: address walkthrough review on live test scoping, twin run comparison advice, and replay consistency --- docs/how-to/author-behavioral-contracts.md | 16 ++++++++-------- 1 file changed, 8 insertions(+), 8 deletions(-) diff --git a/docs/how-to/author-behavioral-contracts.md b/docs/how-to/author-behavioral-contracts.md index c2c0c1c6..820beeca 100644 --- a/docs/how-to/author-behavioral-contracts.md +++ b/docs/how-to/author-behavioral-contracts.md @@ -107,8 +107,8 @@ The numbered sections of this guide correspond to a single, continuous pipeline: │ Caller / Harness Boundary: SUT + Evaluator │ │ (Neither SUT nor LLM runs in this tutorial.) │ │ The tutorial uses a prepared run record │ - │ representing a defective-SUT/evaluator run. │ - │ Neither is executed by this repository. │ + │ representing a defective-SUT/evaluator run; │ + │ neither is executed live by this walkthrough.│ └──────────────────────────────────────────────┘ │ ▼ @@ -138,13 +138,13 @@ Because this tutorial focuses on learning the `eval-quality` contract and verifi 1. **Preflight probing gap (before step 5):** In this tutorial, `eval-quality preflight` does not start the Notes API service or send HTTP requests. The tutorial builder supplies prepared preflight observations representing the responses a harness would collect. It executes `preflightFromObservations`: it plans the legs implied by the contract and supplied probe list, then reduces the prepared observations. In production, when using the TypeScript library, `runPreflight` can actively drive a caller-supplied `EnvironmentProbePort` to send planned legs directly to a running service. 2. **Evaluator and SUT execution gap (before step 6):** - In this tutorial, neither the defective Notes API nor an LLM evaluator runs live. In a live system, your harness executes the defective SUT, presents `sealed-evaluator-brief.json` to the evaluator, captures the resulting tool calls and responses, and seals them into `sealed-run-record.json`. Here, the tutorial uses a prepared sealed run record representing a defective-SUT/evaluator run; the SUT and evaluator are not executed by this repository. The resulting record is committed in `examples/tutorials/walkthrough/sealed-run-record.json` for deterministic replay. + In this tutorial, neither the defective Notes API nor an LLM evaluator runs live. In a live system, your harness executes the defective SUT, presents `sealed-evaluator-brief.json` to the evaluator, captures the resulting tool calls and responses, and seals them into `sealed-run-record.json`. Here, the tutorial uses a prepared sealed run record representing a defective-SUT/evaluator run; the SUT and evaluator are not executed by the commands in this walkthrough (which replays prepared evidence; separate repository tests exercise the Notes SUT live). The resulting record is committed in `examples/tutorials/walkthrough/sealed-run-record.json` for deterministic replay. ### How this maps to a full twin run The full [twin-run model](/explanation/behavioral-evaluation-contracts/#the-twin-run) stress-tests an evaluation by comparing two conditions: a clean system and a mutated system with a seeded defect. -This walkthrough executes **one scored arm** in depth (the defective SUT arm) so you can understand the artifacts and decision rules firsthand. Once you master this sequence, repeating it across both arms is straightforward: you compile and seal the contract once, run preflight reduction on each arm, and score each arm against its own probe ([Next: run a real clean and mutated experiment](#next-run-a-real-clean-and-mutated-experiment)). +This walkthrough replays **one scored arm** in depth (the defective SUT arm) so you can understand the artifacts and decision rules firsthand. Once you master this sequence, repeating it across both arms is straightforward: you compile and seal the contract once, run preflight reduction on each arm, and score each arm against its own probe ([Next: run a real clean and mutated experiment](#next-run-a-real-clean-and-mutated-experiment)). ### Artifact map @@ -613,7 +613,7 @@ Before scoring, inspect the inputs that originate outside the four-stage CLI pip - `findings`: Finding `F-001` reports that the note kept its old title, citing probe `P-001`, oracle `O-001`, and observation `obs-002`. - `oracleDispositions`: Evaluator judgments for each oracle (`violated` for O-001; `held` for O-002, O-003, and O-004). - `evaluatorRecommendation`: Records `FAIL` for the system under test. -- This file is a prepared record representing a defective-SUT and evaluator run; the SUT and evaluator are not executed by this repository. +- This file is a prepared record representing a defective-SUT and evaluator run; the SUT and evaluator are not executed by the commands in this walkthrough. Scoring replays this evidence and does not launch a fresh evaluator. ```text @@ -922,9 +922,9 @@ A clean target does not guarantee an overall PASS verdict, as other contract con On the mutated arm, the oracle the defect targets should resolve `caught`, which in `contract-scoring` mode indicates the contract succeeded. An oracle that resolves `missed` on the mutated arm indicates an unaddressed blind spot. -Compare `scoringVersion` across the two artifacts before comparing anything else in them. -That comparison is yours to make, and the library makes no such check. -[Contract strength](/explanation/contract-strength/) covers `compareDominance`, which is the comparison it does make. +Check `comparabilityKey` and `strength.comparable` before comparing strength. +Inspect `scoringVersion` differences to understand whether fixture, evaluator configuration, mode, or other declared experiment inputs changed. +[Contract strength](/explanation/contract-strength/) covers `compareDominance`, which performs the component-wise dominance comparison and enforces the severity-floor override. ## Two guards