Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions PROGRAM_STATUS.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@

**Coverage: 21/21 expected repositories; duplicates: 0.**

Last verified: `2026-08-31T02:59:01Z`. HEADs are the verified default-branch revisions, not an assumption about later changes.
Last verified: `2026-08-31T12:52:33Z`. HEADs are the verified default-branch revisions, not an assumption about later changes.

## Portfolio totals

Expand Down Expand Up @@ -43,7 +43,7 @@ Program state totals: active 5; waiting 16; complete 0.
| 15 | [agent-recovery-policy](https://github.com/opsle/agent-recovery-policy) | concept | `1b733a111e26` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-recovery-policy/blob/1b733a111e26e0a409fee3b96f627048531daefe/THEORY.md); placeholder only; no automated tests | No shared failure schema, attempt ledger, route evaluator, or comparative fixture set. | After decision evidence and route schemas stabilize, define same-failure convergence on synthetic failures. | `agent-routing-policy`, `agent-state-ledger`, `decision-evidence-protocol` | waiting |
| 16 | [ephemeral-agent-workers](https://github.com/opsle/ephemeral-agent-workers) | concept | `ad96fcfdfac0` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/ephemeral-agent-workers/blob/ad96fcfdfac06d340b5e96d369634980cee78ef4/THEORY.md); placeholder only; no automated tests | Portable authority, claim, and handoff contracts are not ready; no safe synthetic containment harness exists. | Wait for prerequisite contracts, then define a fake worker adapter and destruction receipt without infrastructure changes. | `agent-execution-authorization`, `agent-resource-claims`, `verifiable-agent-handoff` | waiting |
| 17 | [gearbox](https://github.com/opsle/gearbox) | concept | `f3fab9f292cf` | `PROTOTYPED` | [provider-free Python reference core with strict authority-policy admission, exact deterministic argv execution, content-addressed staged helper context, injected one-shot helper transport, passive process waiting, compact results, raw-artifact accounting, fail-closed budgets, and Visible Value receipts](https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/src/opsle_gearbox/core.py); 19 of 19 provider-free automated tests passed locally, in PR #1 CI, and in final-main CI; ruff, shellcheck, actionlint, gitleaks, wheel build, receipt validation, and public raw-locator/hash checks passed | A production-quality bounded helper transport, independently verified isolation and termination, full Context Firewall integration, and a frozen comparative benchmark remain missing. | Freeze a provider-free deterministic-versus-direct baseline and helper-transport conformance corpus before considering any live model/provider run. | `context-firewall`, `decision-evidence-protocol`, `agent-trajectory-profiler`, `agent-routing-policy`, `agent-execution-authorization` | waiting |
| 18 | [affected-verification](https://github.com/opsle/affected-verification) | concept | `641aee9d29a8` | `VERIFIED` | [dependency-free Node.js deterministic planner plus an AV-EXP-001 benchmark-only Git/catalog/source-graph/Vitest adapter, identity-bound SHADOW result validator, complete frozen-oracle harness, explainable skip records, fail-closed uncertainty handling, and opsle.value-receipt.v1 telemetry](https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/REPORT.md); 70 of 70 automated tests, 15 of 15 conformance scenarios, and 6 of 6 determinism checks passed locally, in PR #2 CI, and from a fresh detached worktree at exact main; invalid-state coverage includes target, patch, catalog, selector, adapter, baseline, oracle, scenario, skip-reason, policy, trust-state, and result tampering | A second ecosystem, historical real-change replay, production-quality evidence adapter, and independent qualifying replication remain missing; AV therefore remains OBSERVE/SHADOW and no TRUSTED_BOUNDED change class is authorized. | Preregister and run a second public-repository shadow calibration in a different ecosystem with a meaningful native selector and the same full-catalog oracle discipline. | — | active |
| 18 | [affected-verification](https://github.com/opsle/affected-verification) | concept | `3ff41688dded` | `VERIFIED` | [dependency-free Node.js deterministic planner plus benchmark-only Git/catalog/source-graph adapters for AV-EXP-001 Vitest and AV-EXP-002 Python/pytest-testmon, identity-bound SHADOW result validation, complete frozen-oracle harnesses, explainable skip records, fail-closed uncertainty handling, and opsle.value-receipt.v1 telemetry](https://github.com/opsle/affected-verification/blob/3ff41688dded6e96e65da7cc44fe2608cf86d073/benchmark/av-exp-002/REPORT.md); 86 of 86 automated tests, 15 of 15 conformance scenarios, and 7 of 7 determinism checks passed locally, in PR #3 CI, and from a fresh detached worktree at exact main; the AV-EXP-002 result validator also passed at exact main, and invalid-state coverage includes Python/toolchain, target, patch, catalog, selector state/version/output, dynamic/conftest uncertainty, baseline, oracle, scenario, skip, policy, trust-state, and result tampering | AV-EXP-002 observed a false targeted-sufficiency claim for runtime/subprocess import behavior; historical real-change replay, a production-quality evidence adapter, and independent qualifying replication also remain missing. AV remains OBSERVE/SHADOW and no TRUSTED_BOUNDED change class is authorized. | Preregister and execute a selection-miss repair for AV2-006 that represents runtime/subprocess import uncertainty without changing the preserved AV-EXP-002 result. | — | active |
| 19 | [research](https://github.com/opsle/research) | program infrastructure | `9ee43197880c` | `PROTOTYPED` | [authoritative 21-repository ledger, machine-readable 18-concept theory registry including Affected Verification, canonical theory map, normative Visible Value controls, and provider-free EXP-001 benchmark, launch, one-block coordinator, external four-label LIVE_PROVIDER_RUN authorization, and current catalogue/pricing preflight artifacts with six content-addressed tasks, deterministic oracle, four arm contracts, sealed blinded allocation, exact subject configuration and adapter, exact authorization admission, private boundaries, receipts, mutation tests, and integrity CI](program/THEORY_MAP.md); 88 of 88 repository tests pass locally after deliberate migration to the 21-repository, 18-concept anti-forgetting set, including 13 authorization validations and two byte-identical replays; generated status and registry validation pass | The exact live authorization set remains unconsumed and unreleased, account-specific API entitlement is unverified under the zero-provider-call policy, no immutable dated model snapshot is documented, and the program has no canonical measured concept experiment. | Independently review and release the provider-free live-authorization and catalogue/pricing preflight; do not consume authorization or launch a provider/model subject. | — | active |
| 20 | [site](https://github.com/opsle/site) | program infrastructure | `28ad65be4750` | `PROTOTYPED` | [React/Vinext source implementation with content routes](https://github.com/opsle/site/blob/28ad65be4750dc849976fbf5c9eae9501c6bbb25/README.md); automated build/render tests present; not rerun because this reconciliation kept other repositories read-only | Wait for validated registry data and measured research; deployment requires separate authorization. | After registry merge, add a read-only registry ingestion design without deploying the site. | `research` | waiting |
| 21 | [.github](https://github.com/opsle/.github) | program infrastructure | `01c38e726db7` | `THEORY` | [documentation-only organization profile](https://github.com/opsle/.github/blob/01c38e726db7c3e45059d25fccce55e071e35938/profile/README.md); not applicable to current single Markdown profile; consistency is unverified | No mechanical registry consistency check exists in this repository. | After registry merge, design a read-only consistency check for organization-profile repository links. | `research` | waiting |
Expand Down
115 changes: 114 additions & 1 deletion program/experiments.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"schema_version": 1,
"last_verified_at": "2026-08-31T02:59:01Z",
"last_verified_at": "2026-08-31T12:52:33Z",
"experiments": [
{
"id": "EXP-001",
Expand Down Expand Up @@ -343,6 +343,119 @@
"lifecycle_impact": "PROMOTE_TO_VERIFIED_ONLY: the narrow scoped correctness and safety claims pass meaningful automated and revision-bound benchmark checks, but this run is explicitly capped below BENCHMARK_READY and does not authorize trusted execution.",
"next_task": "Preregister and run a second public-repository shadow calibration in a different ecosystem with a meaningful native selector and the same full-catalog oracle discipline."
},
{
"id": "AV-EXP-002",
"title": "Cross-Ecosystem Minimum Defensible Verification Shadow Calibration",
"status": "RECORDED",
"hypothesis": "On a pinned Python repository with an established ecosystem affected-test selector, Affected Verification preserves every oracle-relevant frozen catalog check while proposing less than FULL and broadening when evidence is incomplete.",
"participating_repositories": [
"affected-verification",
"research"
],
"roles": {
"primary": "affected-verification",
"expected_support": [
"research"
],
"potential_support": []
},
"baseline": "The complete frozen 2,024-check Click verification catalog at 36baa15ff831b939a22bc527cd76ce653ef6f66d, containing 2,016 pytest nodes and eight non-test checks, executed for every scenario and authoritative over all selector predictions.",
"experimental_arms": [
"FULL frozen verification catalog",
"ECOSYSTEM_SELECTOR using pytest-testmon 2.2.0 under its test-selection contract",
"AV_CORE with normalized Git, Python import graph, pytest catalog, verification catalog, and policy evidence",
"AV_WITH_SELECTOR_EVIDENCE with pytest-testmon output as an additional normalized evidence source"
],
"primary_metric": "Selection misses reported individually against checks whose full-catalog outcome changed because of a frozen scenario.",
"secondary_metrics": [
"relevant-check recall and scenario-level misses",
"exact selected and skipped pytest nodes, test files, and non-test checks by compatible class",
"uncertainty broadening and full-verification escalation",
"observed wall-clock telemetry without a causal time-saved claim",
"static-graph versus runtime selector compensation",
"normalized comparison to AV-EXP-001 without aggregating incompatible units"
],
"correctness_gate": "Every frozen catalog check runs for every scenario after selector and AV proposals are frozen; FULL remains authoritative, and a miss is any omitted oracle-relevant failing check.",
"failure_classifications": [
"outside selector contract",
"dependency evidence miss",
"verification-class omission",
"policy omission",
"adapter defect",
"planner defect",
"oracle or harness defect",
"unresolved",
"conservative broadening",
"insufficient evidence"
],
"dataset_fixture_identity": "Click 36baa15ff831b939a22bc527cd76ce653ef6f66d; preregistration commit f8a183c460535f3352fad2fb4990b0c54818d623; catalog sha256:28ed20abf60e7c785052308298dc6ed647b7a20a737513e8fb3c76aa62d9094c; corpus sha256:d5bc43405a5ab0ac34feef6d5fd5df111eace7f1a355400f963f8a7d4399640b; eleven frozen scenarios and patches.",
"model_provider_configuration": "NONE: no model/provider benchmark subject or external provider workload was used; one interactive Codex session used native shell and patch facilities without child agents.",
"run_identities": [
"sha256:5b3f99bfbebd3a0d061651d66adfb5a6aaef899475c6267cb20e4040e6ed5768"
],
"result_artifacts": [
"https://github.com/opsle/affected-verification/blob/3ff41688dded6e96e65da7cc44fe2608cf86d073/benchmark/av-exp-002/REPORT.md",
"https://github.com/opsle/affected-verification/blob/3ff41688dded6e96e65da7cc44fe2608cf86d073/benchmark/av-exp-002/results-v1/summary.json",
"https://github.com/opsle/affected-verification/blob/3ff41688dded6e96e65da7cc44fe2608cf86d073/benchmark/av-exp-002/results-v1/cross-experiment.json",
"https://github.com/opsle/affected-verification/blob/3ff41688dded6e96e65da7cc44fe2608cf86d073/benchmark/av-exp-002/results-v1/evidence-manifest.json"
],
"target": {
"repository": "https://github.com/pallets/click.git",
"sha": "36baa15ff831b939a22bc527cd76ce653ef6f66d",
"license": "BSD-3-Clause"
},
"preregistration": {
"commit_sha": "f8a183c460535f3352fad2fb4990b0c54818d623",
"identity": "sha256:93743e15caee647de1807ee27d35518cd9520394983803e1b93a58e9a2841db6",
"amendment_count": 5,
"comparative_outcomes_observed_before_commit": false
},
"benchmark_result": {
"affected_verification_main_sha": "3ff41688dded6e96e65da7cc44fe2608cf86d073",
"results_commit_sha": "5d126e0ae17557065b55ab84a46f6a3577a49989",
"summary_identity": "sha256:5b3f99bfbebd3a0d061651d66adfb5a6aaef899475c6267cb20e4040e6ed5768",
"evidence_bundle_identity": "sha256:435d8e5356ed6868edfc1747523ff75d1b327389bf14cc2675e867e45f4de705",
"cross_experiment_identity": "sha256:6ac2695ee80c9cb711e756846dad4c7138a3177c11d733456636259c36948da0",
"selector_baseline_identity": "sha256:9a3eb5630a868869179585210e9e6034e0c16039d97874921a09f63ef61aafd8",
"scenario_count": 11,
"synthetic_fault_count": 7,
"synthetic_benign_change_count": 1,
"uncertainty_scenario_count": 3,
"relevant_check_count": 87,
"ecosystem_selector_selected_relevant_check_count": 77,
"ecosystem_selector_missed_relevant_check_count": 10,
"ecosystem_selector_scenario_miss_count": 7,
"av_core_selected_relevant_check_count": 86,
"av_core_missed_relevant_check_count": 1,
"av_core_scenario_miss_count": 1,
"av_core_full_broadening_count": 6,
"av_with_selector_selected_relevant_check_count": 86,
"av_with_selector_missed_relevant_check_count": 1,
"av_with_selector_scenario_miss_count": 1,
"av_with_selector_full_broadening_count": 5,
"av_miss": {
"scenario_id": "AV2-006",
"check_id": "pytest:tests/test_imports.py::test_light_imports",
"classification": "PLANNER_OR_ADAPTER_MISS",
"finding": "A subprocess instrumented runtime imports through the public click package; the static Python graph declared completeness and pytest-testmon also omitted the relevant node."
}
},
"major_findings": [
"Catalog, policy escalation, explicit uncertainty, skip evidence, FULL fallback, shadow classification, and Visible Value generalized without changing the core plan schema.",
"Python dependency evidence, pytest node identity, conftest coupling, selector database lifecycle, and dynamic or subprocess imports required material ecosystem-specific logic.",
"Runtime selector evidence compensated for static dynamic-plugin uncertainty in AV2-009 but could not repair AV2-006 because the selector also omitted the relevant runtime-import test.",
"Both deliberately degraded cases denied aggressive skipping and required FULL."
],
"replication_status": "SAME_HOST_RELEASE_AND_BUNDLE_VERIFICATION_ONLY",
"verdict": "FAIL for the safety hypothesis: AV_CORE and AV_WITH_SELECTOR_EVIDENCE each selected 86/87 oracle-relevant checks and omitted the same runtime/subprocess import test in AV2-006. The controlled shadow calibration itself completed and FULL exposed the miss.",
"blockers": [
"The AV2-006 runtime/subprocess import miss invalidates a positive targeted-sufficiency claim for the frozen corpus.",
"The Python and pytest-testmon adapters are benchmark-only and no historical real-change replay or independent qualifying replication exists.",
"Affected Verification remains OBSERVE/SHADOW; no TRUSTED_BOUNDED class is authorized."
],
"lifecycle_impact": "REMAIN_VERIFIED: the run adds revision-bound falsification evidence and failure modes, but the observed selection miss and explicit run cap do not establish BENCHMARK_READY or EXPERIMENTED lifecycle promotion for the repository.",
"next_task": "Preregister and execute a selection-miss repair for AV2-006 that represents runtime/subprocess import uncertainty without changing the preserved AV-EXP-002 result."
},
{
"id": "LEGACY-001",
"title": "Graphify plus Antigravity semantic adapter integration observation",
Expand Down
Loading