Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions PROGRAM_STATUS.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@

**Coverage: 21/21 expected repositories; duplicates: 0.**

Last verified: `2026-08-31T12:52:33Z`. HEADs are the verified default-branch revisions, not an assumption about later changes.
Last verified: `2026-08-31T16:17:40Z`. HEADs are the verified default-branch revisions, not an assumption about later changes.

## Portfolio totals

Expand Down Expand Up @@ -43,7 +43,7 @@ Program state totals: active 5; waiting 16; complete 0.
| 15 | [agent-recovery-policy](https://github.com/opsle/agent-recovery-policy) | concept | `1b733a111e26` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-recovery-policy/blob/1b733a111e26e0a409fee3b96f627048531daefe/THEORY.md); placeholder only; no automated tests | No shared failure schema, attempt ledger, route evaluator, or comparative fixture set. | After decision evidence and route schemas stabilize, define same-failure convergence on synthetic failures. | `agent-routing-policy`, `agent-state-ledger`, `decision-evidence-protocol` | waiting |
| 16 | [ephemeral-agent-workers](https://github.com/opsle/ephemeral-agent-workers) | concept | `ad96fcfdfac0` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/ephemeral-agent-workers/blob/ad96fcfdfac06d340b5e96d369634980cee78ef4/THEORY.md); placeholder only; no automated tests | Portable authority, claim, and handoff contracts are not ready; no safe synthetic containment harness exists. | Wait for prerequisite contracts, then define a fake worker adapter and destruction receipt without infrastructure changes. | `agent-execution-authorization`, `agent-resource-claims`, `verifiable-agent-handoff` | waiting |
| 17 | [gearbox](https://github.com/opsle/gearbox) | concept | `f3fab9f292cf` | `PROTOTYPED` | [provider-free Python reference core with strict authority-policy admission, exact deterministic argv execution, content-addressed staged helper context, injected one-shot helper transport, passive process waiting, compact results, raw-artifact accounting, fail-closed budgets, and Visible Value receipts](https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/src/opsle_gearbox/core.py); 19 of 19 provider-free automated tests passed locally, in PR #1 CI, and in final-main CI; ruff, shellcheck, actionlint, gitleaks, wheel build, receipt validation, and public raw-locator/hash checks passed | A production-quality bounded helper transport, independently verified isolation and termination, full Context Firewall integration, and a frozen comparative benchmark remain missing. | Freeze a provider-free deterministic-versus-direct baseline and helper-transport conformance corpus before considering any live model/provider run. | `context-firewall`, `decision-evidence-protocol`, `agent-trajectory-profiler`, `agent-routing-policy`, `agent-execution-authorization` | waiting |
| 18 | [affected-verification](https://github.com/opsle/affected-verification) | concept | `3ff41688dded` | `VERIFIED` | [dependency-free Node.js deterministic planner plus benchmark-only Git/catalog/source-graph adapters for AV-EXP-001 Vitest and AV-EXP-002 Python/pytest-testmon, identity-bound SHADOW result validation, complete frozen-oracle harnesses, explainable skip records, fail-closed uncertainty handling, and opsle.value-receipt.v1 telemetry](https://github.com/opsle/affected-verification/blob/3ff41688dded6e96e65da7cc44fe2608cf86d073/benchmark/av-exp-002/REPORT.md); 86 of 86 automated tests, 15 of 15 conformance scenarios, and 7 of 7 determinism checks passed locally, in PR #3 CI, and from a fresh detached worktree at exact main; the AV-EXP-002 result validator also passed at exact main, and invalid-state coverage includes Python/toolchain, target, patch, catalog, selector state/version/output, dynamic/conftest uncertainty, baseline, oracle, scenario, skip, policy, trust-state, and result tampering | AV-EXP-002 observed a false targeted-sufficiency claim for runtime/subprocess import behavior; historical real-change replay, a production-quality evidence adapter, and independent qualifying replication also remain missing. AV remains OBSERVE/SHADOW and no TRUSTED_BOUNDED change class is authorized. | Preregister and execute a selection-miss repair for AV2-006 that represents runtime/subprocess import uncertainty without changing the preserved AV-EXP-002 result. | — | active |
| 18 | [affected-verification](https://github.com/opsle/affected-verification) | concept | `97f490a67337` | `VERIFIED` | [dependency-free Node.js deterministic plan-v2 planner with check-level dependency-completeness states, mechanism and boundary evidence, check-local fail-closed forced selection, explainable skips, bounded deterministic Python boundary inspection, identity-bound SHADOW validation, frozen-oracle repair/replay harnesses, and opsle.value-receipt.v1 safety-cost telemetry](https://github.com/opsle/affected-verification/blob/97f490a67337552fee25757266f3dc034660dca0/benchmark/av-exp-003/REPORT.md); 107 of 107 automated tests, 15 of 15 conformance scenarios, and 10 of 10 determinism checks passed locally, in PR #4 CI, in exact final-main CI, and from a fresh detached worktree at 97f490a67337552fee25757266f3dc034660dca0; the AV-EXP-003 verifier and full deterministic repair reproduction also passed with identical result and regression-matrix identities | The bounded repair does not close opaque boundaries or constitute dynamic analysis; historical real-change replay, a production-quality evidence adapter, and independent qualifying replication remain missing. AV remains OBSERVE/SHADOW and no TRUSTED_BOUNDED change class is authorized. | Preregister a bounded child-process import-tracing evidence-provider experiment to test whether selected opaque checks can regain precision without weakening the fail-closed rule. | — | active |
| 19 | [research](https://github.com/opsle/research) | program infrastructure | `9ee43197880c` | `PROTOTYPED` | [authoritative 21-repository ledger, machine-readable 18-concept theory registry including Affected Verification, canonical theory map, normative Visible Value controls, and provider-free EXP-001 benchmark, launch, one-block coordinator, external four-label LIVE_PROVIDER_RUN authorization, and current catalogue/pricing preflight artifacts with six content-addressed tasks, deterministic oracle, four arm contracts, sealed blinded allocation, exact subject configuration and adapter, exact authorization admission, private boundaries, receipts, mutation tests, and integrity CI](program/THEORY_MAP.md); 88 of 88 repository tests pass locally after deliberate migration to the 21-repository, 18-concept anti-forgetting set, including 13 authorization validations and two byte-identical replays; generated status and registry validation pass | The exact live authorization set remains unconsumed and unreleased, account-specific API entitlement is unverified under the zero-provider-call policy, no immutable dated model snapshot is documented, and the program has no canonical measured concept experiment. | Independently review and release the provider-free live-authorization and catalogue/pricing preflight; do not consume authorization or launch a provider/model subject. | — | active |
| 20 | [site](https://github.com/opsle/site) | program infrastructure | `28ad65be4750` | `PROTOTYPED` | [React/Vinext source implementation with content routes](https://github.com/opsle/site/blob/28ad65be4750dc849976fbf5c9eae9501c6bbb25/README.md); automated build/render tests present; not rerun because this reconciliation kept other repositories read-only | Wait for validated registry data and measured research; deployment requires separate authorization. | After registry merge, add a read-only registry ingestion design without deploying the site. | `research` | waiting |
| 21 | [.github](https://github.com/opsle/.github) | program infrastructure | `01c38e726db7` | `THEORY` | [documentation-only organization profile](https://github.com/opsle/.github/blob/01c38e726db7c3e45059d25fccce55e071e35938/profile/README.md); not applicable to current single Markdown profile; consistency is unverified | No mechanical registry consistency check exists in this repository. | After registry merge, design a read-only consistency check for organization-profile repository links. | `research` | waiting |
Expand Down
111 changes: 110 additions & 1 deletion program/experiments.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"schema_version": 1,
"last_verified_at": "2026-08-31T12:52:33Z",
"last_verified_at": "2026-08-31T16:17:40Z",
"experiments": [
{
"id": "EXP-001",
Expand Down Expand Up @@ -456,6 +456,115 @@
"lifecycle_impact": "REMAIN_VERIFIED: the run adds revision-bound falsification evidence and failure modes, but the observed selection miss and explicit run cap do not establish BENCHMARK_READY or EXPERIMENTED lifecycle promotion for the repository.",
"next_task": "Preregister and execute a selection-miss repair for AV2-006 that represents runtime/subprocess import uncertainty without changing the preserved AV-EXP-002 result."
},
{
"id": "AV-EXP-003",
"title": "Opaque Dependency Boundary Repair",
"status": "RECORDED",
"hypothesis": "Affected Verification may skip a check only when available evidence defends completeness for every declared dependency mechanism capable of connecting the change to that check; an unmodeled, incomplete, unknown, or opaque boundary forces selection of that check unless identified evidence closes it.",
"participating_repositories": [
"affected-verification",
"research"
],
"roles": {
"primary": "affected-verification",
"expected_support": [
"research"
],
"potential_support": []
},
"baseline": "The permanently preserved AV-EXP-002 FAIL and exact frozen AV2-006 Click scenario, plus frozen AV-EXP-001/002 selector/full-oracle evidence and ten preregistered generalized repair cases.",
"experimental_arms": [
"historical AV-EXP-001 and AV-EXP-002 AV selections preserved as the pre-repair baseline",
"Affected Verification plan v2 with check-level dependency completeness and deterministic boundary evidence",
"FULL frozen oracle for every repair case and preserved complete FULL evidence for every prior-corpus replay"
],
"primary_metric": "Zero repaired selection misses, including selection of the exact known AV2-006 check for a generalized dependency-completeness reason.",
"secondary_metrics": [
"prior and repaired test checks or executions selected in compatible exact units",
"additional verification introduced by dependency-safety policy",
"tests still skipped, scenarios remaining targeted, scenarios broadened, and FULL escalations",
"check-level boundary provenance and forced-selection explanations",
"non-test check differences by class"
],
"correctness_gate": "FULL remains authoritative in SHADOW; every adversarial catalog is completely executed, every frozen prior relevant set is replayed, and any newly missed relevant check fails the experiment.",
"failure_classifications": [
"known regression remains omitted",
"new regression miss",
"target-specific special case",
"malformed or unsupported boundary evidence accepted",
"incomplete full oracle",
"unmeasured or hidden precision cost",
"historical result mutation",
"unsafe trust promotion"
],
"dataset_fixture_identity": "Affected Verification final 97f490a67337552fee25757266f3dc034660dca0; AV-EXP-003 preregistration 7aa4d13e42d6a547973d7f2a6b330821145cedc2; Click 36baa15ff831b939a22bc527cd76ce653ef6f66d; result sha256:03b2f7d6a380c84f6a1749531067cf8b87404c879f42380de8f07cce48251519; regression matrix sha256:7260c2d3476a6e78323e75d36c54c8409ea4cb18fa3a8f9a76b5533e1df08615.",
"model_provider_configuration": "NONE: one interactive Codex session used native shell, patch, Git, and Graphify facilities; no Codex children, external model/provider workloads, or production systems were used.",
"run_identities": [
"sha256:03b2f7d6a380c84f6a1749531067cf8b87404c879f42380de8f07cce48251519",
"sha256:7260c2d3476a6e78323e75d36c54c8409ea4cb18fa3a8f9a76b5533e1df08615"
],
"result_artifacts": [
"https://github.com/opsle/affected-verification/blob/97f490a67337552fee25757266f3dc034660dca0/benchmark/av-exp-003/REPORT.md",
"https://github.com/opsle/affected-verification/blob/97f490a67337552fee25757266f3dc034660dca0/benchmark/av-exp-003/results-v1/summary.json",
"https://github.com/opsle/affected-verification/blob/97f490a67337552fee25757266f3dc034660dca0/benchmark/av-exp-003/results-v1/repair-regression-matrix.json",
"https://github.com/opsle/affected-verification/blob/97f490a67337552fee25757266f3dc034660dca0/benchmark/av-exp-003/results-v1/boundary-evidence.json",
"https://github.com/opsle/affected-verification/blob/3ff41688dded6e96e65da7cc44fe2608cf86d073/benchmark/av-exp-002/REPORT.md"
],
"target": {
"repository": "https://github.com/pallets/click.git",
"sha": "36baa15ff831b939a22bc527cd76ce653ef6f66d",
"license": "BSD-3-Clause"
},
"preregistration": {
"commit_sha": "7aa4d13e42d6a547973d7f2a6b330821145cedc2",
"known_failure_outcome_disclosed": true,
"repaired_outcomes_observed_before_commit": false
},
"benchmark_result": {
"affected_verification_main_sha": "97f490a67337552fee25757266f3dc034660dca0",
"result_identity": "sha256:03b2f7d6a380c84f6a1749531067cf8b87404c879f42380de8f07cce48251519",
"regression_matrix_identity": "sha256:7260c2d3476a6e78323e75d36c54c8409ea4cb18fa3a8f9a76b5533e1df08615",
"evidence_manifest_identity": "sha256:932292a4f47372cab963db85772bb3d6c1e4fa536edd7b99f9686b72afd7e93c",
"known_replay": {
"scenario_id": "AV2-006",
"check_id": "pytest:tests/test_imports.py::test_light_imports",
"historical_outcome": "MISS",
"repaired_outcome": "SELECTED",
"additional_test_executions_per_av_arm": 1
},
"adversarial_scenario_count": 10,
"adversarial_repaired_miss_count": 0,
"adversarial_prior_tests_selected": 10,
"adversarial_repaired_tests_selected": 17,
"adversarial_additional_tests_due_to_repair": 7,
"adversarial_tests_still_skipped": 13,
"adversarial_targeted_scenarios_retained": 10,
"adversarial_broadened_scenarios": 7,
"adversarial_full_escalations": 0,
"av_exp_001_new_misses": 0,
"av_exp_001_additional_test_executions": 0,
"av_exp_002_av_core_new_misses": 0,
"av_exp_002_av_core_additional_test_executions": 6,
"av_exp_002_with_selector_new_misses": 0,
"av_exp_002_with_selector_additional_test_executions": 7,
"non_test_check_differences": 0
},
"major_findings": [
"Evidence-source agreement and evidence completeness are independent; static and native selector omission cannot prove irrelevance across an unmodeled boundary.",
"Check-local fail-closed selection repaired the known replay without global FULL broadening; all ten repair scenarios remained targeted.",
"The bounded Python inspector found two open subprocess/child-interpreter checks among 2,016 Click pytest checks and retained provenance without claiming it closes those boundaries.",
"AV-EXP-002 remains a permanent FAIL and no trust promotion follows from repairing its known miss."
],
"replication_status": "SAME_HOST_DETERMINISTIC_REPLAY_ONLY",
"verdict": "PASS only for the preregistered defect-repair claim: the known AV2-006 check is selected, ten generalized cases have zero misses, no frozen prior AV-relevant check becomes newly missed, and exact precision cost is measured. This does not solve dynamic dependencies, prove general safety, or authorize trusted execution.",
"blockers": [
"The static boundary inspector identifies but does not close child-process, runtime import, arbitrary plugin, or reflection boundaries.",
"Historical real-change replay, a production-quality evidence adapter, and independent qualifying replication remain missing.",
"Affected Verification remains OBSERVE/SHADOW; no TRUSTED_BOUNDED class is authorized."
],
"lifecycle_impact": "REMAIN_VERIFIED: the run repairs and measures one defect under SHADOW but is explicitly capped below lifecycle or trust promotion.",
"next_task": "Preregister a bounded child-process import-tracing evidence-provider experiment to test whether selected opaque checks can regain precision without weakening the fail-closed rule."
},
{
"id": "LEGACY-001",
"title": "Graphify plus Antigravity semantic adapter integration observation",
Expand Down
Loading