diff --git a/OPSLE_TASKS_MIGRATION_PLAN.md b/OPSLE_TASKS_MIGRATION_PLAN.md index a8b7fcb..f3c4bd9 100644 --- a/OPSLE_TASKS_MIGRATION_PLAN.md +++ b/OPSLE_TASKS_MIGRATION_PLAN.md @@ -2,6 +2,12 @@ > **NOT AUTHORIZED FOR EXECUTION DURING THIS RUN. PLANNING ONLY.** +Opsle Tasks is the NEXT primary real-world workload after Durable Supervisor +v0.1 is declared and feature-frozen. That workload role does not authorize any +identity, repository, runtime, schema, service, DNS, TLS, provider, or release +migration. The machine-readable program priority remains +`program/registry.json`. + ## Intended future identities - repository: `sneakocom/taslos-tasks` → `opsle/tasks`; diff --git a/OPSLE_TASKS_PUBLIC_RELEASE.md b/OPSLE_TASKS_PUBLIC_RELEASE.md index 35424c3..1d26fd1 100644 --- a/OPSLE_TASKS_PUBLIC_RELEASE.md +++ b/OPSLE_TASKS_PUBLIC_RELEASE.md @@ -2,6 +2,10 @@ > Planning only. Taslos Tasks remains private in its existing location. Publication was not executed. +This release is LATER program work. It does not begin merely because Opsle Tasks +becomes the NEXT Durable Supervisor workload; it requires separate explicit +release authority and the eligibility evidence below. + ## Full-history audit Audit every reachable commit, tag, branch, release asset, PR artifact, LFS object, submodule reference, and archive for credentials, tokens, private keys, connection strings, private addresses, environment secrets, database data, backups, logs, sensitive screenshots, and accidental artifacts. Removing a value from HEAD is insufficient if reachable history retains it. diff --git a/PROGRAM_STATUS.md b/PROGRAM_STATUS.md index ae7db81..6cdd43a 100644 --- a/PROGRAM_STATUS.md +++ b/PROGRAM_STATUS.md @@ -4,49 +4,71 @@ **Coverage: 21/21 expected repositories; duplicates: 0.** -Last verified: `2026-08-31T16:17:40Z`. HEADs are the verified default-branch revisions, not an assumption about later changes. +Last verified: `2026-09-05T15:28:58Z`. HEADs are the verified default-branch revisions, not an assumption about later changes. + +## Program direction + +> What prevents Durable Supervisor from successfully finishing Opsle Tasks with less intelligence, less context, less human involvement, and defensible evidence? + +Current lane: **NOW**. Durable Supervisor v0.1: **IN_PROGRESS** (0/10 stopping criteria satisfied). + +Full generated priority view: `program/PRIORITY.md`. + +### Priority stack + +| Lane | Objective | Repositories | Entry gate | +|---|---|---|---| +| **NOW** | Finish Durable Supervisor v0.1 as a bounded measured system, then freeze feature work. | `durable-supervisor`, `research` | Current program lane. | +| **NEXT** | Use Opsle Tasks as the primary real-world workload and collect integrated measurements without performing the deferred Taslos-to-Opsle migration. | `gearbox`, `context-firewall`, `decision-evidence-protocol`, `agent-trajectory-profiler`, `affected-verification` | Durable Supervisor v0.1 is declared and feature-frozen. | +| **THEN** | Advance an individual concept only when Durable Supervisor or Opsle Tasks evidence demonstrates a concrete need. | `semantic-edit-protocol`, `event-driven-agent-wakeup`, `agent-state-ledger`, `agent-scheduler-runtime`, `verifiable-agent-handoff`, `agent-routing-policy`, `agent-resource-claims`, `agent-discovery-control`, `agent-execution-authorization`, `controlled-agent-acceptance`, `agent-recovery-policy`, `ephemeral-agent-workers` | A qualifying work-item reason and real workload evidence identify the smallest relevant concept boundary. | +| **LATER** | Run controlled experiments, freeze a real-workload benchmark corpus, seek independent replication, and only then consider public product and research-site releases. | `site` | NOW, NEXT, and evidence-triggered THEN work establish a defensible need and release prerequisites are separately authorized. | +| **PARKED** | Retain useful non-priority ideas without turning them into active work. | `.github` | The idea is useful but lacks a qualifying reason to compete with the current objective. | + +### Exact next execution + +In opsle/durable-supervisor, make the selected Gearbox route, routing rationale, and Context Firewall raw-versus-retained accounting operator-visible in each child receipt without broadening Durable Supervisor v0.1. ## Portfolio totals | Lifecycle stage | Count | |---|---:| -| `THEORY` | 12 | +| `THEORY` | 11 | | `SPECIFIED` | 0 | | `PROTOTYPED` | 6 | -| `VERIFIED` | 3 | +| `VERIFIED` | 4 | | `BENCHMARK_READY` | 0 | | `EXPERIMENTED` | 0 | | `REPRODUCED` | 0 | | `DOCUMENTED` | 0 | | `COMPLETE` | 0 | -Program state totals: active 5; waiting 16; complete 0. +Program state totals: active 2; waiting 19; complete 0. ## Repository inventory | # | Repository | Type | HEAD | Stage | Evidence | Blocker | Next task | Dependencies | State | |---:|---|---|---|---|---|---|---|---|---| -| 1 | [agent-trajectory-profiler](https://github.com/opsle/agent-trajectory-profiler) | concept | `0a8966164072` | `VERIFIED` | [dependency-free JavaScript trajectory profiler with value-receipt ingestion, observational run records, class/unit/trust-safe per-run and cumulative summaries, Context Firewall adapter, canonical CLI/API, and content-addressed conformance corpus](https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/src/value-summary.js); 74 of 74 automated tests and 14 of 14 content-addressed measurement fixtures passed locally at the verified HEAD; 5 of 5 pinned exact-revision interoperability cases and 5 determinism tests passed; PR #2 CI passed | EXP-001 exact budgeted authorization, catalogue/pricing preflight, measured experiment, and independent replication remain missing; the one-block coordinator is provider-free qualified and zero model/provider subjects have run. | Provider-free validate one exact four-label live authorization set and catalogue/pricing preflight without launching a subject. | — | active | -| 2 | [semantic-edit-protocol](https://github.com/opsle/semantic-edit-protocol) | concept | `29caad5c0382` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/semantic-edit-protocol/blob/29caad5c03827cde17aabd71c38bc25899413a33/THEORY.md); placeholder only; no automated tests | No executable semantic operation, validator, tests, or benchmark fixtures. | Specify one JavaScript symbol-replacement operation with preconditions, rollback behavior, and conformance cases. | `agent-resource-claims`, `agent-trajectory-profiler` | waiting | -| 3 | [durable-supervisor](https://github.com/opsle/durable-supervisor) | concept | `555ebedb992a` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/durable-supervisor/blob/555ebedb992ac74236bb7da8230b4d6b0489830b/THEORY.md); placeholder only; no automated tests | Portable ledger, scheduler, and wakeup fixtures do not yet exist. | After ledger and wakeup fixtures exist, specify the minimum reconstruction envelope and failure behavior. | `agent-state-ledger`, `agent-scheduler-runtime`, `event-driven-agent-wakeup`, `decision-evidence-protocol` | waiting | -| 4 | [event-driven-agent-wakeup](https://github.com/opsle/event-driven-agent-wakeup) | concept | `a6209860c215` | `PROTOTYPED` | [dependency-free JavaScript state-machine prototype](https://github.com/opsle/event-driven-agent-wakeup/blob/a6209860c2151450cc28ed648bc8c2631c8db7ef/src/runtime.js); 2 of 2 automated tests passed locally at the verified HEAD | Current prototype is in-memory and has no restart or durable-store harness. | Add a deterministic persisted-event fixture covering restart, duplicate delivery, wrong-wait, and timeout cases. | `agent-state-ledger` | waiting | -| 5 | [context-firewall](https://github.com/opsle/context-firewall) | concept | `953c48f1cfd1` | `PROTOTYPED` | [dependency-free deterministic JavaScript TAP-subset reducer with compact evidence packets, separate opsle.value-receipt.v1 sidecars, named operator indicators, payload ceilings, raw escalation, and synthetic conformance corpus](https://github.com/opsle/context-firewall/blob/953c48f1cfd154d6b7ed10b51b87fe54e4df45f2/src/value-receipt.js); 42 of 42 automated tests and 30 of 30 synthetic conformance fixtures passed locally at the verified HEAD; deterministic stdout/sidecar separation and operator stderr tests passed; PR #2 CI passed | EXP-001 exact budgeted authorization, catalogue/pricing preflight, measured correctness experiment, and replication remain missing; the one-block coordinator is provider-free qualified and zero model/provider subjects have run. | Provider-free validate one exact four-label live authorization set and catalogue/pricing preflight without launching a subject. | `decision-evidence-protocol`, `agent-trajectory-profiler` | active | -| 6 | [decision-evidence-protocol](https://github.com/opsle/decision-evidence-protocol) | concept | `b17ae3b41cea` | `VERIFIED` | [dependency-free generic value-receipt validator plus independent Context Firewall packet/source/value-receipt validator, Decision Evidence validation receipts, canonical CLI with named operator indicators, and self-contained conformance suite](https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/src/context-firewall-value.js); 60 of 60 automated tests and 24 of 24 public-safe conformance vectors passed locally at the verified HEAD; 5 of 5 exact-revision interoperability cases, 8 determinism tests, operator separation, source-unverified, and tamper paths passed; PR #2 CI passed | Additional real tool classes, measured subject decision adequacy, EXP-001 exact budgeted authorization and catalogue/pricing preflight, a controlled comparative result, and independent replication are missing. | Provider-free validate one exact four-label live authorization set and catalogue/pricing preflight without launching a subject. | — | active | -| 7 | [agent-state-ledger](https://github.com/opsle/agent-state-ledger) | concept | `acab03b1ff71` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-state-ledger/blob/acab03b1ff7168222552050e21e7553b07d00e7c/THEORY.md); placeholder only; no automated tests | No portable schema, implementation, projection oracle, or replay fixtures. | Specify a minimal append-only event schema and deterministic projection with contradiction fixtures. | — | waiting | -| 8 | [agent-scheduler-runtime](https://github.com/opsle/agent-scheduler-runtime) | concept | `d97cd3c218b2` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-scheduler-runtime/blob/d97cd3c218b20e0b2b0e09873f6a3d15c396b3a0/THEORY.md); placeholder only; no automated tests | Ledger and claim interfaces are not portable or executable in these repositories. | After ledger schema exists, specify deterministic readiness and lease transitions with a fake-clock harness. | `agent-resource-claims`, `agent-state-ledger` | waiting | -| 9 | [verifiable-agent-handoff](https://github.com/opsle/verifiable-agent-handoff) | concept | `399e5cfae943` | `PROTOTYPED` | [dependency-free JavaScript HMAC manifest prototype](https://github.com/opsle/verifiable-agent-handoff/blob/399e5cfae94345affa3f087f0f6eb9e77669d33c/src/seal.js); 3 of 3 automated tests passed locally at the verified HEAD | Prototype authenticates a manifest but does not build, transport, or reconstruct artifacts. | Build a synthetic artifact round-trip harness with tamper, missing-object, and source-destruction cases. | `decision-evidence-protocol` | waiting | -| 10 | [agent-routing-policy](https://github.com/opsle/agent-routing-policy) | concept | `43fc2a72d2c8` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-routing-policy/blob/43fc2a72d2c8494b2dcdca7b5a209de61d8fe2d8/THEORY.md); placeholder only; no automated tests | No normative route schema, evaluator, fixtures, or provider-independent quality evidence. | Specify strict immutable route envelopes and deterministic rejection reasons using synthetic profiles. | — | waiting | -| 11 | [agent-resource-claims](https://github.com/opsle/agent-resource-claims) | concept | `dfe0fbc90c67` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-resource-claims/blob/dfe0fbc90c67ce5ef4256354bb62f1d511b1304c/THEORY.md); placeholder only; no automated tests | No portable resource catalog, claim-set state machine, concurrency tests, or fairness fixtures. | Specify canonical resource identities and atomic claim-set transitions with stale-fence fixtures. | — | waiting | -| 12 | [agent-discovery-control](https://github.com/opsle/agent-discovery-control) | concept | `926547dd9fd7` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-discovery-control/blob/926547dd9fd713990b1d6f1f2e650aa6c0883564/THEORY.md); placeholder only; no automated tests | No durable proposal dataset, similarity oracle, policy evaluator, or convergence harness. | After the ledger schema exists, freeze duplicate/already-satisfied proposal fixtures and expected dispositions. | `agent-state-ledger` | waiting | -| 13 | [agent-execution-authorization](https://github.com/opsle/agent-execution-authorization) | concept | `8fb02e943c83` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-execution-authorization/blob/8fb02e943c83fae121c63185d2f0d0dde8c4260a/THEORY.md); placeholder only; no automated tests | No grant schema, current-authority oracle, revocation model, or adversarial fixtures. | Define a versioned grant binding source authority, current fence, target revision, stage, purpose, and expiry. | `agent-resource-claims`, `agent-state-ledger` | waiting | -| 14 | [controlled-agent-acceptance](https://github.com/opsle/controlled-agent-acceptance) | concept | `2d652adf56e5` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/controlled-agent-acceptance/blob/2d652adf56e53953327d09b1ba9c4a9c3445f052/THEORY.md); placeholder only; no automated tests | Authorization, routing, scheduling, and handoff contracts are not yet executable together. | Wait for prerequisite contracts, then specify an offline one-shot manifest state machine before any provider run. | `agent-execution-authorization`, `agent-routing-policy`, `agent-scheduler-runtime`, `verifiable-agent-handoff` | waiting | -| 15 | [agent-recovery-policy](https://github.com/opsle/agent-recovery-policy) | concept | `1b733a111e26` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-recovery-policy/blob/1b733a111e26e0a409fee3b96f627048531daefe/THEORY.md); placeholder only; no automated tests | No shared failure schema, attempt ledger, route evaluator, or comparative fixture set. | After decision evidence and route schemas stabilize, define same-failure convergence on synthetic failures. | `agent-routing-policy`, `agent-state-ledger`, `decision-evidence-protocol` | waiting | -| 16 | [ephemeral-agent-workers](https://github.com/opsle/ephemeral-agent-workers) | concept | `ad96fcfdfac0` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/ephemeral-agent-workers/blob/ad96fcfdfac06d340b5e96d369634980cee78ef4/THEORY.md); placeholder only; no automated tests | Portable authority, claim, and handoff contracts are not ready; no safe synthetic containment harness exists. | Wait for prerequisite contracts, then define a fake worker adapter and destruction receipt without infrastructure changes. | `agent-execution-authorization`, `agent-resource-claims`, `verifiable-agent-handoff` | waiting | -| 17 | [gearbox](https://github.com/opsle/gearbox) | concept | `f3fab9f292cf` | `PROTOTYPED` | [provider-free Python reference core with strict authority-policy admission, exact deterministic argv execution, content-addressed staged helper context, injected one-shot helper transport, passive process waiting, compact results, raw-artifact accounting, fail-closed budgets, and Visible Value receipts](https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/src/opsle_gearbox/core.py); 19 of 19 provider-free automated tests passed locally, in PR #1 CI, and in final-main CI; ruff, shellcheck, actionlint, gitleaks, wheel build, receipt validation, and public raw-locator/hash checks passed | A production-quality bounded helper transport, independently verified isolation and termination, full Context Firewall integration, and a frozen comparative benchmark remain missing. | Freeze a provider-free deterministic-versus-direct baseline and helper-transport conformance corpus before considering any live model/provider run. | `context-firewall`, `decision-evidence-protocol`, `agent-trajectory-profiler`, `agent-routing-policy`, `agent-execution-authorization` | waiting | -| 18 | [affected-verification](https://github.com/opsle/affected-verification) | concept | `97f490a67337` | `VERIFIED` | [dependency-free Node.js deterministic plan-v2 planner with check-level dependency-completeness states, mechanism and boundary evidence, check-local fail-closed forced selection, explainable skips, bounded deterministic Python boundary inspection, identity-bound SHADOW validation, frozen-oracle repair/replay harnesses, and opsle.value-receipt.v1 safety-cost telemetry](https://github.com/opsle/affected-verification/blob/97f490a67337552fee25757266f3dc034660dca0/benchmark/av-exp-003/REPORT.md); 107 of 107 automated tests, 15 of 15 conformance scenarios, and 10 of 10 determinism checks passed locally, in PR #4 CI, in exact final-main CI, and from a fresh detached worktree at 97f490a67337552fee25757266f3dc034660dca0; the AV-EXP-003 verifier and full deterministic repair reproduction also passed with identical result and regression-matrix identities | The bounded repair does not close opaque boundaries or constitute dynamic analysis; historical real-change replay, a production-quality evidence adapter, and independent qualifying replication remain missing. AV remains OBSERVE/SHADOW and no TRUSTED_BOUNDED change class is authorized. | Preregister a bounded child-process import-tracing evidence-provider experiment to test whether selected opaque checks can regain precision without weakening the fail-closed rule. | — | active | -| 19 | [research](https://github.com/opsle/research) | program infrastructure | `9ee43197880c` | `PROTOTYPED` | [authoritative 21-repository ledger, machine-readable 18-concept theory registry including Affected Verification, canonical theory map, normative Visible Value controls, and provider-free EXP-001 benchmark, launch, one-block coordinator, external four-label LIVE_PROVIDER_RUN authorization, and current catalogue/pricing preflight artifacts with six content-addressed tasks, deterministic oracle, four arm contracts, sealed blinded allocation, exact subject configuration and adapter, exact authorization admission, private boundaries, receipts, mutation tests, and integrity CI](program/THEORY_MAP.md); 88 of 88 repository tests pass locally after deliberate migration to the 21-repository, 18-concept anti-forgetting set, including 13 authorization validations and two byte-identical replays; generated status and registry validation pass | The exact live authorization set remains unconsumed and unreleased, account-specific API entitlement is unverified under the zero-provider-call policy, no immutable dated model snapshot is documented, and the program has no canonical measured concept experiment. | Independently review and release the provider-free live-authorization and catalogue/pricing preflight; do not consume authorization or launch a provider/model subject. | — | active | -| 20 | [site](https://github.com/opsle/site) | program infrastructure | `28ad65be4750` | `PROTOTYPED` | [React/Vinext source implementation with content routes](https://github.com/opsle/site/blob/28ad65be4750dc849976fbf5c9eae9501c6bbb25/README.md); automated build/render tests present; not rerun because this reconciliation kept other repositories read-only | Wait for validated registry data and measured research; deployment requires separate authorization. | After registry merge, add a read-only registry ingestion design without deploying the site. | `research` | waiting | -| 21 | [.github](https://github.com/opsle/.github) | program infrastructure | `01c38e726db7` | `THEORY` | [documentation-only organization profile](https://github.com/opsle/.github/blob/01c38e726db7c3e45059d25fccce55e071e35938/profile/README.md); not applicable to current single Markdown profile; consistency is unverified | No mechanical registry consistency check exists in this repository. | After registry merge, design a read-only consistency check for organization-profile repository links. | `research` | waiting | +| 1 | [agent-trajectory-profiler](https://github.com/opsle/agent-trajectory-profiler) | concept | `0a8966164072` | `VERIFIED` | [dependency-free JavaScript trajectory profiler with value-receipt ingestion, observational run records, class/unit/trust-safe per-run and cumulative summaries, Context Firewall adapter, canonical CLI/API, and content-addressed conformance corpus](https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/src/value-summary.js); 74 of 74 automated tests and 14 of 14 content-addressed measurement fixtures passed locally at the verified HEAD; 5 of 5 pinned exact-revision interoperability cases and 5 determinism tests passed; PR #2 CI passed | EXP-001 exact budgeted authorization, catalogue/pricing preflight, measured experiment, and independent replication remain missing; the one-block coordinator is provider-free qualified and zero model/provider subjects have run. | Wait for a Durable Supervisor or Opsle Tasks workload to require trajectory measurement; do not advance the profiler merely to polish it. | — | waiting | +| 2 | [semantic-edit-protocol](https://github.com/opsle/semantic-edit-protocol) | concept | `29caad5c0382` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/semantic-edit-protocol/blob/29caad5c03827cde17aabd71c38bc25899413a33/THEORY.md); placeholder only; no automated tests | No executable semantic operation, validator, tests, or benchmark fixtures. | Wait until a real Durable Supervisor or Opsle Tasks workload demonstrates a semantic-edit deficiency that simpler bounded edits cannot satisfy. | `agent-resource-claims`, `agent-trajectory-profiler` | waiting | +| 3 | [durable-supervisor](https://github.com/opsle/durable-supervisor) | concept | `1b5ab7631ba6` | `VERIFIED` | [dependency-free Node.js Durable Supervisor v0.1 with persistent supervisor authority, task handoff, discovery, exact Gearbox routing, claims and fencing, detached Runner ownership, Context Firewall packets, acceptance and immutable evaluation, opsled wake delivery, bounded reconstruction, policy controls, and runtime-release fencing](https://github.com/opsle/durable-supervisor/blob/1b5ab7631ba651a32592bbbdab8001865a3baf3d/README.md); PR #2 recorded 141 full-suite tests and canonical validation passing; PR #3 recorded 176 passing with 1 intentional skip; D5 PR #4 passed 6 focused atomicity, operator, and wake checks; projection-reconciliation PR #5 passed 76 focused evaluation, invariant, operator, recovery, and wake checks plus syntax, release-manifest, and canonical validation at the exact head whose tree is main | Durable Supervisor v0.1 still lacks the operator-visible measurement surface, avoidable-intelligence classification, smallest evidence-driven preflight, and another measured foreign-repository workload required by the program stopping criteria. | Make the selected Gearbox route, routing rationale, and Context Firewall raw-versus-retained accounting operator-visible in each child receipt without broadening v0.1. | `agent-state-ledger`, `agent-scheduler-runtime`, `event-driven-agent-wakeup`, `decision-evidence-protocol` | active | +| 4 | [event-driven-agent-wakeup](https://github.com/opsle/event-driven-agent-wakeup) | concept | `a6209860c215` | `PROTOTYPED` | [dependency-free JavaScript state-machine prototype](https://github.com/opsle/event-driven-agent-wakeup/blob/a6209860c2151450cc28ed648bc8c2631c8db7ef/src/runtime.js); 2 of 2 automated tests passed locally at the verified HEAD | Current prototype is in-memory and has no restart or durable-store harness. | Wait for a real Durable Supervisor or Opsle Tasks wake deficiency; advance the standalone concept only if the integrated runtime exposes one. | `agent-state-ledger` | waiting | +| 5 | [context-firewall](https://github.com/opsle/context-firewall) | concept | `953c48f1cfd1` | `PROTOTYPED` | [dependency-free deterministic JavaScript TAP-subset reducer with compact evidence packets, separate opsle.value-receipt.v1 sidecars, named operator indicators, payload ceilings, raw escalation, and synthetic conformance corpus](https://github.com/opsle/context-firewall/blob/953c48f1cfd154d6b7ed10b51b87fe54e4df45f2/src/value-receipt.js); 42 of 42 automated tests and 30 of 30 synthetic conformance fixtures passed locally at the verified HEAD; deterministic stdout/sidecar separation and operator stderr tests passed; PR #2 CI passed | EXP-001 exact budgeted authorization, catalogue/pricing preflight, measured correctness experiment, and replication remain missing; the one-block coordinator is provider-free qualified and zero model/provider subjects have run. | Wait for Durable Supervisor or Opsle Tasks evidence of context overload, unsafe omission, or unsupported input before advancing the standalone reducer. | `decision-evidence-protocol`, `agent-trajectory-profiler` | waiting | +| 6 | [decision-evidence-protocol](https://github.com/opsle/decision-evidence-protocol) | concept | `b17ae3b41cea` | `VERIFIED` | [dependency-free generic value-receipt validator plus independent Context Firewall packet/source/value-receipt validator, Decision Evidence validation receipts, canonical CLI with named operator indicators, and self-contained conformance suite](https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/src/context-firewall-value.js); 60 of 60 automated tests and 24 of 24 public-safe conformance vectors passed locally at the verified HEAD; 5 of 5 exact-revision interoperability cases, 8 determinism tests, operator separation, source-unverified, and tamper paths passed; PR #2 CI passed | Additional real tool classes, measured subject decision adequacy, EXP-001 exact budgeted authorization and catalogue/pricing preflight, a controlled comparative result, and independent replication are missing. | Wait for an integrated Durable Supervisor or Opsle Tasks receipt to expose a decision-evidence gap that blocks a defensible decision. | — | waiting | +| 7 | [agent-state-ledger](https://github.com/opsle/agent-state-ledger) | concept | `acab03b1ff71` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-state-ledger/blob/acab03b1ff7168222552050e21e7553b07d00e7c/THEORY.md); placeholder only; no automated tests | No portable schema, implementation, projection oracle, or replay fixtures. | Wait for a real Durable Supervisor or Opsle Tasks workload to expose a portable durable-state deficiency before advancing this repository. | — | waiting | +| 8 | [agent-scheduler-runtime](https://github.com/opsle/agent-scheduler-runtime) | concept | `d97cd3c218b2` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-scheduler-runtime/blob/d97cd3c218b20e0b2b0e09873f6a3d15c396b3a0/THEORY.md); placeholder only; no automated tests | Ledger and claim interfaces are not portable or executable in these repositories. | Wait for a real workload to expose a scheduler-runtime deficiency not already satisfied by Durable Supervisor's local implementation. | `agent-resource-claims`, `agent-state-ledger` | waiting | +| 9 | [verifiable-agent-handoff](https://github.com/opsle/verifiable-agent-handoff) | concept | `399e5cfae943` | `PROTOTYPED` | [dependency-free JavaScript HMAC manifest prototype](https://github.com/opsle/verifiable-agent-handoff/blob/399e5cfae94345affa3f087f0f6eb9e77669d33c/src/seal.js); 3 of 3 automated tests passed locally at the verified HEAD | Prototype authenticates a manifest but does not build, transport, or reconstruct artifacts. | Wait for a real workload to require evidence survival across an isolation or source-destruction boundary before advancing this protocol. | `decision-evidence-protocol` | waiting | +| 10 | [agent-routing-policy](https://github.com/opsle/agent-routing-policy) | concept | `43fc2a72d2c8` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-routing-policy/blob/43fc2a72d2c8494b2dcdca7b5a209de61d8fe2d8/THEORY.md); placeholder only; no automated tests | No normative route schema, evaluator, fixtures, or provider-independent quality evidence. | Wait for measured Durable Supervisor routing deficiencies before advancing a standalone routing policy. | — | waiting | +| 11 | [agent-resource-claims](https://github.com/opsle/agent-resource-claims) | concept | `dfe0fbc90c67` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-resource-claims/blob/dfe0fbc90c67ce5ef4256354bb62f1d511b1304c/THEORY.md); placeholder only; no automated tests | No portable resource catalog, claim-set state machine, concurrency tests, or fairness fixtures. | Wait for a workload-exposed resource-claim or fencing deficiency not already covered by Durable Supervisor's local claims. | — | waiting | +| 12 | [agent-discovery-control](https://github.com/opsle/agent-discovery-control) | concept | `926547dd9fd7` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-discovery-control/blob/926547dd9fd713990b1d6f1f2e650aa6c0883564/THEORY.md); placeholder only; no automated tests | No durable proposal dataset, similarity oracle, policy evaluator, or convergence harness. | Wait for real duplicate-discovery or already-satisfied work evidence before advancing this repository. | `agent-state-ledger` | waiting | +| 13 | [agent-execution-authorization](https://github.com/opsle/agent-execution-authorization) | concept | `8fb02e943c83` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-execution-authorization/blob/8fb02e943c83fae121c63185d2f0d0dde8c4260a/THEORY.md); placeholder only; no automated tests | No grant schema, current-authority oracle, revocation model, or adversarial fixtures. | Wait for a real execution-authorization gap that blocks the current Durable Supervisor or Opsle Tasks objective. | `agent-resource-claims`, `agent-state-ledger` | waiting | +| 14 | [controlled-agent-acceptance](https://github.com/opsle/controlled-agent-acceptance) | concept | `2d652adf56e5` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/controlled-agent-acceptance/blob/2d652adf56e53953327d09b1ba9c4a9c3445f052/THEORY.md); placeholder only; no automated tests | Authorization, routing, scheduling, and handoff contracts are not yet executable together. | Wait for real acceptance evidence to show a missing portable capability before advancing the standalone concept. | `agent-execution-authorization`, `agent-routing-policy`, `agent-scheduler-runtime`, `verifiable-agent-handoff` | waiting | +| 15 | [agent-recovery-policy](https://github.com/opsle/agent-recovery-policy) | concept | `1b733a111e26` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-recovery-policy/blob/1b733a111e26e0a409fee3b96f627048531daefe/THEORY.md); placeholder only; no automated tests | No shared failure schema, attempt ledger, route evaluator, or comparative fixture set. | Wait for repeated real recovery failures to demonstrate a policy deficiency; do not create speculative recovery work. | `agent-routing-policy`, `agent-state-ledger`, `decision-evidence-protocol` | waiting | +| 16 | [ephemeral-agent-workers](https://github.com/opsle/ephemeral-agent-workers) | concept | `ad96fcfdfac0` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/ephemeral-agent-workers/blob/ad96fcfdfac06d340b5e96d369634980cee78ef4/THEORY.md); placeholder only; no automated tests | Portable authority, claim, and handoff contracts are not ready; no safe synthetic containment harness exists. | Wait for a real workload to require an ephemeral-worker boundary before advancing this repository. | `agent-execution-authorization`, `agent-resource-claims`, `verifiable-agent-handoff` | waiting | +| 17 | [gearbox](https://github.com/opsle/gearbox) | concept | `f3fab9f292cf` | `PROTOTYPED` | [provider-free Python reference core with strict authority-policy admission, exact deterministic argv execution, content-addressed staged helper context, injected one-shot helper transport, passive process waiting, compact results, raw-artifact accounting, fail-closed budgets, and Visible Value receipts](https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/src/opsle_gearbox/core.py); 19 of 19 provider-free automated tests passed locally, in PR #1 CI, and in final-main CI; ruff, shellcheck, actionlint, gitleaks, wheel build, receipt validation, and public raw-locator/hash checks passed | A production-quality bounded helper transport, independently verified isolation and termination, full Context Firewall integration, and a frozen comparative benchmark remain missing. | Support Durable Supervisor receipt visibility and Opsle Tasks workload measurement; advance the standalone Gearbox only if integrated evidence exposes a routing or preflight deficiency. | `context-firewall`, `decision-evidence-protocol`, `agent-trajectory-profiler`, `agent-routing-policy`, `agent-execution-authorization` | waiting | +| 18 | [affected-verification](https://github.com/opsle/affected-verification) | concept | `97f490a67337` | `VERIFIED` | [dependency-free Node.js deterministic plan-v2 planner with check-level dependency-completeness states, mechanism and boundary evidence, check-local fail-closed forced selection, explainable skips, bounded deterministic Python boundary inspection, identity-bound SHADOW validation, frozen-oracle repair/replay harnesses, and opsle.value-receipt.v1 safety-cost telemetry](https://github.com/opsle/affected-verification/blob/97f490a67337552fee25757266f3dc034660dca0/benchmark/av-exp-003/REPORT.md); 107 of 107 automated tests, 15 of 15 conformance scenarios, and 10 of 10 determinism checks passed locally, in PR #4 CI, in exact final-main CI, and from a fresh detached worktree at 97f490a67337552fee25757266f3dc034660dca0; the AV-EXP-003 verifier and full deterministic repair reproduction also passed with identical result and regression-matrix identities | The bounded repair does not close opaque boundaries or constitute dynamic analysis; historical real-change replay, a production-quality evidence adapter, and independent qualifying replication remain missing. AV remains OBSERVE/SHADOW and no TRUSTED_BOUNDED change class is authorized. | Remain OBSERVE/SHADOW and wait for Durable Supervisor-driven Opsle Tasks work to expose a concrete verification-selection deficiency before adding another experiment. | — | waiting | +| 19 | [research](https://github.com/opsle/research) | program infrastructure | `e8a2c36678ac` | `PROTOTYPED` | [authoritative 21-repository portfolio and priority ledger, machine-readable 18-concept theory registry, five-experiment registry, lifecycle and anti-nitpick controls, normative Visible Value target, deterministic generated status and priority views, provider-free EXP-001 preparation, and recorded AV-EXP-001/002/003 shadow evidence with integrity CI](program/registry.json); 90 of 90 repository tests passed locally at the verified default-branch HEAD; the 21-repository and 18-concept registries, five experiment records, generated dashboard, and deterministic experiment evidence checks passed | Durable Supervisor v0.1 measurement and foreign-workload stopping criteria remain open; Opsle Tasks cannot become the primary workload until v0.1 is declared and frozen. | Keep the authoritative priority and portfolio views current while Durable Supervisor completes its ten bounded v0.1 stopping criteria. | — | active | +| 20 | [site](https://github.com/opsle/site) | program infrastructure | `28ad65be4750` | `PROTOTYPED` | [React/Vinext source implementation with content routes](https://github.com/opsle/site/blob/28ad65be4750dc849976fbf5c9eae9501c6bbb25/README.md); automated build/render tests present; not rerun because this reconciliation kept other repositories read-only | Wait for validated registry data and measured research; deployment requires separate authorization. | Remain later until measured research and separate site-release authorization justify registry-derived public content. | `research` | waiting | +| 21 | [.github](https://github.com/opsle/.github) | program infrastructure | `01c38e726db7` | `THEORY` | [documentation-only organization profile](https://github.com/opsle/.github/blob/01c38e726db7c3e45059d25fccce55e071e35938/profile/README.md); not applicable to current single Markdown profile; consistency is unverified | No mechanical registry consistency check exists in this repository. | Remain parked until a broken organization-profile link or external release condition creates a concrete need. | `research` | waiting | ## Theory reconciliation @@ -54,25 +76,19 @@ Concept coverage: 18 canonical concepts; 18 current concept repositories mapped Canonical map: `program/THEORY_MAP.md`. Machine registry: `program/theory-registry.json`. -## Highest-priority workstream - -Independently review and release the provider-free EXP-001 live-authorization and current catalogue/pricing preflight; do not consume authorization or launch any provider/model subject. +## Experiment state `EXP-001` — **PLANNED** — How much context can an AI coding agent safely not see? Blockers: The validated live authorization set remains unconsumed; separate execution authority is required before any provider/model subject launch. Account-specific API entitlement remains unverified under the zero-provider-call policy and requires a fresh fail-closed launch-time catalogue check. No distinct dated gpt-5.6-sol snapshot was publicly listed on 2026-08-30, so immutable model weights cannot be claimed. -## Exact recommended next execution - -In opsle/research, create and provider-free validate one exact four-label LIVE_PROVIDER_RUN authorization set plus a model catalogue/pricing preflight artifact; do not consume authorization or launch a provider/model subject. - ## Mechanical source of truth -- Registry: `program/registry.json` +- Portfolio and priority authority: `program/registry.json` +- Generated priority view: `program/PRIORITY.md` - Experiments: `program/experiments.json` - Theory registry: `program/theory-registry.json` - Theory map: `program/THEORY_MAP.md` - Lifecycle: `program/LIFECYCLE.md` - Operating rules: `program/OPERATING_RULES.md` -- Priority rationale: `program/PRIORITY.md` - Integrity check: `python3 tools/validate_program.py` diff --git a/README.md b/README.md index d126298..6c2c184 100644 --- a/README.md +++ b/README.md @@ -30,8 +30,9 @@ Important mechanisms should remain understandable, falsifiable, benchmarkable, r ## Start here -- [PROGRAM_STATUS.md](PROGRAM_STATUS.md) — generated 20-repository dashboard. -- [program/registry.json](program/registry.json) — authoritative machine-readable ledger. +- [PROGRAM_STATUS.md](PROGRAM_STATUS.md) — generated 21-repository dashboard. +- [program/registry.json](program/registry.json) — authoritative machine-readable portfolio and priority ledger. +- [program/PRIORITY.md](program/PRIORITY.md) — generated NOW / NEXT / THEN / LATER / PARKED view. - [program/THEORY_MAP.md](program/THEORY_MAP.md) — canonical conceptual topology and Gearbox boundary. - [program/theory-registry.json](program/theory-registry.json) — machine-readable concept classifications and dispositions. - [program/experiments.json](program/experiments.json) — canonical experiment registry. @@ -49,7 +50,7 @@ Important mechanisms should remain understandable, falsifiable, benchmarkable, r - [OPSLE_SITE_PLAN.md](OPSLE_SITE_PLAN.md) — future opsle.com content architecture. Validate ledger integrity with `python3 tools/validate_program.py`. Regenerate the -dashboard with `python3 tools/render_program_status.py`. +dashboard and priority view with `python3 tools/render_program_status.py`. ## Product relationship diff --git a/program/OPERATING_RULES.md b/program/OPERATING_RULES.md index 79d7c5f..0716c17 100644 --- a/program/OPERATING_RULES.md +++ b/program/OPERATING_RULES.md @@ -1,10 +1,15 @@ # Program operating rules +`program/registry.json` is authoritative for both portfolio state and the +NOW / NEXT / THEN / LATER / PARKED priority state. Generated Markdown is never +an independent planning authority. + Every Opsle execution must: 1. Read `program/registry.json` before selecting or performing work. 2. Verify the relevant repository default branch and HEAD before relying on recorded state. -3. Work against an explicit repository or experiment objective. +3. Start from the current program lane and operating question before selecting + an explicit repository or experiment objective. 4. Preserve immutable or content-addressed evidence for every material claim. 5. Promote lifecycle state only after satisfying the canonical gate in `program/LIFECYCLE.md`. 6. Update the registry when verified state changes. @@ -22,10 +27,18 @@ Every Opsle execution must: 13. Classify everyday telemetry as observational. Do not relabel an accumulated production corpus as causal or `EXPERIMENTAL` evidence without the controlled method required by `program/VISIBLE_VALUE_CONTRACT.md`. +14. Do not create a work item solely because an implementation can be improved. + New work normally requires a violated invariant, demonstrated defect, + measured inefficiency, missing capability blocking the current program + objective, experiment requirement, security or safety issue, or externally + required release condition. +15. Park cosmetic cleanup, architectural taste, hypothetical robustness, and + speculative future requirements unless qualifying evidence appears. Run `python3 tools/validate_program.py` and `python3 tools/render_program_status.py --check` before committing a registry -change. +change. The renderer checks both `PROGRAM_STATUS.md` and +`program/PRIORITY.md`. ## Portfolio discipline diff --git a/program/PRIORITY.md b/program/PRIORITY.md index 3c62628..9cdabb6 100644 --- a/program/PRIORITY.md +++ b/program/PRIORITY.md @@ -1,99 +1,147 @@ -# Evidence-driven execution order - -The portfolio is a dependency-aware set of parallel workstreams, not a serial -checklist. - -## Workstream 0: canonical Gearbox home - -The 2026-08-29 reconciliation established that Agent Gearbox is a coherent -primary-developer intelligence-and-context transmission capability and that no -current repository owns its irreducible mechanism. Routing, execution -authorization, and resource claims are supporting policies; Durable Supervisor, -ledger, scheduler, wakeup, discovery, and recovery solve autonomous durable -orchestration instead. - -Completed on 2026-08-29 through public `opsle/gearbox` PR #1. The repository now -contains a narrow provider-free prototype, exact Taslos source provenance, -provider-free tests, and revision-bound release evidence. No existing repository -was consolidated, renamed, transferred, archived, or deleted; Taslos Tasks -remained unchanged; zero model/provider subjects ran. - -## Workstream 1: EXP-001 prerequisites and experiment - -1. Context Firewall now emits deterministic Visible Value receipts and a named - operator indicator while keeping its compact packet canonical. -2. Decision Evidence now independently validates those packets and receipts and - emits its own validation receipt and named operator indicator. -3. Agent Trajectory Profiler now ingests compatible receipts and produces - deterministic class/unit/trust-safe per-run and cumulative summaries. -4. Completed through public research PR #9: six content-addressed tasks, the - deterministic correctness oracle, raw plus three Context Firewall arm - contracts, the balanced/blinded allocation method, and the provider-free - exact-revision harness are frozen and qualified. -5. Completed through public research PR #11: the exact OpenAI Responses subject - configuration, medium reasoning effort, no-retry adapter, 10-repetition fixed - sample, 240-label blinded index, 60-block encrypted mapping, seed commitment, - stopping rules, and provider-free verification are preregistered. The - coordinator seed remains outside Git and subject context. Zero model/provider - subjects ran. Add Verifiable Agent Handoff only if a future arm destroys or - isolates the source environment. -6. Completed through public research PR #13: the provider-free one-block - coordinator unseals one mapping outside subject context, validates an exact - four-label authorization set, renders all four arms, prepares empty result - envelopes, retains private artifacts outside model context, and returns only - seed-keyed commitments and exact counts. Secret-backed deterministic replay - used fixture-only authorizations, proved zero canonical arm identifiers in - subject-visible artifacts, and ran zero provider/model subjects. - -These three foundational projects are naturally tested together. The expected -first major measured experiment remains EXP-001 because no repository evidence -establishes a stronger prerequisite experiment. The coordinator prerequisite is -now qualified. The experiment itself must still wait for four exact label-bound -budget authorizations and a current catalogue/pricing preflight under separate -provider/model authority. - -## Workstream 2: durable orchestration - -Develop Agent State Ledger and the portable core of Agent Scheduler Runtime, -then test Event-Driven Agent Wakeup with Durable Supervisor. One restart, -duplicate-event, and reconstruction campaign can exercise all four mechanisms. - -## Workstream 3: bounded execution and verification - -Develop Agent Resource Claims and Agent Execution Authorization before combining -Ephemeral Agent Workers, Verifiable Agent Handoff, and Controlled Agent -Acceptance. This sequence makes lease/fence authority and evidence survival -testable before any real-provider acceptance is considered. - -## Workstream 4: routing and recovery - -Agent Routing Policy and Agent Recovery Policy should wait for durable state, -structured decision evidence, and authorization inputs. Their comparative study -can share failure fixtures and evaluate retry, alternate route, and terminal -decisions without invoking live providers initially. - -## Workstream 5: editing and discovery - -Semantic Edit Protocol should use Agent Trajectory Profiler for correctness-gated -payload/churn measurement and Agent Resource Claims for concurrent-region cases. -Agent Discovery Control should wait for a portable ledger and supervisor fixture -so duplicate convergence and already-satisfied proofs are durable rather than -conversation-local. - -## Workstream 6: affected verification - -Affected Verification is independently verified for its narrow planner and the -AV-EXP-001 Zustand/Vitest shadow calibration, and remains outside Gearbox. Both -AV arms selected 8/8 oracle-relevant checks in the frozen ten-scenario corpus; -the native tests-only arm selected 6/8 and omitted relevant lint and typecheck. -This is one-ecosystem calibration evidence, not general safety or bounded trust. -Its exact next execution is a preregistered second public-repository shadow -calibration in a different ecosystem with a meaningful native selector and the -same full-catalog oracle discipline. +# Evidence-driven program priority + + + +`program/registry.json` is the sole priority authority. Edit the registry, not this file. + +## Operating question + +> What prevents Durable Supervisor from successfully finishing Opsle Tasks with less intelligence, less context, less human involvement, and defensible evidence? + +## Anti-nitpick guardrail + +Do not create a work item solely because an implementation can be improved. + +Normally admit new work only for: + +- violated invariant +- demonstrated defect +- measured inefficiency +- missing capability blocking the current program objective +- experiment requirement +- security or safety issue +- externally required release condition + +Park by default: + +- cosmetic cleanup +- architectural taste +- hypothetical robustness +- speculative future requirement + +## NOW / NEXT / THEN / LATER / PARKED + +| Lane | Objective | Repositories | Entry gate | +|---|---|---|---| +| **NOW** | Finish Durable Supervisor v0.1 as a bounded measured system, then freeze feature work. | `durable-supervisor`, `research` | Current program lane. | +| **NEXT** | Use Opsle Tasks as the primary real-world workload and collect integrated measurements without performing the deferred Taslos-to-Opsle migration. | `gearbox`, `context-firewall`, `decision-evidence-protocol`, `agent-trajectory-profiler`, `affected-verification` | Durable Supervisor v0.1 is declared and feature-frozen. | +| **THEN** | Advance an individual concept only when Durable Supervisor or Opsle Tasks evidence demonstrates a concrete need. | `semantic-edit-protocol`, `event-driven-agent-wakeup`, `agent-state-ledger`, `agent-scheduler-runtime`, `verifiable-agent-handoff`, `agent-routing-policy`, `agent-resource-claims`, `agent-discovery-control`, `agent-execution-authorization`, `controlled-agent-acceptance`, `agent-recovery-policy`, `ephemeral-agent-workers` | A qualifying work-item reason and real workload evidence identify the smallest relevant concept boundary. | +| **LATER** | Run controlled experiments, freeze a real-workload benchmark corpus, seek independent replication, and only then consider public product and research-site releases. | `site` | NOW, NEXT, and evidence-triggered THEN work establish a defensible need and release prerequisites are separately authorized. | +| **PARKED** | Retain useful non-priority ideas without turning them into active work. | `.github` | The idea is useful but lacks a qualifying reason to compete with the current objective. | + +### NOW — Finish Durable Supervisor v0.1 as a bounded measured system, then freeze feature work. + +Entry: Current program lane. + +Exit: Every Durable Supervisor v0.1 stopping criterion is satisfied and the release is explicitly declared and frozen. + +### NEXT — Use Opsle Tasks as the primary real-world workload and collect integrated measurements without performing the deferred Taslos-to-Opsle migration. + +Entry: Durable Supervisor v0.1 is declared and feature-frozen. + +Exit: Durable Supervisor has driven the remaining authorized Opsle Tasks readiness work and produced defensible integrated measurements. + +### THEN — Advance an individual concept only when Durable Supervisor or Opsle Tasks evidence demonstrates a concrete need. + +Entry: A qualifying work-item reason and real workload evidence identify the smallest relevant concept boundary. + +Exit: The demonstrated deficiency is resolved or falsified at its narrowest justified boundary; do not march mechanically through the portfolio. + +### LATER — Run controlled experiments, freeze a real-workload benchmark corpus, seek independent replication, and only then consider public product and research-site releases. + +Entry: NOW, NEXT, and evidence-triggered THEN work establish a defensible need and release prerequisites are separately authorized. + +Exit: Applicable evidence and separate release authorization exist. + +### PARKED — Retain useful non-priority ideas without turning them into active work. + +Entry: The idea is useful but lacks a qualifying reason to compete with the current objective. + +Exit: New evidence supplies a qualifying work-item reason. + +## Durable Supervisor v0.1 stopping criteria + +Status: **IN_PROGRESS**. Foundation: `VERIFIED_ENOUGH_TO_STOP_NITPICKING_UNLESS_REAL_WORKLOAD_EXPOSES_A_DEFECT`. + +Verified main: `1b5ab7631ba651a32592bbbdab8001865a3baf3d`. Runtime: `PAUSED_NO_ACTIVE_TASK_OR_ATTEMPT` as durably recorded at `2026-09-05T12:20:01.551Z`. + +1. **OPEN** — Expose the selected Gearbox route, routing rationale, and Context Firewall accounting in an operator-visible per-child receipt. +2. **OPEN** — Expose each child's model and reasoning effort. +3. **OPEN** — Record actual tokens and cost when provider evidence supplies them, and clearly labeled estimates otherwise. +4. **OPEN** — Measure raw evidence or context versus Context Firewall retained context and report the reduction. +5. **OPEN** — Report first-pass success rate. +6. **OPEN** — Report repair-child or retry rate and the token cost of retries. +7. **OPEN** — Estimate avoidable intelligence consumption by identifying failures whose needed facts were discoverable through deterministic preflight before model execution. +8. **OPEN** — Implement only the smallest useful deterministic preflight or reconnaissance mechanism shown necessary by the avoidable-failure evidence. +9. **OPEN** — Complete another real foreign-repository workload that exercises and preserves the new measurements. +10. **OPEN** — Declare Durable Supervisor v0.1 and freeze feature work except for defects exposed by real workloads. + +## Opsle Tasks boundary + +Opsle Tasks is the NEXT primary real-world workload after Durable Supervisor v0.1. Its current repository remains `sneakocom/taslos-tasks`. + +Measure: Gearbox, Context Firewall, Decision Evidence Protocol, Agent Trajectory Profiler, Affected Verification. + +Without separate authorization, do not: + +- move apps/taslos-tasks +- transfer sneakocom/taslos-tasks +- rename production services +- change schemas merely for rebranding +- public release +- DNS or TLS changes +- launch provider work + +## Evidence-triggered concept activation + +- routing → `agent-routing-policy` +- durable state → `agent-state-ledger` +- retry or deterministic preflight → `gearbox` +- context overload → `context-firewall` +- verification selection → `affected-verification` +- recovery → `agent-recovery-policy` + +## Visible Value target + +Baseline: The configured primary supervisor model and reasoning effort performs all child work itself and receives raw unfiltered evidence. + +Savings require an inspectable same-work baseline or a clearly labeled estimate derived from that baseline; marketing counterfactuals are prohibited. + +- **MEASURED** — The value comes directly from deterministic artifacts or provider usage records and still carries the applicable Visible Value evidence class, normally EXACT or OBSERVED. +- **DERIVED** — The value is a reproducible calculation over identified measured inputs and records its method and applicable Visible Value evidence class. +- **ESTIMATED** — The value carries the ESTIMATED evidence class and names its method, assumptions, pricing source, and uncertainty. +- **UNAVAILABLE** — The value is not currently supported and must remain absent rather than being rendered as zero. + +Per-child receipt: child/task identity; model; reasoning effort; Gearbox route; routing rationale; input tokens; output tokens; raw evidence/context size; Context Firewall retained size; reduction percentage; estimated tokens avoided; estimated cost avoided; duration; attempt number; success/failure; escalation/retry reason; retry potentially avoidable through deterministic preflight. + +Supervisor/run summary: total children; model/effort distribution; total model tokens consumed; work completed deterministically without model use; Context Firewall reduction; estimated token/cost savings; first-pass success rate; repair-child rate; tokens spent on retries; avoidable-intelligence estimate. + +## Later + +- controlled empirical experiments +- frozen real-workload benchmark corpus +- independent replication +- Opsle Tasks public and self-hosted release +- opsle.com research and public site +- hosted Opsle offering + +## Parked + +- Durable Supervisor src/cli.js is 1,570 lines at the verified main SHA and may eventually warrant decomposition; this is a maintenance smell, not a current objective, unless measured work shows material reliability or efficiency harm. +- Background projection repair or reconciliation is not justified merely because explicit retry exists. +- Historical pre-fix cleanup or migration requires evidence of need. +- General architectural polishing remains parked unless a real workload exposes a concrete defect. ## Exact next execution -In `opsle/research`, independently review and release the provider-free exact -four-label `LIVE_PROVIDER_RUN` authorization set and current model -catalogue/pricing preflight branch. Do not consume authorization or launch a -provider/model subject. EXP-001 has no technical dependency on Gearbox. +In opsle/durable-supervisor, make the selected Gearbox route, routing rationale, and Context Firewall raw-versus-retained accounting operator-visible in each child receipt without broadening Durable Supervisor v0.1. diff --git a/program/THEORY_MAP.md b/program/THEORY_MAP.md index 5473e9c..8b2a01e 100644 --- a/program/THEORY_MAP.md +++ b/program/THEORY_MAP.md @@ -438,18 +438,17 @@ Three layers must remain separate: 3. **Experimental hypothesis:** reduced context preserves task correctness under equal fixtures, models, prompts, and correctness gates. -The existing conformance corpora and Prompt 005 dogfood evidence validate -implementation behavior; they are not the EXP-001 dataset or result. The -content-addressed task corpus, correctness oracle, baseline/arms, offline -harness, exact model/provider configuration, and blinded/randomized allocation -remain unfrozen. `EXP-001` stays `PLANNED` with zero run identities and zero -result artifacts. - -The offline benchmark-freeze work remains technically correct and needs no -experiment redesign. The independent public Gearbox publication is complete, -so the provider-free EXP-001 freeze is now the immediate next execution. +The conformance corpora and Prompt 005 dogfood evidence validate implementation +behavior; they are not an EXP-001 result. The content-addressed task corpus, +correctness oracle, baseline and arms, offline harness, exact model/provider +configuration, blinded allocation, coordinator, and provider-free live +preflight are frozen or qualified at their recorded revisions. `EXP-001` stays +`PLANNED` with zero run identities and zero result artifacts. + EXP-001 has no technical dependency on Gearbox, and no provider/model subject is -authorized by the freeze. +authorized by the provider-free preparation. Controlled experiments are now +LATER program work rather than the immediate execution; the authoritative +current priority is the machine state in `program/registry.json`. ## Public Opsle and Taslos Tasks boundary diff --git a/program/registry.json b/program/registry.json index 7e4f40d..c43e0f9 100644 --- a/program/registry.json +++ b/program/registry.json @@ -36,9 +36,105 @@ "provider_model_runs": 0, "repository_consolidations": 0 }, - "current_highest_priority_workstream": "Independently review and release the provider-free EXP-001 live-authorization and current catalogue/pricing preflight; do not consume authorization or launch any provider/model subject.", - "recommended_next_execution": "In opsle/research, create and provider-free validate one exact four-label LIVE_PROVIDER_RUN authorization set plus a model catalogue/pricing preflight artifact; do not consume authorization or launch a provider/model subject.", - "last_verified_at": "2026-08-31T16:17:40Z", + "program_control": { + "schema_version": 1, + "priority_order": ["NOW", "NEXT", "THEN", "LATER", "PARKED"], + "current_lane": "NOW", + "operating_question": "What prevents Durable Supervisor from successfully finishing Opsle Tasks with less intelligence, less context, less human involvement, and defensible evidence?", + "exact_next_execution": "In opsle/durable-supervisor, make the selected Gearbox route, routing rationale, and Context Firewall raw-versus-retained accounting operator-visible in each child receipt without broadening Durable Supervisor v0.1.", + "work_item_admission": { + "rule": "Do not create a work item solely because an implementation can be improved.", + "qualifying_reasons": ["violated invariant", "demonstrated defect", "measured inefficiency", "missing capability blocking the current program objective", "experiment requirement", "security or safety issue", "externally required release condition"], + "parked_by_default": ["cosmetic cleanup", "architectural taste", "hypothetical robustness", "speculative future requirement"] + }, + "lanes": [ + { + "name": "NOW", + "objective": "Finish Durable Supervisor v0.1 as a bounded measured system, then freeze feature work.", + "repositories": ["durable-supervisor", "research"], + "entry_condition": "Current program lane.", + "exit_condition": "Every Durable Supervisor v0.1 stopping criterion is satisfied and the release is explicitly declared and frozen." + }, + { + "name": "NEXT", + "objective": "Use Opsle Tasks as the primary real-world workload and collect integrated measurements without performing the deferred Taslos-to-Opsle migration.", + "repositories": ["gearbox", "context-firewall", "decision-evidence-protocol", "agent-trajectory-profiler", "affected-verification"], + "entry_condition": "Durable Supervisor v0.1 is declared and feature-frozen.", + "exit_condition": "Durable Supervisor has driven the remaining authorized Opsle Tasks readiness work and produced defensible integrated measurements." + }, + { + "name": "THEN", + "objective": "Advance an individual concept only when Durable Supervisor or Opsle Tasks evidence demonstrates a concrete need.", + "repositories": ["semantic-edit-protocol", "event-driven-agent-wakeup", "agent-state-ledger", "agent-scheduler-runtime", "verifiable-agent-handoff", "agent-routing-policy", "agent-resource-claims", "agent-discovery-control", "agent-execution-authorization", "controlled-agent-acceptance", "agent-recovery-policy", "ephemeral-agent-workers"], + "entry_condition": "A qualifying work-item reason and real workload evidence identify the smallest relevant concept boundary.", + "exit_condition": "The demonstrated deficiency is resolved or falsified at its narrowest justified boundary; do not march mechanically through the portfolio." + }, + { + "name": "LATER", + "objective": "Run controlled experiments, freeze a real-workload benchmark corpus, seek independent replication, and only then consider public product and research-site releases.", + "repositories": ["site"], + "entry_condition": "NOW, NEXT, and evidence-triggered THEN work establish a defensible need and release prerequisites are separately authorized.", + "exit_condition": "Applicable evidence and separate release authorization exist." + }, + { + "name": "PARKED", + "objective": "Retain useful non-priority ideas without turning them into active work.", + "repositories": [".github"], + "entry_condition": "The idea is useful but lacks a qualifying reason to compete with the current objective.", + "exit_condition": "New evidence supplies a qualifying work-item reason." + } + ], + "durable_supervisor_v0_1": { + "status": "IN_PROGRESS", + "durability_foundation": "VERIFIED_ENOUGH_TO_STOP_NITPICKING_UNLESS_REAL_WORKLOAD_EXPOSES_A_DEFECT", + "verified_main_sha": "1b5ab7631ba651a32592bbbdab8001865a3baf3d", + "verified_runtime_state": "PAUSED_NO_ACTIVE_TASK_OR_ATTEMPT", + "runtime_state_recorded_at": "2026-09-05T12:20:01.551Z", + "runtime_state_source": "Read-only .opsle/state.json authority showed supervisor_state PAUSED with null active_task_id and active_attempt_id; .opsle/supervisor.json showed AUTHORITATIVE authority.", + "stopping_criteria": [ + {"id": "DS-V0.1-01", "status": "OPEN", "criterion": "Expose the selected Gearbox route, routing rationale, and Context Firewall accounting in an operator-visible per-child receipt."}, + {"id": "DS-V0.1-02", "status": "OPEN", "criterion": "Expose each child's model and reasoning effort."}, + {"id": "DS-V0.1-03", "status": "OPEN", "criterion": "Record actual tokens and cost when provider evidence supplies them, and clearly labeled estimates otherwise."}, + {"id": "DS-V0.1-04", "status": "OPEN", "criterion": "Measure raw evidence or context versus Context Firewall retained context and report the reduction."}, + {"id": "DS-V0.1-05", "status": "OPEN", "criterion": "Report first-pass success rate."}, + {"id": "DS-V0.1-06", "status": "OPEN", "criterion": "Report repair-child or retry rate and the token cost of retries."}, + {"id": "DS-V0.1-07", "status": "OPEN", "criterion": "Estimate avoidable intelligence consumption by identifying failures whose needed facts were discoverable through deterministic preflight before model execution."}, + {"id": "DS-V0.1-08", "status": "OPEN", "criterion": "Implement only the smallest useful deterministic preflight or reconnaissance mechanism shown necessary by the avoidable-failure evidence."}, + {"id": "DS-V0.1-09", "status": "OPEN", "criterion": "Complete another real foreign-repository workload that exercises and preserves the new measurements."}, + {"id": "DS-V0.1-10", "status": "OPEN", "criterion": "Declare Durable Supervisor v0.1 and freeze feature work except for defects exposed by real workloads."} + ] + }, + "opsle_tasks": { + "current_repository": "sneakocom/taslos-tasks", + "future_name": "Opsle Tasks", + "role": "NEXT primary real-world workload after Durable Supervisor v0.1", + "measurements": ["Gearbox", "Context Firewall", "Decision Evidence Protocol", "Agent Trajectory Profiler", "Affected Verification"], + "prohibited_without_separate_authorization": ["move apps/taslos-tasks", "transfer sneakocom/taslos-tasks", "rename production services", "change schemas merely for rebranding", "public release", "DNS or TLS changes", "launch provider work"] + }, + "concept_activation": [ + {"deficiency": "routing", "repositories": ["agent-routing-policy"]}, + {"deficiency": "durable state", "repositories": ["agent-state-ledger"]}, + {"deficiency": "retry or deterministic preflight", "repositories": ["gearbox"]}, + {"deficiency": "context overload", "repositories": ["context-firewall"]}, + {"deficiency": "verification selection", "repositories": ["affected-verification"]}, + {"deficiency": "recovery", "repositories": ["agent-recovery-policy"]} + ], + "later_items": ["controlled empirical experiments", "frozen real-workload benchmark corpus", "independent replication", "Opsle Tasks public and self-hosted release", "opsle.com research and public site", "hosted Opsle offering"], + "parked_items": ["Durable Supervisor src/cli.js is 1,570 lines at the verified main SHA and may eventually warrant decomposition; this is a maintenance smell, not a current objective, unless measured work shows material reliability or efficiency harm.", "Background projection repair or reconciliation is not justified merely because explicit retry exists.", "Historical pre-fix cleanup or migration requires evidence of need.", "General architectural polishing remains parked unless a real workload exposes a concrete defect."] + }, + "visible_value": { + "baseline": "The configured primary supervisor model and reasoning effort performs all child work itself and receives raw unfiltered evidence.", + "baseline_rule": "Savings require an inspectable same-work baseline or a clearly labeled estimate derived from that baseline; marketing counterfactuals are prohibited.", + "value_kinds": [ + {"name": "MEASURED", "definition": "The value comes directly from deterministic artifacts or provider usage records and still carries the applicable Visible Value evidence class, normally EXACT or OBSERVED."}, + {"name": "DERIVED", "definition": "The value is a reproducible calculation over identified measured inputs and records its method and applicable Visible Value evidence class."}, + {"name": "ESTIMATED", "definition": "The value carries the ESTIMATED evidence class and names its method, assumptions, pricing source, and uncertainty."}, + {"name": "UNAVAILABLE", "definition": "The value is not currently supported and must remain absent rather than being rendered as zero."} + ], + "per_child_receipt_fields": ["child/task identity", "model", "reasoning effort", "Gearbox route", "routing rationale", "input tokens", "output tokens", "raw evidence/context size", "Context Firewall retained size", "reduction percentage", "estimated tokens avoided", "estimated cost avoided", "duration", "attempt number", "success/failure", "escalation/retry reason", "retry potentially avoidable through deterministic preflight"], + "supervisor_summary_fields": ["total children", "model/effort distribution", "total model tokens consumed", "work completed deterministically without model use", "Context Firewall reduction", "estimated token/cost savings", "first-pass success rate", "repair-child rate", "tokens spent on retries", "avoidable-intelligence estimate"] + }, + "last_verified_at": "2026-09-05T15:28:58Z", "repositories": [ { "name": "agent-trajectory-profiler", @@ -62,13 +158,13 @@ "dependents": ["context-firewall", "gearbox", "semantic-edit-protocol"], "active_experiment_ids": ["EXP-001"], "blockers": ["EXP-001 exact budgeted authorization, catalogue/pricing preflight, measured experiment, and independent replication remain missing; the one-block coordinator is provider-free qualified and zero model/provider subjects have run."], - "next_task": "Provider-free validate one exact four-label live authorization set and catalogue/pricing preflight without launching a subject.", + "next_task": "Wait for a Durable Supervisor or Opsle Tasks workload to require trajectory measurement; do not advance the profiler merely to polish it.", "evidence": ["https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/src/value-summary.js", "https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/src/run-record.js", "https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/tests/value-summary.test.js", "https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/tools/verify-context-firewall-interop.js", "https://github.com/opsle/agent-trajectory-profiler/pull/2", "https://github.com/opsle/research/blob/04234a65bf36192d63f1dd173c440d45a6604d2b/program/evidence/exp-001-offline-freeze/profile.json", "https://github.com/opsle/research/blob/31848c3f25ff9371055932657e8e2f8ad54cc8c7/program/evidence/exp-001-preregistration/verification-report.json", "https://github.com/opsle/research/blob/9ee43197880c18d4e185cf7e29e02a151d22a12e/program/evidence/exp-001-block-coordinator/qualification-report.json", "https://github.com/opsle/research/pull/11", "https://github.com/opsle/research/pull/13"], "completion_criteria": ["Publish metric semantics and executable profiler.", "Run correctness-gated benchmarks across multiple trajectories.", "Replicate predictive-value claims and document failure modes."], "completion_evidence": [], "completion_status": "INCOMPLETE", - "program_state": "active", - "last_verified_at": "2026-08-29T11:24:05Z" + "program_state": "waiting", + "last_verified_at": "2026-09-05T15:28:58Z" }, { "name": "semantic-edit-protocol", @@ -92,43 +188,43 @@ "dependents": [], "active_experiment_ids": [], "blockers": ["No executable semantic operation, validator, tests, or benchmark fixtures."], - "next_task": "Specify one JavaScript symbol-replacement operation with preconditions, rollback behavior, and conformance cases.", + "next_task": "Wait until a real Durable Supervisor or Opsle Tasks workload demonstrates a semantic-edit deficiency that simpler bounded edits cannot satisfy.", "evidence": ["https://github.com/opsle/semantic-edit-protocol/blob/29caad5c03827cde17aabd71c38bc25899413a33/THEORY.md", "https://github.com/opsle/semantic-edit-protocol/blob/29caad5c03827cde17aabd71c38bc25899413a33/SPEC.md"], "completion_criteria": ["Publish a precise operation contract and executable reference adapter.", "Benchmark against patch and whole-file baselines under identical correctness gates.", "Replicate results and document conflicts and unsupported syntax."], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "waiting", - "last_verified_at": "2026-08-25T05:14:23Z" + "last_verified_at": "2026-09-05T15:28:58Z" }, { "name": "durable-supervisor", "github_url": "https://github.com/opsle/durable-supervisor", "default_branch": "main", - "last_verified_head_sha": "555ebedb992ac74236bb7da8230b4d6b0489830b", + "last_verified_head_sha": "1b5ab7631ba651a32592bbbdab8001865a3baf3d", "project_type": "concept", - "purpose": "Separate supervisor decisions from child lifecycle and reconstruct work from durable state after inactivity.", - "lifecycle_stage": "THEORY", - "implementation_status": "none; placeholder source directory only", + "purpose": "Own autonomous objective progress through durable authority, bounded child execution, evidence reduction, evaluation, pause, wake, and reconstruction across model activations.", + "lifecycle_stage": "VERIFIED", + "implementation_status": "dependency-free Node.js Durable Supervisor v0.1 with persistent supervisor authority, task handoff, discovery, exact Gearbox routing, claims and fencing, detached Runner ownership, Context Firewall packets, acceptance and immutable evaluation, opsled wake delivery, bounded reconstruction, policy controls, and runtime-release fencing", "implementation_requirement": "A durable supervisor/runner reference implementation is required.", - "specification_status": "experimental theory contract; reconstruction and delegation semantics remain incomplete", - "test_status": "placeholder only; no automated tests", - "benchmark_status": "prose plan only; no runnable harness or frozen fixtures", - "measured_experiment_status": "none", - "reproducibility_status": "not available", - "documentation_status": "public theory, preliminary specification, architecture, and benchmark plan present", + "specification_status": "public v0.1 contract defines durable authority, immutable history, task and attempt state, exact routes, claims and fencing, Runner ownership, Context Firewall and acceptance boundaries, evaluation, wake, reconstruction, pause, and runtime compatibility", + "test_status": "PR #2 recorded 141 full-suite tests and canonical validation passing; PR #3 recorded 176 passing with 1 intentional skip; D5 PR #4 passed 6 focused atomicity, operator, and wake checks; projection-reconciliation PR #5 passed 76 focused evaluation, invariant, operator, recovery, and wake checks plus syntax, release-manifest, and canonical validation at the exact head whose tree is main", + "benchmark_status": "bounded self-hosting evidence records meaningful child work, exact route and acceptance artifacts, and measured Context Firewall byte reduction; PR #2 additionally records disposable and real foreign-repository portability plus a live automatic wake, but no frozen comparative baseline or controlled token/cost benchmark exists", + "measured_experiment_status": "operational self-hosting and foreign-repository observations only; no controlled comparative experiment", + "reproducibility_status": "provider-free tests, schema and release checks, reconstruction, portability fixtures, concurrency campaigns, and explicit evaluation replay are reproducible from the public source; latest PRs used focused affected verification rather than a new full-suite run, and no independent replication exists", + "documentation_status": "public v0.1 README, normative specification, architecture, operations guidance, and bounded self-hosting proof document supported behavior and explicit claim limits", "site_publication_status": "GitHub documentation only; no evidence-backed site publication", - "known_limitations": ["Controlled long-horizon experiments, reconstruction thresholds, delegation policy, and continuous-session comparison are missing."], + "known_limitations": ["Operator receipts do not yet expose the full Gearbox route/rationale, child model/effort, actual token/cost, retry cost, or avoidable-intelligence accounting required for v0.1 completion; existing Context Firewall evidence is byte-level, the latest repairs used focused rather than full-suite verification, and comparative benefit and independent replication remain unproven."], "dependencies": ["agent-state-ledger", "agent-scheduler-runtime", "event-driven-agent-wakeup", "decision-evidence-protocol"], "dependents": [], "active_experiment_ids": [], - "blockers": ["Portable ledger, scheduler, and wakeup fixtures do not yet exist."], - "next_task": "After ledger and wakeup fixtures exist, specify the minimum reconstruction envelope and failure behavior.", - "evidence": ["https://github.com/opsle/durable-supervisor/blob/555ebedb992ac74236bb7da8230b4d6b0489830b/THEORY.md", "https://github.com/opsle/durable-supervisor/blob/555ebedb992ac74236bb7da8230b4d6b0489830b/ARCHITECTURE.md"], - "completion_criteria": ["Implement a runner-independent durable supervisor.", "Measure correctness and inference use against a continuous-session baseline.", "Replicate restart and reconstruction behavior."], + "blockers": ["Durable Supervisor v0.1 still lacks the operator-visible measurement surface, avoidable-intelligence classification, smallest evidence-driven preflight, and another measured foreign-repository workload required by the program stopping criteria."], + "next_task": "Make the selected Gearbox route, routing rationale, and Context Firewall raw-versus-retained accounting operator-visible in each child receipt without broadening v0.1.", + "evidence": ["https://github.com/opsle/durable-supervisor/blob/1b5ab7631ba651a32592bbbdab8001865a3baf3d/README.md", "https://github.com/opsle/durable-supervisor/blob/1b5ab7631ba651a32592bbbdab8001865a3baf3d/SPEC.md", "https://github.com/opsle/durable-supervisor/blob/1b5ab7631ba651a32592bbbdab8001865a3baf3d/ARCHITECTURE.md", "https://github.com/opsle/durable-supervisor/blob/1b5ab7631ba651a32592bbbdab8001865a3baf3d/docs/SELF_HOSTING_PROOF.md", "https://github.com/opsle/durable-supervisor/pull/2", "https://github.com/opsle/durable-supervisor/pull/3", "https://github.com/opsle/durable-supervisor/pull/4", "https://github.com/opsle/durable-supervisor/pull/5"], + "completion_criteria": ["Satisfy the ten bounded Durable Supervisor v0.1 stopping criteria in the program-control registry.", "Prove the measurement surface on another real foreign-repository workload.", "Declare v0.1 and freeze feature work except for workload-exposed defects."], "completion_evidence": [], "completion_status": "INCOMPLETE", - "program_state": "waiting", - "last_verified_at": "2026-08-25T05:14:23Z" + "program_state": "active", + "last_verified_at": "2026-09-05T15:28:58Z" }, { "name": "event-driven-agent-wakeup", @@ -152,13 +248,13 @@ "dependents": ["durable-supervisor"], "active_experiment_ids": [], "blockers": ["Current prototype is in-memory and has no restart or durable-store harness."], - "next_task": "Add a deterministic persisted-event fixture covering restart, duplicate delivery, wrong-wait, and timeout cases.", + "next_task": "Wait for a real Durable Supervisor or Opsle Tasks wake deficiency; advance the standalone concept only if the integrated runtime exposes one.", "evidence": ["https://github.com/opsle/event-driven-agent-wakeup/blob/a6209860c2151450cc28ed648bc8c2631c8db7ef/src/runtime.js", "https://github.com/opsle/event-driven-agent-wakeup/blob/a6209860c2151450cc28ed648bc8c2631c8db7ef/tests/runtime.test.js"], "completion_criteria": ["Implement restart-safe durable wait registration and wakeup.", "Benchmark polling and model-turn baselines.", "Replicate event-loss and duplicate-delivery correctness."], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "waiting", - "last_verified_at": "2026-08-25T05:14:23Z" + "last_verified_at": "2026-09-05T15:28:58Z" }, { "name": "context-firewall", @@ -182,13 +278,13 @@ "dependents": ["gearbox"], "active_experiment_ids": ["EXP-001"], "blockers": ["EXP-001 exact budgeted authorization, catalogue/pricing preflight, measured correctness experiment, and replication remain missing; the one-block coordinator is provider-free qualified and zero model/provider subjects have run."], - "next_task": "Provider-free validate one exact four-label live authorization set and catalogue/pricing preflight without launching a subject.", + "next_task": "Wait for Durable Supervisor or Opsle Tasks evidence of context overload, unsafe omission, or unsupported input before advancing the standalone reducer.", "evidence": ["https://github.com/opsle/context-firewall/blob/953c48f1cfd154d6b7ed10b51b87fe54e4df45f2/src/value-receipt.js", "https://github.com/opsle/context-firewall/blob/953c48f1cfd154d6b7ed10b51b87fe54e4df45f2/tests/reducer.test.js", "https://github.com/opsle/context-firewall/blob/953c48f1cfd154d6b7ed10b51b87fe54e4df45f2/fixtures/corpus.js", "https://github.com/opsle/context-firewall/pull/2", "https://github.com/opsle/research/blob/04234a65bf36192d63f1dd173c440d45a6604d2b/program/evidence/exp-001-offline-freeze/context-firewall-value-receipt.json", "https://github.com/opsle/research/blob/31848c3f25ff9371055932657e8e2f8ad54cc8c7/program/evidence/exp-001-preregistration/verification-report.json", "https://github.com/opsle/research/blob/9ee43197880c18d4e185cf7e29e02a151d22a12e/program/evidence/exp-001-block-coordinator/qualification-report.json", "https://github.com/opsle/research/pull/11", "https://github.com/opsle/research/pull/13"], "completion_criteria": ["Publish reducer policy and executable conformance validator.", "Run correctness-gated measured context-reduction experiments.", "Replicate the safe frontier and publish omission failure modes."], "completion_evidence": [], "completion_status": "INCOMPLETE", - "program_state": "active", - "last_verified_at": "2026-08-29T11:24:05Z" + "program_state": "waiting", + "last_verified_at": "2026-09-05T15:28:58Z" }, { "name": "decision-evidence-protocol", @@ -212,13 +308,13 @@ "dependents": ["agent-recovery-policy", "context-firewall", "durable-supervisor", "gearbox", "verifiable-agent-handoff"], "active_experiment_ids": ["EXP-001"], "blockers": ["Additional real tool classes, measured subject decision adequacy, EXP-001 exact budgeted authorization and catalogue/pricing preflight, a controlled comparative result, and independent replication are missing."], - "next_task": "Provider-free validate one exact four-label live authorization set and catalogue/pricing preflight without launching a subject.", + "next_task": "Wait for an integrated Durable Supervisor or Opsle Tasks receipt to expose a decision-evidence gap that blocks a defensible decision.", "evidence": ["https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/src/context-firewall-value.js", "https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/src/value-receipt-v1.js", "https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/tests/value-receipt.test.js", "https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/docs/value-receipt-v1.md", "https://github.com/opsle/decision-evidence-protocol/pull/2", "https://github.com/opsle/research/blob/04234a65bf36192d63f1dd173c440d45a6604d2b/program/evidence/exp-001-offline-freeze/decision-evidence-validation.json", "https://github.com/opsle/research/blob/31848c3f25ff9371055932657e8e2f8ad54cc8c7/program/evidence/exp-001-preregistration/verification-report.json", "https://github.com/opsle/research/blob/9ee43197880c18d4e185cf7e29e02a151d22a12e/program/evidence/exp-001-block-coordinator/qualification-report.json", "https://github.com/opsle/research/pull/11", "https://github.com/opsle/research/pull/13"], "completion_criteria": ["Publish a versioned protocol and comprehensive executable conformance suite.", "Measure adequacy across multiple real tool classes.", "Replicate interoperability and document loss/escalation failure modes."], "completion_evidence": [], "completion_status": "INCOMPLETE", - "program_state": "active", - "last_verified_at": "2026-08-29T11:24:05Z" + "program_state": "waiting", + "last_verified_at": "2026-09-05T15:28:58Z" }, { "name": "agent-state-ledger", @@ -242,13 +338,13 @@ "dependents": ["agent-discovery-control", "agent-execution-authorization", "agent-recovery-policy", "agent-scheduler-runtime", "durable-supervisor", "event-driven-agent-wakeup"], "active_experiment_ids": [], "blockers": ["No portable schema, implementation, projection oracle, or replay fixtures."], - "next_task": "Specify a minimal append-only event schema and deterministic projection with contradiction fixtures.", + "next_task": "Wait for a real Durable Supervisor or Opsle Tasks workload to expose a portable durable-state deficiency before advancing this repository.", "evidence": ["https://github.com/opsle/agent-state-ledger/blob/acab03b1ff7168222552050e21e7553b07d00e7c/THEORY.md", "https://github.com/opsle/agent-state-ledger/blob/acab03b1ff7168222552050e21e7553b07d00e7c/SPEC.md"], "completion_criteria": ["Implement append-only storage and deterministic projection.", "Prove replay/reconstruction against failure fixtures.", "Measure and replicate reconstruction size and correctness."], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "waiting", - "last_verified_at": "2026-08-25T05:14:23Z" + "last_verified_at": "2026-09-05T15:28:58Z" }, { "name": "agent-scheduler-runtime", @@ -272,13 +368,13 @@ "dependents": ["controlled-agent-acceptance", "durable-supervisor"], "active_experiment_ids": [], "blockers": ["Ledger and claim interfaces are not portable or executable in these repositories."], - "next_task": "After ledger schema exists, specify deterministic readiness and lease transitions with a fake-clock harness.", + "next_task": "Wait for a real workload to expose a scheduler-runtime deficiency not already satisfied by Durable Supervisor's local implementation.", "evidence": ["https://github.com/opsle/agent-scheduler-runtime/blob/d97cd3c218b20e0b2b0e09873f6a3d15c396b3a0/THEORY.md", "https://github.com/opsle/agent-scheduler-runtime/blob/d97cd3c218b20e0b2b0e09873f6a3d15c396b3a0/BENCHMARK.md"], "completion_criteria": ["Implement the portable deterministic core and store contract.", "Pass restart, lease, pause, cancellation, duplicate, and fairness campaigns.", "Benchmark and replicate latency and correctness."], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "waiting", - "last_verified_at": "2026-08-25T05:14:23Z" + "last_verified_at": "2026-09-05T15:28:58Z" }, { "name": "verifiable-agent-handoff", @@ -302,13 +398,13 @@ "dependents": ["controlled-agent-acceptance", "ephemeral-agent-workers"], "active_experiment_ids": ["EXP-001"], "blockers": ["Prototype authenticates a manifest but does not build, transport, or reconstruct artifacts."], - "next_task": "Build a synthetic artifact round-trip harness with tamper, missing-object, and source-destruction cases.", + "next_task": "Wait for a real workload to require evidence survival across an isolation or source-destruction boundary before advancing this protocol.", "evidence": ["https://github.com/opsle/verifiable-agent-handoff/blob/399e5cfae94345affa3f087f0f6eb9e77669d33c/src/seal.js", "https://github.com/opsle/verifiable-agent-handoff/blob/399e5cfae94345affa3f087f0f6eb9e77669d33c/tests/seal.test.js"], "completion_criteria": ["Publish a versioned seal/reconstruction contract and executable conformance suite.", "Benchmark full artifact handoff and failure cases.", "Replicate with an independent implementation and environment."], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "waiting", - "last_verified_at": "2026-08-25T05:14:23Z" + "last_verified_at": "2026-09-05T15:28:58Z" }, { "name": "agent-routing-policy", @@ -332,13 +428,13 @@ "dependents": ["agent-recovery-policy", "controlled-agent-acceptance", "gearbox"], "active_experiment_ids": [], "blockers": ["No normative route schema, evaluator, fixtures, or provider-independent quality evidence."], - "next_task": "Specify strict immutable route envelopes and deterministic rejection reasons using synthetic profiles.", + "next_task": "Wait for measured Durable Supervisor routing deficiencies before advancing a standalone routing policy.", "evidence": ["https://github.com/opsle/agent-routing-policy/blob/43fc2a72d2c8494b2dcdca7b5a209de61d8fe2d8/THEORY.md", "https://github.com/opsle/agent-routing-policy/blob/43fc2a72d2c8494b2dcdca7b5a209de61d8fe2d8/SPEC.md"], "completion_criteria": ["Publish route schema and executable policy evaluator.", "Measure strict and adaptive routing against real baselines.", "Replicate quality/cost claims and document unstable routes."], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "waiting", - "last_verified_at": "2026-08-25T05:14:23Z" + "last_verified_at": "2026-09-05T15:28:58Z" }, { "name": "agent-resource-claims", @@ -362,13 +458,13 @@ "dependents": ["agent-execution-authorization", "agent-scheduler-runtime", "ephemeral-agent-workers", "semantic-edit-protocol"], "active_experiment_ids": [], "blockers": ["No portable resource catalog, claim-set state machine, concurrency tests, or fairness fixtures."], - "next_task": "Specify canonical resource identities and atomic claim-set transitions with stale-fence fixtures.", + "next_task": "Wait for a workload-exposed resource-claim or fencing deficiency not already covered by Durable Supervisor's local claims.", "evidence": ["https://github.com/opsle/agent-resource-claims/blob/dfe0fbc90c67ce5ef4256354bb62f1d511b1304c/THEORY.md", "https://github.com/opsle/agent-resource-claims/blob/dfe0fbc90c67ce5ef4256354bb62f1d511b1304c/ARCHITECTURE.md"], "completion_criteria": ["Implement atomic claims, expiry, renewal, takeover, and fencing.", "Pass deadlock, crash, stale-actor, and fairness campaigns.", "Replicate across at least two store/runtime adapters."], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "waiting", - "last_verified_at": "2026-08-25T05:14:23Z" + "last_verified_at": "2026-09-05T15:28:58Z" }, { "name": "agent-discovery-control", @@ -392,13 +488,13 @@ "dependents": [], "active_experiment_ids": [], "blockers": ["No durable proposal dataset, similarity oracle, policy evaluator, or convergence harness."], - "next_task": "After the ledger schema exists, freeze duplicate/already-satisfied proposal fixtures and expected dispositions.", + "next_task": "Wait for real duplicate-discovery or already-satisfied work evidence before advancing this repository.", "evidence": ["https://github.com/opsle/agent-discovery-control/blob/926547dd9fd713990b1d6f1f2e650aa6c0883564/THEORY.md", "https://github.com/opsle/agent-discovery-control/blob/926547dd9fd713990b1d6f1f2e650aa6c0883564/BENCHMARK.md"], "completion_criteria": ["Implement deterministic proposal admission and convergence.", "Benchmark duplicates, storms, and already-satisfied proofs.", "Replicate calibrated policies across objective sizes."], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "waiting", - "last_verified_at": "2026-08-25T05:14:23Z" + "last_verified_at": "2026-09-05T15:28:58Z" }, { "name": "agent-execution-authorization", @@ -422,13 +518,13 @@ "dependents": ["controlled-agent-acceptance", "ephemeral-agent-workers", "gearbox"], "active_experiment_ids": [], "blockers": ["No grant schema, current-authority oracle, revocation model, or adversarial fixtures."], - "next_task": "Define a versioned grant binding source authority, current fence, target revision, stage, purpose, and expiry.", + "next_task": "Wait for a real execution-authorization gap that blocks the current Durable Supervisor or Opsle Tasks objective.", "evidence": ["https://github.com/opsle/agent-execution-authorization/blob/8fb02e943c83fae121c63185d2f0d0dde8c4260a/THEORY.md", "https://github.com/opsle/agent-execution-authorization/blob/8fb02e943c83fae121c63185d2f0d0dde8c4260a/SPEC.md"], "completion_criteria": ["Publish a portable grant schema and executable validator.", "Pass stale, revoked, drifted, replayed, and composed-authority cases.", "Measure and replicate revocation and interoperability behavior."], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "waiting", - "last_verified_at": "2026-08-25T05:14:23Z" + "last_verified_at": "2026-09-05T15:28:58Z" }, { "name": "controlled-agent-acceptance", @@ -452,13 +548,13 @@ "dependents": [], "active_experiment_ids": [], "blockers": ["Authorization, routing, scheduling, and handoff contracts are not yet executable together."], - "next_task": "Wait for prerequisite contracts, then specify an offline one-shot manifest state machine before any provider run.", + "next_task": "Wait for real acceptance evidence to show a missing portable capability before advancing the standalone concept.", "evidence": ["https://github.com/opsle/controlled-agent-acceptance/blob/2d652adf56e53953327d09b1ba9c4a9c3445f052/THEORY.md", "https://github.com/opsle/controlled-agent-acceptance/blob/2d652adf56e53953327d09b1ba9c4a9c3445f052/ARCHITECTURE.md"], "completion_criteria": ["Implement and verify a one-shot fail-closed acceptance controller.", "Run bounded acceptance experiments only under separately authorized exact manifests.", "Replicate interruption, restoration, and residue reconciliation behavior."], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "waiting", - "last_verified_at": "2026-08-25T05:14:23Z" + "last_verified_at": "2026-09-05T15:28:58Z" }, { "name": "agent-recovery-policy", @@ -482,13 +578,13 @@ "dependents": [], "active_experiment_ids": [], "blockers": ["No shared failure schema, attempt ledger, route evaluator, or comparative fixture set."], - "next_task": "After decision evidence and route schemas stabilize, define same-failure convergence on synthetic failures.", + "next_task": "Wait for repeated real recovery failures to demonstrate a policy deficiency; do not create speculative recovery work.", "evidence": ["https://github.com/opsle/agent-recovery-policy/blob/1b733a111e26e0a409fee3b96f627048531daefe/THEORY.md", "https://github.com/opsle/agent-recovery-policy/blob/1b733a111e26e0a409fee3b96f627048531daefe/BENCHMARK.md"], "completion_criteria": ["Publish a typed failure schema and executable bounded policy.", "Compare retry, consultation, alternate route, and terminal baselines.", "Replicate outcome and convergence claims across failure classes."], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "waiting", - "last_verified_at": "2026-08-25T05:14:23Z" + "last_verified_at": "2026-09-05T15:28:58Z" }, { "name": "ephemeral-agent-workers", @@ -512,13 +608,13 @@ "dependents": [], "active_experiment_ids": [], "blockers": ["Portable authority, claim, and handoff contracts are not ready; no safe synthetic containment harness exists."], - "next_task": "Wait for prerequisite contracts, then define a fake worker adapter and destruction receipt without infrastructure changes.", + "next_task": "Wait for a real workload to require an ephemeral-worker boundary before advancing this repository.", "evidence": ["https://github.com/opsle/ephemeral-agent-workers/blob/ad96fcfdfac06d340b5e96d369634980cee78ef4/THEORY.md", "https://github.com/opsle/ephemeral-agent-workers/blob/ad96fcfdfac06d340b5e96d369634980cee78ef4/ARCHITECTURE.md"], "completion_criteria": ["Implement a portable bounded broker/worker adapter and destruction proof.", "Pass containment, credential, process, network, seal, and cleanup campaigns.", "Replicate across isolation technologies with a published threat model."], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "waiting", - "last_verified_at": "2026-08-25T05:14:23Z" + "last_verified_at": "2026-09-05T15:28:58Z" }, { "name": "gearbox", @@ -542,13 +638,13 @@ "dependents": [], "active_experiment_ids": [], "blockers": ["A production-quality bounded helper transport, independently verified isolation and termination, full Context Firewall integration, and a frozen comparative benchmark remain missing."], - "next_task": "Freeze a provider-free deterministic-versus-direct baseline and helper-transport conformance corpus before considering any live model/provider run.", + "next_task": "Support Durable Supervisor receipt visibility and Opsle Tasks workload measurement; advance the standalone Gearbox only if integrated evidence exposes a routing or preflight deficiency.", "evidence": ["https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/src/opsle_gearbox/core.py", "https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/tests/test_core.py", "https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/evidence/release-001/verification.json", "https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/PROVENANCE.md", "program/evidence/gearbox-publication/README.md", "https://github.com/opsle/gearbox/pull/1", "https://github.com/opsle/gearbox/actions/runs/33229696388"], "completion_criteria": ["Publish and verify a production-quality bounded helper transport without absorbing durable orchestration.", "Benchmark deterministic and bounded-cognitive gears against real direct-execution baselines under equal correctness gates.", "Replicate bounded value claims and document transport, isolation, cleanup, escalation, and unsupported-task failure modes."], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "waiting", - "last_verified_at": "2026-08-29T02:53:08Z" + "last_verified_at": "2026-09-05T15:28:58Z" }, { "name": "affected-verification", @@ -572,43 +668,43 @@ "dependents": [], "active_experiment_ids": ["AV-EXP-001", "AV-EXP-002", "AV-EXP-003"], "blockers": ["The bounded repair does not close opaque boundaries or constitute dynamic analysis; historical real-change replay, a production-quality evidence adapter, and independent qualifying replication remain missing. AV remains OBSERVE/SHADOW and no TRUSTED_BOUNDED change class is authorized."], - "next_task": "Preregister a bounded child-process import-tracing evidence-provider experiment to test whether selected opaque checks can regain precision without weakening the fail-closed rule.", + "next_task": "Remain OBSERVE/SHADOW and wait for Durable Supervisor-driven Opsle Tasks work to expose a concrete verification-selection deficiency before adding another experiment.", "evidence": ["https://github.com/opsle/affected-verification/blob/97f490a67337552fee25757266f3dc034660dca0/benchmark/av-exp-003/REPORT.md", "https://github.com/opsle/affected-verification/blob/7aa4d13e42d6a547973d7f2a6b330821145cedc2/benchmark/av-exp-003/preregistration-v1/preregistration.json", "https://github.com/opsle/affected-verification/blob/97f490a67337552fee25757266f3dc034660dca0/benchmark/av-exp-003/results-v1/summary.json", "https://github.com/opsle/affected-verification/blob/97f490a67337552fee25757266f3dc034660dca0/benchmark/av-exp-003/results-v1/repair-regression-matrix.json", "https://github.com/opsle/affected-verification/blob/97f490a67337552fee25757266f3dc034660dca0/benchmark/av-exp-003/results-v1/evidence-manifest.json", "https://github.com/opsle/affected-verification/blob/3ff41688dded6e96e65da7cc44fe2608cf86d073/benchmark/av-exp-002/REPORT.md", "https://github.com/opsle/affected-verification/pull/4", "https://github.com/opsle/affected-verification/actions/runs/33412942072"], "completion_criteria": ["Publish production evidence adapters and independently validate their completeness boundaries.", "Run correctness-first comparisons against full verification and native selectors with durable shadow miss evidence.", "Replicate bounded workload and safety claims on independent repositories or environments."], "completion_evidence": [], "completion_status": "INCOMPLETE", - "program_state": "active", - "last_verified_at": "2026-08-31T16:17:40Z" + "program_state": "waiting", + "last_verified_at": "2026-09-05T15:28:58Z" }, { "name": "research", "github_url": "https://github.com/opsle/research", "default_branch": "main", - "last_verified_head_sha": "9ee43197880c18d4e185cf7e29e02a151d22a12e", + "last_verified_head_sha": "e8a2c36678ac4d4f72a79a56e03e9f896811be02", "project_type": "program infrastructure", "purpose": "Own public research methodology, portfolio reconciliation, experiment records, and the authoritative program ledger.", "lifecycle_stage": "PROTOTYPED", - "implementation_status": "authoritative 21-repository ledger, machine-readable 18-concept theory registry including Affected Verification, canonical theory map, normative Visible Value controls, and provider-free EXP-001 benchmark, launch, one-block coordinator, external four-label LIVE_PROVIDER_RUN authorization, and current catalogue/pricing preflight artifacts with six content-addressed tasks, deterministic oracle, four arm contracts, sealed blinded allocation, exact subject configuration and adapter, exact authorization admission, private boundaries, receipts, mutation tests, and integrity CI", + "implementation_status": "authoritative 21-repository portfolio and priority ledger, machine-readable 18-concept theory registry, five-experiment registry, lifecycle and anti-nitpick controls, normative Visible Value target, deterministic generated status and priority views, provider-free EXP-001 preparation, and recorded AV-EXP-001/002/003 shadow evidence with integrity CI", "implementation_requirement": "Program infrastructure requires a validated registry, generated dashboard, experiment ledger, operating rules, and CI.", - "specification_status": "canonical lifecycle, theory classifications and dispositions, Gearbox and Context Firewall definitions, Gearbox-versus-Durable boundary, Visible Value receipt and measurement classes, operator/model channels, observational corpus, shadow/replay attachment, operating, priority, registry, experiment, and generated-status controls are specified", - "test_status": "88 of 88 repository tests pass locally after deliberate migration to the 21-repository, 18-concept anti-forgetting set, including 13 authorization validations and two byte-identical replays; generated status and registry validation pass", - "benchmark_status": "the EXP-001 provider-free components are frozen at 04234a65bf36192d63f1dd173c440d45a6604d2b, launch controls are preregistered at 31848c3f25ff9371055932657e8e2f8ad54cc8c7, and the one-block coordinator is qualified at 9ee43197880c18d4e185cf7e29e02a151d22a12e; the unreleased branch adds one external exact four-label LIVE_PROVIDER_RUN set and current gpt-5.6-sol catalogue/pricing evidence with zero consumptions, provider/model launches, experiment runs, or results; no causal experiment", - "measured_experiment_status": "one legacy integration observation; no canonical measured concept experiment", - "reproducibility_status": "program and receipt validation, status generation, public dogfood artifacts, exact-revision EXP-001 offline qualification, public allocation checks, adapter self-test, structural preregistration verification, coordinator unit checks, and live-authorization reconstruction from frozen private inputs are reproducible; the live set replay is byte-identical twice, full secret-backed coordinator replay still requires the external seed, and no subject result exists", - "documentation_status": "public canonical theory map, project reconciliation, registered Gearbox boundary and prototype scope, Context Firewall adapter scope, consolidation provenance policy, program documentation, Visible Value semantics, machine controls, evidence, and generated portfolio status are present", + "specification_status": "canonical lifecycle, theory classifications and dispositions, Gearbox and Context Firewall definitions, Gearbox-versus-Durable boundary, Visible Value receipt and measurement classes, operator/model channels, observational corpus, shadow/replay attachment, NOW/NEXT/THEN/LATER/PARKED lanes, work-item admission, registry, experiment, and deterministic generated-view controls are specified", + "test_status": "90 of 90 repository tests passed locally at the verified default-branch HEAD; the 21-repository and 18-concept registries, five experiment records, generated dashboard, and deterministic experiment evidence checks passed", + "benchmark_status": "EXP-001 provider-free components remain frozen and unconsumed; AV-EXP-001/002/003 are recorded correctness-first shadow calibrations, including the preserved AV-EXP-002 miss and bounded AV-EXP-003 repair evidence; no provider/model subject ran", + "measured_experiment_status": "AV-EXP-001, AV-EXP-002, and AV-EXP-003 are recorded shadow experiments; EXP-001 remains planned and unconsumed; no provider/model experiment run exists", + "reproducibility_status": "program and receipt validation, both generated views, public dogfood artifacts, exact-revision EXP-001 provider-free qualification, and AV-EXP-001/002/003 deterministic shadow evidence are reproducible; full secret-backed coordinator replay still requires the external seed, and no provider/model subject result exists", + "documentation_status": "public canonical theory map, project reconciliation, generated priority and portfolio views, registered Gearbox boundary and prototype scope, Context Firewall adapter scope, consolidation provenance policy, program documentation, Visible Value semantics, machine controls, and evidence are present", "site_publication_status": "GitHub research hub only", - "known_limitations": ["The registry cannot self-reference the commit that contains its own SHA, default-branch HEAD verification remains operator-driven, the frozen EXP-001 corpus is Python-only, the exact live authorization set is external and unconsumed, account-specific API entitlement is unverified, the public model ID is not an immutable dated snapshot, site content is not registry-derived, Gearbox comparative evidence is absent, and no canonical concept experiment has run."], + "known_limitations": ["The registry cannot self-reference the commit that contains its own SHA, default-branch HEAD verification remains operator-driven, EXP-001 is still unconsumed, AV remains OBSERVE/SHADOW, the site is not registry-derived, and integrated Durable Supervisor/Opsle Tasks token, cost, first-pass, retry, and avoidable-intelligence measurements do not yet exist."], "dependencies": [], "dependents": [".github", "site"], - "active_experiment_ids": ["EXP-001", "LEGACY-001"], - "blockers": ["The exact live authorization set remains unconsumed and unreleased, account-specific API entitlement is unverified under the zero-provider-call policy, no immutable dated model snapshot is documented, and the program has no canonical measured concept experiment."], - "next_task": "Independently review and release the provider-free live-authorization and catalogue/pricing preflight; do not consume authorization or launch a provider/model subject.", - "evidence": ["program/THEORY_MAP.md", "program/theory-registry.json", "program/evidence/gearbox-publication/README.md", "program/evidence/exp-001-offline-freeze/README.md", "program/evidence/exp-001-preregistration/README.md", "program/evidence/exp-001-block-coordinator/README.md", "program/evidence/exp-001-live-preflight/README.md", "program/evidence/exp-001-live-preflight/model-catalogue.json", "program/evidence/exp-001-live-preflight/pricing-preflight.json", "program/evidence/exp-001-live-preflight/qualification-report.json", "program/evidence/exp-001-live-preflight/value-receipt.json", "https://github.com/opsle/research/blob/04234a65bf36192d63f1dd173c440d45a6604d2b/experiments/exp-001/benchmark.json", "https://github.com/opsle/research/blob/04234a65bf36192d63f1dd173c440d45a6604d2b/program/evidence/exp-001-offline-freeze/qualification-report.json", "https://github.com/opsle/research/blob/31848c3f25ff9371055932657e8e2f8ad54cc8c7/experiments/exp-001/preregistration-v1/preregistration.json", "https://github.com/opsle/research/blob/31848c3f25ff9371055932657e8e2f8ad54cc8c7/program/evidence/exp-001-preregistration/verification-report.json", "https://github.com/opsle/research/blob/9ee43197880c18d4e185cf7e29e02a151d22a12e/experiments/exp-001/coordinator-v1/coordinator.py", "https://github.com/opsle/research/blob/9ee43197880c18d4e185cf7e29e02a151d22a12e/program/evidence/exp-001-block-coordinator/qualification-report.json", "https://github.com/opsle/research/blob/72f9e4a0326d68d3870e2e79ce4e351acb1d8ffa/program/VISIBLE_VALUE_CONTRACT.md", "https://github.com/opsle/research/pull/9", "https://github.com/opsle/research/pull/11", "https://github.com/opsle/research/pull/13"], + "active_experiment_ids": ["EXP-001", "AV-EXP-001", "AV-EXP-002", "AV-EXP-003", "LEGACY-001"], + "blockers": ["Durable Supervisor v0.1 measurement and foreign-workload stopping criteria remain open; Opsle Tasks cannot become the primary workload until v0.1 is declared and frozen."], + "next_task": "Keep the authoritative priority and portfolio views current while Durable Supervisor completes its ten bounded v0.1 stopping criteria.", + "evidence": ["program/registry.json", "PROGRAM_STATUS.md", "program/PRIORITY.md", "program/THEORY_MAP.md", "program/theory-registry.json", "program/evidence/gearbox-publication/README.md", "program/evidence/exp-001-offline-freeze/README.md", "program/evidence/exp-001-preregistration/README.md", "program/evidence/exp-001-block-coordinator/README.md", "program/evidence/exp-001-live-preflight/README.md", "https://github.com/opsle/research/pull/19", "https://github.com/opsle/research/pull/20", "https://github.com/opsle/research/pull/21", "https://github.com/opsle/research/blob/04234a65bf36192d63f1dd173c440d45a6604d2b/experiments/exp-001/benchmark.json", "https://github.com/opsle/research/blob/31848c3f25ff9371055932657e8e2f8ad54cc8c7/experiments/exp-001/preregistration-v1/preregistration.json", "https://github.com/opsle/research/blob/9ee43197880c18d4e185cf7e29e02a151d22a12e/experiments/exp-001/coordinator-v1/coordinator.py", "https://github.com/opsle/research/blob/72f9e4a0326d68d3870e2e79ce4e351acb1d8ffa/program/VISIBLE_VALUE_CONTRACT.md"], "completion_criteria": ["Keep the 21-repository ledger and dashboard mechanically consistent.", "Retain immutable experiment evidence and lifecycle promotion proof.", "Publish program documentation without stale or unsupported claims."], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "active", - "last_verified_at": "2026-08-31T01:29:54Z" + "last_verified_at": "2026-09-05T15:28:58Z" }, { "name": "site", @@ -632,13 +728,13 @@ "dependents": [], "active_experiment_ids": [], "blockers": ["Wait for validated registry data and measured research; deployment requires separate authorization."], - "next_task": "After registry merge, add a read-only registry ingestion design without deploying the site.", + "next_task": "Remain later until measured research and separate site-release authorization justify registry-derived public content.", "evidence": ["https://github.com/opsle/site/blob/28ad65be4750dc849976fbf5c9eae9501c6bbb25/README.md", "https://github.com/opsle/site/blob/28ad65be4750dc849976fbf5c9eae9501c6bbb25/tests/rendered-html.test.mjs"], "completion_criteria": ["Render registry-derived public state without unsupported claims.", "Pass build, render, accessibility, and content-consistency checks.", "Publish only under a separate authorized release with exact revision evidence."], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "waiting", - "last_verified_at": "2026-08-25T05:14:23Z" + "last_verified_at": "2026-09-05T15:28:58Z" }, { "name": ".github", @@ -662,13 +758,13 @@ "dependents": [], "active_experiment_ids": [], "blockers": ["No mechanical registry consistency check exists in this repository."], - "next_task": "After registry merge, design a read-only consistency check for organization-profile repository links.", + "next_task": "Remain parked until a broken organization-profile link or external release condition creates a concrete need.", "evidence": ["https://github.com/opsle/.github/blob/01c38e726db7c3e45059d25fccce55e071e35938/profile/README.md"], "completion_criteria": ["Keep organization identity and repository links consistent with the authoritative registry.", "Automate or verify consistency at exact revisions.", "Document ownership and update workflow."], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "waiting", - "last_verified_at": "2026-08-25T05:14:23Z" + "last_verified_at": "2026-09-05T15:28:58Z" } ] } diff --git a/tests/test_validate_program.py b/tests/test_validate_program.py index 04bb53a..6a6d980 100644 --- a/tests/test_validate_program.py +++ b/tests/test_validate_program.py @@ -8,7 +8,7 @@ ROOT = Path(__file__).resolve().parents[1] sys.path.insert(0, str(ROOT / "tools")) -from render_program_status import render # noqa: E402 +from render_program_status import render, render_priority # noqa: E402 from validate_program import ( # noqa: E402 COMPLETION_GATES, DEFAULT_EXPERIMENTS, @@ -127,6 +127,62 @@ def test_duplicate_repository_fails(self): errors = self.errors_for(registry=registry) self.assertTrue(any("duplicate repositories" in error for error in errors)) + def test_priority_lanes_cover_each_repository_exactly_once(self): + lanes = self.registry["program_control"]["lanes"] + repositories = [name for lane in lanes for name in lane["repositories"]] + self.assertEqual(len(repositories), 21) + self.assertEqual(len(set(repositories)), 21) + + def test_duplicate_priority_repository_fails(self): + registry = copy.deepcopy(self.registry) + registry["program_control"]["lanes"][1]["repositories"].append( + "durable-supervisor" + ) + errors = self.errors_for(registry=registry) + self.assertTrue( + any("duplicate priority repositories" in error for error in errors) + ) + + def test_priority_order_drift_fails(self): + registry = copy.deepcopy(self.registry) + registry["program_control"]["priority_order"] = [ + "NOW", "THEN", "NEXT", "LATER", "PARKED" + ] + errors = self.errors_for(registry=registry) + self.assertTrue(any("priority_order" in error for error in errors)) + + def test_active_repositories_match_now_lane(self): + registry = copy.deepcopy(self.registry) + project = next( + item for item in registry["repositories"] + if item["name"] == "affected-verification" + ) + project["program_state"] = "active" + errors = self.errors_for(registry=registry) + self.assertTrue(any("current priority lane" in error for error in errors)) + + def test_durable_supervisor_stopping_criteria_are_fenced(self): + registry = copy.deepcopy(self.registry) + registry["program_control"]["durable_supervisor_v0_1"][ + "stopping_criteria" + ].pop() + errors = self.errors_for(registry=registry) + self.assertTrue(any("ten ordered stopping criteria" in error for error in errors)) + + def test_anti_nitpick_reasons_are_fenced(self): + registry = copy.deepcopy(self.registry) + registry["program_control"]["work_item_admission"][ + "qualifying_reasons" + ].pop() + errors = self.errors_for(registry=registry) + self.assertTrue(any("qualifying reasons drifted" in error for error in errors)) + + def test_visible_value_fields_are_fenced(self): + registry = copy.deepcopy(self.registry) + registry["visible_value"]["per_child_receipt_fields"].pop() + errors = self.errors_for(registry=registry) + self.assertTrue(any("per-child receipt fields drifted" in error for error in errors)) + def test_malformed_repository_name_reports_error(self): registry = copy.deepcopy(self.registry) registry["repositories"][0]["name"] = ["not", "a", "string"] @@ -210,6 +266,15 @@ def test_nonexistent_experiment_project_fails(self): errors = self.errors_for(experiments=experiments) self.assertTrue(any("references nonexistent project" in error for error in errors)) + def test_experiment_participation_must_be_reciprocal(self): + registry = copy.deepcopy(self.registry) + research = next( + item for item in registry["repositories"] if item["name"] == "research" + ) + research["active_experiment_ids"].remove("AV-EXP-003") + errors = self.errors_for(registry=registry) + self.assertTrue(any("does not reciprocally list" in error for error in errors)) + def test_malformed_experiment_id_reports_error(self): experiments = copy.deepcopy(self.experiments) experiments["experiments"][0]["id"] = ["EXP-001"] @@ -308,9 +373,25 @@ def test_exp001_block_coordinator_release_cannot_drift(self): ] = "0" * 40 errors = self.errors_for(experiments=experiments) self.assertTrue( - any("release SHA must match research" in error for error in errors) + any( + "block coordinator research_release_sha must be" in error + for error in errors + ) ) + def test_block_coordinator_release_is_historical_not_current_head(self): + coordinator = self.experiments["experiments"][0][ + "block_coordinator_qualification" + ] + research = next( + item for item in self.registry["repositories"] + if item["name"] == "research" + ) + self.assertNotEqual( + coordinator["research_release_sha"], research["last_verified_head_sha"] + ) + self.assertEqual(self.errors_for(), []) + def test_exp001_block_coordinator_evidence_hash_cannot_drift(self): experiments = copy.deepcopy(self.experiments) experiments["experiments"][0]["block_coordinator_qualification"][ @@ -386,6 +467,10 @@ def test_dashboard_is_current(self): expected = (ROOT / "PROGRAM_STATUS.md").read_text(encoding="utf-8") self.assertEqual(render(self.registry, self.experiments), expected) + def test_priority_view_is_current(self): + expected = (ROOT / "program" / "PRIORITY.md").read_text(encoding="utf-8") + self.assertEqual(render_priority(self.registry, self.experiments), expected) + if __name__ == "__main__": unittest.main() diff --git a/tools/render_program_status.py b/tools/render_program_status.py index d2406e1..79ce90b 100644 --- a/tools/render_program_status.py +++ b/tools/render_program_status.py @@ -1,5 +1,5 @@ #!/usr/bin/env python3 -"""Render PROGRAM_STATUS.md deterministically from the program registries.""" +"""Render deterministic human views from the Opsle program registries.""" from __future__ import annotations @@ -22,21 +22,164 @@ ROOT = Path(__file__).resolve().parents[1] DEFAULT_OUTPUT = ROOT / "PROGRAM_STATUS.md" +DEFAULT_PRIORITY_OUTPUT = ROOT / "program" / "PRIORITY.md" def _cell(value: str) -> str: return value.replace("|", "\\|").replace("\n", " ") -def render(registry: dict, experiments: dict) -> str: +def _validate_for_render(registry: dict, experiments: dict) -> dict: errors = validate(registry, experiments) theory = load_json(DEFAULT_THEORY_REGISTRY) theory_map_text = DEFAULT_THEORY_MAP.read_text(encoding="utf-8") errors.extend(validate_theory(theory, registry, theory_map_text)) if errors: raise ValueError("cannot render invalid registry: " + "; ".join(errors)) + return theory + + +def _priority_table_lines(control: dict) -> list[str]: + lines = [ + "| Lane | Objective | Repositories | Entry gate |", + "|---|---|---|---|", + ] + for lane in control["lanes"]: + repositories = ", ".join(f"`{name}`" for name in lane["repositories"]) + lines.append( + f"| **{lane['name']}** | {_cell(lane['objective'])} | " + f"{repositories} | {_cell(lane['entry_condition'])} |" + ) + return lines + +def render_priority(registry: dict, experiments: dict) -> str: + _validate_for_render(registry, experiments) + control = registry["program_control"] + admission = control["work_item_admission"] + ds = control["durable_supervisor_v0_1"] + tasks = control["opsle_tasks"] + visible = registry["visible_value"] + lines = [ + "# Evidence-driven program priority", + "", + "", + "", + "`program/registry.json` is the sole priority authority. Edit the registry, not this file.", + "", + "## Operating question", + "", + f"> {control['operating_question']}", + "", + "## Anti-nitpick guardrail", + "", + admission["rule"], + "", + "Normally admit new work only for:", + "", + ] + lines.extend(f"- {reason}" for reason in admission["qualifying_reasons"]) + lines.extend(["", "Park by default:", ""]) + lines.extend(f"- {reason}" for reason in admission["parked_by_default"]) + lines.extend(["", "## NOW / NEXT / THEN / LATER / PARKED", ""]) + lines.extend(_priority_table_lines(control)) + for lane in control["lanes"]: + lines.extend( + [ + "", + f"### {lane['name']} — {lane['objective']}", + "", + f"Entry: {lane['entry_condition']}", + "", + f"Exit: {lane['exit_condition']}", + ] + ) + + lines.extend( + [ + "", + "## Durable Supervisor v0.1 stopping criteria", + "", + f"Status: **{ds['status']}**. Foundation: `{ds['durability_foundation']}`.", + "", + f"Verified main: `{ds['verified_main_sha']}`. Runtime: `{ds['verified_runtime_state']}` " + f"as durably recorded at `{ds['runtime_state_recorded_at']}`.", + "", + ] + ) + lines.extend( + f"{index}. **{item['status']}** — {item['criterion']}" + for index, item in enumerate(ds["stopping_criteria"], 1) + ) + lines.extend( + [ + "", + "## Opsle Tasks boundary", + "", + f"{tasks['future_name']} is the {tasks['role']}. Its current repository remains " + f"`{tasks['current_repository']}`.", + "", + "Measure: " + ", ".join(tasks["measurements"]) + ".", + "", + "Without separate authorization, do not:", + "", + ] + ) + lines.extend(f"- {item}" for item in tasks["prohibited_without_separate_authorization"]) + lines.extend(["", "## Evidence-triggered concept activation", ""]) + lines.extend( + f"- {item['deficiency']} → " + + ", ".join(f"`{name}`" for name in item["repositories"]) + for item in control["concept_activation"] + ) + lines.extend( + [ + "", + "## Visible Value target", + "", + f"Baseline: {visible['baseline']}", + "", + visible["baseline_rule"], + "", + ] + ) + lines.extend( + f"- **{item['name']}** — {item['definition']}" + for item in visible["value_kinds"] + ) + lines.extend( + [ + "", + "Per-child receipt: " + "; ".join(visible["per_child_receipt_fields"]) + ".", + "", + "Supervisor/run summary: " + + "; ".join(visible["supervisor_summary_fields"]) + + ".", + "", + "## Later", + "", + ] + ) + lines.extend(f"- {item}" for item in control["later_items"]) + lines.extend(["", "## Parked", ""]) + lines.extend(f"- {item}" for item in control["parked_items"]) + lines.extend( + [ + "", + "## Exact next execution", + "", + control["exact_next_execution"], + "", + ] + ) + return "\n".join(lines) + + +def render(registry: dict, experiments: dict) -> str: + theory = _validate_for_render(registry, experiments) repositories = registry["repositories"] + control = registry["program_control"] + ds = control["durable_supervisor_v0_1"] stages = Counter(repo["lifecycle_stage"] for repo in repositories) states = Counter(repo["program_state"] for repo in repositories) experiment_by_id = {item["id"]: item for item in experiments["experiments"]} @@ -45,6 +188,9 @@ def render(registry: dict, experiments: dict) -> str: for concept in theory["concepts"] if concept["current_repository"] is not None ] + satisfied = sum( + item["status"] == "SATISFIED" for item in ds["stopping_criteria"] + ) lines = [ "# Opsle program status", "", @@ -56,11 +202,32 @@ def render(registry: dict, experiments: dict) -> str: "", f"Last verified: `{registry['last_verified_at']}`. HEADs are the verified default-branch revisions, not an assumption about later changes.", "", - "## Portfolio totals", + "## Program direction", + "", + f"> {control['operating_question']}", + "", + f"Current lane: **{control['current_lane']}**. Durable Supervisor v0.1: " + f"**{ds['status']}** ({satisfied}/{len(ds['stopping_criteria'])} stopping criteria satisfied).", + "", + "Full generated priority view: `program/PRIORITY.md`.", + "", + "### Priority stack", "", - "| Lifecycle stage | Count |", - "|---|---:|", ] + lines.extend(_priority_table_lines(control)) + lines.extend( + [ + "", + "### Exact next execution", + "", + control["exact_next_execution"], + "", + "## Portfolio totals", + "", + "| Lifecycle stage | Count |", + "|---|---:|", + ] + ) for stage in LIFECYCLE_STAGES: lines.append(f"| `{stage}` | {stages.get(stage, 0)} |") lines.extend( @@ -111,27 +278,21 @@ def render(registry: dict, experiments: dict) -> str: "", "Canonical map: `program/THEORY_MAP.md`. Machine registry: `program/theory-registry.json`.", "", - "## Highest-priority workstream", - "", - registry["current_highest_priority_workstream"], + "## Experiment state", "", f"`EXP-001` — **{exp001['status']}** — {exp001['title']}", "", "Blockers: " + " ".join(exp001["blockers"]), "", - "## Exact recommended next execution", - "", - registry["recommended_next_execution"], - "", "## Mechanical source of truth", "", - "- Registry: `program/registry.json`", + "- Portfolio and priority authority: `program/registry.json`", + "- Generated priority view: `program/PRIORITY.md`", "- Experiments: `program/experiments.json`", "- Theory registry: `program/theory-registry.json`", "- Theory map: `program/THEORY_MAP.md`", "- Lifecycle: `program/LIFECYCLE.md`", "- Operating rules: `program/OPERATING_RULES.md`", - "- Priority rationale: `program/PRIORITY.md`", "- Integrity check: `python3 tools/validate_program.py`", "", ] @@ -143,22 +304,42 @@ def main(argv: list[str] | None = None) -> int: parser = argparse.ArgumentParser(description=__doc__) parser.add_argument("--check", action="store_true") parser.add_argument("--output", type=Path, default=DEFAULT_OUTPUT) + parser.add_argument( + "--priority-output", type=Path, default=DEFAULT_PRIORITY_OUTPUT + ) args = parser.parse_args(argv) registry = load_json(DEFAULT_REGISTRY) experiments = load_json(DEFAULT_EXPERIMENTS) try: content = render(registry, experiments) + priority_content = render_priority(registry, experiments) except ValueError as exc: print(f"FAIL: {exc}", file=sys.stderr) return 1 if args.check: + stale = [] if not args.output.exists() or args.output.read_text(encoding="utf-8") != content: - print(f"FAIL: {args.output} is stale; rerun tools/render_program_status.py", file=sys.stderr) + stale.append(str(args.output)) + if ( + not args.priority_output.exists() + or args.priority_output.read_text(encoding="utf-8") != priority_content + ): + stale.append(str(args.priority_output)) + if stale: + print( + "FAIL: stale generated files: " + + ", ".join(stale) + + "; rerun tools/render_program_status.py", + file=sys.stderr, + ) return 1 - print(f"PASS: {args.output.name} matches the registry") + print( + f"PASS: {args.output.name} and {args.priority_output.name} match the registry" + ) return 0 args.output.write_text(content, encoding="utf-8") - print(f"Rendered {args.output}") + args.priority_output.write_text(priority_content, encoding="utf-8") + print(f"Rendered {args.output} and {args.priority_output}") return 0 diff --git a/tools/validate_program.py b/tools/validate_program.py index f844ea2..66c494d 100644 --- a/tools/validate_program.py +++ b/tools/validate_program.py @@ -55,6 +55,90 @@ "COMPLETE", ) +PRIORITY_LANES = ("NOW", "NEXT", "THEN", "LATER", "PARKED") + +WORK_ITEM_QUALIFYING_REASONS = frozenset( + { + "violated invariant", + "demonstrated defect", + "measured inefficiency", + "missing capability blocking the current program objective", + "experiment requirement", + "security or safety issue", + "externally required release condition", + } +) + +PER_CHILD_VALUE_FIELDS = frozenset( + { + "child/task identity", + "model", + "reasoning effort", + "Gearbox route", + "routing rationale", + "input tokens", + "output tokens", + "raw evidence/context size", + "Context Firewall retained size", + "reduction percentage", + "estimated tokens avoided", + "estimated cost avoided", + "duration", + "attempt number", + "success/failure", + "escalation/retry reason", + "retry potentially avoidable through deterministic preflight", + } +) + +SUPERVISOR_VALUE_FIELDS = frozenset( + { + "total children", + "model/effort distribution", + "total model tokens consumed", + "work completed deterministically without model use", + "Context Firewall reduction", + "estimated token/cost savings", + "first-pass success rate", + "repair-child rate", + "tokens spent on retries", + "avoidable-intelligence estimate", + } +) + +OPSLE_TASKS_MEASUREMENTS = frozenset( + { + "Gearbox", + "Context Firewall", + "Decision Evidence Protocol", + "Agent Trajectory Profiler", + "Affected Verification", + } +) + +OPSLE_TASKS_PROHIBITIONS = frozenset( + { + "move apps/taslos-tasks", + "transfer sneakocom/taslos-tasks", + "rename production services", + "change schemas merely for rebranding", + "public release", + "DNS or TLS changes", + "launch provider work", + } +) + +LATER_ITEMS = frozenset( + { + "controlled empirical experiments", + "frozen real-workload benchmark corpus", + "independent replication", + "Opsle Tasks public and self-hosted release", + "opsle.com research and public site", + "hosted Opsle offering", + } +) + THEORY_CLASSIFICATIONS = frozenset( { "INDEPENDENT_OPSLE_TOOL", @@ -362,8 +446,8 @@ def validate( "theory_registry", "theory_map", "theory_reconciliation", - "current_highest_priority_workstream", - "recommended_next_execution", + "program_control", + "visible_value", "last_verified_at", ): if not registry.get(field): @@ -393,6 +477,208 @@ def validate( if isinstance(repo, dict) and isinstance(repo.get("name"), str) } + control = registry.get("program_control") + if not isinstance(control, dict): + errors.append("registry.program_control must be an object") + else: + if control.get("schema_version") != 1: + errors.append("program_control.schema_version must be 1") + if control.get("priority_order") != list(PRIORITY_LANES): + errors.append("program_control.priority_order must be NOW, NEXT, THEN, LATER, PARKED") + if control.get("current_lane") != "NOW": + errors.append("program_control.current_lane must be NOW") + for field in ("operating_question", "exact_next_execution"): + if not _nonempty_string(control.get(field)): + errors.append(f"program_control.{field} must be a nonempty string") + + admission = control.get("work_item_admission") + if not isinstance(admission, dict): + errors.append("program_control.work_item_admission must be an object") + else: + if not _nonempty_string(admission.get("rule")): + errors.append("program_control.work_item_admission.rule must be nonempty") + reasons = admission.get("qualifying_reasons") + if ( + not _string_list(reasons, allow_empty=False) + or set(reasons) != WORK_ITEM_QUALIFYING_REASONS + ): + errors.append("program_control work-item qualifying reasons drifted") + if not _string_list(admission.get("parked_by_default"), allow_empty=False): + errors.append("program_control parked_by_default must be a nonempty string array") + + lanes = control.get("lanes") + if not isinstance(lanes, list): + errors.append("program_control.lanes must be an array") + lanes = [] + lane_names = [lane.get("name") for lane in lanes if isinstance(lane, dict)] + if lane_names != list(PRIORITY_LANES): + errors.append("program_control lanes must appear once in canonical priority order") + priority_repositories: list[str] = [] + current_lane_repositories: list[str] = [] + for lane in lanes: + if not isinstance(lane, dict): + errors.append("program_control lane entries must be objects") + continue + name = lane.get("name") + for field in ("objective", "entry_condition", "exit_condition"): + if not _nonempty_string(lane.get(field)): + errors.append(f"program_control lane {name}: {field} must be nonempty") + lane_repositories = lane.get("repositories") + if not _string_list(lane_repositories, allow_empty=False): + errors.append(f"program_control lane {name}: repositories must be nonempty strings") + continue + priority_repositories.extend(lane_repositories) + if name == control.get("current_lane"): + current_lane_repositories = lane_repositories + priority_counts = Counter(priority_repositories) + duplicate_priority_repositories = sorted( + name for name, count in priority_counts.items() if count > 1 + ) + if duplicate_priority_repositories: + errors.append( + "program_control duplicate priority repositories: " + + ", ".join(duplicate_priority_repositories) + ) + missing_priority_repositories = sorted(expected - set(priority_repositories)) + unexpected_priority_repositories = sorted(set(priority_repositories) - expected) + if missing_priority_repositories: + errors.append( + "program_control missing priority repositories: " + + ", ".join(missing_priority_repositories) + ) + if unexpected_priority_repositories: + errors.append( + "program_control unexpected priority repositories: " + + ", ".join(unexpected_priority_repositories) + ) + active_repositories = sorted( + name for name, repo in repo_by_name.items() if repo.get("program_state") == "active" + ) + if active_repositories != sorted(current_lane_repositories): + errors.append("active repositories must exactly match the current priority lane") + + ds = control.get("durable_supervisor_v0_1") + if not isinstance(ds, dict): + errors.append("program_control.durable_supervisor_v0_1 must be an object") + else: + if ds.get("status") not in {"IN_PROGRESS", "DECLARED_FROZEN"}: + errors.append("durable_supervisor_v0_1 has invalid status") + if ds.get("verified_runtime_state") != "PAUSED_NO_ACTIVE_TASK_OR_ATTEMPT": + errors.append("durable_supervisor_v0_1 runtime state drifted") + if not _valid_timestamp(ds.get("runtime_state_recorded_at")): + errors.append("durable_supervisor_v0_1 runtime state needs a recorded timestamp") + if not _nonempty_string(ds.get("runtime_state_source")): + errors.append("durable_supervisor_v0_1 runtime observation needs a source") + durable_repository = repo_by_name.get("durable-supervisor", {}) + if ds.get("verified_main_sha") != durable_repository.get("last_verified_head_sha"): + errors.append("durable_supervisor_v0_1 SHA must match the repository ledger") + criteria = ds.get("stopping_criteria") + if not isinstance(criteria, list): + errors.append("durable_supervisor_v0_1.stopping_criteria must be an array") + criteria = [] + expected_ids = [f"DS-V0.1-{index:02d}" for index in range(1, 11)] + if [item.get("id") for item in criteria if isinstance(item, dict)] != expected_ids: + errors.append("durable_supervisor_v0_1 must retain ten ordered stopping criteria") + for item in criteria: + if not isinstance(item, dict): + errors.append("durable_supervisor_v0_1 criteria must be objects") + continue + if item.get("status") not in {"OPEN", "SATISFIED"}: + errors.append(f"{item.get('id')}: invalid stopping-criterion status") + if not _nonempty_string(item.get("criterion")): + errors.append(f"{item.get('id')}: criterion must be nonempty") + + opsle_tasks = control.get("opsle_tasks") + if not isinstance(opsle_tasks, dict): + errors.append("program_control.opsle_tasks must be an object") + else: + if opsle_tasks.get("current_repository") != "sneakocom/taslos-tasks": + errors.append("Opsle Tasks current repository identity drifted") + if "NEXT primary real-world workload" not in str(opsle_tasks.get("role")): + errors.append("Opsle Tasks must remain the NEXT primary real-world workload") + measurements = opsle_tasks.get("measurements") + if ( + not _string_list(measurements, allow_empty=False) + or set(measurements) != OPSLE_TASKS_MEASUREMENTS + ): + errors.append("Opsle Tasks integrated measurements drifted") + prohibitions = opsle_tasks.get("prohibited_without_separate_authorization") + if ( + not _string_list(prohibitions, allow_empty=False) + or set(prohibitions) != OPSLE_TASKS_PROHIBITIONS + ): + errors.append("Opsle Tasks prohibitions drifted") + for field in ("concept_activation", "later_items", "parked_items"): + value = control.get(field) + if not isinstance(value, list) or not value: + errors.append(f"program_control.{field} must be a nonempty array") + activations = control.get("concept_activation") + if isinstance(activations, list): + for index, activation in enumerate(activations): + if not isinstance(activation, dict): + errors.append(f"concept_activation index {index} must be an object") + continue + if not _nonempty_string(activation.get("deficiency")): + errors.append(f"concept_activation index {index} needs a deficiency") + activation_repositories = activation.get("repositories") + if not _string_list(activation_repositories, allow_empty=False): + errors.append( + f"concept_activation index {index} repositories must be nonempty strings" + ) + elif not set(activation_repositories).issubset(expected): + errors.append( + f"concept_activation index {index} references an unknown repository" + ) + later_items = control.get("later_items") + if ( + not _string_list(later_items, allow_empty=False) + or set(later_items) != LATER_ITEMS + ): + errors.append("program_control later items drifted") + parked_items = control.get("parked_items") + if _string_list(parked_items, allow_empty=False): + required_parked_fragments = ( + "src/cli.js", + "Background projection", + "Historical pre-fix", + "architectural polishing", + ) + for fragment in required_parked_fragments: + if not any(fragment in item for item in parked_items): + errors.append(f"program_control parked item missing {fragment}") + + visible_value = registry.get("visible_value") + if not isinstance(visible_value, dict): + errors.append("registry.visible_value must be an object") + else: + for field in ("baseline", "baseline_rule"): + if not _nonempty_string(visible_value.get(field)): + errors.append(f"visible_value.{field} must be a nonempty string") + value_kinds = visible_value.get("value_kinds") + if not isinstance(value_kinds, list): + errors.append("visible_value.value_kinds must be an array") + value_kinds = [] + class_names = [ + item.get("name") for item in value_kinds if isinstance(item, dict) + ] + if class_names != ["MEASURED", "DERIVED", "ESTIMATED", "UNAVAILABLE"]: + errors.append("visible_value value kinds drifted") + for item in value_kinds: + if not isinstance(item, dict) or not _nonempty_string(item.get("definition")): + errors.append("visible_value value kinds require definitions") + child_fields = visible_value.get("per_child_receipt_fields") + if ( + not _string_list(child_fields, allow_empty=False) + or set(child_fields) != PER_CHILD_VALUE_FIELDS + ): + errors.append("visible_value per-child receipt fields drifted") + summary_fields = visible_value.get("supervisor_summary_fields") + if ( + not _string_list(summary_fields, allow_empty=False) + or set(summary_fields) != SUPERVISOR_VALUE_FIELDS + ): + errors.append("visible_value supervisor summary fields drifted") + for index, repo in enumerate(repositories): label = ( repo.get("name") @@ -534,6 +820,11 @@ def validate( continue if project not in expected: errors.append(f"experiment {label}: references nonexistent project {project}") + elif label not in repo_by_name.get(project, {}).get("active_experiment_ids", []): + errors.append( + f"experiment {label}: participating repository {project} " + "does not reciprocally list the experiment" + ) roles = experiment.get("roles", {}) role_projects: list[Any] = [] if isinstance(roles, dict): @@ -731,13 +1022,6 @@ def validate( f"EXP-001 block coordinator {field} must be " f"{expected_value!r}" ) - research_repository = repo_by_name.get("research", {}) - if coordinator.get("research_release_sha") != research_repository.get( - "last_verified_head_sha" - ): - errors.append( - "EXP-001 block coordinator release SHA must match research" - ) artifacts = ( ( "qualification_artifact",