Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions OPSLE_TASKS_MIGRATION_PLAN.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,12 @@

> **NOT AUTHORIZED FOR EXECUTION DURING THIS RUN. PLANNING ONLY.**

Opsle Tasks is the NEXT primary real-world workload after Durable Supervisor
v0.1 is declared and feature-frozen. That workload role does not authorize any
identity, repository, runtime, schema, service, DNS, TLS, provider, or release
migration. The machine-readable program priority remains
`program/registry.json`.

## Intended future identities

- repository: `sneakocom/taslos-tasks` → `opsle/tasks`;
Expand Down
4 changes: 4 additions & 0 deletions OPSLE_TASKS_PUBLIC_RELEASE.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,10 @@

> Planning only. Taslos Tasks remains private in its existing location. Publication was not executed.

This release is LATER program work. It does not begin merely because Opsle Tasks
becomes the NEXT Durable Supervisor workload; it requires separate explicit
release authority and the eligibility evidence below.

## Full-history audit

Audit every reachable commit, tag, branch, release asset, PR artifact, LFS object, submodule reference, and archive for credentials, tokens, private keys, connection strings, private addresses, environment secrets, database data, backups, logs, sensitive screenshots, and accidental artifacts. Removing a value from HEAD is insufficient if reachable history retains it.
Expand Down
84 changes: 50 additions & 34 deletions PROGRAM_STATUS.md

Large diffs are not rendered by default.

7 changes: 4 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,8 +30,9 @@ Important mechanisms should remain understandable, falsifiable, benchmarkable, r

## Start here

- [PROGRAM_STATUS.md](PROGRAM_STATUS.md) — generated 20-repository dashboard.
- [program/registry.json](program/registry.json) — authoritative machine-readable ledger.
- [PROGRAM_STATUS.md](PROGRAM_STATUS.md) — generated 21-repository dashboard.
- [program/registry.json](program/registry.json) — authoritative machine-readable portfolio and priority ledger.
- [program/PRIORITY.md](program/PRIORITY.md) — generated NOW / NEXT / THEN / LATER / PARKED view.
- [program/THEORY_MAP.md](program/THEORY_MAP.md) — canonical conceptual topology and Gearbox boundary.
- [program/theory-registry.json](program/theory-registry.json) — machine-readable concept classifications and dispositions.
- [program/experiments.json](program/experiments.json) — canonical experiment registry.
Expand All @@ -49,7 +50,7 @@ Important mechanisms should remain understandable, falsifiable, benchmarkable, r
- [OPSLE_SITE_PLAN.md](OPSLE_SITE_PLAN.md) — future opsle.com content architecture.

Validate ledger integrity with `python3 tools/validate_program.py`. Regenerate the
dashboard with `python3 tools/render_program_status.py`.
dashboard and priority view with `python3 tools/render_program_status.py`.

## Product relationship

Expand Down
17 changes: 15 additions & 2 deletions program/OPERATING_RULES.md
Original file line number Diff line number Diff line change
@@ -1,10 +1,15 @@
# Program operating rules

`program/registry.json` is authoritative for both portfolio state and the
NOW / NEXT / THEN / LATER / PARKED priority state. Generated Markdown is never
an independent planning authority.

Every Opsle execution must:

1. Read `program/registry.json` before selecting or performing work.
2. Verify the relevant repository default branch and HEAD before relying on recorded state.
3. Work against an explicit repository or experiment objective.
3. Start from the current program lane and operating question before selecting
an explicit repository or experiment objective.
4. Preserve immutable or content-addressed evidence for every material claim.
5. Promote lifecycle state only after satisfying the canonical gate in `program/LIFECYCLE.md`.
6. Update the registry when verified state changes.
Expand All @@ -22,10 +27,18 @@ Every Opsle execution must:
13. Classify everyday telemetry as observational. Do not relabel an accumulated
production corpus as causal or `EXPERIMENTAL` evidence without the controlled
method required by `program/VISIBLE_VALUE_CONTRACT.md`.
14. Do not create a work item solely because an implementation can be improved.
New work normally requires a violated invariant, demonstrated defect,
measured inefficiency, missing capability blocking the current program
objective, experiment requirement, security or safety issue, or externally
required release condition.
15. Park cosmetic cleanup, architectural taste, hypothetical robustness, and
speculative future requirements unless qualifying evidence appears.

Run `python3 tools/validate_program.py` and
`python3 tools/render_program_status.py --check` before committing a registry
change.
change. The renderer checks both `PROGRAM_STATUS.md` and
`program/PRIORITY.md`.

## Portfolio discipline

Expand Down
240 changes: 144 additions & 96 deletions program/PRIORITY.md
Original file line number Diff line number Diff line change
@@ -1,99 +1,147 @@
# Evidence-driven execution order

The portfolio is a dependency-aware set of parallel workstreams, not a serial
checklist.

## Workstream 0: canonical Gearbox home

The 2026-08-29 reconciliation established that Agent Gearbox is a coherent
primary-developer intelligence-and-context transmission capability and that no
current repository owns its irreducible mechanism. Routing, execution
authorization, and resource claims are supporting policies; Durable Supervisor,
ledger, scheduler, wakeup, discovery, and recovery solve autonomous durable
orchestration instead.

Completed on 2026-08-29 through public `opsle/gearbox` PR #1. The repository now
contains a narrow provider-free prototype, exact Taslos source provenance,
provider-free tests, and revision-bound release evidence. No existing repository
was consolidated, renamed, transferred, archived, or deleted; Taslos Tasks
remained unchanged; zero model/provider subjects ran.

## Workstream 1: EXP-001 prerequisites and experiment

1. Context Firewall now emits deterministic Visible Value receipts and a named
operator indicator while keeping its compact packet canonical.
2. Decision Evidence now independently validates those packets and receipts and
emits its own validation receipt and named operator indicator.
3. Agent Trajectory Profiler now ingests compatible receipts and produces
deterministic class/unit/trust-safe per-run and cumulative summaries.
4. Completed through public research PR #9: six content-addressed tasks, the
deterministic correctness oracle, raw plus three Context Firewall arm
contracts, the balanced/blinded allocation method, and the provider-free
exact-revision harness are frozen and qualified.
5. Completed through public research PR #11: the exact OpenAI Responses subject
configuration, medium reasoning effort, no-retry adapter, 10-repetition fixed
sample, 240-label blinded index, 60-block encrypted mapping, seed commitment,
stopping rules, and provider-free verification are preregistered. The
coordinator seed remains outside Git and subject context. Zero model/provider
subjects ran. Add Verifiable Agent Handoff only if a future arm destroys or
isolates the source environment.
6. Completed through public research PR #13: the provider-free one-block
coordinator unseals one mapping outside subject context, validates an exact
four-label authorization set, renders all four arms, prepares empty result
envelopes, retains private artifacts outside model context, and returns only
seed-keyed commitments and exact counts. Secret-backed deterministic replay
used fixture-only authorizations, proved zero canonical arm identifiers in
subject-visible artifacts, and ran zero provider/model subjects.

These three foundational projects are naturally tested together. The expected
first major measured experiment remains EXP-001 because no repository evidence
establishes a stronger prerequisite experiment. The coordinator prerequisite is
now qualified. The experiment itself must still wait for four exact label-bound
budget authorizations and a current catalogue/pricing preflight under separate
provider/model authority.

## Workstream 2: durable orchestration

Develop Agent State Ledger and the portable core of Agent Scheduler Runtime,
then test Event-Driven Agent Wakeup with Durable Supervisor. One restart,
duplicate-event, and reconstruction campaign can exercise all four mechanisms.

## Workstream 3: bounded execution and verification

Develop Agent Resource Claims and Agent Execution Authorization before combining
Ephemeral Agent Workers, Verifiable Agent Handoff, and Controlled Agent
Acceptance. This sequence makes lease/fence authority and evidence survival
testable before any real-provider acceptance is considered.

## Workstream 4: routing and recovery

Agent Routing Policy and Agent Recovery Policy should wait for durable state,
structured decision evidence, and authorization inputs. Their comparative study
can share failure fixtures and evaluate retry, alternate route, and terminal
decisions without invoking live providers initially.

## Workstream 5: editing and discovery

Semantic Edit Protocol should use Agent Trajectory Profiler for correctness-gated
payload/churn measurement and Agent Resource Claims for concurrent-region cases.
Agent Discovery Control should wait for a portable ledger and supervisor fixture
so duplicate convergence and already-satisfied proofs are durable rather than
conversation-local.

## Workstream 6: affected verification

Affected Verification is independently verified for its narrow planner and the
AV-EXP-001 Zustand/Vitest shadow calibration, and remains outside Gearbox. Both
AV arms selected 8/8 oracle-relevant checks in the frozen ten-scenario corpus;
the native tests-only arm selected 6/8 and omitted relevant lint and typecheck.
This is one-ecosystem calibration evidence, not general safety or bounded trust.
Its exact next execution is a preregistered second public-repository shadow
calibration in a different ecosystem with a meaningful native selector and the
same full-catalog oracle discipline.
# Evidence-driven program priority

<!-- Generated by tools/render_program_status.py from program/registry.json. -->

`program/registry.json` is the sole priority authority. Edit the registry, not this file.

## Operating question

> What prevents Durable Supervisor from successfully finishing Opsle Tasks with less intelligence, less context, less human involvement, and defensible evidence?

## Anti-nitpick guardrail

Do not create a work item solely because an implementation can be improved.

Normally admit new work only for:

- violated invariant
- demonstrated defect
- measured inefficiency
- missing capability blocking the current program objective
- experiment requirement
- security or safety issue
- externally required release condition

Park by default:

- cosmetic cleanup
- architectural taste
- hypothetical robustness
- speculative future requirement

## NOW / NEXT / THEN / LATER / PARKED

| Lane | Objective | Repositories | Entry gate |
|---|---|---|---|
| **NOW** | Finish Durable Supervisor v0.1 as a bounded measured system, then freeze feature work. | `durable-supervisor`, `research` | Current program lane. |
| **NEXT** | Use Opsle Tasks as the primary real-world workload and collect integrated measurements without performing the deferred Taslos-to-Opsle migration. | `gearbox`, `context-firewall`, `decision-evidence-protocol`, `agent-trajectory-profiler`, `affected-verification` | Durable Supervisor v0.1 is declared and feature-frozen. |
| **THEN** | Advance an individual concept only when Durable Supervisor or Opsle Tasks evidence demonstrates a concrete need. | `semantic-edit-protocol`, `event-driven-agent-wakeup`, `agent-state-ledger`, `agent-scheduler-runtime`, `verifiable-agent-handoff`, `agent-routing-policy`, `agent-resource-claims`, `agent-discovery-control`, `agent-execution-authorization`, `controlled-agent-acceptance`, `agent-recovery-policy`, `ephemeral-agent-workers` | A qualifying work-item reason and real workload evidence identify the smallest relevant concept boundary. |
| **LATER** | Run controlled experiments, freeze a real-workload benchmark corpus, seek independent replication, and only then consider public product and research-site releases. | `site` | NOW, NEXT, and evidence-triggered THEN work establish a defensible need and release prerequisites are separately authorized. |
| **PARKED** | Retain useful non-priority ideas without turning them into active work. | `.github` | The idea is useful but lacks a qualifying reason to compete with the current objective. |

### NOW — Finish Durable Supervisor v0.1 as a bounded measured system, then freeze feature work.

Entry: Current program lane.

Exit: Every Durable Supervisor v0.1 stopping criterion is satisfied and the release is explicitly declared and frozen.

### NEXT — Use Opsle Tasks as the primary real-world workload and collect integrated measurements without performing the deferred Taslos-to-Opsle migration.

Entry: Durable Supervisor v0.1 is declared and feature-frozen.

Exit: Durable Supervisor has driven the remaining authorized Opsle Tasks readiness work and produced defensible integrated measurements.

### THEN — Advance an individual concept only when Durable Supervisor or Opsle Tasks evidence demonstrates a concrete need.

Entry: A qualifying work-item reason and real workload evidence identify the smallest relevant concept boundary.

Exit: The demonstrated deficiency is resolved or falsified at its narrowest justified boundary; do not march mechanically through the portfolio.

### LATER — Run controlled experiments, freeze a real-workload benchmark corpus, seek independent replication, and only then consider public product and research-site releases.

Entry: NOW, NEXT, and evidence-triggered THEN work establish a defensible need and release prerequisites are separately authorized.

Exit: Applicable evidence and separate release authorization exist.

### PARKED — Retain useful non-priority ideas without turning them into active work.

Entry: The idea is useful but lacks a qualifying reason to compete with the current objective.

Exit: New evidence supplies a qualifying work-item reason.

## Durable Supervisor v0.1 stopping criteria

Status: **IN_PROGRESS**. Foundation: `VERIFIED_ENOUGH_TO_STOP_NITPICKING_UNLESS_REAL_WORKLOAD_EXPOSES_A_DEFECT`.

Verified main: `1b5ab7631ba651a32592bbbdab8001865a3baf3d`. Runtime: `PAUSED_NO_ACTIVE_TASK_OR_ATTEMPT` as durably recorded at `2026-09-05T12:20:01.551Z`.

1. **OPEN** — Expose the selected Gearbox route, routing rationale, and Context Firewall accounting in an operator-visible per-child receipt.
2. **OPEN** — Expose each child's model and reasoning effort.
3. **OPEN** — Record actual tokens and cost when provider evidence supplies them, and clearly labeled estimates otherwise.
4. **OPEN** — Measure raw evidence or context versus Context Firewall retained context and report the reduction.
5. **OPEN** — Report first-pass success rate.
6. **OPEN** — Report repair-child or retry rate and the token cost of retries.
7. **OPEN** — Estimate avoidable intelligence consumption by identifying failures whose needed facts were discoverable through deterministic preflight before model execution.
8. **OPEN** — Implement only the smallest useful deterministic preflight or reconnaissance mechanism shown necessary by the avoidable-failure evidence.
9. **OPEN** — Complete another real foreign-repository workload that exercises and preserves the new measurements.
10. **OPEN** — Declare Durable Supervisor v0.1 and freeze feature work except for defects exposed by real workloads.

## Opsle Tasks boundary

Opsle Tasks is the NEXT primary real-world workload after Durable Supervisor v0.1. Its current repository remains `sneakocom/taslos-tasks`.

Measure: Gearbox, Context Firewall, Decision Evidence Protocol, Agent Trajectory Profiler, Affected Verification.

Without separate authorization, do not:

- move apps/taslos-tasks
- transfer sneakocom/taslos-tasks
- rename production services
- change schemas merely for rebranding
- public release
- DNS or TLS changes
- launch provider work

## Evidence-triggered concept activation

- routing → `agent-routing-policy`
- durable state → `agent-state-ledger`
- retry or deterministic preflight → `gearbox`
- context overload → `context-firewall`
- verification selection → `affected-verification`
- recovery → `agent-recovery-policy`

## Visible Value target

Baseline: The configured primary supervisor model and reasoning effort performs all child work itself and receives raw unfiltered evidence.

Savings require an inspectable same-work baseline or a clearly labeled estimate derived from that baseline; marketing counterfactuals are prohibited.

- **MEASURED** — The value comes directly from deterministic artifacts or provider usage records and still carries the applicable Visible Value evidence class, normally EXACT or OBSERVED.
- **DERIVED** — The value is a reproducible calculation over identified measured inputs and records its method and applicable Visible Value evidence class.
- **ESTIMATED** — The value carries the ESTIMATED evidence class and names its method, assumptions, pricing source, and uncertainty.
- **UNAVAILABLE** — The value is not currently supported and must remain absent rather than being rendered as zero.

Per-child receipt: child/task identity; model; reasoning effort; Gearbox route; routing rationale; input tokens; output tokens; raw evidence/context size; Context Firewall retained size; reduction percentage; estimated tokens avoided; estimated cost avoided; duration; attempt number; success/failure; escalation/retry reason; retry potentially avoidable through deterministic preflight.

Supervisor/run summary: total children; model/effort distribution; total model tokens consumed; work completed deterministically without model use; Context Firewall reduction; estimated token/cost savings; first-pass success rate; repair-child rate; tokens spent on retries; avoidable-intelligence estimate.

## Later

- controlled empirical experiments
- frozen real-workload benchmark corpus
- independent replication
- Opsle Tasks public and self-hosted release
- opsle.com research and public site
- hosted Opsle offering

## Parked

- Durable Supervisor src/cli.js is 1,570 lines at the verified main SHA and may eventually warrant decomposition; this is a maintenance smell, not a current objective, unless measured work shows material reliability or efficiency harm.
- Background projection repair or reconciliation is not justified merely because explicit retry exists.
- Historical pre-fix cleanup or migration requires evidence of need.
- General architectural polishing remains parked unless a real workload exposes a concrete defect.

## Exact next execution

In `opsle/research`, independently review and release the provider-free exact
four-label `LIVE_PROVIDER_RUN` authorization set and current model
catalogue/pricing preflight branch. Do not consume authorization or launch a
provider/model subject. EXP-001 has no technical dependency on Gearbox.
In opsle/durable-supervisor, make the selected Gearbox route, routing rationale, and Context Firewall raw-versus-retained accounting operator-visible in each child receipt without broadening Durable Supervisor v0.1.
Loading