Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions PROGRAM_STATUS.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,7 +43,7 @@ Program state totals: active 4; waiting 16; complete 0.
| 15 | [agent-recovery-policy](https://github.com/opsle/agent-recovery-policy) | concept | `1b733a111e26` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-recovery-policy/blob/1b733a111e26e0a409fee3b96f627048531daefe/THEORY.md); placeholder only; no automated tests | No shared failure schema, attempt ledger, route evaluator, or comparative fixture set. | After decision evidence and route schemas stabilize, define same-failure convergence on synthetic failures. | `agent-routing-policy`, `agent-state-ledger`, `decision-evidence-protocol` | waiting |
| 16 | [ephemeral-agent-workers](https://github.com/opsle/ephemeral-agent-workers) | concept | `ad96fcfdfac0` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/ephemeral-agent-workers/blob/ad96fcfdfac06d340b5e96d369634980cee78ef4/THEORY.md); placeholder only; no automated tests | Portable authority, claim, and handoff contracts are not ready; no safe synthetic containment harness exists. | Wait for prerequisite contracts, then define a fake worker adapter and destruction receipt without infrastructure changes. | `agent-execution-authorization`, `agent-resource-claims`, `verifiable-agent-handoff` | waiting |
| 17 | [gearbox](https://github.com/opsle/gearbox) | concept | `f3fab9f292cf` | `PROTOTYPED` | [provider-free Python reference core with strict authority-policy admission, exact deterministic argv execution, content-addressed staged helper context, injected one-shot helper transport, passive process waiting, compact results, raw-artifact accounting, fail-closed budgets, and Visible Value receipts](https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/src/opsle_gearbox/core.py); 19 of 19 provider-free automated tests passed locally, in PR #1 CI, and in final-main CI; ruff, shellcheck, actionlint, gitleaks, wheel build, receipt validation, and public raw-locator/hash checks passed | A production-quality bounded helper transport, independently verified isolation and termination, full Context Firewall integration, and a frozen comparative benchmark remain missing. | Freeze a provider-free deterministic-versus-direct baseline and helper-transport conformance corpus before considering any live model/provider run. | `context-firewall`, `decision-evidence-protocol`, `agent-trajectory-profiler`, `agent-routing-policy`, `agent-execution-authorization` | waiting |
| 18 | [research](https://github.com/opsle/research) | program infrastructure | `9ee43197880c` | `PROTOTYPED` | [authoritative 20-repository ledger, machine-readable 17-concept theory registry, canonical theory map, normative Visible Value controls, and provider-free EXP-001 benchmark, launch, and one-block coordinator freezes with six content-addressed tasks, deterministic oracle, four arm contracts, sealed blinded allocation, exact subject configuration and adapter, exact four-authorization admission, private result-envelope templates, exact-revision interoperability, receipts, mutation tests, and integrity CI](program/THEORY_MAP.md); 75 of 75 program-registry, theory-registry, theory-map-hash, generated-status, repository anti-forgetting, EXP-001 freeze/preregistration/coordinator, Gearbox evidence, and Visible Value tests passed locally; PR #13 and exact main CI validated the provider-free coordinator release, and the offline harness replayed all 48 arms | EXP-001 lacks an exact live budgeted authorization set and catalogue/pricing preflight; its one-block coordinator is provider-free qualified, zero model/provider subjects have run, and the program has no canonical measured concept experiment. | Create and provider-free validate one exact four-label LIVE_PROVIDER_RUN authorization set plus a model catalogue/pricing preflight artifact; do not launch a provider/model subject. | — | active |
| 18 | [research](https://github.com/opsle/research) | program infrastructure | `9ee43197880c` | `PROTOTYPED` | [authoritative 20-repository ledger, machine-readable 17-concept theory registry, canonical theory map, normative Visible Value controls, and provider-free EXP-001 benchmark, launch, one-block coordinator, external four-label LIVE_PROVIDER_RUN authorization, and current catalogue/pricing preflight artifacts with six content-addressed tasks, deterministic oracle, four arm contracts, sealed blinded allocation, exact subject configuration and adapter, exact authorization admission, private boundaries, receipts, mutation tests, and integrity CI](program/THEORY_MAP.md); 88 of 88 repository tests pass on the unreleased provider-free live-preflight branch, including 13 authorization validations and two byte-identical replays; PR #13 and exact main CI remain the released coordinator baseline | The exact live authorization set remains unconsumed and unreleased, account-specific API entitlement is unverified under the zero-provider-call policy, no immutable dated model snapshot is documented, and the program has no canonical measured concept experiment. | Independently review and release the provider-free live-authorization and catalogue/pricing preflight; do not consume authorization or launch a provider/model subject. | — | active |
| 19 | [site](https://github.com/opsle/site) | program infrastructure | `28ad65be4750` | `PROTOTYPED` | [React/Vinext source implementation with content routes](https://github.com/opsle/site/blob/28ad65be4750dc849976fbf5c9eae9501c6bbb25/README.md); automated build/render tests present; not rerun because this reconciliation kept other repositories read-only | Wait for validated registry data and measured research; deployment requires separate authorization. | After registry merge, add a read-only registry ingestion design without deploying the site. | `research` | waiting |
| 20 | [.github](https://github.com/opsle/.github) | program infrastructure | `01c38e726db7` | `THEORY` | [documentation-only organization profile](https://github.com/opsle/.github/blob/01c38e726db7c3e45059d25fccce55e071e35938/profile/README.md); not applicable to current single Markdown profile; consistency is unverified | No mechanical registry consistency check exists in this repository. | After registry merge, design a read-only consistency check for organization-profile repository links. | `research` | waiting |

Expand All @@ -55,11 +55,11 @@ Canonical map: `program/THEORY_MAP.md`. Machine registry: `program/theory-regist

## Highest-priority workstream

Prepare EXP-001's final non-launch admission boundary: one exact four-label LIVE_PROVIDER_RUN authorization set plus a model catalogue/pricing preflight artifact, validated without launching any provider/model subject.
Independently review and release the provider-free EXP-001 live-authorization and current catalogue/pricing preflight; do not consume authorization or launch any provider/model subject.

`EXP-001` — **PLANNED** — How much context can an AI coding agent safely not see?

Blockers: No machine-readable four-label budgeted authorization set exists; zero subject/provider runs are authorized or recorded. No distinct dated gpt-5.6-sol snapshot was publicly listed on 2026-08-29, so immutable model weights cannot be claimed and launch requires a catalogue-drift preflight.
Blockers: The validated live authorization set remains unconsumed; separate execution authority is required before any provider/model subject launch. Account-specific API entitlement remains unverified under the zero-provider-call policy and requires a fresh fail-closed launch-time catalogue check. No distinct dated gpt-5.6-sol snapshot was publicly listed on 2026-08-30, so immutable model weights cannot be claimed.

## Exact recommended next execution

Expand Down
7 changes: 7 additions & 0 deletions experiments/exp-001/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -64,6 +64,13 @@ commitments. Its committed qualification evidence is in
Qualification uses non-launchable fixture authorizations and therefore neither
authorizes nor executes a provider request.

The unreleased provider-free live admission preparation is recorded in
[`live-preflight-v1/`](live-preflight-v1/). Its public evidence is in
[`program/evidence/exp-001-live-preflight/`](../../program/evidence/exp-001-live-preflight/).
The actual exact four-label `LIVE_PROVIDER_RUN` set remains external to Git and
outside every subject context. It is validated but unconsumed; this preparation
does not authorize lifecycle advancement or launch a subject.

## Provider-free verification

The harness fails closed unless the three pinned public Opsle dependencies are
Expand Down
29 changes: 29 additions & 0 deletions experiments/exp-001/live-preflight-v1/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
# EXP-001 live authorization and catalogue preflight v1

Status: `PROVIDER_FREE_VALIDATION_ONLY`.

This directory implements the provider-free creation and validation of one
external four-label `LIVE_PROVIDER_RUN` authorization set and the current model
catalogue/pricing preflight for the exact preregistered `gpt-5.6-sol`
configuration.

The live authorization set is intentionally outside Git and outside every
subject context. The public evidence contains its auditable set ID and content
identity but not the block ID, subject labels, per-label authorization IDs,
allocation seed, arm mapping, raw evidence, or result envelopes.

`preflight.py create` is exclusive-create: it refuses an existing destination,
creates four unique label-bound records, marks every record `UNCONSUMED`, and
does not expose any provider transport or execution path. `preflight.py qualify`
validates without mutating the set, reproduces it twice from frozen inputs, and
executes the required fail-closed negative cases in temporary copies.

The catalogue contains only the exact preregistered candidate. Other providers
or models are not eligible without a versioned preregistration amendment. Model
availability and prices come from official OpenAI API documentation; account
entitlement remains unresolved because this task permits zero provider calls.

This preparation is not an EXP-001 run or result. It consumes no authorization,
launches no subject, creates no result envelope, advances no lifecycle state,
and makes no correctness, token, cost, latency, savings, or causal-benefit
claim.
79 changes: 79 additions & 0 deletions experiments/exp-001/live-preflight-v1/model-catalogue.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,79 @@
{
"protocol_version": "opsle.exp001.model-catalogue/v1",
"experiment_id": "EXP-001",
"purpose": "Current provider-free catalogue preflight for the exact preregistered EXP-001 live block configuration; this artifact does not authorize or launch any provider/model subject.",
"retrieved_at": "2026-08-30T18:49:10Z",
"catalogue_scope": "Only the exact preregistered candidate is eligible. Any provider or model substitution requires a versioned preregistration amendment and is outside this catalogue.",
"eligible_candidates": [
{
"provider": "OpenAI API",
"provider_id": "openai",
"model_id": "gpt-5.6-sol",
"model_family": "GPT-5.6",
"wire_api": "Responses API",
"endpoint": "https://api.openai.com/v1/responses",
"availability": "DOCUMENTED_API_MODEL_ACCOUNT_ACCESS_UNVERIFIED",
"availability_detail": "The official model page documents gpt-5.6-sol on v1/responses and lists API rate tiers. This provider-free task did not query account-specific entitlement.",
"pricing_basis": "USD_PER_1M_TEXT_TOKENS_STANDARD_SHORT_CONTEXT",
"input_price_usd": 4.0,
"cached_input_price_usd": 0.4,
"cache_write_price_usd": 5.0,
"output_price_usd": 20.0,
"long_context_input_price_usd": 8.0,
"long_context_cached_input_price_usd": 0.8,
"long_context_cache_write_price_usd": 10.0,
"long_context_output_price_usd": 30.0,
"long_context_threshold_input_tokens": 272000,
"context_window_tokens": 1050000,
"max_output_tokens": 128000,
"subscription_api_distinction": "Direct API usage pricing; ChatGPT or Codex subscription access is not used and does not authorize this experiment.",
"snapshot_id": null,
"pricing_status": "AUTHORITATIVELY_DOCUMENTED_FOR_DIRECT_API_STANDARD_TIER",
"pricing_promotion": "Official documentation states promotional pricing is available at least through 2026-11-21; a fresh launch-time check remains mandatory.",
"sources": [
{
"publisher": "OpenAI",
"title": "GPT-5.6 Sol Model",
"url": "https://developers.openai.com/api/docs/models/gpt-5.6-sol",
"retrieved_at": "2026-08-30T18:49:10Z",
"supports": [
"exact model identifier",
"model family",
"Responses API availability",
"context and output limits",
"short-context prices",
"long-context threshold",
"snapshot listing"
]
},
{
"publisher": "OpenAI",
"title": "OpenAI API Pricing",
"url": "https://developers.openai.com/api/docs/pricing",
"retrieved_at": "2026-08-30T18:49:10Z",
"supports": [
"standard short-context token prices",
"standard long-context token prices",
"cached-input and cache-write prices",
"promotional pricing horizon"
]
}
],
"uncertainty_or_unavailable_fields": [
{
"field": "account_specific_api_entitlement",
"value": null,
"status": "UNRESOLVED_ZERO_PROVIDER_CALL_POLICY"
},
{
"field": "distinct_dated_snapshot_id",
"value": null,
"status": "UNAVAILABLE_OFFICIAL_PAGE_LISTS_ONLY_PUBLIC_MODEL_ID"
}
]
}
],
"authorization_effect": "NONE",
"provider_model_launch_count": 0,
"identity": "sha256:334a48850dacf3d9ad5d5c9088287fecc78925e64e10f974cb63575b74592b3c"
}
Loading