diff --git a/SEO.md b/SEO.md index 0da3066..9c48515 100644 --- a/SEO.md +++ b/SEO.md @@ -49,6 +49,7 @@ that page's frontmatter, never here. | `docs/production/migrations` | ctrlrun schema migration | Migrations run at open, forward only, with no flag that opens a database un-migrated. | | `docs/production/recovery` | agent crashed mid action | A restarted process repairs nothing and cannot know the holder is dead. | | `docs/production/receipt-integrity` | verify receipt chain | Run ctrlrun receipts --verify-chain and read the six names it can report. | +| `docs/production/anchoring` | anchor receipt chain outside database | Anchor the chain's head where your database's writer cannot reach it, and what that does not prove. | | `docs/production/soak` | ctrlrun soak test results | One published run, its measured duration, and the exit criterion it does not meet. | | `docs/production/operations` | ctrlrun monitoring | Watch how many effects are sitting in an unknown outcome that nobody has answered. | | `docs/mcp/overview` | MCP gateway human approval | CTRLRun works with MCP in four ways. | diff --git a/docs.json b/docs.json index 50da4a4..d1cb573 100644 --- a/docs.json +++ b/docs.json @@ -81,6 +81,7 @@ "docs/production/migrations", "docs/production/recovery", "docs/production/receipt-integrity", + "docs/production/anchoring", "docs/production/soak", "docs/production/operations" ] @@ -485,6 +486,11 @@ "destination": "/docs/production/receipt-integrity", "permanent": true }, + { + "source": "/production/anchoring", + "destination": "/docs/production/anchoring", + "permanent": true + }, { "source": "/production/soak", "destination": "/docs/production/soak", diff --git a/docs.mdx b/docs.mdx index 9ffb2d3..8353ae4 100644 --- a/docs.mdx +++ b/docs.mdx @@ -219,8 +219,8 @@ the framework's own interrupt, and a framework with no such primitive does not n {/* generated from the suite, pyproject and the soak (mdx) — run the generator */} - **Version 0.10.0**, on [PyPI](https://pypi.org/project/ctrlrun/), Python 3.11 and later, tested on 3.11 to 3.14. -- **6,080 tests**, every version specified before it was written and every requirement mutation-tested. -- **27 guarantees you can check in your own setup**, with `ctrlrun verify` against your policy, on your store's backend, in a scratch store it creates. +- **6,158 tests**, every version specified before it was written and every requirement mutation-tested. +- **29 guarantees you can check in your own setup**, with `ctrlrun verify` against your policy, on your store's backend, in a scratch store it creates. - **One host: a file.** SQLite, no server, no ops. **Many hosts: Postgres**, the same guarantees, graded by the same suite. - **Soaked for 20m 0s on postgres**: 889,735 actions, 0 unattributed ambiguous outcomes, positive control fired. Nothing here establishes what only accumulates over days. [What it does not establish](https://ctrlrun.dev/docs/production/soak). - **Each receipt carries the hash of the one before it**, so an alteration is detected and named. diff --git a/docs/CLAIMS.md b/docs/CLAIMS.md index a52d6ce..c6f2bcf 100644 --- a/docs/CLAIMS.md +++ b/docs/CLAIMS.md @@ -25,17 +25,17 @@ by its quoted claim, and `tests/test_docs_audit.py` fails if a named row is not | Claim | Code | Proof | |---|---|---| -| "The last check before an AI agent does something it can't undo." | `Control.execute` — `control.py:1304` — resolves the principal, evaluates authority and policy, consumes the approval and reserves the effect key **before** the executor runs; nothing in the wrapper calls the function first | `test_T1_a_blind_retry_is_refused_and_never_reaches_the_remote`, `test_T3_the_fake_remote_is_called_exactly_once` | +| "The last check before an AI agent does something it can't undo." | `Control.execute` — `control.py:1318` — resolves the principal, evaluates authority and policy, consumes the approval and reserves the effect key **before** the executor runs; nothing in the wrapper calls the function first | `test_T1_a_blind_retry_is_refused_and_never_reaches_the_remote`, `test_T3_the_fake_remote_is_called_exactly_once` | | "Autonomy belongs to the action, not the agent." | `Policy.evaluate(action)` — `policy.py:657` — passes only the action's **name and arguments** to `_ActionPolicy.evaluate` (`policy.py:657`), whose signature has no principal in it. A rule cannot read who is acting even by accident. `agent_eq` and `user_eq` are refused at load by `RESERVED_ARGUMENTS` (`policy.py:253`) rather than silently matching nothing. | `test_T6_an_action_name_is_matched_exactly`, `test_a_condition_naming_an_action_field_is_refused_at_load` | -| "A consequential action happens at most once, exactly as approved, and leaves a receipt — and when the outcome is unknown, CTRLRun says so instead of guessing." | At most once: `plan_reservation` — `effect.py:250`. Exactly as approved: the approval is bound to `action_hash` and consumed with the reservation — `_authorize_and_reserve` — `state.py:1198`. Or not at all: a refusal raises before the executor — `Control.execute` — `control.py:1304`. Says so instead of guessing: only `NotExecuted` maps to `FAILED` — `_outcome` — `control.py:2209` — and everything else is `AMBIGUOUS`. A receipt: `Receipt` — `receipt.py:252`. **This sentence read *happens once … or not at all* until 0.6**, a two-way disjunction that excluded the third outcome the product exists for: a lost reply is neither, and the README's own first section says so. | `test_T3_exactly_one_agent_reserves_and_seven_are_blocked`, `test_T2_a_mutated_action_presenting_the_approval_raises_ApprovalMismatch`, `test_T1_a_lost_response_leaves_the_effect_ambiguous`, `test_T11_every_demo_receipt_carries_every_field_in_the_spec` | +| "A consequential action happens at most once, exactly as approved, and leaves a receipt — and when the outcome is unknown, CTRLRun says so instead of guessing." | At most once: `plan_reservation` — `effect.py:250`. Exactly as approved: the approval is bound to `action_hash` and consumed with the reservation — `_authorize_and_reserve` — `state.py:1255`. Or not at all: a refusal raises before the executor — `Control.execute` — `control.py:1318`. Says so instead of guessing: only `NotExecuted` maps to `FAILED` — `_outcome` — `control.py:2223` — and everything else is `AMBIGUOUS`. A receipt: `Receipt` — `receipt.py:252`. **This sentence read *happens once … or not at all* until 0.6**, a two-way disjunction that excluded the third outcome the product exists for: a lost reply is neither, and the README's own first section says so. | `test_T3_exactly_one_agent_reserves_and_seven_are_blocked`, `test_T2_a_mutated_action_presenting_the_approval_raises_ApprovalMismatch`, `test_T1_a_lost_response_leaves_the_effect_ambiguous`, `test_T11_every_demo_receipt_carries_every_field_in_the_spec` | | "A Python library that sits between the decision to act and the call that acts." | `@protect` — `control.py` — wraps the callable that acts, and `Control.execute` runs every check before invoking it. The category noun was on `docs.mdx` and in `pyproject.toml`'s `description` and nowhere in the README until 0.6, so a reader had to infer what CTRLRun **is** from three slogans. | `test_the_header_carries_the_fixed_copy_and_the_five_badges`, `test_T1_a_blind_retry_is_refused_and_never_reaches_the_remote` | -| "Runs in production on a single file, or on Postgres across hosts" | SQLite: `SQLiteStateStore` reserves inside the `BEGIN IMMEDIATE` of `_authorize_and_reserve` — `state.py:1478` — which is a write lock on the file and holds across OS processes. Postgres: `PostgresStateStore` over `UNIQUE(effect_key)` with `INSERT … ON CONFLICT DO NOTHING` and checked row counts (SPEC-v0.6 §4.2), the same `StateStore` protocol, extended by nothing | `test_T3_exactly_one_agent_reserves_and_seven_are_blocked` (8 OS processes, both backends), `test_T141_the_shipped_backends_pass`, `test_T154_postgres_passes_the_store_conformance_suite` | +| "Runs in production on a single file, or on Postgres across hosts" | SQLite: `SQLiteStateStore` reserves inside the `BEGIN IMMEDIATE` of `_authorize_and_reserve` — `state.py:1535` — which is a write lock on the file and holds across OS processes. Postgres: `PostgresStateStore` over `UNIQUE(effect_key)` with `INSERT … ON CONFLICT DO NOTHING` and checked row counts (SPEC-v0.6 §4.2), the same `StateStore` protocol, extended by nothing | `test_T3_exactly_one_agent_reserves_and_seven_are_blocked` (8 OS processes, both backends), `test_T141_the_shipped_backends_pass`, `test_T154_postgres_passes_the_store_conformance_suite` | ## The refund that happened twice | Claim | Code | Proof | |---|---|---| -| "A lost reply is `AMBIGUOUS`, never `FAILED`, and a retry against an `AMBIGUOUS` effect is refused — until a human, or a `reconcile` hook, says what happened." | Only `NotExecuted` maps to `FAILED` — `_outcome` — `control.py:1311`; a retry against an `AMBIGUOUS` key is refused by `plan_reservation` — `effect.py:250`; the two things permitted to move the record on and nothing else — `resolve` — `cli/main.py:591` — and `Control._reconciled` — `control.py:2863` | `test_T1_a_lost_response_leaves_the_effect_ambiguous`, `test_T1_a_blind_retry_is_refused_and_never_reaches_the_remote`, `test_T160_there_is_no_reaper`, `test_T13_a_hook_answering_not_executed_moves_the_record_to_failed` | +| "A lost reply is `AMBIGUOUS`, never `FAILED`, and a retry against an `AMBIGUOUS` effect is refused — until a human, or a `reconcile` hook, says what happened." | Only `NotExecuted` maps to `FAILED` — `_outcome` — `control.py:1325`; a retry against an `AMBIGUOUS` key is refused by `plan_reservation` — `effect.py:250`; the two things permitted to move the record on and nothing else — `resolve` — `cli/main.py:743` — and `Control._reconciled` — `control.py:2877` | `test_T1_a_lost_response_leaves_the_effect_ambiguous`, `test_T1_a_blind_retry_is_refused_and_never_reaches_the_remote`, `test_T160_there_is_no_reaper`, `test_T13_a_hook_answering_not_executed_moves_the_record_to_failed` | | "The customer is refunded twice, and nothing in the stack noticed." — said of a stack without CTRLRun; the demo runs the same sequence with it, and counts the calls the remote received | `ctrlrun demo` scenario 1, which retries against a fake remote that counts its calls and prints the count | `test_T3_the_fake_remote_is_called_exactly_once`, `test_T1_a_blind_retry_is_refused_and_never_reaches_the_remote` | ## Protect your first action @@ -50,24 +50,24 @@ by its quoted claim, and `tests/test_docs_audit.py` fails if a named row is not | Claim | Code | Proof | |---|---|---| -| "A lost reply is `AMBIGUOUS`, never `FAILED`, and a retry against an `AMBIGUOUS` effect is refused." | Only `NotExecuted` maps to `FAILED` — `_outcome` — `control.py:2209`; `plan_reservation` — `effect.py:250` | `test_T1_a_lost_response_leaves_the_effect_ambiguous`, `test_T1_a_blind_retry_is_refused_and_never_reaches_the_remote` | -| "reserved atomically across processes and hosts; one worker wins" | `reserve_effect` — `state.py:639`; `PostgresStateStore.reserve_effect` — `postgres.py:742` | `test_T3_exactly_one_agent_reserves_and_seven_are_blocked`, `test_T154_postgres_passes_the_store_conformance_suite` | -| "bound to the hash of the exact action a human saw, used once, and refused for anything else" | `_authorize_and_reserve` — `state.py:1198` | `test_T2_a_mutated_action_presenting_the_approval_raises_ApprovalMismatch`, `test_T4_replaying_the_approval_raises_ApprovalMismatch_with_reason_consumed` | +| "A lost reply is `AMBIGUOUS`, never `FAILED`, and a retry against an `AMBIGUOUS` effect is refused." | Only `NotExecuted` maps to `FAILED` — `_outcome` — `control.py:2223`; `plan_reservation` — `effect.py:250` | `test_T1_a_lost_response_leaves_the_effect_ambiguous`, `test_T1_a_blind_retry_is_refused_and_never_reaches_the_remote` | +| "reserved atomically across processes and hosts; one worker wins" | `reserve_effect` — `state.py:641`; `PostgresStateStore.reserve_effect` — `postgres.py:752` | `test_T3_exactly_one_agent_reserves_and_seven_are_blocked`, `test_T154_postgres_passes_the_store_conformance_suite` | +| "bound to the hash of the exact action a human saw, used once, and refused for anything else" | `_authorize_and_reserve` — `state.py:1255` | `test_T2_a_mutated_action_presenting_the_approval_raises_ApprovalMismatch`, `test_T4_replaying_the_approval_raises_ApprovalMismatch_with_reason_consumed` | | "An action the policy does not list is denied" | `Policy.evaluate` — `policy.py:657` | `test_T6_unknown_action_is_denied_with_reason_unknown_action` | -| "Authority first ... then policy" / "authority first" | `Control.execute` evaluates authority before policy and a denial appends `AUTHORITY_DENIED` and never `POLICY_EVALUATED` — `control.py:1304` | `test_T74_a_denial_leaves_no_pending_approval_request` | +| "Authority first ... then policy" / "authority first" | `Control.execute` evaluates authority before policy and a denial appends `AUTHORITY_DENIED` and never `POLICY_EVALUATED` — `control.py:1318` | `test_T74_a_denial_leaves_no_pending_approval_request` | | "Neither axis reads the agent's instructions" | `Policy.evaluate` — `policy.py:657` — sees the action's name and arguments; `Authority.evaluate` — `authority.py:1296` — sees the action and the principal; neither is handed a prompt, a message or a tool result | `test_T6_an_action_name_is_matched_exactly`, `test_T67_a_principal_with_no_grant_is_denied` | | "canonical arguments (sorted keys, no floats) ... Its SHA-256 is the action hash" | `canonicalize` / `action_hash` — `action.py`; `float` refused at any depth — `action.py:81` | `test_T7_canonical_form_is_exactly_the_specified_serialization`, `test_T7_nested_dicts_are_sorted_recursively` | -| "The approval is single-use, expires, and matches nothing but that exact action." | `_authorize_and_reserve` — `state.py:1198` — checks expiry at consumption | `test_T5_expiry_is_checked_at_consumption_not_only_at_grant`, `test_T4_replaying_the_approval_raises_ApprovalMismatch_with_reason_consumed` | -| "Only `NotExecuted`, raised by you, means `FAILED`." | `_outcome` — `control.py:2209`; `NotExecuted` — `errors.py:159` | `test_T1_a_lost_response_leaves_the_effect_ambiguous` | +| "The approval is single-use, expires, and matches nothing but that exact action." | `_authorize_and_reserve` — `state.py:1255` — checks expiry at consumption | `test_T5_expiry_is_checked_at_consumption_not_only_at_grant`, `test_T4_replaying_the_approval_raises_ApprovalMismatch_with_reason_consumed` | +| "Only `NotExecuted`, raised by you, means `FAILED`." | `_outcome` — `control.py:2223`; `NotExecuted` — `errors.py:159` | `test_T1_a_lost_response_leaves_the_effect_ambiguous` | | "the hash of the policy that decided it, chained to the receipt before it" | `Policy.policy_hash` — `policy.py:795`; `prev_hash`, `GENESIS_HASH` for the first — `receipt.py:127` | `test_T172_every_receipt_carries_the_hash_and_the_declared_version`, `test_T164_an_altered_receipt_is_content_altered_at_its_seq` | ## Three ways to use it | Claim | Code | Proof | |---|---|---| -| "You probably do not need an adapter" | Three ways in, and `@protect` (`control.py:4939`) covers this process while the gateway covers MCP — an adapter buys only the interrupt | `test_T139_the_adapter_section_says_when_you_do_not_need_one_up_front` | -| "`ctrlrun init` writes a starter" | `init` — `cli/main.py:357` | CI's `package` job runs `ctrlrun init` from the wheel and asserts `ctrlrun.yaml` exists | -| "The human runs `ctrlrun approve ` and the agent calls again inside `ctrlrun.with_approval(request_id)`" | `approve` — `cli/main.py:381`; `with_approval` — `control.py:392`; `ApprovalRequired` (`errors.py:90`) carries `request_id` | `test_T2_a_mutated_action_presenting_the_approval_raises_ApprovalMismatch` (the granted path first), `test_T4_replaying_the_approval_raises_ApprovalMismatch_with_reason_consumed` | +| "You probably do not need an adapter" | Three ways in, and `@protect` (`control.py:4969`) covers this process while the gateway covers MCP — an adapter buys only the interrupt | `test_T139_the_adapter_section_says_when_you_do_not_need_one_up_front` | +| "`ctrlrun init` writes a starter" | `init` — `cli/main.py:367` | CI's `package` job runs `ctrlrun init` from the wheel and asserts `ctrlrun.yaml` exists | +| "The human runs `ctrlrun approve ` and the agent calls again inside `ctrlrun.with_approval(request_id)`" | `approve` — `cli/main.py:391`; `with_approval` — `control.py:446`; `ApprovalRequired` (`errors.py:90`) carries `request_id` | `test_T2_a_mutated_action_presenting_the_approval_raises_ApprovalMismatch` (the granted path first), `test_T4_replaying_the_approval_raises_ApprovalMismatch_with_reason_consumed` | | "No agent changes" | `INTERCEPTED_METHOD` is `tools/call` and every other method is relayed unchanged — `gateway/mcp.py:42` | `test_a_non_intercepted_method_is_relayed_with_no_ctrlrun_outcome` | | "Point the MCP client at the gateway instead of at the tool server" | `Gateway.handle` — `gateway/server.py:399`; `serve` — `gateway/__init__.py:45` | `test_T19_the_upstream_receives_the_canonical_arguments` | | "Tools become actions named `mcp..`" | `Gateway._intercept` — `gateway/server.py:442` | `test_T19_the_action_is_named_for_the_alias_and_the_tool` | @@ -89,26 +89,26 @@ by its quoted claim, and `tests/test_docs_audit.py` fails if a named row is not | "Unknown actions are denied; there is no default-allow." | `Policy.evaluate` — `policy.py:657` | `test_T6_unknown_action_is_denied_with_reason_unknown_action` | | "Amounts are integer minor units; floats are rejected outright" | `float` refused at any depth — `action.py:81` | `test_T7_canonical_form_is_exactly_the_specified_serialization` | | "The policy cannot see who is asking — deliberately, since v0.1" | `Policy.evaluate` still takes only the action's name and arguments; `RESERVED_ARGUMENTS` — `policy.py:657` — refuses `agent_eq` and every other principal-addressing condition at load, in a document of **every** schema version | `test_T74b_a_reserved_name_in_a_policy_rule_is_a_load_error`, `test_T74b_a_reserved_name_in_a_grant_constraint_is_a_load_error` | -| "the second axis, `authority:`" | `Authority.evaluate` — `authority.py:1288`; `Control._authority_result` — `control.py:1114` | `test_T67_a_principal_with_no_grant_is_denied` | +| "the second axis, `authority:`" | `Authority.evaluate` — `authority.py:1288`; `Control._authority_result` — `control.py:1128` | `test_T67_a_principal_with_no_grant_is_denied` | | "opt-in, and then fail-closed" | `_optional_authority` returns `None` for a document with no section — `control.py`; `Control.authority is None` is v0.2 behaviour exactly | `test_T66_a_document_with_no_authority_section_leaves_control_authority_none`, `test_T66_no_authority_event_is_appended_without_a_section`, and T66's session-wide guard in `tests/conftest.py` | | "every principal needs a grant and no grant means denied" | `NO_AUTHORITY` — the fail-closed default of `Authority.evaluate` (`authority.py:86`), reached for reads and for actions with no effect key alike | `test_T67_an_action_the_policy_allows_outright_still_needs_a_grant` | | "A grant carries no `decision:`" | `_GRANT_KEYS` — `authority.py` — is a closed set that does not contain `decision` | `test_T73b_grant_refuses_what_the_loader_refuses` | | "combine as the **stricter of the two**" | `Control.evaluate` returns the combined result — `control.py`; a denial on either axis is a denial | `test_T70_the_stricter_of_the_two_wins` | -| "narrow it at runtime with `ctrlrun delegate`" | `Control.delegate` — `control.py:4151`; `Authority.plan_delegation` — `authority.py:1549`; `ctrlrun delegate` — `cli/main.py:1096` | `test_t75_the_delegation_authorizes_an_action_within_its_limits` | +| "narrow it at runtime with `ctrlrun delegate`" | `Control.delegate` — `control.py:4165`; `Authority.plan_delegation` — `authority.py:1549`; `ctrlrun delegate` — `cli/main.py:1270` | `test_t75_the_delegation_authorizes_an_action_within_its_limits` | | "provably a subset of its parent on every dimension, at creation and again at every evaluation" | `contained_dimension` — `authority.py:986` — runs from `plan_delegation` (`authority.py:1549`) **and** from the chain walk in `Authority.evaluate` (`authority.py:1296`) | `test_t76_each_dimension_violated_alone`, `test_t77b_a_narrowed_parent_narrows_its_children` | -| "provably a subset of its parent on every dimension, at creation and again at every evaluation; a consequence budget is consumed inside the reservation's own transaction, and a rolling window bounds what may start rather than recalling what already did" | Containment as in the row above. The budget: `check_charges` — `state.py:571` — is evaluated inside `reserve_effect`'s own transaction on all three backends, and `_charges_for` — `authority.py:1404` — charges every ancestor in the chain. The window is rolling and bounds the next reserve only: `_spent` sums `[now - window, now]` and nothing reads it again after a reservation is taken | `test_T408_a_charge_and_its_reservation_are_one_transaction`, `test_T409_N_processes_racing_one_budget_spend_at_most_the_limit`, `test_T412b_every_ancestor_is_charged_through_a_real_chain`, `test_T408c_the_rolling_window_forgets` | +| "provably a subset of its parent on every dimension, at creation and again at every evaluation; a consequence budget is consumed inside the reservation's own transaction, and a rolling window bounds what may start rather than recalling what already did" | Containment as in the row above. The budget: `check_charges` — `state.py:573` — is evaluated inside `reserve_effect`'s own transaction on all three backends, and `_charges_for` — `authority.py:1404` — charges every ancestor in the chain. The window is rolling and bounds the next reserve only: `_spent` sums `[now - window, now]` and nothing reads it again after a reservation is taken | `test_T408_a_charge_and_its_reservation_are_one_transaction`, `test_T409_N_processes_racing_one_budget_spend_at_most_the_limit`, `test_T412b_every_ancestor_is_charged_through_a_real_chain`, `test_T408c_the_rolling_window_forgets` | | "omitting a dimension the parent constrains is rejected rather than inherited" | `contained_dimension` treats an absent child dimension as unconstrained and therefore wider — `authority.py:986`; the subject half is `_subject_contained` (`authority.py:1065`) | `test_t81_omission_is_not_unlimited`, `test_T73b_a_subject_addressed_to_every_principal_is_refused`, `test_t76_each_dimension_violated_alone` | -| "`ctrlrun revoke` cuts a chain of any depth with one write" | `Control.revoke` — `control.py:4488` — writes one row — `revoke_delegation` — `state.py:837` and visits no children; every evaluation walks to the root | `test_t78_a_revoked_parent_denies_its_grandchild`, `test_put_delegation_is_never_an_upsert` | -| "`mode: observe` … records what *would* have been blocked, without blocking anything" | `_parse_mode` — `policy.py:778`; `Control._observed` — `control.py:1694`; `_WouldHave` — `receipt.py:346`; `ReceiptResult.OBSERVED` — `receipt.py:258` | `test_T82_observe_executes_what_enforce_would_deny`, `test_T83_a_duplicate_is_recorded_and_still_runs` | +| "`ctrlrun revoke` cuts a chain of any depth with one write" | `Control.revoke` — `control.py:4505` — writes one row — `revoke_delegation` — `state.py:839` and visits no children; every evaluation walks to the root | `test_t78_a_revoked_parent_denies_its_grandchild`, `test_put_delegation_is_never_an_upsert` | +| "`mode: observe` … records what *would* have been blocked, without blocking anything" | `_parse_mode` — `policy.py:778`; `Control._observed` — `control.py:1708`; `_WouldHave` — `receipt.py:346`; `ReceiptResult.OBSERVED` — `receipt.py:258` | `test_T82_observe_executes_what_enforce_would_deny`, `test_T83_a_duplicate_is_recorded_and_still_runs` | | "One top-level line" | `mode:` is refused anywhere but the top level — `reject_nested_mode`, `policy.py:778` | `test_T84_mode_is_refused_anywhere_but_the_top_level` | -| "`ctrlrun stats` gives you the numbers" | `stats` — `cli/main.py:917`; counted from `would_have.blocked_reason` and nothing else | `test_T86_stats_counts_what_observe_mode_recorded`, `test_T86_stats_reaches_no_network` | -| "It is not a dry run: it executes" | `_observed` runs the executor on every path, including the ones enforce mode would have refused — `control.py:1694` | `test_T82_observe_executes_what_enforce_would_deny`, `test_T83_an_executor_that_fails_on_a_held_key_still_writes_the_record` | +| "`ctrlrun stats` gives you the numbers" | `stats` — `cli/main.py:1074`; counted from `would_have.blocked_reason` and nothing else | `test_T86_stats_counts_what_observe_mode_recorded`, `test_T86_stats_reaches_no_network` | +| "It is not a dry run: it executes" | `_observed` runs the executor on every path, including the ones enforce mode would have refused — `control.py:1708` | `test_T82_observe_executes_what_enforce_would_deny`, `test_T83_an_executor_that_fails_on_a_held_key_still_writes_the_record` | ## Prove it holds in your setup | Claim | Code | Proof | |---|---|---| -| "runs the kernel's own failure scenarios against the configuration in front of it" | `ctrlrun.verify.run` — `verify/__init__.py:156`; the eleven guarantees — `GUARANTEES` — `verify/guarantees.py:51`; the scenarios — `verify/scenarios.py` | `test_T100_the_authority_example_passes_every_non_authority_guarantee` (11/11), `test_T100_a_v1_document_with_no_templates_and_no_grants` | +| "runs the kernel's own failure scenarios against the configuration in front of it" | `ctrlrun.verify.run` — `verify/__init__.py:156`; the eleven guarantees — `GUARANTEES` — `verify/guarantees.py:55`; the scenarios — `verify/scenarios.py` | `test_T100_the_authority_example_passes_every_non_authority_guarantee` (11/11), `test_T100_a_v1_document_with_no_templates_and_no_grants` | | "in a scratch store, with fake executors, and no network" | One scratch store per guarantee under a temporary directory — `verify/scenarios.py`, `Engine.control`; `state_path()` is never called and `Control.from_file()` is never used | `test_T103_the_operators_store_is_byte_identical_before_and_after`, `test_T103_a_store_that_does_not_exist_is_not_created`, `test_T107_a_full_run_completes_with_no_network` | | "Your `.ctrlrun/state.db` is byte-identical before and after" | The scratch path is a `tempfile.mkdtemp` removed in a `finally` — `verify/__init__.py` | `test_T103_the_operators_store_is_byte_identical_before_and_after` (SHA-256 and `st_mtime_ns`), `test_T103_CTRLRUN_STATE_is_not_read_and_not_created` | | "Not applicable is not a pass" | `Report.applicable` is passes plus failures — `verify/report.py`; every N/A reason is a statement about the document — `verify/guarantees.py` | `test_T101_a_policy_with_no_approve_rule_makes_G1_and_G2_not_applicable`, `test_T102_a_policy_with_no_effect_templates_makes_G3_G4_and_G5_not_applicable` | @@ -133,47 +133,47 @@ keeps it honest: ## The capability matrix Rendered from `capabilities.yaml`; the six rows are the six groups of the verify -catalogue, `GUARANTEES` (`verify/guarantees.py:51`). +catalogue, `GUARANTEES` (`verify/guarantees.py:55`). | Claim | Code | Proof | |---|---|---| -| "An approval is bound to the exact action; a mutated or replayed one is refused." | `action_hash` — `action.py`; the approval record stores it and `_authorize_and_reserve` compares it — `state.py:667`; single use is the `granted → consumed` transition in the same `BEGIN IMMEDIATE` | `test_T2_a_mutated_action_presenting_the_approval_raises_ApprovalMismatch`, `test_T4_replaying_the_approval_raises_ApprovalMismatch_with_reason_consumed`, `test_T5_expiry_is_checked_at_consumption_not_only_at_grant` | -| "One logical effect happens at most once, across threads, processes and hosts." | `reserve_effect` — `state.py:639`, decided inside the `BEGIN IMMEDIATE` of `_authorize_and_reserve` (`state.py:1198`) against `effect_key TEXT PRIMARY KEY` (`migrations.py:109`; `COLLATE "C"` on Postgres, §4.4) | `test_T3_exactly_one_agent_reserves_and_seven_are_blocked` (8 OS processes, both backends), `test_T3_the_fake_remote_is_called_exactly_once` | -| "An unknown outcome is AMBIGUOUS, never FAILED, and blocks a blind retry." | Only `NotExecuted` maps to `FAILED` — `_outcome` — `control.py:2209`. Every other exception, timeouts included, yields `AMBIGUOUS`. A retry against an `AMBIGUOUS` key is refused — `effect.py:174`, the one place `plan_reservation` decides it for every store | `test_T1_a_blind_retry_writes_a_blocked_receipt`, `test_T1_the_ambiguous_record_survives_the_blocked_retry`, `test_T1_a_lost_response_leaves_the_effect_ambiguous` | -| "An unknown action, a missing policy or a missing principal is denied." | Unknown action: `Policy.evaluate` — `policy.py:657` — answers `deny` for a name the document does not list. Missing or malformed policy: `Policy.from_file` — `policy.py:814` — raises `PolicyError`, and there is no `Control` without a policy. Missing principal: `_refuse_no_principal` — `control.py:4821` | `test_T6_unknown_action_raises_ActionDenied_with_reason_unknown_action`, `test_missing_policy_file_is_a_policy_error`, `test_malformed_policy_document_is_a_policy_error`, `test_T62_a_declining_provider_with_no_context_is_no_principal` | +| "An approval is bound to the exact action; a mutated or replayed one is refused." | `action_hash` — `action.py`; the approval record stores it and `_authorize_and_reserve` compares it — `state.py:669`; single use is the `granted → consumed` transition in the same `BEGIN IMMEDIATE` | `test_T2_a_mutated_action_presenting_the_approval_raises_ApprovalMismatch`, `test_T4_replaying_the_approval_raises_ApprovalMismatch_with_reason_consumed`, `test_T5_expiry_is_checked_at_consumption_not_only_at_grant` | +| "One logical effect happens at most once, across threads, processes and hosts." | `reserve_effect` — `state.py:641`, decided inside the `BEGIN IMMEDIATE` of `_authorize_and_reserve` (`state.py:1255`) against `effect_key TEXT PRIMARY KEY` (`migrations.py:109`; `COLLATE "C"` on Postgres, §4.4) | `test_T3_exactly_one_agent_reserves_and_seven_are_blocked` (8 OS processes, both backends), `test_T3_the_fake_remote_is_called_exactly_once` | +| "An unknown outcome is AMBIGUOUS, never FAILED, and blocks a blind retry." | Only `NotExecuted` maps to `FAILED` — `_outcome` — `control.py:2223`. Every other exception, timeouts included, yields `AMBIGUOUS`. A retry against an `AMBIGUOUS` key is refused — `effect.py:174`, the one place `plan_reservation` decides it for every store | `test_T1_a_blind_retry_writes_a_blocked_receipt`, `test_T1_the_ambiguous_record_survives_the_blocked_retry`, `test_T1_a_lost_response_leaves_the_effect_ambiguous` | +| "An unknown action, a missing policy or a missing principal is denied." | Unknown action: `Policy.evaluate` — `policy.py:657` — answers `deny` for a name the document does not list. Missing or malformed policy: `Policy.from_file` — `policy.py:814` — raises `PolicyError`, and there is no `Control` without a policy. Missing principal: `_refuse_no_principal` — `control.py:4851` | `test_T6_unknown_action_raises_ActionDenied_with_reason_unknown_action`, `test_missing_policy_file_is_a_policy_error`, `test_malformed_policy_document_is_a_policy_error`, `test_T62_a_declining_provider_with_no_context_is_no_principal` | | "With authority on, every principal needs a grant, and delegation cannot widen one." | `NO_AUTHORITY` — the fail-closed default of `Authority.evaluate` (`authority.py:86`); `contained_dimension` — `authority.py:986` — runs from `plan_delegation` (`authority.py:1549`) and from the chain walk in `Authority.evaluate` | `test_T67_a_principal_with_no_grant_is_denied`, `test_t76_each_dimension_violated_alone` | -| "Every executed action leaves a portable JSON receipt" | `ReceiptResult` — `receipt.py:242`; `Event` — `receipt.py:308`; the store is authoritative — `append_event` — `state.py:848`; the JSONL export — `JSONLEventSink` — `receipt.py:819` | `test_T11_every_demo_receipt_carries_every_field_in_the_spec`, `test_T11_every_demo_receipt_parses_back_into_a_Receipt` | +| "Every executed action leaves a portable JSON receipt" | `ReceiptResult` — `receipt.py:242`; `Event` — `receipt.py:308`; the store is authoritative — `append_event` — `state.py:850`; the JSONL export — `JSONLEventSink` — `receipt.py:934` | `test_T11_every_demo_receipt_carries_every_field_in_the_spec`, `test_T11_every_demo_receipt_parses_back_into_a_Receipt` | ## What it guarantees | Claim | Code | Proof | |---|---|---| -| "On SQLite that is `BEGIN IMMEDIATE`" | `_authorize_and_reserve` — `state.py:1198` | `test_T3_exactly_one_agent_reserves_and_seven_are_blocked` | -| "a unique index on the effect key and compare-and-set updates whose row counts are checked" | `reserve_effect` — `postgres.py:742` — `INSERT … ON CONFLICT DO NOTHING` against `effect_key TEXT PRIMARY KEY COLLATE "C"` (`migrations.py:109`) | `test_T3_exactly_one_agent_reserves_and_seven_are_blocked` (8 OS processes, both backends) | -| "Same `StateStore` protocol, extended by nothing" | `PostgresStateStore.reserve_effect` — `postgres.py:742` — and every other method implement `v0.1 §5.3`'s frozen protocol; the decisions stay in `plan_reservation` (`effect.py:250`) | `test_T154_postgres_passes_the_store_conformance_suite` | +| "On SQLite that is `BEGIN IMMEDIATE`" | `_authorize_and_reserve` — `state.py:1255` | `test_T3_exactly_one_agent_reserves_and_seven_are_blocked` | +| "a unique index on the effect key and compare-and-set updates whose row counts are checked" | `reserve_effect` — `postgres.py:752` — `INSERT … ON CONFLICT DO NOTHING` against `effect_key TEXT PRIMARY KEY COLLATE "C"` (`migrations.py:109`) | `test_T3_exactly_one_agent_reserves_and_seven_are_blocked` (8 OS processes, both backends) | +| "Same `StateStore` protocol, extended by nothing" | `PostgresStateStore.reserve_effect` — `postgres.py:752` — and every other method implement `v0.1 §5.3`'s frozen protocol; the decisions stay in `plan_reservation` (`effect.py:250`) | `test_T154_postgres_passes_the_store_conformance_suite` | | "graded by the suite written for SQLite" | `ctrlrun.conformance.store.run` — `conformance/store/__init__.py:55` | `test_T140_every_fixture_fails_the_suite_named_for_it` | -| "It will not *knowingly* execute the same logical effect twice, and will never treat an unknown outcome as a failure." | `plan_reservation` — `effect.py:250` (refuse retry on `AMBIGUOUS`) and `_outcome` — `control.py:2209` (only `NotExecuted` → `FAILED`) | `test_T1_a_blind_retry_is_refused_and_never_reaches_the_remote`, `test_T1_a_lost_response_leaves_the_effect_ambiguous` | -| "a lost connection during `COMMIT` ... are `AMBIGUOUS`" | `_resolve_lost_insert` — `postgres.py:933`; `_resolve_lost_update` — `postgres.py:1559`; only `NotExecuted` maps to `FAILED` — `_outcome` — `control.py:2209` | `test_T155_a_connection_killed_during_commit_is_resolved_by_the_re_read`, `test_T155_no_effect_is_ever_recorded_failed_by_a_lost_commit` | -| "the store re-reads the row to find out which" | The six branches, named and logged — `A2_LANDED` — `postgres.py:142` | `test_T155b_a_landed_commit_on_a_transition_is_seen_as_landed`, `test_T155d_a_commit_the_server_never_received_retries_the_insert` | -| "A crashed worker's effect stays `AMBIGUOUS` until a human runs `ctrlrun resolve` or a `reconcile` hook asks the remote what happened" | An expired lease is `AMBIGUOUS` and nothing sweeps it — `LEASE_EXPIRED` — `effect.py:174`; who resolved it — `resolved_by` — `effect.py:210`; `resolve` — `cli/main.py:591` | `test_T159_ambiguous_survives_a_restart_and_still_refuses_a_blind_retry`, `test_T160_there_is_no_reaper`, `test_T161_a_human_resolution_records_who` | -| "the only thing besides a human permitted to move a record out of `AMBIGUOUS`" | `Control._reconciled` — `control.py:2863`; `RECONCILED_STATES` — `effect.py` | `test_T13_a_hook_answering_not_executed_moves_the_record_to_failed`, `test_T14_a_hook_answering_committed_refuses_the_retry_as_a_duplicate` | +| "It will not *knowingly* execute the same logical effect twice, and will never treat an unknown outcome as a failure." | `plan_reservation` — `effect.py:250` (refuse retry on `AMBIGUOUS`) and `_outcome` — `control.py:2223` (only `NotExecuted` → `FAILED`) | `test_T1_a_blind_retry_is_refused_and_never_reaches_the_remote`, `test_T1_a_lost_response_leaves_the_effect_ambiguous` | +| "a lost connection during `COMMIT` ... are `AMBIGUOUS`" | `_resolve_lost_insert` — `postgres.py:943`; `_resolve_lost_update` — `postgres.py:1569`; only `NotExecuted` maps to `FAILED` — `_outcome` — `control.py:2223` | `test_T155_a_connection_killed_during_commit_is_resolved_by_the_re_read`, `test_T155_no_effect_is_ever_recorded_failed_by_a_lost_commit` | +| "the store re-reads the row to find out which" | The six branches, named and logged — `A2_LANDED` — `postgres.py:152` | `test_T155b_a_landed_commit_on_a_transition_is_seen_as_landed`, `test_T155d_a_commit_the_server_never_received_retries_the_insert` | +| "A crashed worker's effect stays `AMBIGUOUS` until a human runs `ctrlrun resolve` or a `reconcile` hook asks the remote what happened" | An expired lease is `AMBIGUOUS` and nothing sweeps it — `LEASE_EXPIRED` — `effect.py:174`; who resolved it — `resolved_by` — `effect.py:210`; `resolve` — `cli/main.py:743` | `test_T159_ambiguous_survives_a_restart_and_still_refuses_a_blind_retry`, `test_T160_there_is_no_reaper`, `test_T161_a_human_resolution_records_who` | +| "the only thing besides a human permitted to move a record out of `AMBIGUOUS`" | `Control._reconciled` — `control.py:2877`; `RECONCILED_STATES` — `effect.py` | `test_T13_a_hook_answering_not_executed_moves_the_record_to_failed`, `test_T14_a_hook_answering_committed_refuses_the_retry_as_a_duplicate` | | "and only in the direction its answer points" | `"unknown"` is absent from `RECONCILED_STATES` — `effect.py` | `test_T15_a_hook_that_cannot_answer_leaves_the_record_ambiguous` | -| "Unknown action, missing policy, malformed policy, missing principal, missing or mismatched approval and inconsistent state are all `deny`." | `Policy.evaluate` — `policy.py:657`; `Policy.from_file` — `policy.py:814`; `_refuse_no_principal` — `control.py:4821`; `_authorize_and_reserve` — `state.py:1198` | `test_T6_unknown_action_raises_ActionDenied_with_reason_unknown_action`, `test_malformed_policy_document_is_a_policy_error`, `test_T62_a_declining_provider_with_no_context_is_no_principal`, `test_T2_a_mutated_action_presenting_the_approval_raises_ApprovalMismatch` | +| "Unknown action, missing policy, malformed policy, missing principal, missing or mismatched approval and inconsistent state are all `deny`." | `Policy.evaluate` — `policy.py:657`; `Policy.from_file` — `policy.py:814`; `_refuse_no_principal` — `control.py:4851`; `_authorize_and_reserve` — `state.py:1255` | `test_T6_unknown_action_raises_ActionDenied_with_reason_unknown_action`, `test_malformed_policy_document_is_a_policy_error`, `test_T62_a_declining_provider_with_no_context_is_no_principal`, `test_T2_a_mutated_action_presenting_the_approval_raises_ApprovalMismatch` | | "No flag makes a consequential action permissive by default" | There is no such option on `Control`, on `@protect`, on the CLI or in the policy schema's closed key sets — `_TOP_LEVEL_KEYS` — `policy.py:129` | `test_T84_mode_is_refused_anywhere_but_the_top_level`, `test_T101b_zero_applicable_guarantees_is_not_a_pass` | -| "With `authority:` on, every principal needs a grant, delegation cannot widen one, and `ctrlrun revoke` cuts a chain with one write." | `Authority.evaluate` — `authority.py:1288`; `contained_dimension` — `authority.py:986`; `Control.revoke` — `control.py:4488` | `test_T67_a_principal_with_no_grant_is_denied`, `test_t76_each_dimension_violated_alone`, `test_t78_a_revoked_parent_denies_its_grandchild` | +| "With `authority:` on, every principal needs a grant, delegation cannot widen one, and `ctrlrun revoke` cuts a chain with one write." | `Authority.evaluate` — `authority.py:1288`; `contained_dimension` — `authority.py:986`; `Control.revoke` — `control.py:4505` | `test_T67_a_principal_with_no_grant_is_denied`, `test_t76_each_dimension_violated_alone`, `test_t78_a_revoked_parent_denies_its_grandchild` | | "verifies a bearer token against a JWKS or a pinned key" | `JWTIdentityProvider._verified` — `jwt_identity.py:210`; the algorithm comes from the configured list and never from the token | `test_T88_a_valid_token_becomes_a_principal`, `test_T89_every_invalid_token_is_refused_by_cause` | | "maps the verified claims onto a principal" | `_principal` — `jwt_identity.py` — copies only the claims named in `claim_names` | `test_T88_only_the_named_claims_reach_the_principal` | | "`pip install \"ctrlrun[identity]\"`" | `identity = ["pyjwt[crypto]>=2.8"]` in `pyproject.toml`; imported lazily by `_jwt()` — `jwt_identity.py` | `test_T92_constructing_without_the_extra_names_the_install_command`, `test_T92_importing_ctrlrun_pulls_in_no_jwt_module` | | "CTRLRun issues no credential and defines no identity format" | There is no minting, signing or issuing code path in the package: `jwt_identity.py` calls `decode` and never `encode` | `test_the_package_never_encodes_a_token` | -| "every receipt records which policy decided it" | `Policy.policy_hash` — `policy.py:795`, over `_canonical_policy` — `policy.py:1043`; carried into the receipt by `_record` — `control.py:4712` | `test_T172_every_receipt_carries_the_hash_and_the_declared_version`, `test_T172_two_policies_sharing_a_version_string_are_told_apart_by_the_hash` | +| "every receipt records which policy decided it" | `Policy.policy_hash` — `policy.py:795`, over `_canonical_policy` — `policy.py:1043`; carried into the receipt by `_record` — `control.py:4742` | `test_T172_every_receipt_carries_the_hash_and_the_declared_version`, `test_T172_two_policies_sharing_a_version_string_are_told_apart_by_the_hash` | | "the policy's declared `version:` and a hash of its canonical content" | `version:` is recorded and never authoritative; `policy_hash` is what tells two documents apart — `policy.py:783` | `test_T171_the_declared_version_alone_does_not_change_the_hash`, `test_T171_comments_key_order_and_whitespace_do_not_change_the_hash` | -| "the approval is re-checked against the policy in force at execution" | `Control.execute` — `control.py:1304`; `_spend_unneeded_approval` — `control.py:2803` | `test_T173_the_DENY_row_refuses_and_leaves_the_approval_granted`, `test_T173_the_ALLOW_row_invalidates_the_approval_it_did_not_need` | -| "Each receipt carries the hash of the one before it" | `Receipt.chain_hash` — `receipt.py:535`; `prev_hash` — `receipt.py:454`; `GENESIS_HASH` — `receipt.py:127`; `put_receipt` takes the head row's lock first — `postgres.py:2094` | `test_T164_an_altered_receipt_is_content_altered_at_its_seq`, `test_T164_reordering_two_receipts_is_detected_either_way` | -| "`ctrlrun receipts --verify-chain` reports it by `seq`" | `verify_chain` — `receipt.py:937`; the six names — `CHAIN_BREAKS` — `receipt.py:874` | `test_the_verify_chain_flag_reports_a_break_by_seq_and_by_name`, `test_verify_chain_reads_a_postgres_store_through_store_url` | -| "migrations are automatic at open, forward-only" | `migrate` — `migrations.py:635`, called from both stores' constructors; `HEAD` — `migrations.py:416` | `test_T147_a_v05_database_migrates_and_keeps_every_row`, `test_T150_reopening_does_not_rerun` | -| "An older binary against a newer schema refuses immediately" | `_refuse` — `migrations.py:561`; `SchemaMismatch` — `errors.py` | `test_T148_an_older_binary_refuses_a_newer_database`, `test_T148_no_other_table_is_read_before_the_refusal` | +| "the approval is re-checked against the policy in force at execution" | `Control.execute` — `control.py:1318`; `_spend_unneeded_approval` — `control.py:2817` | `test_T173_the_DENY_row_refuses_and_leaves_the_approval_granted`, `test_T173_the_ALLOW_row_invalidates_the_approval_it_did_not_need` | +| "Each receipt carries the hash of the one before it" | `Receipt.chain_hash` — `receipt.py:535`; `prev_hash` — `receipt.py:454`; `GENESIS_HASH` — `receipt.py:127`; `put_receipt` takes the head row's lock first — `postgres.py:2104` | `test_T164_an_altered_receipt_is_content_altered_at_its_seq`, `test_T164_reordering_two_receipts_is_detected_either_way` | +| "`ctrlrun receipts --verify-chain` reports it by `seq`" | `verify_chain` — `receipt.py:1052`; the six names — `CHAIN_BREAKS` — `receipt.py:989` | `test_the_verify_chain_flag_reports_a_break_by_seq_and_by_name`, `test_verify_chain_reads_a_postgres_store_through_store_url` | +| "migrations are automatic at open, forward-only" | `migrate` — `migrations.py:723`, called from both stores' constructors; `HEAD` — `migrations.py:504` | `test_T147_a_v05_database_migrates_and_keeps_every_row`, `test_T150_reopening_does_not_rerun` | +| "An older binary against a newer schema refuses immediately" | `_refuse` — `migrations.py:649`; `SchemaMismatch` — `errors.py` | `test_T148_an_older_binary_refuses_a_newer_database`, `test_T148_no_other_table_is_read_before_the_refusal` | | "Releases carry PyPI provenance attestations from GitHub Actions" | `.github/workflows/publish.yml` — `pypa/gh-action-pypi-publish` pinned at v1.14.2, which generates and uploads PEP 740 attestations by default since v1.11.0 (its release notes, read 2026-09-06), with no `attestations: false`; the `pypi` job's only permission is `id-token: write` | `test_the_publish_workflow_attests_through_trusted_publishing`, `test_every_action_is_pinned_to_a_commit` | -| "`ctrlrun approve`, `deny`, `resolve`, `inspect`, `receipts` and `stats` work from the shell against any store" | `approve` — `cli/main.py:381`; `receipts` — `cli/main.py:459`; `effects` — `cli/main.py:533`; `resolve` — `cli/main.py:591`; `inspect` — `cli/main.py:635`; `stats` — `cli/main.py:917`; every one takes `--store-url` (SPEC-v0.6 §9.4) | `test_T10_resolve_failed_permits_a_retry`, `test_T18_inspect_json_emits_the_inspection_schema`, `test_T86_stats_counts_what_observe_mode_recorded`, `test_verify_chain_reads_a_postgres_store_through_store_url` | +| "`ctrlrun approve`, `deny`, `resolve`, `inspect`, `receipts` and `stats` work from the shell against any store" | `approve` — `cli/main.py:391`; `receipts` — `cli/main.py:469`; `effects` — `cli/main.py:685`; `resolve` — `cli/main.py:743`; `inspect` — `cli/main.py:787`; `stats` — `cli/main.py:1074`; every one takes `--store-url` (SPEC-v0.6 §9.4) | `test_T10_resolve_failed_permits_a_retry`, `test_T18_inspect_json_emits_the_inspection_schema`, `test_T86_stats_counts_what_observe_mode_recorded`, `test_verify_chain_reads_a_postgres_store_through_store_url` | | "`WebhookApprovalProvider` sends an approval request to a webhook, such as Slack, and takes the answer back through the same grant calls" | `WebhookApprovalProvider` — `webhook.py:143` — one signed POST on `APPROVAL_REQUESTED`; the inbound answer lands through `grant_approval` / `deny_approval` like the CLI's | `test_T27_the_outbound_post_carries_a_signature_over_the_exact_bytes_sent`, `test_T27_the_payload_carries_what_the_spec_names` | | "one OpenTelemetry span per action, one span event per step" | `OTelEventSink` — `otel.py:47` | `test_T29_one_action_produces_one_span_named_for_the_action`, `test_T29_every_event_becomes_a_span_event_named_by_its_type` | | "argument values stay out of it unless you ask for them" | `OTelEventSink(arguments=...)` — `otel.py:47` | `test_T29_argument_values_are_not_attributes_by_default` | @@ -201,7 +201,7 @@ The README also makes negative claims. They matter as much as the positive ones. | Claim | Code | Proof | |---|---|---| | "the same `StateStore` protocol, extended by nothing, graded by the suite written for SQLite rather than one written for it" | `PostgresStateStore` — `postgres.py` — satisfies `StateStore` and adds no method (SPEC-v0.6 §9.1); `ctrlrun.conformance.store.SUITES` is the SQLite suite, run against both | `test_T141_the_shipped_backends_pass`, `test_T154_postgres_passes_the_store_conformance_suite` | -| "automatic at open and forward-only, with no flag that opens a database un-migrated. An older binary against a newer schema refuses immediately." | `migrate` — `migrations.py:635` — called from both stores' constructors; `_refuse` — `migrations.py:561` — raises `SchemaMismatch` on a newer `user_version` | `test_T147_a_v05_database_migrates_and_keeps_every_row`, `test_T148_an_older_binary_refuses_a_newer_database`, `test_T152b_no_flag_opens_a_database_without_migrating` | +| "automatic at open and forward-only, with no flag that opens a database un-migrated. An older binary against a newer schema refuses immediately." | `migrate` — `migrations.py:723` — called from both stores' constructors; `_refuse` — `migrations.py:649` — raises `SchemaMismatch` on a newer `user_version` | `test_T147_a_v05_database_migrates_and_keeps_every_row`, `test_T148_an_older_binary_refuses_a_newer_database`, `test_T152b_no_flag_opens_a_database_without_migrating` | | "an edit, a deletion from the middle or a reordering is detected and named by `seq`" | `verify_chain` — `receipt.py:452` — and the six break names in `CHAIN_BREAKS` | `test_T164_an_altered_receipt_is_content_altered_at_its_seq`, `test_T164_reordering_two_receipts_is_detected_either_way`, `test_the_verify_chain_flag_reports_a_break_by_seq_and_by_name` | | "It detects **alteration**, which is not authorship: receipts are not signed." | No signing code, and a release scan keeps the vocabulary out | `test_T180_the_release_documents_do_not_blur_alteration_and_authorship` | | "every receipt records the policy that decided it, so a receipt from six months ago says what the rules were" | `Policy.policy_hash` — `policy.py:795` — over the parsed decision inputs, recorded on the receipt | `test_T172_every_receipt_carries_the_hash_and_the_declared_version`, `test_T171_any_decision_input_changes_the_hash`, `test_T171_the_declared_version_alone_does_not_change_the_hash` | @@ -228,7 +228,7 @@ restating the code; the ones that are new to the site carry their own code and p | `concepts/outcomes-and-ambiguous` | the outcome table; only a human or a reconcile hook moves a record on, and only in the direction the answer points; nothing sweeps; a lost `COMMIT` on Postgres is `AMBIGUOUS` | the matrix row "An unknown outcome is AMBIGUOUS…", the reconciliation rows, "A crashed worker's effect stays `AMBIGUOUS`…" and the Postgres rows above; `test_T160_there_is_no_reaper` | | `concepts/receipts-and-evidence` | the receipt's fields, the JSONL sink, the policy hash and version, the chain and what it does not prove | the matrix row "Every executed action leaves a portable JSON receipt", the receipt-chain and policy-versioning rows above, and `test_T11_every_demo_receipt_carries_every_field_in_the_spec` | | `concepts/authority-and-delegation` | opt-in then fail-closed, no `decision:` on a grant, stricter of the two, containment at creation and at every evaluation, omission rejected, one-write revocation, identity consumed | the authority rows under "Write down what the agent may do" and "What it guarantees" above | -| `concepts/observe-mode` | executes, records `would_have`, one top-level line, counted by `ctrlrun stats`, never asks a human | the observe-mode rows above; `_observed` — `control.py:4764` | +| `concepts/observe-mode` | executes, records `would_have`, one top-level line, counted by `ctrlrun stats`, never asks a human | the observe-mode rows above; `_observed` — `control.py:4794` | | `concepts/fail-closed` | the refusal table, one exception per row | the matrix row "An unknown action, a missing policy or a missing principal is denied.", `ActionDenied` — `errors.py:31`, `DuplicateEffect` — `errors.py:143`, `AmbiguousEffect` — `errors.py:143`, and `test_a_policy_deny_is_denied_the_same_way_as_an_unknown_action` | ## The docs site: Production @@ -239,7 +239,7 @@ sentence that rots quietly. | Page | Claim | Proved by | |---|---|---| -| `production/index` | the readiness block — version, test count, guarantee count, the two stores, the soak, the chain, the licence | rendered by `tools/docs_audit/render_readiness.py` from `pyproject.toml`, `pytest --collect-only`, the `GUARANTEES` catalogue — `verify/guarantees.py:51` — and `research/soak/results/`; `test_the_readiness_block_is_the_generators_in_every_place_it_appears` asserts the same block in the README, the docs home and this page, and `test_the_readiness_block_refuses_a_shrunken_suite_and_accepts_a_grown_one` makes the count a floor | +| `production/index` | the readiness block — version, test count, guarantee count, the two stores, the soak, the chain, the licence | rendered by `tools/docs_audit/render_readiness.py` from `pyproject.toml`, `pytest --collect-only`, the `GUARANTEES` catalogue — `verify/guarantees.py:55` — and `research/soak/results/`; `test_the_readiness_block_is_the_generators_in_every_place_it_appears` asserts the same block in the README, the docs home and this page, and `test_the_readiness_block_refuses_a_shrunken_suite_and_accepts_a_grown_one` makes the count a floor | | `production/index` | the **Not yet** list: no external security audit, no third-party review of the kernel, no sector packs | stated rather than measured, because nothing in a repository can measure an absence. A fourth line — *no soak of the length the roadmap asks for* — was **derived** from the published run until `SPEC-v0.6.md` §8.1 removed the duration from the criterion on 2026-09-07, which removed the thing being derived; the run's own duration is still printed on the soak line above the list. The list lives inside the generated block so it cannot be scrolled past. `test_the_not_yet_list_is_inside_the_block_and_not_below_it`, `test_the_not_yet_list_is_the_constant_and_derives_nothing_from_the_soak` and `test_the_readiness_block_does_not_report_the_soak_as_an_unmet_gate` assert all of it; removing a stated line is its own pull request with the row that makes the new sentence true | | `production/index` | "SQLite is the default and it is production-grade on one host… Postgres is for many hosts" | the header row above; `test_the_first_line_of_the_section_says_which_store_and_why` asserts the order, because Postgres first would tell a reader with one host something false | | `production/how-reservation-works` | the two rows: an exception before `COMMIT` is a failed write; one during it is unknown and is re-read | SPEC-v0.6 §4.3 Tables A, A1 and A2; `test_T155_a_connection_killed_during_commit_is_resolved_by_the_re_read`, `test_T155e_a_commit_the_server_never_received_re_issues_the_update`, `test_T155c_the_re_read_identity_check_is_not_an_action_id_match`, `test_T156_a_failed_re_read_refuses_to_proceed`; `test_the_two_rows_of_the_lost_commit_are_not_merged` asserts the page keeps them apart | diff --git a/docs/OWASP-AGENTIC-TOP10.md b/docs/OWASP-AGENTIC-TOP10.md index 4cd1bb2..0698274 100644 --- a/docs/OWASP-AGENTIC-TOP10.md +++ b/docs/OWASP-AGENTIC-TOP10.md @@ -74,7 +74,7 @@ mechanism, not the entry. | **G8** expired authority refused | A grant is authority only until its `expires_at`; after that the action it covered is denied, by name. | `ASI03:2026`, `ASI10:2026` | Authority is evaluated on every action against the clock, not at the start of a session, so an agent still running after its grant lapsed is denied on its next proposal. | | **G9** delegation cannot escalate | A delegated grant is valid only if it is provably a subset of its parent on every dimension — and a child that **drops** a dimension its parent constrains is rejected rather than treated as unconstrained. | `ASI03:2026`, `ASI10:2026` | Containment is checked at creation and again on every evaluation by walking the chain to its root, and omission is never inheritance — so an agent handed authority cannot mint itself more of it, and a revocation anywhere in the chain cuts everything beneath it. | | **G10** unknown exception is ambiguous | `NotExecuted` is the only outcome that means "the remote did nothing". Everything else, timeouts included, is `AMBIGUOUS`. | `ASI08:2026` | The mapping from an executor's exception to an outcome is asymmetric on purpose: a timeout is not a failure, so a framework's retry-on-error cannot be the thing that decides whether money moved twice. | -| **G11** an altered receipt is detected | Each receipt carries the hash of the one before it. Altering, deleting or reordering one breaks the chain, and the break is reported by name — `content_altered`, `hash_missing`, `link_broken`, `missing`, `head_mismatch`, `unchained` — and by `seq`. | `ASI09:2026` (partly) | The evidence an operator reads after an incident is the thing an attacker who got that far has the most reason to edit. This does not stop them: it makes **changing what a receipt says, while keeping the receipts after it**, cost a rewrite of all of them plus the head, rather than one statement. What it does not close is the end of the log — erasing a suffix, or appending to it, each cost two statements and are undetected, because the head is a row in the same database and not an external anchor. v0.6 has no anchor and claims none. It is **not** a signature and says nothing about who wrote the log; somebody who can rewrite every row including the head recomputes the chain and it verifies, and `THREAT_MODEL.md` still lists a malicious administrator as out of scope. | +| **G11** an altered receipt is detected | Each receipt carries the hash of the one before it. Altering, deleting or reordering one breaks the chain, and the break is reported by name — `content_altered`, `hash_missing`, `link_broken`, `missing`, `head_mismatch`, `unchained` — and by `seq`. | `ASI09:2026` (partly) | The evidence an operator reads after an incident is the thing an attacker who got that far has the most reason to edit. This does not stop them: it makes **changing what a receipt says, while keeping the receipts after it**, cost a rewrite of all of them plus the head, rather than one statement. What it does not close is the end of the log — erasing a suffix, or appending to it, each cost two statements and are undetected, because the head is a row in the same database and not an external anchor. **`G28` closes the erasing half since 0.11 and not the appending half**; `G11` itself is unchanged by it, and deliberately so. It is **not** a signature and says nothing about who wrote the log; somebody who can rewrite every row including the head recomputes the chain and it verifies, and `THREAT_MODEL.md` still lists a malicious administrator as out of scope. | | **G12** a byte written is ambiguous | `ctrlrun.transport` classifies a transport failure from evidence rather than from an exception type: it raises `NotExecuted` only where a connection it opened was handed no request byte, in an executor run that had offered none. After one byte, a reset, a read timeout, a reused connection or a second connection in the same run is the original exception, and the effect is `AMBIGUOUS`. | `ASI08:2026`, `ASI09:2026` (partly) | `FAILED` versus `AMBIGUOUS` is the one decision this library exists to get right, and the kernel does not make it: an executor does. Until v0.7 the correct rule lived only behind `ctrlrun[gateway]`, so the surface most people use had a docstring and no implementation, and the obvious hand-written classifier maps `ConnectionResetError` to "nothing happened" after the whole request reached the remote. That is a licence to act twice, written by a well-meaning integrator. What this closes is the absence of a correct rule in core. What it does not close: the register sees only this library's own sends, so an executor that sends part of the effect through another transport and then uses the classifier can be handed a claim that is true of these connections and false of the effect. No parameter, attribute or environment variable widens what counts as `FAILED`. | | **G13** clock divergence is named | The kernel measures how far this host's clock disagrees with the store's, on a store that has a clock of its own, and reports the divergence as `CLOCK_SKEW_DETECTED` when it passes the configured threshold. It changes no decision: every lease is still compared against the application clock exactly as before. | `ASI08:2026` (partly) | Once the store is shared, each host brings its own clock, and the failure is fail-closed and therefore quiet. A host running ahead sees a live lease as expired and marks `AMBIGUOUS` a record whose real holder is mid-flight and about to succeed; a host running behind refuses for longer than it should. Neither says why. What this closes is the silence, not the skew: an operator reading a receipt learns that two clocks disagreed and by how much, so the response to the first failure is not itself the second one. It does not synchronize anything, and a skew below the threshold is not reported. | | **G14** token changes across a renewal | `ctrlrun.idempotency_token()` answers inside an executor with a token derived from `(effect_key, attempt)`: stable within one attempt, including across a resume, and different after a renewal. Send it to a provider as its idempotency key. | `ASI08:2026` (partly) | A provider handed the effect key alone would answer the one retry the kernel permits, permitted *because the executor proved nothing happened*, with the cached failure of the attempt that failed. A token that moves with the attempt keeps a provider's cache from becoming a second source of stale outcomes. What it is for is reconciliation, a deterministic handle to ask a provider what became of an attempt whose outcome is unknown; it does not make a retry safe, and after `AMBIGUOUS` the kernel still refuses one. It is unique only as far as the operator's effect keys are, and nothing here checks two stores sharing a provider account. | @@ -91,6 +91,8 @@ mechanism, not the entry. | **G25** a hop narrows or it is refused | Authority handed to a second agent is a **subset** of what the first agent held, on every dimension, checked when the hop is created and again at every evaluation under it. An action proposed under a hop is decided against **that hop's grant alone**, with no fallback to anything else the receiving agent holds, so a hop can only ever narrow. | `ASI07:2026` (partly), `ASI03:2026` (partly), `ASI10:2026` (partly) | Before v0.10 a second agent acted under its own grants and the first agent's limits were a convention. Delegation existed but stopped at the process. What this does **not** do is compel a receiving agent to present the hop it was given: an agent that holds a root grant of its own can act under that instead, and the deployment rule that closes it is that an agent which only ever acts on handed-over work holds no root grant. `ctrlrun scan` names the principals that do. | | **G26** a hop is named on both sides | The receipt of an action taken under a hop names the hop, and so does the record of the hop's creation, so an action and the delegation that authorised it are joined from either end without inference. | `ASI07:2026` (partly), `ASI10:2026` (partly), `ASI03:2026` (partly) | Attribution across a hop was reconstruction before this: a reader had to match timestamps and principals and hope. It is evidence, not prevention, and it is on the receipt rather than in a log that can be rotated away. A hop created *inside* an action is not named on its creator's receipt; the `DELEGATION_CREATED` event carries the `action_id` and holds the join instead. | | **G27** a swapped upstream is denied | An action entry may pin the upstream it authorises, by the SHA-256 of the server's leaf certificate or by the hash of a tool's advertised schema. A server that is not the pinned one, or a tool whose schema moved under an approved action name, is refused `upstream_mismatch`; an upstream nothing observed is refused `upstream_unverified` and never admitted. | `ASI02:2026` (partly), `ASI07:2026` (partly) | An approved action name is a name, and until v0.10 nothing checked that the thing answering to it was the thing that was approved. Enforced by `ctrlrun gateway`, the surface that holds the connection: at startup, at the decision, and at the TLS handshake. In-process there is no upstream to observe, so a pinned action refuses on every call, which is fail-closed and is why `ctrlrun verify` skips a pinned action unless a scenario asks for it by name. This is one slice of a supply chain and not the category: nothing here inspects a package, a model, a build or a signature chain. | +| **G28** truncation past an anchor fails | The chain's head is a row in the same database, so erasing the end of the log and updating that row is two statements and the chain reports itself intact. An anchor records the pair the head holds (`seq` and the hash at it) through a provider **you** supply, outside the store, and anything at or below an anchored `seq` can then no longer be removed or altered without the anchored pair failing to reproduce. Reported as `anchor_broken`, `anchor_missing` or `anchor_repudiated`, in the anchor's own report. | `ASI09:2026` (partly) | **An anchor freezes a prefix, and the limits are the point.** An **append is not detected**: a forged receipt lands above every anchored `seq`, so nothing stops reproducing and the next anchor freezes it like any other. Receipts written and erased entirely between two anchors are not detected either. It is **not a signature** and says nothing about who wrote the log, and an administrator who rewrites everything before the next anchor is still out of scope. The window you are exposed to is `(last anchored seq, current head]`, and its size is your choice of interval: that is the number to tune, and the number to quote instead of any sentence about tamper-evidence. CTRLRun ships **no** anchor provider, because an RFC 3161 client is a network client; the anchor is worth exactly what the record you point it at is worth, and one in the same directory as the database is worth nothing. | +| **G31** five receipt schemas verify | One receipt chain can hold every schema version a store has been written under, and `ctrlrun verify` walks it end to end, hash by hash, **each row hashed by the rule its own version wrote**. A store kept since v0.6 holds five: `v3`, `v4`, `v5`, `v6`, `v7`. A receipt whose schema label this binary does not know is **named** and is not reported as a break. | `ASI09:2026` (partly) | The receipt chain has been checkable since v0.6, but nothing proved it stayed checkable **across an upgrade**, which is the only interesting case: a chain that verifies on the day it is written and stops verifying two releases later is evidence with a shelf life nobody stated. The proof is built from the released wheels rather than from fixtures, because a fixture is this build's opinion of what 0.6 wrote. What it does not close is everything `G11` does not close: this is about whether the record can still be **read and recomputed** years later, not about who wrote it, and an administrator who rewrites every row including the head is still out of scope. Nor does it make a future version's fields intelligible: an unknown label is named so an operator knows which row this binary could not fully interpret, which is a different sentence from *this row was tampered with*. | --- diff --git a/docs/OWASP-SOLUTIONS-LANDSCAPE.md b/docs/OWASP-SOLUTIONS-LANDSCAPE.md index e501791..c6d9faf 100644 --- a/docs/OWASP-SOLUTIONS-LANDSCAPE.md +++ b/docs/OWASP-SOLUTIONS-LANDSCAPE.md @@ -101,7 +101,7 @@ example is theirs and is not a claim that CTRLRun uses it. | Apply PII & Sensitive data masking injected into agent components | No | Out of scope. Raw resource state never reaches a receipt, because a precondition fingerprint is hashed through the canonicalizer, but that is evidence hygiene, not masking. | none | | Apply differential privacy or obfuscation on sensitive data injected into agent memory | No | Out of scope. | none | | Data masking on structured data | No | Out of scope. | none | -| Agent Action Audit | Yes | Every consequential action leaves a receipt naming the principal, the action hash, the decision, the approval and the outcome; `ctrlrun inspect ` reads one; the chain detects alteration (`G11`) and, with the external anchor, truncation and append. | v0.1, v0.6, v0.11 | +| Agent Action Audit | Yes | Every consequential action leaves a receipt naming the principal, the action hash, the decision, the approval and the outcome; `ctrlrun inspect ` reads one; the chain detects alteration (`G11`) and, with the external anchor, truncation (`G28`). An **append is not detected**: it lands above every anchored `seq`, so no anchored pair stops reproducing. | v0.1, v0.6, v0.11 | ### Test & Evaluate diff --git a/docs/ROADMAP.md b/docs/ROADMAP.md index 6f831e8..2b4ec01 100644 --- a/docs/ROADMAP.md +++ b/docs/ROADMAP.md @@ -375,12 +375,12 @@ Standards: A2A, as code. No conformance claim. One question: can the record be trusted after the fact, and kept? -- **An external anchor for the receipt chain.** The chain detects alteration and says on every page that it does not detect truncation or append — both measured at two statements, undetected, because the head is a row in the same database. v0.11 anchors the head outside the database at an interval (an RFC 3161 timestamp, or an equivalent the operator supplies) so that **anything at or below an anchored `seq` can no longer be removed or altered** without the anchored pair failing to reproduce. An anchor freezes a prefix: **an append is not detected**, because an appended row lands above every anchored `seq`, and nor is a receipt created and destroyed entirely between two anchors. This sentence said "erased or appended" until 2026-09-14, when `SPEC-v0.11.md`'s review ran the cases; §2.4 there is the table, and the named kinds are the anchor's own, not the six chain-break kinds. No keys of its own: it consumes a timestamp and issues nothing, which is why it is here and signing is not. +- **An external anchor for the receipt chain.** The chain detects alteration and says on every page that it does not detect truncation or append — both measured at two statements, undetected, because the head is a row in the same database. v0.11 anchors the head outside the database at an interval (an RFC 3161 timestamp, or an equivalent the operator supplies) so that **anything at or below an anchored `seq` can no longer be removed or altered** without the anchored pair failing to reproduce. An anchor freezes a prefix: **an append is not detected**, because an appended row lands above every anchored `seq`, and nor is a receipt created and destroyed entirely between two anchors. This sentence said "erased or appended" until 2026-09-14, when `SPEC-v0.11.md`'s review ran the cases; §2.4 there is the table, and the named kinds are the anchor's own, not the six chain-break kinds. No keys of its own: it consumes a timestamp and issues nothing, which is why it is here and signing is not. **Built by item 2 on 2026-09-14**, with `G28` grading it and `ctrlrun anchor` running it. CTRLRun ships **no** provider: an RFC 3161 client is a network client, so the operator supplies four calls (`make`, `check`, `latest`, `since`) and `examples/anchored-chain/` shows the smallest one that works. `since()` is the call a review added and the reason the design holds: with `make` and `check` alone, the record of *which* anchors exist lived in CTRLRun's own table, so deleting the newest row there left the older anchor reproducing and the truncation invisible, at a cost of one more statement. - **Retention and legal hold.** There is no retention policy today and `docs/postgres.md` says so, in the same breath as the reason one is hard to write: deleting receipts from the middle or the end of the chain is detected as a break by design. v0.11 pays that debt: a chain-preserving prune that leaves a checkpoint receipt verifiable across the gap, and a hold that refuses to prune, both recorded as receipts themselves. **v0.9 adds a second growing table and states the invariant rather than the command**: the budget ledger only grows, and `SPEC-v0.9.md` §7.3 says that rows older than the longest window on any budget of a grant cannot affect a future decision, so somebody else's archiving is safe. One caveat travels with it, because the invariant is about decisions and not about evidence: an `AMBIGUOUS` effect older than that window still **holds** a charge the operator surfaces display, so an archiver on a live ledger excludes un-released rows. `ctrlrun stats` reports the row count so the growth is visible before it matters. - **Enforcement coverage.** From events already written: policy entries never exercised, gateway tools never routed, `@protect` actions never seen. The runtime half of `ctrlrun scan`, under the same rule — a clean result is not a verdict, no score, no percentage, no badge. -- **One chain, several receipt schemas.** `ctrlrun.receipt/v7` is the schema today, and the rule since `SPEC-v0.3.md` §12.2 is that every reader upgrades before any writer switches, so an older receipt on disk still parses. v0.8 (the verified approver; the grant id under break-glass), v0.9 (budget consumption) and v0.10 (the hop) each add fields and each bump the version, so a chain kept from v0.6 across them holds **five receipt schema versions**: `v3`, which 0.6 wrote, `v4`, which v0.7 added, and `v5`, `v6` and `v7` after it. This sentence has now gone stale twice and is corrected here rather than quietly both times. It said *three shapes* and named `v3` as the schema today, before v0.7's precondition fields bumped it; v0.7's release pass fixed that and left *four* and `v4`, which v0.10's hop field made wrong again. A count of versions in a document is a number that goes stale at every release, which is the argument for reading `receipt.py`'s constants instead. And nothing yet proves that `verify` walks it end to end, hash by hash, each receipt hashed by the rule its own version wrote. v0.11 proves it, here, because this is the milestone about whether the record can be trusted after the fact. No new field: the version string already exists. What is new is the test, and the rule that a receipt whose version the binary does not know is *named* and not reported as a break — which is the same distinction v0.6 §3.2 draws for a `schema_version` row the binary does not know. Added 2026-09-10. +- **One chain, several receipt schemas.** `ctrlrun.receipt/v7` is the schema today, and the rule since `SPEC-v0.3.md` §12.2 is that every reader upgrades before any writer switches, so an older receipt on disk still parses. v0.8 (the verified approver; the grant id under break-glass), v0.9 (budget consumption) and v0.10 (the hop) each add fields and each bump the version, so a chain kept from v0.6 across them holds **five receipt schema versions**: `v3`, which 0.6 wrote, `v4`, which v0.7 added, and `v5`, `v6` and `v7` after it. This sentence has now gone stale twice and is corrected here rather than quietly both times. It said *three shapes* and named `v3` as the schema today, before v0.7's precondition fields bumped it; v0.7's release pass fixed that and left *four* and `v4`, which v0.10's hop field made wrong again. A count of versions in a document is a number that goes stale at every release, which is the argument for reading `receipt.py`'s constants instead. And nothing yet proved that `verify` walks it end to end, hash by hash, each receipt hashed by the rule its own version wrote. **v0.11's item 4 proves it and `G31` grades it, since 2026-09-14.** No new field: the version string already existed. What is new is the proof, and the rule that a receipt whose version the binary does not know is *named* and not reported as a break — the same distinction v0.6 §3.2 draws for a `schema_version` row the binary does not know. **The proof is built from the released wheels rather than from fixtures** (`scripts/five_schema_chain.py`): five environments, `pip install ctrlrun==0.6.1`, `0.7.0`, `0.8.0`, `0.9.0`, `0.10.0`, one store, then this build verifies across the whole thing, because a fixture is only this build's opinion of what 0.6 wrote. One thing that proof got wrong first is worth keeping: run with `PYTHONPATH=src`, the variable is inherited by every child, so all five "released wheels" imported the build under test and the run reported **one** schema version while looking exactly like a pass. The script now strips it and checks, per release, that the interpreter ran from that release's own environment. Added 2026-09-10, proved 2026-09-14. -- **A malformed value in a receipt row blinds every reader of the chain, and one `UPDATE` is enough.** Found while building v0.7's item 5, deferred there with a written decision, and named here because it is the evidence surface and this is the evidence milestone. A receipt whose *schema label* is unknown, and a receipt carrying an *added key*, are each reported at their `seq` and leave every other row readable. A malformed **value** of a key the schema declares is not: a float among a receipt's `controls` raises out of `Receipt.from_dict`, so `ctrlrun receipts`, `receipts --verify-chain`, `ctrlrun inspect`, `ctrlrun stats` and `G11` all stop together, and a single tampered row hides the whole document rather than naming itself. 0.6.1 behaves the same way and v0.7 neither introduced nor widened it. Fixing it needs one of two things, and both are amendments rather than patches: a new name in `CHAIN_BREAKS`, which is a closed set on a `SPEC-v0.6.md` §6.5 surface, or a reader that walks raw rows and reports per row without constructing a `Receipt` at all. `SPEC-v0.7.md` §12.5 carries the argument. Added 2026-09-12. +- **A malformed value in a receipt row blinded every reader of the chain, and one `UPDATE` was enough. Closed by v0.11's item 1 on 2026-09-14.** Found while building v0.7's item 5, deferred there with a written decision, and named here because it is the evidence surface and this is the evidence milestone. A receipt whose *schema label* is unknown, and a receipt carrying an *added key*, were each already reported at their `seq` and left every other row readable. A malformed **value** of a key the schema declares was not: a float among a receipt's `controls` raised out of `Receipt.from_dict`, so `ctrlrun receipts`, `receipts --verify-chain`, `ctrlrun inspect`, `ctrlrun stats` and the operator MCP server's `receipts` and `stats` tools all stopped together, and a single tampered row hid the whole document rather than naming itself. 0.6.1 behaved the same way and v0.7 neither introduced nor widened it. **The fix is `SPEC-v0.7.md` §12.5's second candidate**, a reader that reports per row: `StateStore.receipts()` hands back a `ctrlrun.receipt.UnreadableReceipt` for a row it cannot construct, naming its `seq` and the type of what refused it. `CHAIN_BREAKS` did **not** grow, and §12.5's first candidate is declined with a reason in `SPEC-v0.11.md` §5.1: `content_altered` already names a document that cannot be canonicalized, so a second name would be two names for one break. **Two corrections this entry earned by being implemented.** It listed `G11` among the readers that stop; `ctrlrun verify` grades `G11` against a scratch store it creates and fills itself, which no `UPDATE` reaches, so `G11` was never blinded by an operator's tampered row. And it omitted the operator MCP server, which is a *network* surface: the same one statement took out the remote console as well as the terminal. Added 2026-09-12, closed 2026-09-14. **Does not close.** Authorship. An anchor proves the log existed in this form at that time; it does not prove who wrote it, and a malicious administrator who rewrites everything before the next anchor is still out of scope. Signed receipts stay off the roadmap for the reason `SPEC-v0.6.md` §11 gives. diff --git a/docs/THREAT_MODEL.md b/docs/THREAT_MODEL.md index 49ee477..31f9c96 100644 --- a/docs/THREAT_MODEL.md +++ b/docs/THREAT_MODEL.md @@ -178,7 +178,13 @@ certified, not audited. - **Effect key templates do not escape placeholder values.** A template is literal text with values substituted in, so `refund:{tenant}:{payment_id}` resolves `tenant="acme:evil", payment_id="p1"` and `tenant="acme", payment_id="evil:p1"` to the same key. Arguments come from the agent, which this model treats as untrusted, so a crafted argument can make two distinct logical effects share one identity. The consequence is a refusal, not a double execution — the second attempt is blocked as a duplicate — so this costs availability, not correctness, and it fails in the safe direction. Until values are escaped, put the untrusted placeholder last, or use a delimiter the value cannot contain. - Single-host reservation only (SQLite). Multi-host needs Postgres (v0.6). - Approver identity is free text; no authentication of the approver (v0.3). -- Receipts are not signed, and they are not signed after v0.6 either. v0.6 adds a **hash chain** (`SPEC-v0.6.md` §6): each receipt carries the hash of the one before it, with `seq` inside the hashed content, so a partial tamper is detected and named — an `UPDATE` on one row, a `DELETE` from the middle, a reordering. What that closes is **alteration that keeps the receipts after it**: changing what receipt *n* says while leaving the rest in place costs a rewrite of all of them plus the head, rather than one statement. **Not a truncation at the end, and not an append.** Two earlier versions of this line claimed the first; a review measured both at **two statements, undetected** — delete the rows and rewind the head, or insert a well-formed row and advance it. The head is a row in the same database as the receipts, so it raises the cost of *forgetting* and not the cost of erasing; an anchor outside the database is what would close that, and v0.6 has none. What it does **not** close is authorship, and it does not close a database admin who can rewrite every row including the chain head: such an adversary recomputes the chain and it verifies. The malicious-administrator line above is unchanged; v0.6 narrows it rather than removing it. Nor does the chain prove that every action wrote a receipt — a receipt whose write failed leaves no gap in `seq` and is invisible to the chain by construction; the events log is where that is reconciled. +- Receipts are not signed, and they are not signed after v0.6 either. v0.6 adds a **hash chain** (`SPEC-v0.6.md` §6): each receipt carries the hash of the one before it, with `seq` inside the hashed content, so a partial tamper is detected and named — an `UPDATE` on one row, a `DELETE` from the middle, a reordering. What that closes is **alteration that keeps the receipts after it**: changing what receipt *n* says while leaving the rest in place costs a rewrite of all of them plus the head, rather than one statement. **Not a truncation at the end, and not an append.** Two earlier versions of this line claimed the first; a review measured both at **two statements, undetected** — delete the rows and rewind the head, or insert a well-formed row and advance it. The head is a row in the same database as the receipts, so it raises the cost of *forgetting* and not the cost of erasing; an anchor outside the database is what closes that, and **v0.11 adds one**. + +- **What the anchor changes, and exactly how far** (`SPEC-v0.11.md` §2.4, §3). An anchor records the pair the head holds, a `seq` and the hash at it, through a provider the operator supplies and CTRLRun does not ship. It **freezes a prefix**: anything at or below an anchored `seq` can no longer be removed or altered without the anchored pair failing to reproduce, and that is decided by the operator's own record rather than by a row in the database under suspicion. So the truncation measured above at two statements is now named `anchor_broken`, and so is the administrator who rewrites every row **including the head**, for everything at or below an anchored `seq`: the malicious-administrator line is narrowed again rather than removed. + + **An append is still not detected**, and this is the half most likely to be misread. A forged receipt lands at head + 1, above every anchored `seq`, so nothing stops reproducing, and a later anchor freezes the forged chain as readily as an honest one. Receipts written and erased entirely **between** two anchors are not detected either, because they were never at or below an anchored `seq`. An earlier version of the roadmap said the anchor closed "a suffix erased or appended"; a review ran both cases and it closes only the first. + + The window an operator is exposed to is `(last anchored seq, current head]`, and its size is their choice of interval. **That is the number to tune and the number to quote**, in place of any sentence about tamper-evidence. And an anchor is worth exactly what the record behind it is worth: a provider pointed at a file in the same directory as the database is worth nothing, which is the sentence this document already uses of a revocation feed. The anchor mints nothing — no key, no signature — so it still says nothing about **who** wrote the log. What it does **not** close is authorship, and it does not close a database admin who can rewrite every row including the chain head: such an adversary recomputes the chain and it verifies. The malicious-administrator line above is unchanged; v0.6 narrows it rather than removing it. Nor does the chain prove that every action wrote a receipt — a receipt whose write failed leaves no gap in `seq` and is invisible to the chain by construction; the events log is where that is reconciled. - No reconciliation; AMBIGUOUS always needs a human (v0.2 adds the executor `reconcile` hook). - The decorator can be bypassed by code that doesn't use it. diff --git a/docs/cookbook/verify-in-github-actions.mdx b/docs/cookbook/verify-in-github-actions.mdx index f22f8a2..b308b0e 100644 --- a/docs/cookbook/verify-in-github-actions.mdx +++ b/docs/cookbook/verify-in-github-actions.mdx @@ -62,8 +62,8 @@ by tag where you want a ref nobody can move. The agent sees nothing; this is the operator's check. The build sees: ```text -CTRLRun verify — ctrlrun 0.10.0, catalogue ctrlrun.guarantees/v6 -policy ./ctrlrun.yaml (ctrlrun.policy/v2, mode: enforce) +CTRLRun verify — ctrlrun 0.10.0, catalogue ctrlrun.guarantees/v7 +policy /Users/arpanghoshal/ctrlrun-project/wt/v11-i2/examples/cookbook/verify-in-github-actions/ctrlrun.yaml (ctrlrun.policy/v2, mode: enforce) authority none store sqlite, scratch (created and destroyed for this run) @@ -105,12 +105,14 @@ G24 grant refused off its task N/A no authority section G25 a hop narrows or it is refused N/A no authority section G26 a hop is named on both sides N/A no authority section G27 a swapped upstream is denied N/A no action entry pins an upstream +G28 truncation past an anchor fails PASS k8s.delete_namespace +G31 five receipt schemas verify PASS k8s.delete_namespace (a token is unique only as far as your effect keys are: two stores sharing a provider account must not produce the same effect-key string for different effects, and nothing here can check that) -15/15 declared guarantees pass. 12 not applicable: G8, G9, G13, G15, G17, G19, G22, G23, G24, G25, G26, G27. +17/17 declared guarantees pass. 12 not applicable: G8, G9, G13, G15, G17, G19, G22, G23, G24, G25, G26, G27. ``` The first line is on stderr, from G7's own scenario driving an action with no principal — the diff --git a/docs/guides/verify-in-ci.mdx b/docs/guides/verify-in-ci.mdx index fe657fe..62f04ed 100644 --- a/docs/guides/verify-in-ci.mdx +++ b/docs/guides/verify-in-ci.mdx @@ -32,8 +32,8 @@ guarantees pass. ``` ```text - CTRLRun verify — ctrlrun 0.10.0, catalogue ctrlrun.guarantees/v6 - policy ./ctrlrun.yaml (ctrlrun.policy/v2, mode: enforce) + CTRLRun verify — ctrlrun 0.10.0, catalogue ctrlrun.guarantees/v7 + policy /Users/arpanghoshal/ctrlrun-project/wt/v11-i2/examples/cookbook/verify-in-github-actions/ctrlrun.yaml (ctrlrun.policy/v2, mode: enforce) authority none store sqlite, scratch (created and destroyed for this run) @@ -75,12 +75,14 @@ guarantees pass. G25 a hop narrows or it is refused N/A no authority section G26 a hop is named on both sides N/A no authority section G27 a swapped upstream is denied N/A no action entry pins an upstream + G28 truncation past an anchor fails PASS k8s.delete_namespace + G31 five receipt schemas verify PASS k8s.delete_namespace (a token is unique only as far as your effect keys are: two stores sharing a provider account must not produce the same effect-key string for different effects, and nothing here can check that) - 15/15 declared guarantees pass. 12 not applicable: G8, G9, G13, G15, G17, G19, G22, G23, G24, G25, G26, G27. + 17/17 declared guarantees pass. 12 not applicable: G8, G9, G13, G15, G17, G19, G22, G23, G24, G25, G26, G27. ``` The first line is on **stderr**, from G7's own scenario: an action with no principal is diff --git a/docs/production/anchoring.mdx b/docs/production/anchoring.mdx new file mode 100644 index 0000000..d8d5372 --- /dev/null +++ b/docs/production/anchoring.mdx @@ -0,0 +1,88 @@ +--- +title: "Anchoring the receipt chain" +description: "Erasing the end of the receipt log costs two SQL statements. Anchor the head outside your database, and see what that proves and what it does not." +--- + +The receipt chain detects **alteration**: edit a receipt and its hash no longer matches. It does +not detect **truncation**, and that has been written down since 0.6 rather than discovered. The +reason is structural, and worth seeing: + +```sql +DELETE FROM receipts WHERE seq > 3; +UPDATE receipt_chain SET seq = 3, hash = ''; +``` + +Two statements. `ctrlrun receipts --verify-chain` then reports the store intact, because the head +that would have caught it is a row in the same database as the receipts. + +An **anchor** records the pair the head holds, a `seq` and the hash at it, somewhere your +database's writer does not control. + +## Running it + +```bash +ctrlrun anchor --provider yourpkg.anchors:provider # on a schedule +ctrlrun anchor --provider yourpkg.anchors:provider --verify # in the job after a restore +``` + +CTRLRun **ships no provider**. An RFC 3161 client is a network client, and this library's core is +the standard library plus `pyyaml` and `click`. You supply one with four calls: `make`, `check`, +`latest` and `since`. + +A real one is a timestamp authority, a transparency log, or an append-only bucket in another +account with different credentials. **An anchor is worth exactly what the record behind it is +worth**, and one pointed at a file beside your database is worth nothing. +`examples/anchored-chain/` has a minimal provider and runs both halves. + +## What it proves + +Anything **at or below** an anchored `seq` can no longer be removed or altered without the +anchored pair failing to reproduce. That includes the administrator who rewrites every row +*including* the head, which the chain alone cannot catch, because the operator's own record is +what decides rather than a row in the database under suspicion. + +The verification asks your provider what it holds **before** reading CTRLRun's local table, so +deleting rows from that table does not remove the question: a store whose anchor cache was emptied +reports `anchor_missing`, which is a break. + +## What this does not do + +- **It does not detect an append.** A forged receipt lands above every anchored `seq`, so nothing + stops reproducing, and the next anchor freezes it as readily as an honest one. +- **It does not detect receipts written and erased between two anchors.** They were never at or + below an anchored `seq`. +- **It does not tell you who wrote a receipt.** An anchor mints nothing and is not a signature, + and an administrator who rewrites everything before the next anchor is out of scope. + +**The window you are exposed to is `(last anchored seq, current head]`.** Its size is your choice +of interval. That is the number to tune, and the number to quote. + +## The three names + +| Name | What it means | +|---|---| +| `anchor_broken` | an anchored pair does not reproduce: the `seq` is absent, or it hashes differently. This is the one that means tamper | +| `anchor_missing` | you anchor and the store holds no anchor at all, or your provider names one the local table lacks | +| `anchor_repudiated` | your provider answered **no**: it does not recognise a pair the local table claims it anchored | + +They are reported in the **anchor's own** report and never in the chain's. `G11` grades the hash +chain every deployment has; `G28` grades the anchor, is opt-in, and is `N/A` with a reason where +you do not configure one. + +**An unreachable provider is neither.** It reports `unavailable`, and `G28` is `N/A` with that +reason. Refusing to act when you cannot ask is fail-closed; reporting tampering when you cannot ask +is a false positive, and a briefly unreachable timestamp authority must not look like a truncation. + +**Verified by** `T530` for the truncation this exists to catch, with the chain reporting the same +store intact as its negative control; `T530b` for an administrator who rewrites every row +including the head; `T531`, which runs a forged **append** and requires both reports to stay clean, +so the limit above is a tested property rather than a sentence; `T532` for the attacker who +deletes the local anchor row as well; `T533` for `G11` being untouched; and `T534` for a provider +that is unreachable, which is `unavailable` and never a break. + +## Next + +- [Receipt integrity](/docs/production/receipt-integrity): the chain, and the six names it reports. +- [The threat model](/docs/THREAT_MODEL): what stays open. +- [Operations](/docs/production/operations): where these checks belong in a schedule. +- [Get started](/docs/get-started/quickstart) · [Why](/docs/why). diff --git a/docs/production/index.mdx b/docs/production/index.mdx index 20df822..4a3f255 100644 --- a/docs/production/index.mdx +++ b/docs/production/index.mdx @@ -28,8 +28,8 @@ need. `test_the_first_line_of_the_section_says_which_store_and_why` asserts the {/* generated from the suite, pyproject and the soak (full) — run the generator */} - **Version 0.10.0**, on [PyPI](https://pypi.org/project/ctrlrun/), Python 3.11 and later, tested on 3.11 to 3.14. -- **6,080 tests**, every version specified before it was written and every requirement mutation-tested. [Read more](/docs/how-this-is-built). -- **27 guarantees you can check in your own setup**, with `ctrlrun verify` against your policy, on your store's backend, in a scratch store it creates. [Read more](/docs/security/verify-guarantees). +- **6,158 tests**, every version specified before it was written and every requirement mutation-tested. [Read more](/docs/how-this-is-built). +- **29 guarantees you can check in your own setup**, with `ctrlrun verify` against your policy, on your store's backend, in a scratch store it creates. [Read more](/docs/security/verify-guarantees). - **One host: a file.** SQLite, no server, no ops. **Many hosts: Postgres**, the same guarantees, graded by the same suite. [Read more](/docs/production/postgres). - **Soaked for 20m 0s on postgres**: 889,735 actions, 0 unattributed ambiguous outcomes, positive control fired. Nothing here establishes what only accumulates over days. [Read more](/docs/production/soak). - **Each receipt carries the hash of the one before it**, so an alteration is detected and named. [Read more](/docs/production/receipt-integrity). diff --git a/docs/production/receipt-integrity.mdx b/docs/production/receipt-integrity.mdx index e3a1384..7399aff 100644 --- a/docs/production/receipt-integrity.mdx +++ b/docs/production/receipt-integrity.mdx @@ -19,7 +19,7 @@ database directly. | Name | What it means | What to do | |---|---|---| -| `content_altered` | receipt *n* no longer hashes to its stored hash | the row was edited; the stored hash says what it was, the row says what it is now | +| `content_altered` | receipt *n* no longer hashes to its stored hash, **or cannot be read back as a receipt at all** | the row was edited; the stored hash says what it was, the row says what it is now | | `hash_missing` | receipt *n* has a position and no stored hash | the hash column was cleared; nothing can be compared, and this is a failure rather than a skip | | `link_broken` | receipt *n*'s link does not match what *n-1* hashes to | a neighbour changed, or two rows were reordered | | `missing` | a gap in `seq` | a receipt was deleted from the middle | @@ -31,6 +31,18 @@ column destroyed every independent copy of every hash, and an earlier reader **s comparison it could no longer make and called the chain intact. A check that cannot be evaluated is not a check that passed. +## One bad row costs one row + +One declared key set to the wrong type is enough to damage a row past reading, and until 0.11 +**one `UPDATE` blinded every reader at once**. Now the row is named where it sits: + +``` +seq 2 ctr_b9db23088ad0… UNREADABLE this row could not be read (InvalidArgument) +``` + +`--verify-chain` reports `content_altered` there, and `stats` counts the rest as +`unreadable receipts` rather than dropping them from a total. + `unchained` is never a pass either. Rows from before the chain have no position, so they are read separately and reported with their count, and the summary says how many of how many were verified. Folding them into a green count would be the same false green in a new costume. @@ -48,6 +60,12 @@ record is already committed, so a retry of that key is refused as a duplicate ra twice; the head is unadvanced and there is no gap in `seq`. You get an action that ran, an exception you must handle, and a log with no hole in it. +## Truncation + +Erasing the **end** of the log is two statements and this command reports the store intact +afterwards, because the head that would catch it is a row in the same database. +[Anchoring](/docs/production/anchoring) puts that head out of the writer's reach. + ## What this does not do - **It does not tell you who wrote a receipt.** Alteration is not authorship, it does not survive @@ -61,7 +79,7 @@ exception you must handle, and a log with no hole in it. created, because a verification tool with a side effect on the thing it verifies is refused. `--verify-chain` is the one that reads yours. -**Verified by** `T164` — six tamper cases, each asserted on the **name** and the `seq` and not +**Verified by** `T510` and `T511` for the readers that no longer blind, `T164` — six tamper cases, each asserted on the **name** and the `seq` and not merely on "invalid", the reordering case included — `T165` for the positive control without which every one of those rows would pass against a detector that always says broken, `T167` for truncation caught by the head and only by the head, `T168` for the unchained rows that are never diff --git a/docs/reference/api/Control.mdx b/docs/reference/api/Control.mdx index fc33906..19443c9 100644 --- a/docs/reference/api/Control.mdx +++ b/docs/reference/api/Control.mdx @@ -5,7 +5,7 @@ description: "Policy, state and evidence composed around a single action (SPEC-v {/* generated by tools/docs_audit/render_api.py from the docstrings — edit the docstring, never this page */} -`ctrlrun.Control` — class, defined at `src/ctrlrun/control.py:822` +`ctrlrun.Control` — class, defined at `src/ctrlrun/control.py:824` ```python from ctrlrun import Control @@ -14,7 +14,7 @@ from ctrlrun import Control ```python class Control - def __init__(policy: Policy, store: StateStore, approvals: ApprovalProvider | None = None, *, clock: Callable[[], datetime] = _utc_now, approval_ttl: timedelta = DEFAULT_APPROVAL_TTL, lease: timedelta = DEFAULT_LEASE, sinks: Sequence[EventSink] = (), suspend_timeout: timedelta = DEFAULT_SUSPEND_TIMEOUT, identity: IdentityProvider | None = None, authority: Authority | None = None, environment: str | None = None, approver_identity: ApproverIdentity | None = None, require_approved_policy: bool = False, upstream: str | None = None) + def __init__(policy: Policy, store: StateStore, approvals: ApprovalProvider | None = None, *, clock: Callable[[], datetime] = _utc_now, approval_ttl: timedelta = DEFAULT_APPROVAL_TTL, lease: timedelta = DEFAULT_LEASE, sinks: Sequence[EventSink] = (), suspend_timeout: timedelta = DEFAULT_SUSPEND_TIMEOUT, identity: IdentityProvider | None = None, authority: Authority | None = None, environment: str | None = None, approver_identity: ApproverIdentity | None = None, require_approved_policy: bool = False, upstream: str | None = None, anchor: AnchorProvider | None = None) ``` Policy, state and evidence composed around a single action (SPEC-v0.1 §8). diff --git a/docs/reference/api/DelegationRecord.mdx b/docs/reference/api/DelegationRecord.mdx index 682a49b..6437814 100644 --- a/docs/reference/api/DelegationRecord.mdx +++ b/docs/reference/api/DelegationRecord.mdx @@ -5,7 +5,7 @@ description: "One row of the `delegations` table (SPEC-v0.3 §5.2)." {/* generated by tools/docs_audit/render_api.py from the docstrings — edit the docstring, never this page */} -`ctrlrun.DelegationRecord` — class, defined at `src/ctrlrun/state.py:386` +`ctrlrun.DelegationRecord` — class, defined at `src/ctrlrun/state.py:388` ```python from ctrlrun import DelegationRecord diff --git a/docs/reference/api/EventSink.mdx b/docs/reference/api/EventSink.mdx index 517562f..df54bd9 100644 --- a/docs/reference/api/EventSink.mdx +++ b/docs/reference/api/EventSink.mdx @@ -5,7 +5,7 @@ description: "Somewhere a copy of every `Event` and `Receipt` goes (SPEC-v0.2 § {/* generated by tools/docs_audit/render_api.py from the docstrings — edit the docstring, never this page */} -`ctrlrun.EventSink` — class, defined at `src/ctrlrun/receipt.py:800` +`ctrlrun.EventSink` — class, defined at `src/ctrlrun/receipt.py:915` ```python from ctrlrun import EventSink diff --git a/docs/reference/api/InMemoryStateStore.mdx b/docs/reference/api/InMemoryStateStore.mdx index b5b2d63..f66b5b9 100644 --- a/docs/reference/api/InMemoryStateStore.mdx +++ b/docs/reference/api/InMemoryStateStore.mdx @@ -5,7 +5,7 @@ description: "Everything held in process memory: for tests and `ctrlrun demo`." {/* generated by tools/docs_audit/render_api.py from the docstrings — edit the docstring, never this page */} -`ctrlrun.InMemoryStateStore` — class, defined at `src/ctrlrun/state.py:970` +`ctrlrun.InMemoryStateStore` — class, defined at `src/ctrlrun/state.py:1009` ```python from ctrlrun import InMemoryStateStore diff --git a/docs/reference/api/JSONLEventSink.mdx b/docs/reference/api/JSONLEventSink.mdx index 4ea2266..0e3e899 100644 --- a/docs/reference/api/JSONLEventSink.mdx +++ b/docs/reference/api/JSONLEventSink.mdx @@ -5,7 +5,7 @@ description: "The JSONL half of the evidence: two append-only files in one direc {/* generated by tools/docs_audit/render_api.py from the docstrings — edit the docstring, never this page */} -`ctrlrun.JSONLEventSink` — class, defined at `src/ctrlrun/receipt.py:819` +`ctrlrun.JSONLEventSink` — class, defined at `src/ctrlrun/receipt.py:934` ```python from ctrlrun import JSONLEventSink diff --git a/docs/reference/api/SQLiteStateStore.mdx b/docs/reference/api/SQLiteStateStore.mdx index ce0af8b..c8bff6c 100644 --- a/docs/reference/api/SQLiteStateStore.mdx +++ b/docs/reference/api/SQLiteStateStore.mdx @@ -5,7 +5,7 @@ description: "Approvals, effects and evidence in one SQLite file (ARCHITECTURE {/* generated by tools/docs_audit/render_api.py from the docstrings — edit the docstring, never this page */} -`ctrlrun.SQLiteStateStore` — class, defined at `src/ctrlrun/state.py:1478` +`ctrlrun.SQLiteStateStore` — class, defined at `src/ctrlrun/state.py:1535` ```python from ctrlrun import SQLiteStateStore diff --git a/docs/reference/api/StateStore.mdx b/docs/reference/api/StateStore.mdx index f18316e..13838ac 100644 --- a/docs/reference/api/StateStore.mdx +++ b/docs/reference/api/StateStore.mdx @@ -5,7 +5,7 @@ description: "Durable state behind a `Control` (SPEC-v0.1 §5.3): approvals, eff {/* generated by tools/docs_audit/render_api.py from the docstrings — edit the docstring, never this page */} -`ctrlrun.StateStore` — class, defined at `src/ctrlrun/state.py:636` +`ctrlrun.StateStore` — class, defined at `src/ctrlrun/state.py:638` ```python from ctrlrun import StateStore @@ -39,8 +39,11 @@ class StateStore(ApprovalStore, Protocol) def append_event(event: Event) -> Event def put_receipt(receipt: Receipt) -> Receipt def chain_head() -> tuple[int, str] | None + def put_anchor(anchor: Anchor) -> None + def anchors() -> tuple[Anchor, ...] + def checkpoint() -> tuple[int, str] | None def events() -> tuple[Event, ...] - def receipts() -> tuple[Receipt, ...] + def receipts() -> tuple[Receipt | UnreadableReceipt, ...] def close() -> None ``` diff --git a/docs/reference/api/context.mdx b/docs/reference/api/context.mdx index 4cb2fa3..80e97f3 100644 --- a/docs/reference/api/context.mdx +++ b/docs/reference/api/context.mdx @@ -5,7 +5,7 @@ description: "Bind the principal for calls made inside the block." {/* generated by tools/docs_audit/render_api.py from the docstrings — edit the docstring, never this page */} -`ctrlrun.context` — function, defined at `src/ctrlrun/control.py:420` +`ctrlrun.context` — function, defined at `src/ctrlrun/control.py:422` ```python from ctrlrun import context diff --git a/docs/reference/api/idempotency_token.mdx b/docs/reference/api/idempotency_token.mdx index 413f193..2118891 100644 --- a/docs/reference/api/idempotency_token.mdx +++ b/docs/reference/api/idempotency_token.mdx @@ -5,7 +5,7 @@ description: "The provider idempotency token for the attempt this executor is ru {/* generated by tools/docs_audit/render_api.py from the docstrings — edit the docstring, never this page */} -`ctrlrun.idempotency_token` — function, defined at `src/ctrlrun/control.py:388` +`ctrlrun.idempotency_token` — function, defined at `src/ctrlrun/control.py:390` ```python from ctrlrun import idempotency_token diff --git a/docs/reference/api/postgres-PostgresStateStore.mdx b/docs/reference/api/postgres-PostgresStateStore.mdx index 9e91e77..d8eb8a3 100644 --- a/docs/reference/api/postgres-PostgresStateStore.mdx +++ b/docs/reference/api/postgres-PostgresStateStore.mdx @@ -5,7 +5,7 @@ description: "Approvals, effects and evidence in a Postgres schema (SPEC-v0.6 § {/* generated by tools/docs_audit/render_api.py from the docstrings — edit the docstring, never this page */} -`ctrlrun.postgres.PostgresStateStore` — class, defined at `src/ctrlrun/postgres.py:335` +`ctrlrun.postgres.PostgresStateStore` — class, defined at `src/ctrlrun/postgres.py:345` ```python from ctrlrun.postgres import PostgresStateStore diff --git a/docs/reference/api/protect.mdx b/docs/reference/api/protect.mdx index 261ea15..5122897 100644 --- a/docs/reference/api/protect.mdx +++ b/docs/reference/api/protect.mdx @@ -5,7 +5,7 @@ description: "Bind a function to an action name: every call becomes a decided, r {/* generated by tools/docs_audit/render_api.py from the docstrings — edit the docstring, never this page */} -`ctrlrun.protect` — function, defined at `src/ctrlrun/control.py:4939` +`ctrlrun.protect` — function, defined at `src/ctrlrun/control.py:4969` ```python from ctrlrun import protect diff --git a/docs/reference/api/state-Charge.mdx b/docs/reference/api/state-Charge.mdx index 7f828af..a6f39ae 100644 --- a/docs/reference/api/state-Charge.mdx +++ b/docs/reference/api/state-Charge.mdx @@ -5,7 +5,7 @@ description: "What one reservation spends against one grant's budget (SPEC-v0.9 {/* generated by tools/docs_audit/render_api.py from the docstrings — edit the docstring, never this page */} -`ctrlrun.state.Charge` — class, defined at `src/ctrlrun/state.py:513` +`ctrlrun.state.Charge` — class, defined at `src/ctrlrun/state.py:515` ```python from ctrlrun.state import Charge diff --git a/docs/reference/api/state-Consumption.mdx b/docs/reference/api/state-Consumption.mdx index 5c2d3aa..77fa872 100644 --- a/docs/reference/api/state-Consumption.mdx +++ b/docs/reference/api/state-Consumption.mdx @@ -5,7 +5,7 @@ description: "One ledger row, as `consumptions()` hands it back (SPEC-v0.9 §3.3 {/* generated by tools/docs_audit/render_api.py from the docstrings — edit the docstring, never this page */} -`ctrlrun.state.Consumption` — class, defined at `src/ctrlrun/state.py:553` +`ctrlrun.state.Consumption` — class, defined at `src/ctrlrun/state.py:555` ```python from ctrlrun.state import Consumption diff --git a/docs/reference/api/state-check_charges.mdx b/docs/reference/api/state-check_charges.mdx index 352b812..67c3329 100644 --- a/docs/reference/api/state-check_charges.mdx +++ b/docs/reference/api/state-check_charges.mdx @@ -5,7 +5,7 @@ description: "SPEC-v0.9 §3.3.1's predicate, in one place so three backends cann {/* generated by tools/docs_audit/render_api.py from the docstrings — edit the docstring, never this page */} -`ctrlrun.state.check_charges` — function, defined at `src/ctrlrun/state.py:571` +`ctrlrun.state.check_charges` — function, defined at `src/ctrlrun/state.py:573` ```python from ctrlrun.state import check_charges diff --git a/docs/reference/api/with_approval.mdx b/docs/reference/api/with_approval.mdx index 366818b..155305e 100644 --- a/docs/reference/api/with_approval.mdx +++ b/docs/reference/api/with_approval.mdx @@ -5,7 +5,7 @@ description: "Present a granted approval to the calls made inside the block (SPE {/* generated by tools/docs_audit/render_api.py from the docstrings — edit the docstring, never this page */} -`ctrlrun.with_approval` — function, defined at `src/ctrlrun/control.py:443` +`ctrlrun.with_approval` — function, defined at `src/ctrlrun/control.py:445` ```python from ctrlrun import with_approval diff --git a/docs/reference/cli.mdx b/docs/reference/cli.mdx index 8dfb504..7e65e23 100644 --- a/docs/reference/cli.mdx +++ b/docs/reference/cli.mdx @@ -22,6 +22,7 @@ Options: --help Show this message and exit. Commands: + anchor Anchor this store's chain head outside the store, or check... approve Grant a pending approval request. delegate Create a delegated grant beneath an existing one. demo Run the five scenarios, in process, with no network. @@ -110,6 +111,41 @@ Options: --help Show this message and exit. ``` +## ctrlrun anchor + +```text +Usage: ctrlrun anchor [OPTIONS] + + Anchor this store's chain head outside the store, or check the anchors already + made. + + The chain detects alteration. It does not detect truncation or append, because + the head that would catch them is a row in the same database (SPEC-v0.6 §6.4). + An anchor records the pair the head holds somewhere your database's writer + does not control. + + **What it proves:** anything at or below an anchored seq can no longer be + removed or altered without the anchored pair failing to reproduce. **What it + does not:** an append is not detected, because it lands above every anchored + seq; nor are receipts created and destroyed between two anchors; nor who wrote + any of it. The window you are exposed to is (last anchored seq, current head], + and its size is your choice of interval. + +Options: + --provider MODULE:ATTR Your anchor provider (SPEC-v0.11 §3.2). CTRLRun + ships none. [required] + --verify Check every anchor the provider holds against + this chain, and make none. + --kind [interval|checkpoint] interval (the scheduled anchor) or checkpoint (a + prune's). + --json Print the report as JSON. + --store-url TEXT The store to open. Default: $CTRLRUN_STORE_URL, + else the SQLite database beside the policy + (.ctrlrun/state.db, or wherever $CTRLRUN_STATE + points). + --help Show this message and exit. +``` + ## ctrlrun effects ```text diff --git a/docs/verify.md b/docs/verify.md index 47de744..170ad64 100644 --- a/docs/verify.md +++ b/docs/verify.md @@ -14,7 +14,7 @@ what could not be tested at all. ```console $ ctrlrun verify -CTRLRun verify — ctrlrun 0.10.0, catalogue ctrlrun.guarantees/v6 +CTRLRun verify — ctrlrun 0.10.0, catalogue ctrlrun.guarantees/v7 policy examples/authority/payments.yaml (ctrlrun.policy/v7, mode: enforce) authority same document, 3 grants store sqlite, scratch (created and destroyed for this run) @@ -61,12 +61,14 @@ G24 grant refused off its task PASS head-of-support G25 a hop narrows or it is refused PASS head-of-support G26 a hop is named on both sides PASS head-of-support G27 a swapped upstream is denied N/A no action entry pins an upstream +G28 truncation past an anchor fails PASS stripe.refund +G31 five receipt schemas verify PASS stripe.refund (a token is unique only as far as your effect keys are: two stores sharing a provider account must not produce the same effect-key string for different effects, and nothing here can check that) -24/24 declared guarantees pass. 3 not applicable: G13, G15, G27. +26/26 declared guarantees pass. 3 not applicable: G13, G15, G27. ``` It reads the policy document — `$CTRLRUN_CONFIG`, else `./ctrlrun.yaml` — and the authority @@ -139,13 +141,13 @@ G14 token changes across a renewal N/A no action declares an `effect:` temp G15 renewal past the ceiling refused N/A no action verify can drive to allow or approve declares both `effect:` and `max_attempts` -8/8 declared guarantees pass. 8 not applicable: G3, G4, G5, G8, G9, G13, G14, G15. +9/9 declared guarantees pass. 8 not applicable: G3, G4, G5, G8, G9, G13, G14, G15. ``` -That run is `8/8`, never `16/16`. There is no flag that folds an N/A into the count, and there +That run is `9/9`, never `17/17`. There is no flag that folds an N/A into the count, and there will not be one: a number that counts guarantees nobody exercised is a number that means nothing. Point the same document at Postgres and G13 becomes a graded `PASS`, so the run reads -`9/9` with seven not applicable: the denominator moves with what the setup can actually +`10/10` with seven not applicable: the denominator moves with what the setup can actually exercise, which is the whole idea. An N/A is always a statement about your **document**, derived from it. A scenario verify could diff --git a/generated/readiness.full.mdx b/generated/readiness.full.mdx index f797016..e4f03f3 100644 --- a/generated/readiness.full.mdx +++ b/generated/readiness.full.mdx @@ -1,7 +1,7 @@ {/* generated from the suite, pyproject and the soak (full) — run the generator */} - **Version 0.10.0**, on [PyPI](https://pypi.org/project/ctrlrun/), Python 3.11 and later, tested on 3.11 to 3.14. -- **6,080 tests**, every version specified before it was written and every requirement mutation-tested. [Read more](/docs/how-this-is-built). -- **27 guarantees you can check in your own setup**, with `ctrlrun verify` against your policy, on your store's backend, in a scratch store it creates. [Read more](/docs/security/verify-guarantees). +- **6,158 tests**, every version specified before it was written and every requirement mutation-tested. [Read more](/docs/how-this-is-built). +- **29 guarantees you can check in your own setup**, with `ctrlrun verify` against your policy, on your store's backend, in a scratch store it creates. [Read more](/docs/security/verify-guarantees). - **One host: a file.** SQLite, no server, no ops. **Many hosts: Postgres**, the same guarantees, graded by the same suite. [Read more](/docs/production/postgres). - **Soaked for 20m 0s on postgres**: 889,735 actions, 0 unattributed ambiguous outcomes, positive control fired. Nothing here establishes what only accumulates over days. [Read more](/docs/production/soak). - **Each receipt carries the hash of the one before it**, so an alteration is detected and named. [Read more](/docs/production/receipt-integrity). diff --git a/generated/readiness.json b/generated/readiness.json index 24cdd23..9ac78fa 100644 --- a/generated/readiness.json +++ b/generated/readiness.json @@ -1,5 +1,5 @@ { - "guarantees": 27, + "guarantees": 29, "python": { "floor": "3.11", "tested": [ @@ -18,6 +18,6 @@ "positive_control": true, "unexplained": 0 }, - "tests": 6080, + "tests": 6158, "version": "0.10.0" } diff --git a/generated/readiness.mdx b/generated/readiness.mdx index 5d095a0..ce46e0a 100644 --- a/generated/readiness.mdx +++ b/generated/readiness.mdx @@ -1,7 +1,7 @@ {/* generated from the suite, pyproject and the soak (mdx) — run the generator */} - **Version 0.10.0**, on [PyPI](https://pypi.org/project/ctrlrun/), Python 3.11 and later, tested on 3.11 to 3.14. -- **6,080 tests**, every version specified before it was written and every requirement mutation-tested. -- **27 guarantees you can check in your own setup**, with `ctrlrun verify` against your policy, on your store's backend, in a scratch store it creates. +- **6,158 tests**, every version specified before it was written and every requirement mutation-tested. +- **29 guarantees you can check in your own setup**, with `ctrlrun verify` against your policy, on your store's backend, in a scratch store it creates. - **One host: a file.** SQLite, no server, no ops. **Many hosts: Postgres**, the same guarantees, graded by the same suite. - **Soaked for 20m 0s on postgres**: 889,735 actions, 0 unattributed ambiguous outcomes, positive control fired. Nothing here establishes what only accumulates over days. [What it does not establish](https://ctrlrun.dev/docs/production/soak). - **Each receipt carries the hash of the one before it**, so an alteration is detected and named. diff --git a/generated/readiness.readme.md b/generated/readiness.readme.md index 2539a25..dc404c0 100644 --- a/generated/readiness.readme.md +++ b/generated/readiness.readme.md @@ -1,7 +1,7 @@ - **Version 0.10.0**, on [PyPI](https://pypi.org/project/ctrlrun/), Python 3.11 and later, tested on 3.11 to 3.14. -- **6,080 tests**, every version specified before it was written and every requirement mutation-tested. -- **27 guarantees you can check in your own setup**, with `ctrlrun verify` against your policy, on your store's backend, in a scratch store it creates. +- **6,158 tests**, every version specified before it was written and every requirement mutation-tested. +- **29 guarantees you can check in your own setup**, with `ctrlrun verify` against your policy, on your store's backend, in a scratch store it creates. - **One host: a file.** SQLite, no server, no ops. **Many hosts: Postgres**, the same guarantees, graded by the same suite. - **Soaked for 20m 0s on postgres**: 889,735 actions, 0 unattributed ambiguous outcomes, positive control fired. Nothing here establishes what only accumulates over days. [What it does not establish](https://ctrlrun.dev/docs/production/soak). - **Each receipt carries the hash of the one before it**, so an alteration is detected and named. diff --git a/tests/test_docs_production.py b/tests/test_docs_production.py index 7ec37b9..69d88d3 100644 --- a/tests/test_docs_production.py +++ b/tests/test_docs_production.py @@ -132,6 +132,10 @@ def test_the_section_exists_and_has_a_page_for_each_thing_that_breaks(): "migrations", "recovery", "receipt-integrity", + # SPEC-v0.11 §3. Its own page rather than a section of receipt-integrity: that page is a + # "what to run" page on a 900-word budget, and the anchor's limits need as much room as + # its claim, which is what the budget exists to force a decision about. + "anchoring", "soak", "operations", } @@ -198,6 +202,11 @@ def test_every_production_page_cites_acceptance_tests_that_exist(page: Path): # unwatched. `T180` allow-lists the same sentence, one line at a time, in # `tests/test_release_v0_6.py`. "## what this does not do - **it does not tell you who wrote a receipt.** alteration is not authorship, it does not survive an administrator who can rewrite every row including the head, and it vouches for nothing that was never recorded.", # noqa: E501 + # `anchoring.mdx`, SPEC-v0.11 §2.4. The anchor is the first thing in this project a + # reader could mistake for tamper-proofing, so its page says what it is not in the same + # breath, and this is that sentence. Added deliberately, which is what the allow-list is + # for. + "- **it does not tell you who wrote a receipt.** an anchor mints nothing and is not a signature, and an administrator who rewrites everything before the next anchor is out of scope.", # noqa: E501 } ) diff --git a/tests/test_release_documents.py b/tests/test_release_documents.py index f8d09ef..61ddb0a 100644 --- a/tests/test_release_documents.py +++ b/tests/test_release_documents.py @@ -102,6 +102,9 @@ def _load(name: str) -> set[str]: "- **`docs/docs/THREAT_MODEL.md`'s \"Receipts are not signed; a database admin can alter history", # noqa: E501 "which half it does not: **truncation at the end**, authorship, an adversary who can rewrite", # noqa: E501 '- Receipts are not signed. A database administrator can alter history. **This line read "(v0.6)" until v0.6 was built, and that was a promise v0.6 does not keep**: v0.6 adds a hash chain, which detects alteration and is not evidence of authorship, and it does not stop an administrator who can rewrite every row including the chain head. Signing is out of scope (`SPEC-v0.6.md` §11).', # noqa: E501 + # SPEC-v0.11 §1.1 rule 1 — the anchor consumes a timestamp and issues nothing, + # which is a sentence about what CTRLRun does NOT do and has to say the word. + "no revocation, no signing. Signing stays off the roadmap for the reason `SPEC-v0.6.md` §11", # noqa: E501 ), "docs/postgres.md": ( "Receipts are not signed, alteration is not authorship, and the chain is not tamper-proof", @@ -110,7 +113,9 @@ def _load(name: str) -> set[str]: "- **It does not tell you who wrote a receipt.** Alteration is not authorship, it does not survive", # noqa: E501 ), "docs/THREAT_MODEL.md": ( - "- Receipts are not signed, and they are not signed after v0.6 either. v0.6 adds a **hash chain** (`SPEC-v0.6.md` §6): each receipt carries the hash of the one before it, with `seq` inside the hashed content, so a partial tamper is detected and named — an `UPDATE` on one row, a `DELETE` from the middle, a reordering. What that closes is **alteration that keeps the receipts after it**: changing what receipt *n* says while leaving the rest in place costs a rewrite of all of them plus the head, rather than one statement. **Not a truncation at the end, and not an append.** Two earlier versions of this line claimed the first; a review measured both at **two statements, undetected** — delete the rows and rewind the head, or insert a well-formed row and advance it. The head is a row in the same database as the receipts, so it raises the cost of *forgetting* and not the cost of erasing; an anchor outside the database is what would close that, and v0.6 has none. What it does **not** close is authorship, and it does not close a database admin who can rewrite every row including the chain head: such an adversary recomputes the chain and it verifies. The malicious-administrator line above is unchanged; v0.6 narrows it rather than removing it. Nor does the chain prove that every action wrote a receipt — a receipt whose write failed leaves no gap in `seq` and is invisible to the chain by construction; the events log is where that is reconciled.", # noqa: E501 + "- Receipts are not signed, and they are not signed after v0.6 either. v0.6 adds a **hash chain** (`SPEC-v0.6.md` §6): each receipt carries the hash of the one before it, with `seq` inside the hashed content, so a partial tamper is detected and named — an `UPDATE` on one row, a `DELETE` from the middle, a reordering. What that closes is **alteration that keeps the receipts after it**: changing what receipt *n* says while leaving the rest in place costs a rewrite of all of them plus the head, rather than one statement. **Not a truncation at the end, and not an append.** Two earlier versions of this line claimed the first; a review measured both at **two statements, undetected** — delete the rows and rewind the head, or insert a well-formed row and advance it. The head is a row in the same database as the receipts, so it raises the cost of *forgetting* and not the cost of erasing; an anchor outside the database is what closes that, and **v0.11 adds one**.", # noqa: E501 + # SPEC-v0.11 §2.4 — the anchor's bounded claim, on the page that states it. + "The window an operator is exposed to is `(last anchored seq, current head]`, and its size is their choice of interval. **That is the number to tune and the number to quote**, in place of any sentence about tamper-evidence. And an anchor is worth exactly what the record behind it is worth: a provider pointed at a file in the same directory as the database is worth nothing, which is the sentence this document already uses of a revocation feed. The anchor mints nothing — no key, no signature — so it still says nothing about **who** wrote the log. What it does **not** close is authorship, and it does not close a database admin who can rewrite every row including the chain head: such an adversary recomputes the chain and it verifies. The malicious-administrator line above is unchanged; v0.6 narrows it rather than removing it. Nor does the chain prove that every action wrote a receipt — a receipt whose write failed leaves no gap in `seq` and is invisible to the chain by construction; the events log is where that is reconciled.", # noqa: E501 ), } diff --git a/tests/test_verify_page.py b/tests/test_verify_page.py index ee22303..3d2393d 100644 --- a/tests/test_verify_page.py +++ b/tests/test_verify_page.py @@ -87,7 +87,9 @@ def test_the_verify_page_says_what_not_applicable_means(): page = " ".join(_repository_file(VERIFY_DOC).split()) assert "Not applicable is not a pass" in page - assert "never `16/16`" in page + # 17/17 since v0.11 item 4: `G31` is applicable wherever a chain can be built, so the + # illustrated run is 9 passing over 8 not applicable rather than 8 over 8. + assert "never `17/17`" in page assert "no flag that folds an N/A into the count" in page assert "declared guarantees pass" in page