diff --git a/.gitignore b/.gitignore index 7f920fecd..578acd5c8 100644 --- a/.gitignore +++ b/.gitignore @@ -243,3 +243,10 @@ examples/hivemind/src/hivemind-telemetry.jsonl /hivemind-governed-transcript.jsonl /hivemind-audit.jsonl /hivemind-governed-audit.jsonl + +# Vigilant (docs/book) compiles these from gateway.go / shield.rs at run time +# (modules/synthesizer.naab builds into Vigilant/bin/). The two copies that were +# committed were Android/aarch64 builds nothing read -- ~9.4 MB of executables. +docs/book/verification/ch0_full_projects/Vigilant/bin/ +docs/book/verification/ch0_full_projects/Vigilant/proxy/gateway_vessel +docs/book/verification/ch0_full_projects/Vigilant/scanner/shield_vessel diff --git a/docs/book/verification/ch0_full_projects/Vigilant/proxy/gateway_vessel b/docs/book/verification/ch0_full_projects/Vigilant/proxy/gateway_vessel deleted file mode 100755 index 4e2d0a571..000000000 Binary files a/docs/book/verification/ch0_full_projects/Vigilant/proxy/gateway_vessel and /dev/null differ diff --git a/docs/book/verification/ch0_full_projects/Vigilant/scanner/shield_vessel b/docs/book/verification/ch0_full_projects/Vigilant/scanner/shield_vessel deleted file mode 100755 index 0727175b1..000000000 Binary files a/docs/book/verification/ch0_full_projects/Vigilant/scanner/shield_vessel and /dev/null differ diff --git a/docs/open-investigations.md b/docs/open-investigations.md index d59ea9cc6..467e342a2 100644 --- a/docs/open-investigations.md +++ b/docs/open-investigations.md @@ -902,7 +902,7 @@ campaign doc's live-status column. | C1b | **Confirmed again by the 2x2 (2026-08-18)** — arms at `0.03` were outcome-identical to arms at `0.0`. This row predated that experiment and would have pre-empted it; it was not consulted. **Healing is inert at every shipped rate.** 25 configs set `coherence_natural_healing`, including both templates (0.02) and living-script v1/v2 (0.03). Varying only the rate on v3: 0.0 and 0.03 give **identical** results to the quarantine; 0.25 restores `elevated → normal`. `heal_factor = 1/(1+signals)` halves an already-small number against penalties of 0.08–0.15. | **addressed; R3 withdrawn** | Bound (R1/R2), channel-agnostic booking (S1/S2) and ratchet (R4) shipped. **R3 (raise the rate) is withdrawn on evidence**: the rate at which healing suppresses escalation is damage-relative, not absolute — HIGH is lost at 0.50 in v3 and at 0.15 in the same scenario with two signals disabled. No global constant is correct. **Re-examined 2026-08-25 and UPHELD, with the cliff number corrected**: on two fresh profiles neither 0.13 nor 0.15 suppressed anything, and suppression first appeared between 0.20 and 0.30. The withdrawal now rests on the light profile suppressing ELEVATED at 0.30 — not on HIGH, which that fixture cannot test because it never reaches HIGH with healing off (a vacuous control, stated rather than glossed). See `docs/proposal-bounded-coherence-healing.md`. | | C1a | **`elevated → normal` is unreachable once coherence floors** — **premise updated 2026-08-26 (#174)**: a behaviour-only recovery path now EXISTS but ships off. R5 relative healing (`context_drift.coherence_healing_damage_fraction`, default 0.0) returns a floored agent to 1.0000/normal at `f=0.5`, verified against a no-healing control on the same fixture. So the row is no longer "cannot", it is "can, if enabled, and only while no signal is firing" — see the S17 row below for why that last clause bites. Whether it should default on is the open decision. Original text: With `coherence_natural_healing` at its default 0.0, coherence never climbs back, and `coherence_prox` alone holds composite above `elevated_threshold` — 18 quiet turns in the v3 run left it pinned at 0.5125 vs a 0.35 threshold. | **precondition narrowed (2026-09-06)** | **Measured 2026-09-06** on shipped defaults (`adaptive_baseline_enabled: true`): coherence floors at 0.275, NOT 0.0 — the precondition this row assumes (coherence at 0.0) is not reached by ordinary varied work (25 turns, `gemini-flash-lite-latest`). Level stayed NORMAL throughout; escalation did not occur. **This does NOT prove de-escalation is reachable** — ARM A never escalated, so it produced zero de-escalation events. If any workload reaches ELEVATED on shipped defaults, this row's arithmetic (composite floor 0.45 > `elevated_threshold` 0.40) is unchanged. The resolution is an **accidental side effect** of #176 (the S17 `persona_baseline_adaptive` flip, 2026-08-29) — nobody designed it, and no test asserts coherence stays above 0.0 on varied work. A future S17 change could silently re-expose this row. Remains valid for `adaptive_baseline_enabled: false`. See `investigate/deescalation/FINDINGS.md`. **See C1c: live evidence now generalises this from `elevated → normal` to de-escalation from ANY level.** | | C1e | **S17's frozen warm-up baseline makes an agent pay forever for correct work — and it silently invalidated an acceptance gate.** `persona_baseline_mean` is set ONCE at baseline completion and never updated, and the test is `abs()` on a SINGLE response, so an agent whose later responses differ in keyword count from its warm-up turns fires every turn indefinitely. Measured on the `test_coherence_recovery.sh` fixture (#174): turns 18-40, `persona_fingerprint=0.0500`, every turn, alone, while the agent produced correct on-mandate work. **Exact numbers (2026-08-29, from decision snapshots):** baseline established at turn 6 as mean **7.167**, stddev **1.863**; only **2 distinct values across 40 turns** (`-1.0` then frozen), confirming the set-once guard `persona_baseline_mean < 0.0`. With `persona_deviation_factor` 2.0 the firing band is `|kw - 7.167| > 3.73`, i.e. anything below 3.4 or **above 10.9** keywords. The warm-up turns are terse ("Ledger reconcile part N") averaging 7.2 keywords; the recovery turns are richer ("quarterly totals recomputed and the balance recorded against source ledger section N") and clear 10.9 — so **the agent is punished permanently for producing MORE, scored identically to a collapse, because the test is `abs()` on a SINGLE response.** Note a hypothesis that did NOT survive measurement: the `stddev < 1.0` floor was suspected as the hair-trigger, but the observed stddev was 1.863 and never reached the floor. This evidence supports the DIRECTIONAL fix (b) over naive re-baselining (a) — re-deriving the baseline would also have to decide whether growth counts as drift, which is the actual question. **The consequence was not just noise.** It meant the acceptance gate `test_recovery_default.sh` (RD-03/RD-04, merged in #172) could only ever be passed by a mechanism that OUT-HEALS a live signal — which is suppression: absolute healing at 0.30 pins coherence at 1.0000 for 14 consecutive turns while the signal fires on every one. The gate was red for a reason nobody had checked, and its recorded justification was wrong. | **addressed (DEFAULT ON since 2026-08-29)** | Fixed in #176, default flipped after measurement. **The flip is evidence-backed, not preference:** the frozen baseline inflicts a CERTAIN, UNIVERSAL false penalty (22 of 22 post-drift turns on correct work). The cost is a BOUNDED low-rate hole, measured across 8 arms with no other signal firing in any (so S17 was measured alone) — drift in extra keywords/turn, S17 firings off->on: 0.25: 24->2 (~8% retained, the boiling-frog case); 0.5: 28->12; 1.0: 29->12; 2.0: 30->13. Detection degrades only below ~0.5 kw/turn, never reaches zero, and is flat above. The fixture was ENGINEERED as the worst case (on-mandate vocabulary so nothing else could fire and freeze the baseline), so 24->2 is an UPPER BOUND. Full suite 441/441 0 unexpected; leak 874/0. **LIVE A/B RESULTS (2026-08-29), including one that falsified a claim in this row.** (a) `living-script_extended` vs Gemini: the DEVELOPER agent, at IDENTICAL turn counts (45 v 45) with its non-S17 signals flat (circular 3->3, validation_outcome 4->4, entity_consistency 4->3), fired S17 **16 -> 3**. That is the clean evidence. The run's AGGREGATE (27->10) does NOT survive its own control: arm B was truncated by timeout (108 v 121 analyzed turns) and `semantic_stability` fell FURTHER than S17 relative to that (0.25 v 0.41 of expected), so the aggregate drop is not attributable to the fix. Temperature was not pinned (prompt omission). (b) `living-script_v3` vs a local 0.5B model, greedy + fixed seed: P2 control **perfect** — every signal byte-identical across arms, 39/39 turns. S17 fired **0 times in both arms**, so it says NOTHING about the firing rate. But it revealed what the row and the code comments got WRONG: **this is not a pure leniency change.** On uniform responses the re-derived stddev NARROWS (observed 2.67 -> 1.00, the floor), tightening the 2-sigma band from +/-5.34 to +/-2.00 keywords — a 12-keyword deviation fires with the fix ON and not OFF. The mechanism RECALIBRATES; direction is workload-dependent. **Follow-up check, which shrinks that concern:** the stddev floor of 1.0 is applied in BOTH paths (frozen establishment AND re-derivation — 2 sites), so the minimum +/-2 keyword band PREDATES the fix and is reachable without it whenever warm-up turns are uniform. The fix changes where the band is CENTRED (current behaviour vs warm-up), not how narrow it can get. The v3 comparison was a wide frozen band (varied warm-up) against a tracked narrow one — that is the warm-up being unrepresentative, i.e. the defect. The real open question, PRE-EXISTING and not caused by this change, is whether +/-2 keywords (deviation_factor 2.0 x stddev floor 1.0) is the right minimum tolerance at all. **Anything needing a persistently-firing S17 must now pin `persona_baseline_adaptive: false`** — `test_relative_healing.sh` does, because RH-10 had been relying on this defect to supply its live signal. `context_drift.thresholds.persona_baseline_adaptive` (default **true** since the 2026-08-29 flip this row's own status column records — the "default false" and "ships off" wording elsewhere in this row was written before the flip and was left standing, making the row read as self-contradictory; `governance.h`, `persona_baseline_adaptive = true`; ratcheted false->true as a loosening) re-derives the baseline from the rolling window at the end of any turn where NO OTHER signal fired. Adaptation is gated on independent evidence of health: a degrading agent trips other signals, the baseline freezes, S17 keeps paying. Excluding S17 from its own gate is load-bearing -- while S17 is the only thing firing no turn is ever clean, so the signal would hold its own baseline hostage and the fix would silently do nothing. Gates in `tests/governance_v4/test_persona_adaptive.sh`, 4/4: PF-02 (control) OFF fires 22/22 post-drift turns; PF-03 ON fires 2/22; **PF-04 (positive control) ON fires 10x during genuine drift, identical to OFF's 10x -- detection reduced by exactly zero.** Both candidate fixes named in this row were REJECTED on measurement: directional-only blinds S17 to rambling, naive re-baselining recreates the S9 problem. **Residual risk, open:** an agent whose ONLY symptom is slow persona drift, with no other signal ever firing, would be followed by the baseline -- the boiling-frog case. That was why it shipped off initially; it was flipped ON after the eight-arm measurement above, and the residual risk is accepted rather than avoided. | Known and previously recorded as deliberately-not-changed (`docs/governance-campaign-findings.md`), but that decision predates this evidence. Reopening on new evidence, not on preference. Two candidate fixes, neither traced yet: re-derive the baseline on a rolling window (S9 does this and is deliberately frozen for the opposite reason — so the asymmetry needs its own argument), or make the deviation test directional so producing MORE is not penalised like collapse. Anything here must not weaken S17's ability to catch a genuine persona shift; a positive control for that is mandatory. **DEPENDENCY — read before changing anything here (added 2026-09-06).** The obvious remedy for this row is to touch `adaptive_baseline_enabled`, because that flag is what gates `baseline_complete` and therefore S17 entirely. It was flipped to true FOR THIS ROW (#176) — and that flip is now the only thing keeping the C1a/C1c/C1d premise unreachable. #203 measured the connection: with the flag off, ordinary varied work floors coherence to 0.0 in three turns and the engine sticks at HIGH; with it on, the same work stays clear of the floor. **Nobody decided that.** The de-escalation consequence was a side effect of an S17 fix, not a design choice, and until `tests/governance_v4/test_coherence_floor_precondition.sh` there was nothing asserting it. So: a change here that reverts, narrows or conditions that flag reverts the C1 behaviour too. C1R-01/02/03 will fail and tell you — that is what they are for — but the coupling belongs in your design, not in your surprise. | -| C1f | **S5's frozen baseline is set from THREE turns, and two lines of startup code condemn a run.** Sibling of C1e, found by tracing the same "frozen baseline" shape across all four signals that hold one. `initial_entropy` is set ONCE, from the first half of the sliding window the first time it reaches `vocab_contraction_window` (6) turns — i.e. from turns 0-2 — and never updated. It is the earliest-set of the four (fixed BEFORE `baseline_complete`) and the **only one of the four not exposed in `snapshotState()`**, so unlike `persona_baseline_mean` / `thinking_baseline_mean` / `input_tokens_baseline_mean` it cannot be reconstructed from preserved evidence. Note also that `turn_types` is built from **RuntimeEventType**, not response vocabulary — the signal's name is misleading; nothing an agent *says* can move it. A pure `agent.send()` loop is `{AS, AR}` forever (entropy 1.0, and `1.0 < 0.6` is false) so S5 can never fire; add k distinct non-agent event types **in turn 0 alone** and the frozen baseline rises past the trip point: k=1 → `initial_entropy` 1.449, threshold 0.869, silent; **k=2 → 1.811, threshold 1.087, fires — and then fires every remaining turn.** Measured on shipped defaults, 40 turns, identical prompts and identical stub responses in every arm: control S5 0/40 and coherence **1.0000** at turn 39; k=3 arm S5 **35/40** and coherence **0.0000 from turn 12**, escalating `normal → elevated` at turn 12 while the control never leaves NORMAL. **Adaptive baselining does not absorb it** — S5 never fires *during* the baseline window, so its expected rate is 0 and every later firing pays full weight 0.15. With `output_admissibility` on (threshold 0.70, streak 5) the same pair diverges to: control completes 40/40; k=3 arm **terminated at turn 12 by `QUARANTINE_STREAK_EXCEEDED`** after 5 × `OUTPUT_INADMISSIBLE`. Two lines of ordinary orchestration — an env read and a base64 call before the first send — is the whole difference. | **partially addressed; ships off** | Pinned as-is by `tests/governance_v4/test_vocab_baseline.sh` (VC-00..VC-04), which asserts CURRENT behaviour so a fix reads as a deliberate gate change; VC-00 is the vacuity guard and VC-01 is the lower half of the trip point, without which VC-02 locates no boundary. **Not claimed:** no shipped example is known to be affected — `living-script_extended` does one `file.read` before its first `agent.create`, which is k=1, *below* the trip point. **Fixed 2026-09-04, default OFF.** Candidate 1 (expose `initial_entropy` in `snapshotState()`) had ALREADY shipped — this row was stale on that point, and VC-07 pins it. Candidate 2 shipped as `context_drift.thresholds.entropy_baseline_adaptive` (default false, ratcheted false->true as a loosening), mirroring `persona_baseline_adaptive`: re-derive `initial_entropy` from the current early half on any turn clean APART FROM S5 — the same self-exclusion S17 needs, because while S5 is the only signal firing no turn is ever clean and a baseline gated on total cleanliness could never move. **Measured, 25 turns, identical response stream:** startup artifact 20 firings/coherence 0.0000 -> 5 firings/0.2500; genuine narrowing (varied for 10 turns then stops, `since_last_check` feed) detected on IDENTICAL turns 20-25 with identical coherence 0.8610, while the baseline demonstrably moved (2.000 frozen vs 1.918 adaptive) — so detection cost is zero and the arm is not vacuous. Pinned by VC-08/VC-09/VC-10, where VC-10 is the vacuity guard without which 'detection unchanged' cannot be distinguished from 'the mechanism never ran'. **The residual 5 firings are structural, not a half-fix:** while turn 0 is still inside the sliding window's early half that entropy is genuinely present. **Why it ships OFF despite a measured zero detection cost** — the standard #176 used to flip S17's equivalent ON: this is ONE narrowing fixture against #176's eight arms at different drift rates, AND under the default feed the re-derived baseline converges to 1.000 where S5 can never fire again, so defaulting it on would silently retire S5 for every existing config. Defensible (under the default feed S5 can only ever fire on the startup artifact) but too large a claim for one fixture. Candidate 3 (directional comparison) remains untraced and is now unnecessary. B6 did NOT have to be settled first, as this row predicted — its fix only had to be AVAILABLE, to make the positive control expressible. **Do not re-derive the "S5 is blind under the default feed" result as a new finding.** An external defaults audit (2026-09-05) reported it as a critical coupling between `event_feed: "turn_bucket"` and `entropy_baseline_adaptive`, on the stated mechanism that the gate is `initial_entropy > 1.0` so an entropy of 1.000 fails it. That constant does not exist: the gate is `state.initial_entropy > config_->thresholds.entropy_min_initial`, default **0.5** (`governance.h`), and 1.000 clears it. The blindness is real but is the SECOND clause and has nothing to do with adaptive or with the feed — `per_turn_types` is a set per turn, so a tool-less agent is `{AS, AR}` in both halves of the window, `recent < initial * entropy_contraction_ratio` is `1.0 < 0.6`, and it is false with the flag ON, OFF, and under `since_last_check` alike (that selector changes WHICH events arrive, not their type variety — a tool-less agent raises no other types to deliver). Pinned as VC-00, the control gate of `test_vocab_baseline.sh`, since before the flag existed. +| C1f | **S5's frozen baseline is set from THREE turns, and two lines of startup code condemn a run.** Sibling of C1e, found by tracing the same "frozen baseline" shape across all four signals that hold one. `initial_entropy` is set ONCE, from the first half of the sliding window the first time it reaches `vocab_contraction_window` (6) turns — i.e. from turns 0-2 — and never updated. It is the earliest-set of the four (fixed BEFORE `baseline_complete`) and the **only one of the four not exposed in `snapshotState()`**, so unlike `persona_baseline_mean` / `thinking_baseline_mean` / `input_tokens_baseline_mean` it cannot be reconstructed from preserved evidence. Note also that `turn_types` is built from **RuntimeEventType**, not response vocabulary — the signal's name is misleading; nothing an agent *says* can move it. A pure `agent.send()` loop is `{AS, AR}` forever (entropy 1.0, and `1.0 < 0.6` is false) so S5 can never fire; add k distinct non-agent event types **in turn 0 alone** and the frozen baseline rises past the trip point: k=1 → `initial_entropy` 1.449, threshold 0.869, silent; **k=2 → 1.811, threshold 1.087, fires — and then fires every remaining turn.** Measured on shipped defaults, 40 turns, identical prompts and identical stub responses in every arm: control S5 0/40 and coherence **1.0000** at turn 39; k=3 arm S5 **35/40** and coherence **0.0000 from turn 12**, escalating `normal → elevated` at turn 12 while the control never leaves NORMAL. **Adaptive baselining does not absorb it** — S5 never fires *during* the baseline window, so its expected rate is 0 and every later firing pays full weight 0.15. With `output_admissibility` on (threshold 0.70, streak 5) the same pair diverges to: control completes 40/40; k=3 arm **terminated at turn 12 by `QUARANTINE_STREAK_EXCEEDED`** after 5 × `OUTPUT_INADMISSIBLE`. Two lines of ordinary orchestration — an env read and a base64 call before the first send — is the whole difference. **First real-world case (2026-09-29, repo-sentinel dogfood round 5, NAAb `80c8cd9`, live Gemini):** the sentinel reviewer does git and file work before its first send, so turns 0-2 carry non-agent event types. In all four runs of the `adaptive_baseline_window: 5` arm — two clean fixtures, two adversarial — the reviewer held 1.0000 through turn 5, then paid S5 0.15 on every turn from turn 6 (`vocab_contraction_window` = 6, which is why the first firing did not move when the operator changed the baseline window from 3 to 5), reaching **0.556667 and OUTPUT_INADMISSIBLE at turn 8 in all four, identical to six decimals** whatever the model wrote. The same trajectory in the clean control is what the round-4 write-up reported as a false positive (F-008). Disabling S5 for the reviewer (`context_drift_signals: {vocabulary_contraction: false}`, arm B) removed every quarantine in both clean runs. Every analyzed row in all eight runs reconciled from telemetry (`coherence_adjustments`, #263), so the trajectory is the engine's arithmetic, not a reporting artefact. | **partially addressed; ships off** | Pinned as-is by `tests/governance_v4/test_vocab_baseline.sh` (VC-00..VC-04), which asserts CURRENT behaviour so a fix reads as a deliberate gate change; VC-00 is the vacuity guard and VC-01 is the lower half of the trip point, without which VC-02 locates no boundary. **Not claimed** (as written 2026-09-04, SUPERSEDED — see the repo-sentinel case in the first column): no shipped example is known to be affected — `living-script_extended` does one `file.read` before its first `agent.create`, which is k=1, *below* the trip point. **Fixed 2026-09-04, default OFF.** Candidate 1 (expose `initial_entropy` in `snapshotState()`) had ALREADY shipped — this row was stale on that point, and VC-07 pins it. Candidate 2 shipped as `context_drift.thresholds.entropy_baseline_adaptive` (default false, ratcheted false->true as a loosening), mirroring `persona_baseline_adaptive`: re-derive `initial_entropy` from the current early half on any turn clean APART FROM S5 — the same self-exclusion S17 needs, because while S5 is the only signal firing no turn is ever clean and a baseline gated on total cleanliness could never move. **Measured, 25 turns, identical response stream:** startup artifact 20 firings/coherence 0.0000 -> 5 firings/0.2500; genuine narrowing (varied for 10 turns then stops, `since_last_check` feed) detected on IDENTICAL turns 20-25 with identical coherence 0.8610, while the baseline demonstrably moved (2.000 frozen vs 1.918 adaptive) — so detection cost is zero and the arm is not vacuous. Pinned by VC-08/VC-09/VC-10, where VC-10 is the vacuity guard without which 'detection unchanged' cannot be distinguished from 'the mechanism never ran'. **The residual 5 firings are structural, not a half-fix:** while turn 0 is still inside the sliding window's early half that entropy is genuinely present. **Why it ships OFF despite a measured zero detection cost** — the standard #176 used to flip S17's equivalent ON: this is ONE narrowing fixture against #176's eight arms at different drift rates, AND under the default feed the re-derived baseline converges to 1.000 where S5 can never fire again, so defaulting it on would silently retire S5 for every existing config. Defensible (under the default feed S5 can only ever fire on the startup artifact) but too large a claim for one fixture. Candidate 3 (directional comparison) remains untraced and is now unnecessary. B6 did NOT have to be settled first, as this row predicted — its fix only had to be AVAILABLE, to make the positive control expressible. **Do not re-derive the "S5 is blind under the default feed" result as a new finding.** An external defaults audit (2026-09-05) reported it as a critical coupling between `event_feed: "turn_bucket"` and `entropy_baseline_adaptive`, on the stated mechanism that the gate is `initial_entropy > 1.0` so an entropy of 1.000 fails it. That constant does not exist: the gate is `state.initial_entropy > config_->thresholds.entropy_min_initial`, default **0.5** (`governance.h`), and 1.000 clears it. The blindness is real but is the SECOND clause and has nothing to do with adaptive or with the feed — `per_turn_types` is a set per turn, so a tool-less agent is `{AS, AR}` in both halves of the window, `recent < initial * entropy_contraction_ratio` is `1.0 < 0.6`, and it is false with the flag ON, OFF, and under `since_last_check` alike (that selector changes WHICH events arrive, not their type variety — a tool-less agent raises no other types to deliver). Pinned as VC-00, the control gate of `test_vocab_baseline.sh`, since before the flag existed. | C1d | **A compliant agent doing ordinary work floors its coherence in six turns — and it is NOT about code.** v3's DESIGN (1-3) and IMPLEMENT (4-6) are ordinary well-specified on-mandate work. Code arm (keyed runs 4/5/6): coherence 0.74/0.64/0.56 at turn 3, floored turn 7-8. **Prose arm (keyed, same config via `prose_overlay.json`, only the task differs): 0.46 at turn 3, floored turn 5** — outside the code arm's three-run spread and two turns earlier. Signals, code vs prose: `semantic_stability` 29/36 vs **35/36**, `plan_drift` 14/36 vs **29/36**, `entity_consistency` 34/36 vs 35/36. | **precondition narrowed (2026-09-06)** | Ordinary varied work of ANY modality trips these thresholds and prose trips them harder, so the suspects are the **0.25 defaults** on `semantic_stability` and `entity_consistency` — NOT the code-aware extractor, which appears to be helping code. **Do not move a threshold on this**: it identifies the suspect, not the value. Two caveats stand — n=1 for prose against n=3 for code, and response length is unmatched (prose input baseline 1717 vs 1220), so modality is not isolated from verbosity. Settle by matching response length across arms and running prose twice. **Update (run 24 + the 2x2):** the suspects now have a MECHANISM, which moves this off "threshold value unknown". Both signals score inter-turn *variation*, so they are anti-correlated with repetition drift: measured 67/67/**0**/37% (S15) and 33/33/**12**/32% (S10) across DESIGN/IMPLEMENT/DRIFT_PRESSURE/RECOVERY — firing LESS during induced drift than during correct work — and across 8 consecutive byte-identical responses in the 2x2 neither fired once. In recovery they fire on a period of exactly 6, matching the scenario's task rotation. That is a defect in KIND, not in value, and no threshold move fixes it: raising the threshold makes them fire more on correct work, lowering it makes them blinder to repetition. See `governance-campaign-findings.md`. **Measured 2026-09-06** on shipped defaults: coherence floors at 0.275, not 0.0 — the six-turn floor does not occur (S10/S15 absorbed by adaptive baseline). Threshold decision remains open; the measurement narrows the precondition, not the defect in KIND. See `investigate/deescalation/FINDINGS.md`. | | C1c | **De-escalation is effectively unreachable for a real agent once coherence floors — measured, not inferred.** Composite reduces to `0.35·(1−coherence) + 0.10·depth + 0.25·min(1, signals/4)`, exact against every row of three keyed runs. With coherence at 0.0 the first term is permanently 0.35, so leaving HIGH (hold 0.55) requires **≤1 signal for 3 consecutive turns** — two signals is 0.575. Across keyed runs 4, 5 and 6 there are **84 post-escalation worker turns; 4 had ≤1 signal, and never two consecutively.** `entity_consistency` fired 34/36 and `semantic_stability` 29/36 in run 6. At ~5% of turns being quiet enough, three consecutive is ~0.01%: it does not happen at any run length. | **precondition narrowed (2026-09-06)** | This is no longer "the scenario didn't give the mechanism its calm". A real model's responses fire 2+ CDD signals per turn as a matter of course, so the calm the hysteresis requires is not a thing real traffic produces. The decision is whether that is intended. Options, all of them loosenings and none obviously right: raise the per-level hold margin so 2 signals sits below it; count calm on a *rate* rather than a consecutive run; or accept that de-escalation requires healing (which R3 withdrew as unsettable globally, and R5 supersedes). **Do not "fix" this by tuning v3's prompts** — four prompt revisions across runs 3-6 each moved which signal fired without changing that two fire. **Measured 2026-09-06** on shipped defaults: escalation did not occur (25 turns, coherence floor 0.275, level NORMAL throughout). ARM A never escalated, so this is NOT evidence that de-escalation is reachable — no de-escalation events were produced. Control arm (baseline OFF) reproduced the register findings exactly (coherence 0.0 turn 3, ELEVATED turn 4, HIGH turn 9, 20 turns before timeout). If any workload reaches ELEVATED on shipped defaults, this row's reasoning is unchanged. The resolution is accidental (#176, 2026-08-29); see C1a's blocker. See `investigate/deescalation/FINDINGS.md`. | | C2 | **HIGH and CRITICAL levels** | **HIGH live-confirmed, CRITICAL open** | HIGH reached keylessly at turn 5 on 2026-08-09, then **live against a real model on keyed run 2** (ELEVATED turn 8, HIGH turn 10, `drift_worker` the pressure handle at both) — see E6. CRITICAL was deliberately tuned out of reach in both. | diff --git a/run-all-tests.sh b/run-all-tests.sh index 981dbcf8f..1c37092e6 100755 --- a/run-all-tests.sh +++ b/run-all-tests.sh @@ -2824,9 +2824,11 @@ GOVV4_SWEEP_SKIP=( # It states what "fixed" means for the semantic signals: no signal may fire # as often on correct work as on drift (C1), correct work must not exhaust # the coherence budget (C2), restating the mandate must not outscore doing - # the task (C3), and every drift arm must still be caught (C4). On the - # engine as of this entry it fails all four while all five of its own - # vacuity checks pass, so the failures are about the engine. + # the task (C3), and every drift arm must still be caught (C4). When this + # entry was written it failed all four while all five of its own vacuity + # checks passed, so the failures are about the engine. As of 80c8cd9 + # (2026-09-29) C1 and C2 pass and C3/C4 still fail -- the gate stays out of + # the sweep until all four pass. # # Deleting this line is the act of wiring it in. Do that in the SAME commit # that makes it pass -- not before, or the suite stops being a signal.