Skip to content

CDD: reconcile coherence from telemetry; analyze the response after a failed call - #263

Merged
b-macker merged 1 commit into
masterfrom
claude/naab-inadmissible-action-prevention-4cmn1m
Sep 28, 2026
Merged

b-macker merged 1 commit into
masterfrom
claude/naab-inadmissible-action-prevention-4cmn1m

Conversation

@b-macker

Copy link
Copy Markdown
Owner

Summary

Checking the repo-sentinel F-008 handoff (the raw telemetry for both Stage 4 runs) turned up two CDD defects and one repo-hygiene fix.

1. Coherence moved without any telemetry

F-008 reported that the listed penalties never matched the coherence drops. CDD's arithmetic was correct: each drop was exactly the penalties minus natural healing (0.03 / (1 + signals_fired)), to four decimals on all six turns. The healing just wasn't reported anywhere. Three other movements weren't reported either:

  • temporal decay
  • the clamp at 0, which discards any penalty beyond 0
  • recoverCoherence(), called on a passed step-up challenge or a failed pipeline stage

CDD_TURN now carries a coherence_adjustments field, so every analyzed row can be reconciled from telemetry alone:

coherence = previous analyzed coherence - penalties + validation_recovery + sum(coherence_adjustments)
  • Contents: temporal_decay=-, natural_healing=+, floor_absorbed=+, recovery=+.
  • Decay and recovery happen between analyzed turns. They build up in DriftState.pending_* and are reported on the next analyzed row.
  • It's a separate field on purpose. living-script_v3/report.py, its gate registry and test_signal_contract.sh all read a non-empty penalties_detail as "a signal paid this turn", and healing lands on nearly every turn after damage.
  • validation_recovery now reports the credit actually received. It's capped at 1.0, so it can be less than the configured amount.

2. The response after a failed API call was never analyzed

Writing the reconciliation test exposed this one.

  • Cause: a retry-exhausted failure is reported to checkContextDrift() at the same turn number the next real response will carry, because a failed call doesn't advance the turn. That analysis took the turn's slot, and the next response then hit the interval check.
  • Effect: after any failed call (a 429 is the usual case), the next response was scored by none of the 23 signals, while its CDD_TURN still said analyzed:"true".

The fix:

  • With exclude_infrastructure_errors on (the default): the infrastructure error returns before any analysis, and before the event-feed watermark moves. It feeds no signal, so nothing is lost.
  • With it off: the error still feeds repeated_failures, so the analysis runs as before, then hands the turn's slot back.
  • analyzed label: it now comes from a per-handle analysis counter instead of comparing turn numbers.

3. Untrack examples/hivemind/src/hivemind-telemetry.jsonl (10 MB)

It's the live output file of the hivemind configs, so every run appended to a tracked file, and nothing reads it. It's added to .gitignore; local copies are untouched.

Test Plan

  • New tests/governance_v4/test_coherence_reconcile.sh: 7/7, registered in run-all-tests.sh.
    • RC-01: all 12 analyzed rows reconcile (worst residual 0.0001, from 4-decimal rounding).
    • RC-02: the run exercised every adjustment kind.
    • RC-03: negative control; without the new field, 11 rows fail to reconcile.
    • RC-04: penalties_detail stays signal-only.
    • IA-01/IA-02: under both exclude_infrastructure_errors settings, a verbatim repeat sent after a failed call fires response_repetition.
    • IA-03: the same repeat with no failure before it fires too, so the fixture can fire S21 at all.
  • Against the One implementation of string slice/substring/replace; check inline code passed through process.run #262 build:
    • RC-01 and RC-02 fail: 11 of 12 rows don't reconcile.
    • IA-01 and IA-02 fail: the repeat escapes CDD while labelled analyzed=true.
    • IA-03 still passes.
  • Leak check: 874/874.
  • Still running, to be posted here: the related CDD suites (semantic signals, validation signal, signal contract, challenge fail path) and the full suite. test_event_feed.sh has already passed.

🤖 Generated with Claude Code

https://claude.ai/code/session_01ELUfjXZvx8kzXo1UJjrAhC


Generated by Claude Code

… failed call

Found checking the repo-sentinel F-008 handoff.

1. Coherence moved without telemetry. Natural healing, temporal decay, the
   clamp at 0 and recoverCoherence() (passed step-up, failed pipeline stage)
   were reported nowhere, so listed penalties never matched the drop and the
   dogfood reported CDD's arithmetic as broken. It was exact once healing was
   added back. CDD_TURN now carries coherence_adjustments
   (temporal_decay/natural_healing/floor_absorbed/recovery), kept out of
   penalties_detail because consumers read a non-empty penalties_detail as "a
   signal paid". validation_recovery reports the credit received.

2. The response after a retry-exhausted API failure was never analyzed. The
   failure was analyzed at the same turn number the next response carries,
   took the interval slot, and the response was skipped by all 23 signals
   while its CDD_TURN said analyzed:"true". Infrastructure errors now return
   before analysis when exclude_infrastructure_errors is on (default: they
   feed no signal), and hand the slot back when it is off. The analyzed
   label now comes from an analysis counter, not a turn-number comparison.

3. Untrack examples/hivemind/src/hivemind-telemetry.jsonl (10 MB). It is the
   live output target of the hivemind configs, so every run appended to a
   tracked file; nothing reads it.

Test: tests/governance_v4/test_coherence_reconcile.sh, 7/7. The #262 build
fails RC-01/RC-02 (11 of 12 rows do not reconcile) and IA-01/IA-02 (the
repeated response escapes S21 while labelled analyzed); IA-03 is the control.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELUfjXZvx8kzXo1UJjrAhC
@github-actions

Copy link
Copy Markdown

NAAb Governance Report

Metric Count
Files checked 16
Passed 16
Failed 0

All governance checks passed!

Generated by NAAb Governance Engine v4.0

Copy link
Copy Markdown
Owner Author

Local results at 9aadb73:

CI on 9aadb73 is green, including build-windows.


Generated by Claude Code

@b-macker
b-macker marked this pull request as ready for review September 28, 2026 22:29
@b-macker
b-macker merged commit 80c8cd9 into master Sep 28, 2026
23 checks passed
@b-macker
b-macker deleted the claude/naab-inadmissible-action-prevention-4cmn1m branch September 28, 2026 22:29
b-macker pushed a commit that referenced this pull request Sep 29, 2026
…ries

- docs/open-investigations.md C1f: the repo-sentinel dogfood (round 5, live
  Gemini, NAAb 80c8cd9) is the first real-world case of S5's startup-frozen
  entropy baseline. All four runs of one arm (clean and adversarial fixtures)
  reach 0.556667 and OUTPUT_INADMISSIBLE at turn 8, identical to six
  decimals; S5 first fires at turn 6 (vocab_contraction_window) whatever the
  baseline window. Every analyzed row reconciled from telemetry (#263). The
  row's "no shipped example is known to be affected" is marked superseded.
- run-all-tests.sh: the test_signal_contract.sh skip note said the gate fails
  all four criteria; C1/C2 now pass and C3/C4 still fail.
- Untrack docs/book/.../Vigilant/{proxy/gateway_vessel,scanner/shield_vessel}
  (~9.4 MB, Android aarch64). Nothing reads them: synthesizer.naab builds
  into Vigilant/bin/ from the gateway.go / shield.rs sources beside them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELUfjXZvx8kzXo1UJjrAhC
b-macker added a commit that referenced this pull request Sep 29, 2026
…ries (#264)

- docs/open-investigations.md C1f: the repo-sentinel dogfood (round 5, live
  Gemini, NAAb 80c8cd9) is the first real-world case of S5's startup-frozen
  entropy baseline. All four runs of one arm (clean and adversarial fixtures)
  reach 0.556667 and OUTPUT_INADMISSIBLE at turn 8, identical to six
  decimals; S5 first fires at turn 6 (vocab_contraction_window) whatever the
  baseline window. Every analyzed row reconciled from telemetry (#263). The
  row's "no shipped example is known to be affected" is marked superseded.
- run-all-tests.sh: the test_signal_contract.sh skip note said the gate fails
  all four criteria; C1/C2 now pass and C3/C4 still fail.
- Untrack docs/book/.../Vigilant/{proxy/gateway_vessel,scanner/shield_vessel}
  (~9.4 MB, Android aarch64). Nothing reads them: synthesizer.naab builds
  into Vigilant/bin/ from the gateway.go / shield.rs sources beside them.


Claude-Session: https://claude.ai/code/session_01ELUfjXZvx8kzXo1UJjrAhC

Co-authored-by: Claude <noreply@anthropic.com>
b-macker pushed a commit that referenced this pull request Sep 29, 2026
- docs/governance-campaign-findings.md: new section for rounds 3-5 of the
  repo-sentinel dogfood (live Gemini). coherence_adjustments (#263) is
  confirmed live: 171 analyzed rows across 8 runs, 0 unreconciled. The
  post-failure analysis fix (#263) stays stub-only: round 5 retried on 503s
  but no send was shown to exhaust its retries. C1f's defect is confirmed
  live; its new default (#265) is stub-only until a round runs on it. Also
  records that round 4's "adversarial caught earlier" did not reproduce
  (two failed validations in eight runs), and that round 3's evidence was
  deleted by the project's run.sh before it could be checked.
- docs/open-investigations.md C1f: status "partially addressed; ships off"
  -> "fixed; on by default since #265".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELUfjXZvx8kzXo1UJjrAhC
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants