Skip to content

Cap every check_results_ writer, fix r22/r32 probes, unpin python3.12 - #256

Merged
b-macker merged 3 commits into
masterfrom
claude/naab-inadmissible-action-prevention-4cmn1m
Sep 26, 2026
Merged

b-macker merged 3 commits into
masterfrom
claude/naab-inadmissible-action-prevention-4cmn1m

Conversation

@b-macker

@b-macker b-macker commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

Summary

Bundles portability item 7 with diagnosis and registration of the two held-back security suites, test_r22_fixes.sh and test_r32_fixes.sh. Both suites were failing because their probes were broken, and both broken probes were hiding something.

Changes

  • V-GOV-024 cap gap (real defect): check_results_ was capped at MAX_CHECK_RESULTS only in recordPass() and the violation path. The ten pass2.* writers in the post-execution audit (up to ~7 per polyglot block) and the polyglot_optimization writer had no cap. Every writer now goes through capCheckResultsLocked(), which does a single pass and evicts preflight entries only as a last resort.
  • r32 T4: the old check, grep -A3 push_back | grep erase, missed the real eviction because it was four lines down, and it could never notice the uncapped sites. It now attributes each writer to its enclosing function and requires that function to apply the cap. It has a self-control. Run against the old source it names all 12 uncapped sites; against the new source it names none.
  • r22 V-GOV-018: the fixtures predate the rule that mode: enforce upgrades the sandbox to standard. Under that sandbox every <<shell block was refused (exit 1), so the two "allowed" arms failed and the two "blocked" arms passed without ever reaching per-agent policy. The fixtures now set sandbox_level: elevated. Blocked arms require the governance refusal (exit 3), and allowed arms require the shell output marker. If the platform has no shell executor, as on Windows, the allowed arms are skipped and reported as UNMEASURABLE.
  • r22 V-GOV-017 symlink arms: naab-gov scan never prints file contents, so checking the output for the secret could not fail. The arms now compare scanned-file counts: a symlink that is not followed counts as 0 files, and a positive control checks that a regular .naab file counts as 1. Symlink targets are .naab files, so an extension filter cannot explain a zero. /etc/hostname, which does not exist on Windows, is no longer used. A platform that cannot create symlinks reports SKIP.
  • r22 other probes: if naab-gov is missing, the suite now stops and reports it as UNMEASURABLE. The "normal file" arm grepped for "", which always matches, and it had no discoverable govern.json, so require-governance refused to run it. It now has its own config and expects 42. Scans run from the work dir, so quality-report.* no longer lands in the caller's cwd.
  • CI: naab-gov is now built in both ci.yml Linux jobs and in windows.yml's build-linux, which are all the Linux jobs that run run-all-tests.sh.
  • paths.cpp (item 7): when CMake did not find Python, the fallback returned a hardcoded python3.12 path. It now picks the highest python3.N that contains Python.h under $PREFIX/include, /usr/local/include and /usr/include, and keeps the old path as a last resort.
  • run-all-tests.sh: r22 and r32 are registered, and the old exclusion comment is replaced with the diagnosis.

Test Plan

  • test_r22_fixes.sh 9/9 and test_r32_fixes.sh 9/9 locally; CI is green on all platforms, including build-windows.
  • The python_include_dir() fallback was compiled standalone. It chose 3.13 over 3.9 (numeric ordering), skipped a directory with no Python.h, and fell through correctly when $PREFIX did not exist.
  • test_coverage_visibility.sh and test_evidence_chain.sh pass; leak check 874/0.
  • Full suite in the dev container: 447 tests, 381 passed, 2 unexpected failures. This matches the master baseline; both failures come from a stray root /govern.json that exists only in the dev container. CI has none.

Not changed, and worth a follow-up: naab-gov scan is very slow. A Debug build takes about 35 s/MB on NUL-filled input and more than 5 min on 1 MB of short lines. Because of that, r22's 11–15 MB capped-read arm takes about 6 minutes in Debug.

🤖 Generated with Claude Code

https://claude.ai/code/session_01ELUfjXZvx8kzXo1UJjrAhC

V-GOV-024: check_results_ was capped at MAX_CHECK_RESULTS only in
recordPass() and the violation path. The ten pass2.* writers in the
post-execution audit (up to ~7 per polyglot block) and the
polyglot_optimization writer had no cap. All writers now go through
capCheckResultsLocked(), a single-pass, preflight-aware eviction.

test_r32_fixes.sh T4 grepped for an eviction within three lines of ANY
push_back: it missed the real one (four lines down) and could never see the
uncapped sites. It now attributes each writer to its function and requires
a cap, with a self-control; on the old source it names all 12 sites.

test_r22_fixes.sh: fixtures predated the enforce-mode sandbox upgrade, so
every <<shell was refused by the sandbox -- the "allowed" arms failed and
the "blocked" arms passed without reaching per-agent policy. Fixtures now
set sandbox_level: elevated; blocked arms require exit 3, allowed arms
require the marker. A missing naab-gov is UNMEASURABLE (it made the
symlink-leak arms pass for free), the "normal file" arm no longer greps
for "", and the Linux CI jobs now build naab-gov. Both suites registered.

paths.cpp: when CMake did not find Python, python_include_dir() returned a
hardcoded python3.12 path. It now picks the highest python3.N holding
Python.h under $PREFIX/include, /usr/local/include, /usr/include.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELUfjXZvx8kzXo1UJjrAhC
@github-actions

Copy link
Copy Markdown

NAAb Governance Report

Metric Count
Files checked 16
Passed 16
Failed 0

All governance checks passed!

Generated by NAAb Governance Engine v4.0

build-windows showed three more broken probes in test_r22_fixes.sh:

- The symlink arms grepped scan output for the secret, but naab-gov scan
  never prints file contents, so they could not fail. They now use the
  scanned-file count (a rejected symlink is not counted), with a positive
  control that a regular .naab file IS counted, and a .naab target so an
  extension filter cannot explain the zero. /etc/hostname (absent on
  Windows) is gone; a platform that cannot create symlinks is a SKIP.
- The "normal file" arm had no discoverable govern.json, so the default
  require-governance refused to run. It passed locally only because of a
  stray config at /. It now gets its own.
- The shell "allowed" arms report UNMEASURABLE when the platform has no
  shell executor, instead of failing as a governance block.

Scans now run from the work dir so quality-report.* is not written into
the caller's cwd.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELUfjXZvx8kzXo1UJjrAhC
build-linux runs run-all-tests.sh but built only naab-lang and libnaab.
test_r22_fixes.sh's new usability guard correctly refused to report its
scanner arms as passing without the binary; before the guard, this job
would have passed them without a scan ever running.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELUfjXZvx8kzXo1UJjrAhC
@b-macker
b-macker marked this pull request as ready for review September 26, 2026 16:13
@b-macker
b-macker merged commit 7f1f1cf into master Sep 26, 2026
24 checks passed
@b-macker
b-macker deleted the claude/naab-inadmissible-action-prevention-4cmn1m branch September 26, 2026 16:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants