Repository navigation
Cap every check_results_ writer, fix r22/r32 probes, unpin python3.12 - #256
Merged
b-macker merged 3 commits intoSep 26, 2026
Merged
Conversation
V-GOV-024: check_results_ was capped at MAX_CHECK_RESULTS only in recordPass() and the violation path. The ten pass2.* writers in the post-execution audit (up to ~7 per polyglot block) and the polyglot_optimization writer had no cap. All writers now go through capCheckResultsLocked(), a single-pass, preflight-aware eviction. test_r32_fixes.sh T4 grepped for an eviction within three lines of ANY push_back: it missed the real one (four lines down) and could never see the uncapped sites. It now attributes each writer to its function and requires a cap, with a self-control; on the old source it names all 12 sites. test_r22_fixes.sh: fixtures predated the enforce-mode sandbox upgrade, so every <<shell was refused by the sandbox -- the "allowed" arms failed and the "blocked" arms passed without reaching per-agent policy. Fixtures now set sandbox_level: elevated; blocked arms require exit 3, allowed arms require the marker. A missing naab-gov is UNMEASURABLE (it made the symlink-leak arms pass for free), the "normal file" arm no longer greps for "", and the Linux CI jobs now build naab-gov. Both suites registered. paths.cpp: when CMake did not find Python, python_include_dir() returned a hardcoded python3.12 path. It now picks the highest python3.N holding Python.h under $PREFIX/include, /usr/local/include, /usr/include. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ELUfjXZvx8kzXo1UJjrAhC
NAAb Governance Report
All governance checks passed! Generated by NAAb Governance Engine v4.0 |
build-windows showed three more broken probes in test_r22_fixes.sh: - The symlink arms grepped scan output for the secret, but naab-gov scan never prints file contents, so they could not fail. They now use the scanned-file count (a rejected symlink is not counted), with a positive control that a regular .naab file IS counted, and a .naab target so an extension filter cannot explain the zero. /etc/hostname (absent on Windows) is gone; a platform that cannot create symlinks is a SKIP. - The "normal file" arm had no discoverable govern.json, so the default require-governance refused to run. It passed locally only because of a stray config at /. It now gets its own. - The shell "allowed" arms report UNMEASURABLE when the platform has no shell executor, instead of failing as a governance block. Scans now run from the work dir so quality-report.* is not written into the caller's cwd. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ELUfjXZvx8kzXo1UJjrAhC
build-linux runs run-all-tests.sh but built only naab-lang and libnaab. test_r22_fixes.sh's new usability guard correctly refused to report its scanner arms as passing without the binary; before the guard, this job would have passed them without a scan ever running. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ELUfjXZvx8kzXo1UJjrAhC
b-macker
marked this pull request as ready for review
September 26, 2026 16:13
b-macker
deleted the
claude/naab-inadmissible-action-prevention-4cmn1m
branch
September 26, 2026 16:13
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Bundles portability item 7 with diagnosis and registration of the two held-back security suites,
test_r22_fixes.shandtest_r32_fixes.sh. Both suites were failing because their probes were broken, and both broken probes were hiding something.Changes
check_results_was capped atMAX_CHECK_RESULTSonly inrecordPass()and the violation path. The tenpass2.*writers in the post-execution audit (up to ~7 per polyglot block) and thepolyglot_optimizationwriter had no cap. Every writer now goes throughcapCheckResultsLocked(), which does a single pass and evicts preflight entries only as a last resort.grep -A3 push_back | grep erase, missed the real eviction because it was four lines down, and it could never notice the uncapped sites. It now attributes each writer to its enclosing function and requires that function to apply the cap. It has a self-control. Run against the old source it names all 12 uncapped sites; against the new source it names none.mode: enforceupgrades the sandbox tostandard. Under that sandbox every<<shellblock was refused (exit 1), so the two "allowed" arms failed and the two "blocked" arms passed without ever reaching per-agent policy. The fixtures now setsandbox_level: elevated. Blocked arms require the governance refusal (exit 3), and allowed arms require the shell output marker. If the platform has no shell executor, as on Windows, the allowed arms are skipped and reported as UNMEASURABLE.naab-gov scannever prints file contents, so checking the output for the secret could not fail. The arms now compare scanned-file counts: a symlink that is not followed counts as 0 files, and a positive control checks that a regular.naabfile counts as 1. Symlink targets are.naabfiles, so an extension filter cannot explain a zero./etc/hostname, which does not exist on Windows, is no longer used. A platform that cannot create symlinks reports SKIP.naab-govis missing, the suite now stops and reports it as UNMEASURABLE. The "normal file" arm grepped for"", which always matches, and it had no discoverablegovern.json, so require-governance refused to run it. It now has its own config and expects42. Scans run from the work dir, soquality-report.*no longer lands in the caller's cwd.naab-govis now built in bothci.ymlLinux jobs and inwindows.yml'sbuild-linux, which are all the Linux jobs that runrun-all-tests.sh.python3.12path. It now picks the highestpython3.Nthat containsPython.hunder$PREFIX/include,/usr/local/includeand/usr/include, and keeps the old path as a last resort.Test Plan
test_r22_fixes.sh9/9 andtest_r32_fixes.sh9/9 locally; CI is green on all platforms, including build-windows.python_include_dir()fallback was compiled standalone. It chose 3.13 over 3.9 (numeric ordering), skipped a directory with noPython.h, and fell through correctly when$PREFIXdid not exist.test_coverage_visibility.shandtest_evidence_chain.shpass; leak check 874/0./govern.jsonthat exists only in the dev container. CI has none.Not changed, and worth a follow-up:
naab-gov scanis very slow. A Debug build takes about 35 s/MB on NUL-filled input and more than 5 min on 1 MB of short lines. Because of that, r22's 11–15 MB capped-read arm takes about 6 minutes in Debug.🤖 Generated with Claude Code
https://claude.ai/code/session_01ELUfjXZvx8kzXo1UJjrAhC