Skip to content

One implementation of string slice/substring/replace; check inline code passed through process.run - #262

Merged
b-macker merged 3 commits into
masterfrom
claude/naab-inadmissible-action-prevention-4cmn1m
Sep 28, 2026
Merged

b-macker merged 3 commits into
masterfrom
claude/naab-inadmissible-action-prevention-4cmn1m

Conversation

@b-macker

@b-macker b-macker commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Summary

Round 3 of dogfooding (the repo-sentinel project) turned up two problems. This PR fixes both.

1. String methods disagreed between engines, and one hung

Round 3 reported F-007, the string.slice hint, as "not fixed". It had run a build without #261, which does suggest substring. Checking it turned up something bigger:

  • s.slice(...) worked as a method, while string.slice(...) didn't exist as a function.
  • The method forms of three string operations gave different results in the VM and the tree-walker.

Each operation had four hand-written copies: VM method dispatch, two tree-walker method paths, and the string module.

Call VM Tree-walker
s.replace("-", "+") a+b-c (first match only) a+b+c
s.substring(3, 1) "" "lo" (a length calculation underflows)
s.slice(-3) llo hello (treated as substring)
"ab".replace("", "+") +ab hangs

The last row is the serious one. The tree-walker's replace didn't guard against an empty pattern. An empty string matches at every position, so it kept inserting forever while memory grew. --timeout can't stop it, because the loop never returns to the interpreter, and the REST API runs the tree-walker.

2. Inline code in process.run skipped the code checks

repo-sentinel reported every C++ file as 0 bytes. The cause was a Python helper that swallowed every error:

except Exception:
    return 0

The project's own code_quality.no_incomplete_logic (HARD) blocks that exact code in a <<python>> block and in codegen.run(). repo-sentinel ran it as process.run("python3", ["-c", code]) instead, and it went through with no check at all. NAAb already had the rule that would have caught the bug.

Changes

  • New include/naab/string_ops.h with substring, slice and replaceAll, called from all four places:
    • The string module functions already agreed with each other in every case, so their behaviour is the reference.
    • replace replaces every occurrence; an empty pattern leaves the string unchanged.
    • substring clamps out-of-range indices and returns "" for an empty or reversed range.
    • slice follows JavaScript: negative indices count from the end.
  • string.slice(s, start[, end]) is now a module function, behaving the same as the method.
  • string.substr keeps its hint pointing to substring, because JavaScript's substr takes a length rather than an end index.
  • process.run inline-code check (src/stdlib/process_impl.cpp): when the command is an interpreter given inline code, that code goes through checkPolyglotBlock(), the same check <<python>> blocks and codegen.run() use. Recognised:
    • Interpreters and flags:
      • python/pypy -c
      • node -e/-p/--eval/--print, also --eval=
      • ruby -e
      • perl -e/-E
      • php -r
      • sh/bash/dash/zsh/ksh -c
    • Also matched through a directory prefix, .exe, a version suffix (python3.12), and combined short flags when the code flag comes last (-Ic, -ec).
    • Side effect: languages.allowed/blocked now apply to that inline code too.
    • Not covered: a script file (python3 s.py) isn't scanned, and only the outer interpreter is recognised.
  • CLAUDE.md: notes for both changes.

Test Plan

  • String fixes:
    • New differential probe tests/differential/corpus/string_methods.naab. The corpus only exercised the module form of these functions, which is why none of this showed up. Differential run: 57 passed, 0 failed.
    • tests/parser/test_dogfood_hints.sh: 21/21 pass. S-06 checks string.slice works; E-01 checks empty-pattern replace under a 10 s limit.
    • Against the previous build, S-06 and E-01 fail in both engines; E-01 on the tree-walker fails by being killed at 10 s.
  • process.run check:
    • New tests/security/test_process_run_inline_gate.sh: 14/14 across both engines.
    • Every blocking case checks that the code's marker file was never written, not just the exit code.
    • Against the previous build it fails only the two blocking cases (PI-01, PI-03), in both engines.
    • Controls:
      • PI-00: the <<python>> block is blocked under this config, so the config does exercise the check.
      • PI-02, PI-04, PI-05: harmless python3 -c, harmless sh -c and echo -c <text> still run.
      • PI-06: a script file is out of scope and still runs.
  • Existing callers: the python3 -c and sh -c calls in living-script, living-script_v2, living-script_extended, hivemind and hivemind_governed run unchanged under their own govern.json, on both builds.
  • Leak check: 874 passed, 0 failed. The new code adds no error text of its own; it passes on checkPolyglotBlock's messages.
  • Full suite: a fresh run including the new test is pending. I'll post the result here.

🤖 Generated with Claude Code

https://claude.ai/code/session_01ELUfjXZvx8kzXo1UJjrAhC

…er hang

Dogfood round 3 reported string.slice "not fixed" (it ran a build without
#261), but checking it found that s.slice(...) worked as a METHOD while
string.slice(...) did not exist -- and that the method forms disagreed
between engines. Four copies (VM method dispatch, two tree-walker method
paths, the string module); measured on "hello" / "a-b-c":

  s.replace("-", "+")   VM a+b-c (first only)   tree-walker a+b+c
  s.substring(3, 1)     VM ""                   tree-walker "lo" (wrapped size_t)
  s.slice(-3)           VM llo                  tree-walker hello (aliased to substring)

and "ab".replace("", "+") hung the tree-walker: find("") matches at every
position, so it inserted forever with growing memory. --timeout cannot stop
it (the loop never returns to the interpreter), and the REST API runs the
tree-walker. The VM gave "+ab".

The module functions agreed with each other in every case, so they are the
reference: include/naab/string_ops.h holds substring, slice (JavaScript
semantics) and replaceAll (empty pattern = unchanged), and all four sites
call it. string.slice(s, start[, end]) is added as a module function.

Tests: tests/differential/corpus/string_methods.naab (the corpus only had
module-form string probes, which is why none of this surfaced) and
test_dogfood_hints.sh S-06 (string.slice) and E-01 (the hang, both engines,
10s limit). Against the previous build S-06 and E-01 fail in both engines,
E-01/tree-walk by being killed at 10s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELUfjXZvx8kzXo1UJjrAhC
@github-actions

Copy link
Copy Markdown

NAAb Governance Report

Metric Count
Files checked 16
Passed 16
Failed 0

All governance checks passed!

Generated by NAAb Governance Engine v4.0

process.run("python3", ["-c", code]) is a <<python>> block by another name,
but it skipped every code check the block and codegen.run() get. Found
dogfooding repo-sentinel: a size helper that swallowed every error
(`except Exception: return 0`) and reported every C++ file as 0 bytes. The
project's own no_incomplete_logic (HARD) blocks that code as a block and in
codegen.run(); through process.run it ran silently.

process.run now recognises python/pypy -c, node -e/-p/--eval/--print,
ruby -e, perl -e/-E, php -r and sh/bash/dash/zsh/ksh -c (through a directory
prefix, .exe, a version suffix, and clustered short flags) and passes the
code to checkPolyglotBlock(). A script file is not inline code and is not
scanned; only the outer interpreter is recognised.

Test: tests/security/test_process_run_inline_gate.sh, 14/14 on both engines;
the prior build fails exactly PI-01 and PI-03 on each. The python -c and
sh -c calls in living-script{,_v2,_extended} and hivemind{,_governed} run
unchanged under their own configs on both builds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELUfjXZvx8kzXo1UJjrAhC
@b-macker b-macker changed the title One implementation of string slice/substring/replace; fix a tree-walker hang One implementation of string slice/substring/replace; check inline code passed through process.run Sep 28, 2026

Copy link
Copy Markdown
Owner Author

Local results at 325b765 (includes the process.run inline-code gate):

  • Full suite: 447 total, 381 passed, 52 error-behavior, 12 needs-tree-walk, 2 missing-executor, 2 unexpected.
    • The 2 unexpected are test_require_governance_gov007.sh and test_platform_fixes.sh.
    • Both are caused by a stray /govern.json on this machine, and the pre-change baseline (ce0ca24) had the identical result.
    • The newly registered test_process_run_inline_gate.sh passed.
  • Error-message leak check: 874 passed, 0 failed.
  • Differential v2: 57 passed, 0 failed.

CI on 325b765 is green, including build-windows.


Generated by Claude Code

@b-macker
b-macker marked this pull request as ready for review September 28, 2026 20:02
@b-macker
b-macker merged commit 4bb77dc into master Sep 28, 2026
24 checks passed
@b-macker
b-macker deleted the claude/naab-inadmissible-action-prevention-4cmn1m branch September 28, 2026 20:02
b-macker pushed a commit that referenced this pull request Sep 28, 2026
… failed call

Found checking the repo-sentinel F-008 handoff.

1. Coherence moved without telemetry. Natural healing, temporal decay, the
   clamp at 0 and recoverCoherence() (passed step-up, failed pipeline stage)
   were reported nowhere, so listed penalties never matched the drop and the
   dogfood reported CDD's arithmetic as broken. It was exact once healing was
   added back. CDD_TURN now carries coherence_adjustments
   (temporal_decay/natural_healing/floor_absorbed/recovery), kept out of
   penalties_detail because consumers read a non-empty penalties_detail as "a
   signal paid". validation_recovery reports the credit received.

2. The response after a retry-exhausted API failure was never analyzed. The
   failure was analyzed at the same turn number the next response carries,
   took the interval slot, and the response was skipped by all 23 signals
   while its CDD_TURN said analyzed:"true". Infrastructure errors now return
   before analysis when exclude_infrastructure_errors is on (default: they
   feed no signal), and hand the slot back when it is off. The analyzed
   label now comes from an analysis counter, not a turn-number comparison.

3. Untrack examples/hivemind/src/hivemind-telemetry.jsonl (10 MB). It is the
   live output target of the hivemind configs, so every run appended to a
   tracked file; nothing reads it.

Test: tests/governance_v4/test_coherence_reconcile.sh, 7/7. The #262 build
fails RC-01/RC-02 (11 of 12 rows do not reconcile) and IA-01/IA-02 (the
repeated response escapes S21 while labelled analyzed); IA-03 is the control.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELUfjXZvx8kzXo1UJjrAhC
b-macker added a commit that referenced this pull request Sep 28, 2026
… failed call (#263)

Found checking the repo-sentinel F-008 handoff.

1. Coherence moved without telemetry. Natural healing, temporal decay, the
   clamp at 0 and recoverCoherence() (passed step-up, failed pipeline stage)
   were reported nowhere, so listed penalties never matched the drop and the
   dogfood reported CDD's arithmetic as broken. It was exact once healing was
   added back. CDD_TURN now carries coherence_adjustments
   (temporal_decay/natural_healing/floor_absorbed/recovery), kept out of
   penalties_detail because consumers read a non-empty penalties_detail as "a
   signal paid". validation_recovery reports the credit received.

2. The response after a retry-exhausted API failure was never analyzed. The
   failure was analyzed at the same turn number the next response carries,
   took the interval slot, and the response was skipped by all 23 signals
   while its CDD_TURN said analyzed:"true". Infrastructure errors now return
   before analysis when exclude_infrastructure_errors is on (default: they
   feed no signal), and hand the slot back when it is off. The analyzed
   label now comes from an analysis counter, not a turn-number comparison.

3. Untrack examples/hivemind/src/hivemind-telemetry.jsonl (10 MB). It is the
   live output target of the hivemind configs, so every run appended to a
   tracked file; nothing reads it.

Test: tests/governance_v4/test_coherence_reconcile.sh, 7/7. The #262 build
fails RC-01/RC-02 (11 of 12 rows do not reconcile) and IA-01/IA-02 (the
repeated response escapes S21 while labelled analyzed); IA-03 is the control.


Claude-Session: https://claude.ai/code/session_01ELUfjXZvx8kzXo1UJjrAhC

Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants