Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -256,6 +256,7 @@ include/naab/ All headers
- **Subprocess containment** (`src/runtime/subprocess_helpers.h/cpp`): `SubprocessContainment` struct applied to all polyglot child processes via `execute_subprocess_with_pipes()`. The memory budget is `RLIMIT_DATA` (private writable memory), with `RLIMIT_AS` only as a backstop ceiling (4x the budget, at least 2 GB) for shared anonymous mappings, which `RLIMIT_DATA` does not count. It used to be `RLIMIT_AS` at the budget, which counts untouched reservations: node 22 needs 512-768 MB of address space just to start, so `process.run("node", ...)` died with "Failed to reserve virtual memory" under `elevated`. Test: `tests/security/test_child_memory_limit.sh` (M-02/M-03 are why the fix is not "drop the limit"). 5-layer defense: RLIMIT_NPROC=0 (blocks fork/subprocess), PATH restriction, env scrubbing, timeout (SIGKILL), allow_exec/allow_fork flags. `SubprocessContainment::fromCurrentSandbox()` reads current sandbox level: `standard` → fork+exec blocked, `elevated` → fork allowed, `unrestricted` → no restrictions. Pre-execution source scanning (governance checks) catches static patterns; containment catches runtime-constructed commands that evade static analysis.

### Stdlib Notable
- **String `slice` / `substring` / `replace` have ONE implementation**, `include/naab/string_ops.h`, used by the VM's method dispatch, both tree-walker method paths and the `string` module. There were four copies and they disagreed: `s.replace` replaced only the FIRST match on the VM (all on the tree-walker and in `string.replace`), the tree-walker aliased `slice` to `substring` (so `s.slice(-3)` was `hello` there and `llo` on the VM), and its reversed-range `substring` returned the rest of the string through a wrapped `size_t`. Worst, the tree-walker's `replace` had no empty-pattern guard: `"ab".replace("", "+")` inserted forever with growing memory, which `--timeout` cannot interrupt, and REST runs the tree-walker. `string.slice(s, start[, end])` now exists with the method's JavaScript semantics. The differential corpus had only module-form string probes; `tests/differential/corpus/string_methods.naab` covers the method forms, and `tests/parser/test_dogfood_hints.sh` E-01 guards the hang.
- `array.sort()` mutates in place, `array.sorted()` returns a new sorted array (non-mutating)
- Both support optional comparator: `arr.sorted(fn(a, b) { return a - b })`
- `sorted()` is registered in `hasFunction()` but NOT in `isMutatingFunction()`
Expand Down Expand Up @@ -399,6 +400,7 @@ Always run `bash tests/security/test_error_msg_leaks.sh` after changing any erro
- **Agent responses auto-strip markdown fences** — `agent.send()` strips ` ```json ... ``` ` wrapping when the entire response is a single fenced block. Only strips when fence is at start/end (with optional whitespace). Does not strip partial fences or multiple fence blocks.
- **Blocked turns never enter history (split commit)** — CDD/OA-blocked `agent.send()` turns (including catchable DETECT-level blocks) do NOT append to the handle's message history; the tracker/exposure accounting still commits. Quarantine/attest disposition is controlled by `output_admissibility.inadmissible_history`.
- **Scanner regexes must be bounded — an unbounded one crashes instead of deciding.** `SECRET_PATTERNS` in `governance_engine.cpp` carried `-----BEGIN[\s\S]*PRIVATE KEY-----`, whose greedy `[\s\S]*` runs to end of input then backtracks one position at a time; libstdc++ executes that recursively (`_Executor::_M_dfs`), so ~130k frames exhausted the stack and any input with `-----BEGIN` plus roughly 30KB of following text died with SIGSEGV. A crash is worse than a false negative: `code_quality.no_secrets` is HARD, and a crash renders NO verdict — it neither blocks nor passes, and exits with a signal rather than the documented exit 3. Scope, because the call-site list is easy to over-read: `checkSecrets()` returns at its first line unless `code_quality.no_secrets` is enabled, and that key defaults **false** — with it unset the regex never runs and the same 30KB input exits 0. Where it IS enabled the exposure is real and untrusted, since the call sites cover the agent prompt, response content, tool arguments and tool results, so model output could terminate the interpreter. The surrounding `catch (const std::regex_error&)` cannot help — stack exhaustion is not an exception. Fixed by bounding to `[A-Z0-9 ]{0,40}`, which still matches every real PEM header. When adding to `SECRET_PATTERNS`, `DANGEROUS_PATTERNS_DB` or `PLACEHOLDER_PATTERNS_DB`, never use an unbounded `.*` / `[\s\S]*` between two required literals. Test: `tests/security/test_secret_scan_redos.sh` (Group B is the positive control — deleting the pattern also stops the crash).
- **Inline interpreter code in `process.run` gets the polyglot-block checks.** `process.run("python3", ["-c", code])` is a `<<python>>` block by another name, and it used to skip every code check the block and `codegen.run()` get — measured building repo-sentinel: `except Exception: return 0` was a HARD `no_incomplete_logic` block in both and ran silently here, reporting every C++ file as 0 bytes. `inlineInterpreterCode()` in `process_impl.cpp` recognises python/pypy (`-c`), node (`-e`/`-p`/`--eval`/`--print`, also `--eval=`), ruby (`-e`), perl (`-e`/`-E`), php (`-r`) and sh/bash/dash/zsh/ksh (`-c`) — through a directory prefix, `.exe` and a version suffix (`python3.12`), with clustered short flags when the code flag is last (`-Ic`, `-ec`) — and passes the code to `checkPolyglotBlock()`, so `languages.allowed`/`blocked` now apply to it too. **Scope**: a script FILE (`python3 s.py`) is not inline code and is not scanned, and only the OUTER interpreter is recognised (`sh -c "python3 -c ..."` is checked as shell). Test: `tests/security/test_process_run_inline_gate.sh` (PI-00 is the reference that the config exercises the check at all; PI-02/04/05 are the controls a refuse-everything gate would fail).
- **A `process.run` pipeline is governed by BSD and taint, NOT by the agent layer.** Scripts that reach a model through `process.run` instead of `agent.send()` get zero agent governance: no CDD (all 23 signals are fed from `AGENT_RESPONSE`, emitted only by `agent_impl.cpp`), no output admissibility, no step-up challenges, no pulse, no transcript (all 22 hook points are inside `agentSend()`), and no response secret/PII scan. What DOES still apply is taint tracking and behavioural sequence detection — `vm.cpp` emits `PROCESS_EXEC`/`FILE_READ`/`FILE_WRITE`/`ENV_READ`/`ENV_WRITE`/`ENCODE`/`DECODE` from ordinary stdlib calls, gated only on `behavioral_sequences.enabled` (default **false**). Setting `context_drift.enabled` on such a config governs nothing and prints a warning; setting both silences the warning while still scoring nothing. Worked example, with the trap of narrowing taint sinks to unblock a script documented rather than repeated: `examples/hivemind_governed/` (test: `tests/governance_v4/test_hivemind_governed.sh`).
- **A test's OUTPUT CHANNEL is part of the instrument and needs its own control.** A test that runs a tool, captures its stdout and compares against a baseline has two things that can be wrong: the tool's answer, and the path from that answer to the comparison. The second breaks only on platforms nobody runs locally, and it has bitten three times in two days across two sessions (`test_hivemind_governed.sh`, `test_coverage_visibility.sh`, `test_state_field_screen.sh` twice) — never retained from the previous fix. Two shapes. **Encoding**: Python encodes stdout with the LOCALE's encoding, which is cp1252 under MSYS2 — it writes U+2014 as the single byte `0x97` and a UTF-8 reader dies; under a bare C locale the writer dies instead. A traceback announces itself. **Newline**: `print()` translates `\n` to `\r\n` on Windows, and command substitution strips the trailing newline but not the embedded CRs — so the captured string compares unequal to a baseline it renders IDENTICALLY to. Nothing announces itself: `test_state_field_screen.sh` reported `flagged set drifted from the pinned baseline` above two character-for-character identical lists. The fix for that half is making the difference VISIBLE (`sed -n l`), more than any comparison change — an assertion that cannot show why it failed gets believed over the code. **Reproducing the encoding half on Linux, the detail that costs the most time: `LC_ALL=C` does NOT reproduce it** — PEP 538 coerces the C locale back to UTF-8; use `PYTHONUTF8=0 PYTHONCOERCECLOCALE=0 LC_ALL=C`. The newline half reproduces by injecting `sys.stdout.reconfigure(newline="\r\n")`. Prevention in the tool: print ASCII only, write bytes (`sys.stdout.buffer.write`) for anything a test parses, and give every `open()` an explicit `encoding=` AND `errors=`. Controls: `tests/helpers/encoding_controls.sh` (`enc_has_non_ascii`, `enc_strip_cr`, `enc_escaped_diff`, `enc_self_test`) — both failure shapes reproduce on Linux, so this class is testable without a Windows runner. **The renderer itself must not use an external text tool.** `enc_escaped_diff` originally piped through `sed -n l` and reddened build-windows a THIRD time — the assertion that it surfaces a CR passed on Linux and failed under MSYS2. Two hypotheses fit that evidence and they demand different fixes (sed rendering CR as octal `\015`, or sed stripping a trailing CR as part of a CRLF terminator, in which case nothing renders and widening the grep fixes nothing), so the renderer stopped asking: it now escapes via bash parameter expansion, where no line-ending convention gets a vote. Note POSIX *does* list `\r` among `l`'s named escapes, so the octal hypothesis needs a non-conforming sed — which is why `enc_platform_probe` records what the platform actually does on every self-test failure, rather than leaving the next person to re-derive it from a runner they do not have. **A third shape, and it needs no exotic platform: `set -o pipefail` plus `cmd | grep -q`.** `grep -q` exits the instant it matches and closes the pipe; the producer takes SIGPIPE; `pipefail` then makes the pipeline status non-zero. So the pipeline reports FAILURE precisely when the pattern IS present. In a security suite that inverts every verdict in the safe-looking direction — a successful read reads as "refused", and every assertion expecting a refusal passes for free. Capture into a variable and match with `case`, which also leaves you holding the output to print when an assertion fails. **A fourth shape, and it caught the guard rather than the code: handing a PATH to a helper that does not share the shell's path vocabulary.** `test_path_precedence.sh` validated each generated `govern.json` with `python3 -c "...json.load(open('$W/govern.json'))"`. Under MSYS2 `python3` is a NATIVE Windows build and cannot open an MSYS `/tmp/...` path, so the validator meant to catch a broken fixture became a broken probe — it reported `fixture is broken` eleven times while all eleven real assertions passed in the same run. Feed the helper BYTES, not a path (`python3 -c "...json.load(sys.stdin)" < "$file"`): the shell does the open, and no path crosses the boundary. The general rule is the same one that produced the relative-path fixtures two commits earlier — whenever a shell hands a native program a filename, ask which vocabulary that name is in.
- **A polyglot block that cannot run ABORTS the program, so one unusable executor fails every arm after it.** `test_function_capability_gates.sh` put six one-action functions in a single program, one of them a `<<shell>>` block. On the Windows runner the shell executor is unusable, the program died there, and `flushGroupedAdvisories()` (clean-exit only, at the time) never emitted the summary every arm greps — so 11 of 12 arms failed, including the success-expecting controls. The one that passed, FG-11, is the only arm reading the per-occurrence message instead of the summary, and that asymmetry is what identified the cause: a uniform failure with a single odd survivor points at the shared output path, not at the subject. Two fixes, and both are the general lesson: put anything platform-dependent in its OWN program so it cannot take unrelated assertions down with it, and give it a **viability probe asked with the action GRANTED** — so the only thing that can fail is the executor — reporting UNMEASURABLE rather than a governance FAILURE. Verified by forcing the probe to fail: 10 pass, 2 skip, 0 fail. Note the first attempt to simulate that was itself a broken probe: copying the suite to `/tmp` made `SCRIPT_DIR/../../build/naab-lang` resolve nowhere and every arm returned `rc=127`; a suite must be simulated from its own directory.
Expand Down
66 changes: 66 additions & 0 deletions include/naab/string_ops.h
Original file line number Diff line number Diff line change
@@ -0,0 +1,66 @@
// String operations shared by BOTH engines' method forms (`s.slice(...)`) and
// the string module (`string.slice(s, ...)`).
//
// Each existed as three hand-written copies -- VM callBuiltinMethod, two
// tree-walker method paths in call_dispatch.cpp, and the module -- and the
// copies disagreed. Measured on "hello" / "a-b-c":
//
// s.replace("-", "+") VM "a+b-c" (first only) tree-walker "a+b+c"
// s.substring(3, 1) VM "" tree-walker "lo"
// s.slice(-3) VM "llo" tree-walker "hello"
//
// The module functions agreed with each other in every case, so they are the
// reference: replace replaces every occurrence, substring clamps and yields ""
// for an empty or reversed range, slice follows JavaScript (negative indices
// count from the end, reversed range -> ""). Positions are byte offsets, as
// they always were.

#pragma once

#include <string>

namespace naab {
namespace strops {

// substring(start[, end]): start < 0 -> 0; end clamped to the length;
// end <= start (including a negative end) -> "".
inline std::string substring(const std::string& s, int start, bool has_end, int end) {
int len = static_cast<int>(s.size());
if (start < 0) start = 0;
if (start >= len) return "";
if (!has_end) return s.substr(static_cast<size_t>(start));
if (end > len) end = len;
if (end <= start) return "";
return s.substr(static_cast<size_t>(start), static_cast<size_t>(end - start));
}

// slice(start[, end]), JavaScript semantics: a negative index counts from the
// end; out-of-range indices clamp; an empty or reversed range yields "".
inline std::string slice(const std::string& s, int start, bool has_end, int end) {
int len = static_cast<int>(s.size());
if (start < 0) start += len;
if (start < 0) start = 0;
if (start > len) start = len;
if (!has_end) end = len;
if (end < 0) end += len;
if (end < 0) end = 0;
if (end > len) end = len;
if (start >= end) return "";
return s.substr(static_cast<size_t>(start), static_cast<size_t>(end - start));
}

// replace(from, to): every occurrence. An empty `from` returns s unchanged
// (replacing "" everywhere would never terminate).
inline std::string replaceAll(const std::string& s, const std::string& from, const std::string& to) {
if (from.empty()) return s;
std::string out = s;
size_t pos = 0;
while ((pos = out.find(from, pos)) != std::string::npos) {
out.replace(pos, from.size(), to);
pos += to.size();
}
return out;
}

} // namespace strops
} // namespace naab
17 changes: 17 additions & 0 deletions run-all-tests.sh
Original file line number Diff line number Diff line change
Expand Up @@ -1649,6 +1649,23 @@ else
echo " test_path_policy_reach.sh: not found, skipping"
fi

echo ""
echo "═══════════════════════════════════════════════════════════"
echo " process.run Inline Code Gate (python -c, sh -c, ...)"
echo "═══════════════════════════════════════════════════════════"
echo ""
PROC_INLINE_SCRIPT="tests/security/test_process_run_inline_gate.sh"
if [ -f "$PROC_INLINE_SCRIPT" ]; then
if run_shell_test "$PROC_INLINE_SCRIPT" 2>&1; then
echo " test_process_run_inline_gate.sh: ALL PASSED"
else
FAILED=$((FAILED + 1))
FAILED_TESTS+=("test_process_run_inline_gate.sh")
fi
else
echo " test_process_run_inline_gate.sh: not found, skipping"
fi

# --- Signed govern.json vs Package Operations (F39) ---
echo ""
echo "═══════════════════════════════════════════════════════════"
Expand Down
Loading
Loading