Make --timeout reach Python, http and nested blocks; run unit tests in CI - #259
Merged
b-macker merged 4 commits intoSep 28, 2026
Merged
Conversation
…n CI
--timeout is a flag the interpreter polls. Three ways a script outran it,
all measured with --timeout 3:
- A <<python>> busy loop ran to completion (20.1s). Embedded CPython had
no interrupt path. The timer thread now queues a pending call that raises
TimeoutError and re-queues itself while the timeout stands, so a broad
except cannot swallow it. On CPython 3.11 queuing alone did nothing
(Py_AddPendingCall from a non-main thread does not flag the eval loop);
a brief GIL acquire makes the running thread recompute it. Not on
Android (PyGILState_Ensure on a foreign thread is the CFI crash).
- One <<javascript>> expression, or codegen.run, removed --timeout for the
rest of the script: every executor's ScopedTimeout re-armed the single
per-thread timer and cleared it on exit, and the cancel counter was
process-wide, so a scope on another thread cancelled the main thread's
timer (and one REST request could cancel another's). ScopedTimeout now
nests: it only tightens the deadline in force and restores it on exit.
0 means "no limit of its own" -- it armed a zero-second timer that
killed every subprocess under the UNRESTRICTED preset.
- http.get(url, {}, 0) to a SYN-dropping host was still connecting at 20s.
A curl progress callback now aborts on timeout, and timeout_ms <= 0
falls back to the 30s default.
Known limits, recorded not fixed: Python blocked in C (time.sleep) and
Python on worker threads are not interrupted.
tests/security/test_timeout_reach.sh: 9 arms. On the old code the control
passes and the other 8 fail.
naab_unit_tests now runs in CI via tests/unit/run_unit_tests.sh. 559 pass;
30 known failures are listed with their reasons in known_failures.txt. The
runner fails on an unlisted failure, on a listed test that does not exist,
and on a listed test that starts passing, so the list has to shrink as
fixes land. All three failure modes were checked.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELUfjXZvx8kzXo1UJjrAhC
NAAb Governance Report
All governance checks passed! Generated by NAAb Governance Engine v4.0 |
build-windows: H-01 stopped at 10s, not 3s. The progress-callback abort was not reached during connect on the Windows curl, so the request ran to curl's 10s connect timeout. Capping CURLOPT_TIMEOUT_MS (and through it the connect timeout) at the time left before the ScopedTimeout deadline is exact on every platform; the callback stays as the second line. With the callback disabled locally, the cap alone stops H-01 at 3s. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ELUfjXZvx8kzXo1UJjrAhC
…he sandbox
From building repo-sentinel in NAAb with a governed agent pipeline:
- dict.get() miss hints (F-002): a counting loop over 50 file paths
printed 49 hint lines, most a "did you mean" pointing at a sibling path.
The three copies of the hint code are now one helper that stops after 3
hints per run and then says how to mark an expected miss (a default
argument or d.has()), which already prints nothing.
- Ternary hint (F-004): it existed but fired only when '?' started an
expression. In a dict literal, parentheses or a call argument, expect()
reported "Expected '}'" (plus a missing-brace diagnosis) or "Expected
')'" instead. expect() now gives the if-expression hint for any stray '?'.
- Node under the sandbox (found checking F-005, whose argv report is node's
own `--` handling, not NAAb): children got RLIMIT_AS at the memory
budget, and node 22 needs 512-768 MB of address space just to start, so
process.run("node", ...) died with "Failed to reserve virtual memory"
under `elevated`. The budget now applies to RLIMIT_DATA, with an
RLIMIT_AS ceiling (4x, at least 2 GB) as the backstop for shared
anonymous memory, which RLIMIT_DATA does not count.
tests/parser/test_dogfood_hints.sh (10 arms: old build fails 5, its 5
controls pass) and tests/security/test_child_memory_limit.sh (4 arms: old
build fails M-01 only -- M-02/M-03 assert the limit still bounds private
and shared memory, M-04 is their control). Both pass on the new build.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ELUfjXZvx8kzXo1UJjrAhC
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ELUfjXZvx8kzXo1UJjrAhC
Owner
Author
|
Local results on the head commit
Since the PR was opened, two commits were added from the repo-sentinel dogfooding findings:
On an earlier local run, Generated by Claude Code |
b-macker
marked this pull request as ready for review
September 28, 2026 01:57
b-macker
deleted the
claude/naab-inadmissible-action-prevention-4cmn1m
branch
September 28, 2026 01:57
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
--timeoutis a flag the interpreter checks between steps. Work that runs inside the process without checking it only noticed the timeout after it finished on its own. This PR fixes three ways a script could run past--timeout. All three were measured with--timeout 3.<<python>>busy loop (20 s)try: ... except Exception: passwhile trueafter one<<javascript>>expressionwhile trueaftercodegen.run(...)http.get(url, {}, 0)to a host that never answersThe JavaScript/
codegencase was found while fixing the Python one. It is the most serious of the three: any script that ran a single JavaScript block had no time limit for the rest of the run.This PR also runs the GoogleTest unit tests in CI, which was the rest of item 8.
Changes
resource_limits.h/.cpp):ScopedTimeout, and the JS, shell, subprocess and C++ executors wrap each block in another one.ScopedTimeoutnow records the deadline already in force. It can only shorten it, never extend it, and restores it on exit.0now means "no limit of its own". Before, it set a zero-second timer that killed every subprocess as soon as it started under theUNRESTRICTEDsandbox preset (defect Add tab completion to REPL #4 indocs/unit-test-findings.md).codegen.runset and cleared the timeout by hand; it now usesScopedTimeout.python_c_wrapper.c,python_interpreter_manager.cpp):Py_AddPendingCall) that raisesTimeoutErrorin the running code.http_impl.cpp):timeout_msof 0 or less (libcurl's "never time out") now uses the 30 s default.ci.yml,tests/unit/run_unit_tests.sh,tests/unit/known_failures.txt):docs/unit-test-findings.md.AsyncCallbackPooldeadlocks and use-after-free) are not run at all.docs/unit-test-findings.md(items 2a, 2c, 2d, and the CI section) and a CLAUDE.md note on nesting timeouts.Known limits, recorded but not fixed:
time.sleepisn't interrupted, because CPython only runs the callback between bytecodes.time.sleep(15)took 15 s on both the old and new code.Test Plan
tests/security/test_timeout_reach.sh, 9 checks, all passing, measured on elapsed time rather than error text:codegen.runcall must still be stopped.run_unit_tests.shreports PASS locally in about 4 s. I checked all three ways it can fail by editing the list: dropping an entry, adding a test that passes, and adding a typo. Each made it fail.🤖 Generated with Claude Code
https://claude.ai/code/session_01ELUfjXZvx8kzXo1UJjrAhC
Generated by Claude Code