Conversation
- The agent can no longer reach its own approvals through the control center:
requests are refused (403 [AGENT_BROWSER]) while the agent's browser has a tab
on this control center (matched by Host / any local or tunnel address of the
port); such tabs are sent back to about:blank in a QodeX-launched browser, and
control-center login cookies (qx_ctl*, any port) are swept out of the agent's
profile after each ?k= login, on eviction and every 20s.
- Live-view navigate refuses this control center and any QodeX control link
(local/LAN/tunnel URL with a k= token): opening one in the agent's browser
planted the login cookie there ([CONTROL_CENTER_URL]).
- `qodex control` ignored Ctrl+C in the real CLI: src/index.ts imports the tool
registry, whose process-registry adds a non-exiting SIGINT listener. The
command now installs (from the start) an always-exiting SIGINT handler with a
bounded graceful stop; a second Ctrl+C exits at once.
- ?k= login with a percent-encoded key (%6B=) kept the token in the bounce target
and bounced forever; stripTokenFromUrl now compares decoded keys.
- maskSecrets masks ?k=/&k= control tokens (e.g. mission live URLs) in text
published to viewers/channels.
- A control-center takeover is handed back to the agent after 5 min with no
dashboard connected (takeoverReleaseMs; 0 disables), with a bus notice —
unattended runs no longer wait forever behind a closed tab.
- Control action errors map to proper statuses: *_NOT_FOUND 404,
BAD_REQUEST/INVALID_* /*_BAD_* 400, *_NOT_PENDING/CONFLICT 409 (was always 500).
- missions-bridge: concurrent registerMissionControl calls raced into two event
bridges (every mission event published twice) — now one shared setup; disposers
are idempotent (a double release dropped other holders' actions);
missions.resolveApproval threw nothing on failure (HTTP 200 + {ok:false}, the
dashboard dropped the card) — now throws [APPROVAL_NOT_FOUND]/[APPROVAL_BAD_ANSWER]/
[APPROVAL_NOT_PENDING]; a missions DB that cannot open no longer fails
registration (and with it /control).
- LAN rebind: explicitly-undefined options no longer wipe kept settings.
- Dashboard: emoji, AltGr and macOS Option characters are typed as text (were sent
as unknown key combos); 404/409 on a mission approval removes the card from the
right list; the URL bar follows only the ACTIVE tab (Module A publishes
'navigated' for every tab).
- Tests: real-Chromium end-to-end with the real QodexBrowserManager (takeover-only
launch via navigate, CDP screencast frames, frame-coordinate clicks, typing and
key names, frame.jpg, hand-back resumes waiters, dashboard driving the agent
browser, close/relaunch, agent self-approval attempt), subprocess Ctrl+C test,
missions-bridge tests, and server regressions for each fix.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Desktop control (src/tools/computer/**), found by review and verified
against real Chromium under Xvfb where possible:
- Clipboard paste is now abort-safe. pasteText restores the previous
clipboard in a finally, outside the aborted signal, and clears it when
the old contents couldn't be read. Before, cancelling mid-paste left the
typed text (possibly a password) on the clipboard (reproduced under Xvfb
with xclip). On Windows, a killed paste script now clears the clipboard
if it still holds our text.
- X11 'auto' typing pastes non-ASCII text. `xdotool type` of Persian
types wrong letters but still exits 0 ("سلسم سنیا" for "سلام دنیا",
3 of 4 runs, at any --delay), so the old fallback-on-failure never
fired. An aborted `xdotool type` no longer falls through to the
clipboard.
- X11/Wayland typed text now goes over stdin (`--file -`), not argv, so
it no longer shows in ps or /proc/*/cmdline. Older tools fall back to
argv.
- CRLF is normalized before typing (xdotool pressed Return twice per
"\r\n").
- Window titles are attacker-controlled (a browser window's title is the
page <title>). Output Sentinel does not fence (screenshot result and
notes, out-of-bounds coordinate errors, screen_info, WINDOW_NOT_FOUND
lists, backend errors) now only shows safeTitle(): truncated, invisible
and bidi characters stripped, and withheld if scanInjection flags it.
- computer_use_agent passes ctx.askUser to the sub-agent runner, like
`task` does, so its approvals reach the caller's human or channel.
- computer_use_open: coerceArgs turns host:port targets
("evil.example:8443/x", "localhost:3000") into the http:// URLs the
backends open, so Sentinel's blocked, allowed and private-network
domain policy applies. Before, they were reviewed as a local app. The
tool also refuses QodeX's own secrets (vault, vault key, ~/.qodex/.env,
browser profiles; file: URLs and symlinks too), and screenshot paths
can't point into them.
- computer_use_click takes an optional `element`, auto-filled in
coerceArgs from the last computer_use_locate hit at that point, so the
review and the activity log can say what is clicked.
- computer_use_locate stops waiting on the vision model when the call is
cancelled (vision_analyze ignores the signal).
- The locate parser accepts Qwen2.5-VL bbox_2d / point_2d and Qwen2-VL
<|box_start|> tokens (qwen2.5vl is the recommended local vision model).
- computer_use_scroll with only one of x/y is an error instead of
scrolling under the pointer.
- The default screenshots dir is created and tightened to 0700.
Tests: new test/desktop-hardening.test.ts (23 cases); desktop-backends
updated for stdin typing and auto-paste. tsc clean; full suite
158 files / 1916 tests pass.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Approval bypasses - Foreground `mission start/resume --foreground --yes` on a TTY marked the worker as interactive, so Sentinel routed CRITICAL actions (purchase, payment, send, credential) through the step's askUser, which auto-answers "yes" in --yes mode: purchases were auto-approved. 'auto' mode no longer counts as a human at askUser; the terminal becomes an approval channel (TerminalApprovalChannel) racing the mission queue/control/Telegram. - An agent inside a mission could approve its own requests or inject "user" steering notes via its shell (`qodex mission approve|steer`): refused when QODEX_MISSION_ID is inherited ([MISSION_APPROVAL_FORBIDDEN] / [MISSION_STEER_FORBIDDEN]). - The live control-center URL (token = approval authority) was shown to the model by mission_status/inline mission_start and printed into the agent-readable worker log: token redacted there (CLI status for the human still shows it, except from inside a mission). - mission_start ran no permission check at all (a prompt injection could start detached autonomous work): now goes through PermissionEngine/askUser like the equivalent shell command (allow / autoReject / ask). Worker pid ownership - runMission never released its pid: after an inline (TUI) mission paused or failed, `mission cancel` from another terminal/Telegram SIGTERMed the TUI, and a cancel from the same process left it "Cancelling…" forever. Pids are now released at finalize, only ACTIVE missions are signalled, and shared inline hosts (worker_shared) are never signalled. - PID reuse after a crash/reboot made a stranger process look like the worker (stuck "running", SIGTERM to an unrelated process). The worker's start identity (/proc/<pid>/stat starttime, column pid_start) is recorded and checked (isWorkerAlive). - Racing resumes could start two workers (spawn overwrote the pid; the worker check-then-set was not atomic): claimWorker() is a compare-and-set; a losing spawn reports [MISSION_BUSY] and a losing worker exits 3 without failing the mission. Untrusted output / secrets - Step results were pasted raw into the next step's prompt and the report prompt (injection propagation with user authority): fenced as data with an injection banner. mission_status is untrustedOutput; inline mission_start fences the mission text. - Tool-result excerpts, notices, step/report excerpts, milestones, approval prompts (incl. Telegram/control rows), the --yes audit trail, worklog and notifications are secret-masked (maskSecrets); step results/report keep the real values. Correctness - `qodex mission start/status/list/resume` lost --yes, --model and --json to the root program's identically named global options (commander), so `--yes` and scheduled routines silently ran in 'ask' mode on the default model and --json printed text. Subcommands now read optsWithGlobals. - A deleted/expired DB approval whose options have no 'n*' answer left the tool waiting forever: the poller withdraws it with broker.cancel(). - Approving a pending approval of a dead worker reported success into the void: answer paths reconcile first (approval expires, mission paused). - Planner capping folded overflow steps into the last step but dropped their dependencies (folded work could run before what it needed). - Schedule id prefix lookup used unescaped LIKE (`schedule rm %%%%` matched an arbitrary entry). - Mission worker and schedule run logs (goal, output, report) are 0600 in a 0700 dir. Tests: test/missions-review.test.ts (22 tests, all failing before the fix), plus a real-process E2E (detached worker against a hanging local model, duplicate worker refused, SIGTERM cancel, resume, schedule tick). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
- '# Your Computer' is for the top-level agent only; sub-agents get a focused role brief (subagent-compressed prompt 7.1k -> 6.4k of 6.5k) - shorter shared browser param descriptions (ref/selector/element/ timeout/snapshot are repeated across ~20 tools) - toolTokensNormal covers the five new relevance-gated tool families (measured ~44k over 163 tools, same ~13% headroom as before); the per-turn gated budget is unchanged and still passes Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Guard bypasses and integrity gaps found after integration, each with a failing test first (test/sentinel-review, test/sentinel-browser on real Chromium via ToolRegistry, test/vault-frames on real Chromium): - afterTool trusted a text prefix to decide "already fenced": page text starting with "<untrusted_content ...>" (browser_get_text, evaluate, clipboard) skipped both the injection scan and the fence. Now decided by result metadata only; banner excerpts can no longer carry markup. - web_fetch / web_search / remote http_request and MCP tool output is now scanned and fenced too (localhost / LAN dev servers stay unfenced). - browser_fill_secret judged a field's frame by its URL: an about:blank child of a cross-origin ad iframe (or a sandboxed srcdoc frame) "was" the page and received the password (reproduced in Chromium). Now uses the document's origin (window.origin); opaque/unknown about: frames refused. - Vault: tenants on shared hosting (me.github.io, bucket.s3.amazonaws.com) now match exactly; "evil.bucket.s3.amazonaws.com" is someone else's. - A scripted click routed around the purchase guard: browser_evaluate ".click()/requestSubmit()/dispatchEvent" and javascript: URLs (which run in the page) are now judged like a click on what they select (selector words + the described target elements; checkout/gateway pages). Space on a focused button is an activation (press Space on "Place order"). - browser_fill_form fields addressed by selector were never described or classified (password by selector slipped through); target-less browser_type now describes the focused element; selector words are used when an element can't be described. - Self-approval: shell `qodex mission approve|deny`, `qodex control`, `qodex telegram pair|setup|unpair`, `mission start --yes` (also typed via computer_use_type / clipboard) need a human; any use of the in-process control center (port, token, tunnel) is hard-blocked; ?k= control tokens are masked in every tool result and in the audit log (mission_status exposed the worker's live URL with its token). - Approval trust stores are write-protected like config.yaml: ~/.qodex/.env, sessions.db (SQL writes to mission_approvals), channels/ (Telegram pairing) and sentinel/ (audit). Integrity rules ignore autoApprove. - tar / zip / rsync / cp -r / grep -r / s3_sync of ~/.qodex as a whole are blocked like the vault key and profiles themselves. - workflow_run review: matched the real workflow format (upload `files`, tab-new URLs, param defaults, vault-backed secret params = fixed high) and found workflows by normalizeWorkflowName (Persian/spaced names were "contents unknown"). - browser_dialog accept reads the pending dialog text (confirm "Delete your account?" / "Confirm purchase?"). - Unreadable Sentinel config falls back to the defaults (was: allow all); the audit trail recreates its directory if it is removed mid-session. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Replay (src/workflows/replay.ts)
- Sentinel/vault saw a different element than replay acted on: text/label
matches (and the recorded selector) pick the first VISIBLE match, but the
guard got the bare selector and describeSelector() resolves .first() -- a
hidden look-alike first laundered a payment click (reproduced on real
Chromium with the real Sentinel). The selector handed to the guard and the
vault is now pinned with `>> nth=i`; records keep the reusable selector.
- Recorded selectors are resolved with page.locator (the manager's locator
is .first(), hiding extra / hidden matches); a precise candidate that now
matches several visible elements yields to one that matches exactly one.
- Fail closed: without a ToolContext the guard was skipped entirely; now a
deny-by-default context is used. The default guard is Sentinel (static
import); if it can't be obtained every consequential step is refused.
- Never act after cancel / during a human takeover: abort + takeover are
re-checked right before every page action (an approval can take minutes).
- Secret params no longer leak: composite values ("{{user}}:{{pass}}") are
treated as secret, secrets in names/URLs are scrubbed from guard args,
action records and reports (incl. URL-encoded forms), and a secret param in
a navigation URL is refused.
- start_step past the last step replays nothing instead of being clamped back
onto the last step (the resume hint for a hand-finished last step would
repeat e.g. "Place order").
- Uploads: QodeX-home check also against the symlink-resolved home.
- Final URL moved inside the untrusted-data fence; title timer cleared.
Recorder (src/workflows/recorder.ts)
- The capture nonce leaked through window property names
(__qxRecInstalled_<nonce>): any page could list them and forge steps
(reproduced: a page injected a "Delete account" click). Names now use a
one-way tag; the binding is never re-read from the page global.
- Only trusted input/change events are captured (pages can't author fills by
dispatching synthetic events); input from cross-origin iframes is ignored.
- Enter in a text field recorded fill+press+fill+click (change event and the
implicit-submission click); now one fill + one press.
- Mixed recordings: echo matching rewritten (value-aware, echoes of instant
actions must precede the agent's record, submit:true also echoes Enter,
Enter on the focused field matched by key) and navigations are attributed
to the agent action by its echo time -- previously the wrong capture was
dropped (the human's own input lost) and Enter was duplicated.
- Typed values are parameterized into later URLs only as whole query values /
path segments, never into the host ("shop" on shop.example made the host a
param, so replay navigated to another site).
- A slow capture install that outlived its recording left the binding/init
script in the context forever; now disposed (generation-checked).
- Auto-detected vault fills (no ref/selector) keep a replayable step.
Skills / tools
- Generated SKILL.md (trusted, auto-injected instructions) carried
page-derived labels and the model's description unscanned: workflows whose
text trips Sentinel's / the skill installer's injection scanners get no
skill; step labels are fenced as data with backticks neutralized.
- workflow_show output is untrustedOutput; workflow_run accepts params: null;
skill allowed-tools cover every step kind.
Tests: test/workflows-hardening.test.ts (unit regressions),
test/workflows-e2e.test.ts (real QodexBrowserManager + browser_* tools +
real Sentinel on Chromium), chromium forge/synthetic-event test.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sub-agent isolation (fresh child loop per run), cancellation-bound askUser, run-mode enforcement at execution, wall budget that only fires on stalls, real-browser core tests. Tests aligned with the headless fail-safe policy. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
runSubagent now runs each child on a fresh AgentLoop, so the parent-instance stub of buildInitialMessages no longer reached the child. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
…ping) - control center: an unexpected request error no longer echoes its text to the client (logged only) - mission log tail reads size and data through one descriptor - scheduler lock steal renames atomically and puts back a fresh lock it moved by mistake (two ticks stealing at once could both run) - Windows browser-profile probe opens the lockfile directly - Telegram plain-text fallback strips tags until stable Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
… initial values) - scheduler lock: age, inode and holder pid read through one descriptor - mission log: no exists-check before the open - Telegram plain-text fallback drops stray angle brackets before decoding - control center never stringifies a thrown non-Error into a response - mission_start confirmation: no dead initial values Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
| let stat: fsSync.Stats | null = null; | ||
| let holder: number | null = null; | ||
| const existing = await fs.open(lockPath, 'r').catch(() => null); | ||
| if (!existing) return fs.open(lockPath, 'wx').catch(() => null); // released meanwhile |
Owner
Author
There was a problem hiding this comment.
This is the same lock path as the thread above. The open flagged here is open(lockPath, 'wx'), an exclusive create, used when the lock disappears between the first EEXIST and the inspection. It can't be raced into a double hold, so I'm leaving it as is.
Generated by Claude Code
…branch Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
…fixes) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ept) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rigin Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- R1: process-registry / live-registry no longer install module-load SIGINT
listeners that never exit (importing the tool registry disabled Node's
default Ctrl+C exit for every subcommand); cleanup stays on 'exit'.
qodex control keeps its always-exiting graceful-stop handler (comment
fixed); qodex telegram start now always installs its own exiting handler
(graceful stop confirms the update offset) instead of only when another
listener existed. telegram-command test updated for that.
- R2: dev_server_* schemas put .describe() before .optional() so the
descriptions reach the JSON schema; dev_server_start.env is a
[{key, value}] array (z.record was advertised as a string), with
coerceArgs still accepting a {NAME: value} object or its JSON string.
- R3: ToolRegistry redacts args (redactForAudit, hideTyped) in the
'Executing tool' debug log and in the ARGUMENT_VALIDATION_ERROR echo, so
typed text / password values / bodies never reach qodex.log or the
transcript verbatim.
- R4: remember refuses facts with high-severity Sentinel injection findings
([MEMORY_REFUSED], nothing stored) — same bar as the loop's load filter.
- R5: gather / fanout / orchestrate declare timeoutSeconds = 2400 like task,
so they aren't killed at the global tool timeout and their run time is
excused from the parent's wall budget.
- R6: the agent loop's cancel listener settles the tool race for every
tool, not only no-timeout ones: Ctrl+C no longer waits on a tool that
ignores ctx.signal; the timeout rejects before aborting.
- R7: mission_status / mission_list and the background_job / dev_server
status pollers are state-dependent (repeat guard keyed by result hash).
- R8: TUI tool summaries lift Sentinel's injection banner off the top of a
fenced result (shown clipped after the tool's own lines); the fence
itself was already stripped for display.
- R9: telegram.apiBase, telegram.botTokenEnv and control.host are honored
only from ~/.qodex/config.yaml; a project .qodex/config.yaml setting them
is ignored with a warning.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
…6), no unsafe-inline Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…omium; stale login error fix Fake manager: http (non-localhost) page and wrong-site rotation refused, failed typing rolls back a new entry / restores the previous password, value scrubbed. Real Chromium on local pages: sign-up fills password + confirm and saves the entry, change-password rotates into new + confirm and keeps the old one, maxlength respected, login-only form refused; browser_login classic, identifier-first and TOTP flows, wrong url / redirect to another site refused with nothing filled, one failed attempt then LOGIN_HALTED until the entry changes, a payment-like submit button refused, submit:false. Every test checks that no password, seed or code reaches a result, progress event, action record or the console. Fix: an error message already on the page before submitting (a script-driven form after an earlier attempt) was read as this attempt failing; it now only counts when the form is still there at the deadline. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Mount /api/secrets and /api/vault (full token only; loopback or https; sealed ECDH+AES-GCM over tunnels; plain-http LAN refused) and register the control center as a secret surface while it runs. Dashboard: a request card that seals the typed login in the page, and a vault panel (masked list, add, change password, edit sites/login URL, confirmed delete). Real-Chromium tests cover local and tunnel sealing and assert no secret reaches responses, bus or approvals. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
… target field The guard only gathered the page URL and the target element for browser_* tools, so the approval prompt for vault_generate_and_fill named neither the site nor the field. It now gets both (ref / selector described like browser_fill_secret). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
… the running shell, unattended auto-ask timeout, project instructions first|all Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen # Conflicts: # src/cli/ui.tsx # src/sentinel/guard.ts
….1, model aliases, /checkup prompt-audit, build-eval / hillclimb skills Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen # Conflicts: # src/index.ts
…ault_generate_and_fill A successful vault_generate_and_fill (and browser_login, explicitly) is evidence of a real-world action for the completion gate; a failed login is not. The recorder turns browser_login's auto-detected vault fills into steps on the field's autocomplete token (username / email / current-password / one-time-code) instead of skipping the username step. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
The chat got 'unpaired' while the approval channel was still registered, so an approval raised in that window could still be routed to it (the pairing test caught the gap under load). Same order as /pair: channel state first, then the confirmation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
…ogin and vault_request_login
Tool relevance: a vault/login family (vault_*, browser_login,
browser_fill_secret) keyed on the user's own credentials and 2FA wording,
EN + FA ('log me in with my saved password', 'رمز عبورم', 'گاوصندوق', 'کد دو
مرحله') — never on a bare 'password', so coding tasks about password hashing
stay lean. Budgets unchanged: tool definitions 48143 / 50000 (173 tools),
compressed prompt 6326 / 6500.
Prompts: the Credentials line, the web-task addendum, the browser role and
browser_agent now lead with browser_login (one field: browser_fill_secret;
sign-up: vault_generate_and_fill). vault_request_login is named only when the
tool is registered (system prompt) or as 'if you have it'. The browser
sub-agent may use vault_generate_and_fill / vault_request_login. /vault lists
the edit / rotate / import commands.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
session.ts gains additive human-input observers (before a human click / Enter under takeover, main-frame navigation, takeover switch). src/vault/capture.ts reads the submitted login host-side from the owning frame (iframes followed by focus / hit point, origin checked against the frame URL), keeps it in memory only, and after the next navigation leaves the login form (or at takeover end) asks 'Save the login for <host> (user <masked>)?' via ApprovalBroker — terminal and control center — then adds or merge-patches the vault entry. Two-step logins keep the step-1 username; an identical saved login is not asked again. Installed by the control center and the TUI (local asker). Real-Chromium tests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
…d full run Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
… wait is not missed IDLE only reports EXISTS for mail that arrives after the SELECT, so a message landing between the watcher's last check and the start of the wait sat until the next wake-up (up to the IDLE timeout, 25 min by default). The watcher now passes the next UID it has not seen; the IMAP transport compares it with UIDNEXT right after the SELECT (the fake does the same) and returns at once. Regression test delivers a message exactly in that window. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
…-gated project mods, loop integration, qodex mod CLI Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen # Conflicts: # src/cli/modes/headless.ts # src/cli/slash-catalog.ts # src/index.ts
…UI, built-in context-bar, you-should-know, sample-hello, /mod new, docs/MODS.md Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen # Conflicts: # src/cli/slash-catalog.ts # src/cli/ui.tsx
…ion, CSV import, vault_generate_and_fill, browser_login Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen # Conflicts: # src/sentinel/policy.ts
…rompt and a sealed control-center form, save-login capture, vault panel Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen # Conflicts: # src/agent/tool-relevance.ts # src/control/dashboard.ts # src/control/server.ts # src/sentinel/policy.ts # src/vault/tools.ts
…watcher loop tests wait long enough under load A page opened from a scoped hand-off link loaded the password panel script, which polled /api/secrets (refused for scoped links, but the page should not ask). In hand-off mode the panels stay hidden and nothing polls. The watcher loop tests now wait up to 10 s (still far below the 60 s IDLE timeout they prove is not hit). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
…ringifying the promise
String(page.title()) gave '[object Promise]' — the title signal ('Just a moment…') was never
read when the in-page probe failed — and left the promise's rejection unhandled when the page
closed (seen as an unhandled rejection in the full suite).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
…eady aborted signal A stop() landing while the watcher updated its state (or while the IMAP IDLE connection was being set up) aborted the signal before the wait added its abort listener; that listener never fires, so the wait ran to its timeout (60 s in tests, 25 min against a real server) and stop() hung. The loop now re-checks the signal right before waiting, and the IMAP transport checks it after connecting. Deterministic regression test. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
…arity (README, CHANGELOG, guides, Persian summary) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
…, vault, CAPTCHA hand-off and mods The previous docs commit only carried the MAIL.md move and the vault/CAPTCHA guide; this adds the README sections and slash-command list, the CHANGELOG entry, the standing-grants/watcher/rules section of docs/MAIL.md and the Persian bullets in docs/AGENT_PLATFORM.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
…eeds the pixels browser.lean (auto | on | off, default auto = only a headless browser QodeX launched; QODEX_BROWSER_LEAN=1|0 forces it). One CDP session per tab enables the Fetch domain for the Image / Font / Media resource types only and fails them with BlockedByClient: the DOM, scripts, styles, XHR, forms, cookies and the HTTP cache are untouched (context.route() turns the cache off for every request: a cacheable script was fetched 3x over 3 loads with it, once with this). Never on the user's own Chrome (cdpUrl) or on loopback / LAN / non-http pages (dev servers). It switches itself off for the rest of the session once pixels matter: browser_screenshot, browser_pdf (both say what to reload), the live view, a takeover or a bot check. browser_status shows it; the request log names it and the console drops the per-image "Failed to load resource" noise. 12-photo page, headless: ~0 MB vs 5 MB downloaded, ~270 ms vs ~720 ms for launch + load, ~100-180 MB less Chromium RSS. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
…temp file, file races, dead /mod case) CI "Test" was red on the last three heads: test/vault-entry-tui.test.ts read Ink's output, which Ink keeps to itself until unmount when CI is set (is-in-ci). The test now renders with debug: true like the other Ink tests. That exposed a real race in the prompt itself: Ink attaches key handlers in a deferred effect, so a key typed right after Enter could reach the previous field's handler — a fast "pw⏎pw⏎" became a false "passwords do not match". The form now routes every key through refs for the current field and text, and a pasted value ending in Enter moves on. Two new tests fail on the old component (type-ahead, paste) and pass now; 8/8 runs under CI=true. CodeQL: - wayland: measure the screen through a private mkdtemp dir, not a predictable name in the shared temp dir (symlink pre-creation). - mods $.fs.read and `qodex vault import`: size check and read through one file handle (the file checked is the file read). - recorder: the flush expression uses the hex-checked tag as a plain identifier instead of JSON-in-code. - remove the unreachable second `case 'mod'` (the tested modsmith one stays), unused vault/tools imports, a dead initial value in the mail-grant check, a redundant `!pendingPrompt`, and clear typed secrets in place. - tests: escape the reflected fixture value, one-handle reads, plain string matches instead of unanchored host regexes, explicit regex grouping. Full suite under CI=true: 344 files, 4463 tests pass. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
…bstring check in page code)
CodeQL read `textContent.includes('shop.example.com')` as URL sanitization.
The test now waits for a non-empty row and asserts the host and the masked
username with toContain.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
…identity (opt-in)
An honest alternative to detection evasion: with browser.botAuth on, QodeX
signs its own requests with an Ed25519 key so a site can recognise the agent
("this is QodeX, acting for its user") and choose to let it through. It never
hides that the browser is automated — no fingerprint spoofing, no hiding
navigator.webdriver, no synthetic human input.
- src/tools/browser/bot-auth.ts (pure): Ed25519 key management, RFC 7638 JWK
thumbprint key id, RFC 9421 HTTP Message Signatures with the web-bot-auth tag
over @authority + signature-agent, cached per authority, and the public-key
directory (JWK Set) to publish. Unit tests verify the signature with the
public key and reject a tampered authority.
- session.ts setupBotAuth(): a context.route signs same-site document / xhr /
fetch on public hosts only; launched browser only (never the user's own
Chrome over CDP); loopback / LAN pages are exempt (reuses the lean predicate).
status().botAuth + botAuthSigner().
- config browser.botAuth (off by default; true/false/on/off/object;
QODEX_BROWSER_BOT_AUTH=1|0); `qodex browser bot-auth --init|--directory` and a
line in `qodex browser status` / browser_status.
- The private key lives in ~/.qodex/browser/bot-auth/ (0600); Sentinel adds it
to the protected paths so the agent can never read or change it.
- Real-Chromium test: the server receives a valid signature on the document and
the fetch (not the image), the public key verifies it, and loopback stays
unsigned by default.
Docs: VAULT_AND_CAPTCHA.md (EN + FA), README, CHANGELOG, AGENT_PLATFORM.md (FA).
Full suite under CI=true: 345 files, 4477 tests pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This turns QodeX from a coding CLI into a general autonomous agent that is local-first and model-agnostic, with Persian supported throughout. Every module was built on its own branch, integrated, tested end to end and merged with
main.src/tools/browser)refs. Set-of-marks screenshots, downloads and uploads, dialogs. Attaching to your own Chrome over CDP.browser_agent. Passwords masked in snapshots. Lean mode (browser.lean, defaultauto, headless only) skips images, fonts and media while nobody needs the pixels; the HTTP cache keeps working.src/tools/computer)src/sentinel)src/missions)src/control)src/workflows)src/channels/telegram)/stop, mail rules and grants from your phone.src/security)manual/edits/auto(Shift+Tab,--auto,approval.defaultMode). Auto runs everything inside the project. It still asks for Sentinel-critical actions and for destructive actions outside the project, and writes to agent instruction files ask in every mode. In a measured 43-step scenario, prompts dropped from 23 to 6./goalkeeps working until a check passes or evidence is cited./stopis an emergency stop (TUI, control center, Telegram)./learnturns the task just finished into a skill. Scheduled monitors (--continuity --notify-on-change).src/mail,src/grants)mail_*tools;mail_sendalways asks. Standing reply grants are created only by a human and only cover a same-thread reply to the original sender. An IMAP IDLE watcher, plus rules that start a task with the email fenced as data. docs/MAIL.mdsrc/vault)browser_login(identifier-first, 2FA).vault_generate_and_fill.vault_request_login: the human types the login into a masked TUI prompt or a sealed control-center form, so the agent never sees it. "Save this login?" after a human login, and a vault panel. docs/VAULT_AND_CAPTCHA.md[CHALLENGE_HUMAN_ONLY]).browser.stealthis now off by default.context-barandyou-should-know(docs/MODS.md). Also: a wrap-up allowance at budget caps, send-now (Ctrl+Enter) that keeps a running shell command in the background, an auto-mode ask timeout, `firstReview and hardening
stop()waiting out the IDLE timeout./unpairconfirming before dropping its approval channel.CIis set) led to this one.$.fs.readandqodex vault import, code construction in the recorder, a dead duplicatecase 'mod', the mail attachment TOCTOU, HTML double unescape, and the earlier findings.await, a null value that only clears a field, test-only temp files, and atomicwxlock creates.Test plan
npx tsc --noEmitis clean.CI=true npx vitest run: 344 files and 4,463 tests pass locally. On GitHub, Test, Build and Analyze are green on 25a872d.🤖 Generated with Claude Code
https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen