Skip to content

feat: QodeX agent platform — browser, desktop, missions, Sentinel, auto mode, mail, password vault, CAPTCHA hand-off, mods - #110

Draft
QodeXcli wants to merge 277 commits into
mainfrom
claude/qodex-agent-upgrade-0eng77
Draft

QodeXcli wants to merge 277 commits into
mainfrom
claude/qodex-agent-upgrade-0eng77

Conversation

@QodeXcli

@QodeXcli QodeXcli commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

Summary

This turns QodeX from a coding CLI into a general autonomous agent that is local-first and model-agnostic, with Persian supported throughout. Every module was built on its own branch, integrated, tested end to end and merged with main.

Module What it adds
Dedicated QodeX Browser (src/tools/browser) A persistent profile, so logins survive restarts. Multiple tabs. Accessibility snapshots with refs. Set-of-marks screenshots, downloads and uploads, dialogs. Attaching to your own Chrome over CDP. browser_agent. Passwords masked in snapshots. Lean mode (browser.lean, default auto, headless only) skips images, fonts and media while nobody needs the pixels; the HTTP cache keeps working.
Desktop control (src/tools/computer) macOS, X11, Wayland and Windows backends: click, type, key, scroll, drag, clipboard, open-app, vision locate.
Sentinel (src/sentinel) A guard on every tool call. Purchases, payments, credentials and sending always need an explicit human; nothing bypasses it. Untrusted output is fenced and scanned for prompt injection (English and Persian).
Missions (src/missions) Long-running, detached, resumable goals with cross-process approvals and a live control center per worker.
Control center (src/control) A token-gated web UI: live browser view, human takeover, approvals, steering, timeline, emergency stop.
Workflows (src/workflows) Record agent, human or mixed demonstrations; replay with self-healing selectors under Sentinel.
Telegram (src/channels/telegram) Approvals, missions, status, /stop, mail rules and grants from your phone.
Real auto mode (src/security) Approval modes manual / edits / auto (Shift+Tab, --auto, approval.defaultMode). Auto runs everything inside the project. It still asks for Sentinel-critical actions and for destructive actions outside the project, and writes to agent instruction files ask in every mode. In a measured 43-step scenario, prompts dropped from 23 to 6.
Hermes-style features /goal keeps working until a check passes or evidence is cited. /stop is an emergency stop (TUI, control center, Telegram). /learn turns the task just finished into a skill. Scheduled monitors (--continuity --notify-on-change).
Mail (src/mail, src/grants) IMAP/SMTP accounts (secrets encrypted with the vault key) and mail_* tools; mail_send always asks. Standing reply grants are created only by a human and only cover a same-thread reply to the original sender. An IMAP IDLE watcher, plus rules that start a task with the email fenced as data. docs/MAIL.md
Password vault (src/vault) The key can live in the OS keychain (macOS Keychain, Secret Service, Windows DPAPI). Edit, rotate and CSV import. browser_login (identifier-first, 2FA). vault_generate_and_fill. vault_request_login: the human types the login into a masked TUI prompt or a sealed control-center form, so the agent never sees it. "Save this login?" after a human login, and a vault panel. docs/VAULT_AND_CAPTCHA.md
CAPTCHA hand-off, never solving Bot checks are detected and self-clearing ones are waited out. The rest go to the human: a Telegram card with a cropped screenshot and a scoped, short-lived live-view link, a phone hand-off mode that relays the human's own press-and-hold or drag, and auto-resume. The agent can never act on a challenge ([CHALLENGE_HUMAN_ONLY]). browser.stealth is now off by default.
Claude Code parity Mods: hooks into QodeX itself, compatible with Claude Code mods, with trust-gated project mods and the built-ins context-bar and you-should-know (docs/MODS.md). Also: a wrap-up allowance at budget caps, send-now (Ctrl+Enter) that keeps a running shell command in the background, an auto-mode ask timeout, `first

Review and hardening

  • Each module had an adversarial review that fixed real bugs: Sentinel bypasses, vault leaks, mission auto-approval, payment laundering in replay, token leaks.
  • I reviewed the last batch of branches myself while merging. That found and fixed:
    • A recorder echo-matching bug that showed up only on fast machines.
    • The mail watcher missing mail that arrived between a check and the IDLE wait, and stop() waiting out the IDLE timeout.
    • Telegram /unpair confirming before dropping its approval channel.
    • Challenge detection stringifying the page-title promise.
    • A hand-off page polling the secret routes.
    • The terminal's secure login prompt putting keys typed right after Enter into the previous field, which gave a false "passwords do not match". The CI-only failure of its tests (Ink hides frames when CI is set) led to this one.
  • CodeQL:
    • Fixed: a predictable temp file in the Wayland backend, the check-then-read races in mods $.fs.read and qodex vault import, code construction in the recorder, a dead duplicate case 'mod', the mail attachment TOCTOU, HTML double unescape, and the earlier findings.
    • Every remaining alert has a reply on its thread explaining why it is a false positive: env-derived values printed only through redaction, deliberate double checks across await, a null value that only clears a field, test-only temp files, and atomic wx lock creates.
    • The CodeQL gate needs a maintainer to dismiss those alerts in the Security tab. I can't do that from here.

Test plan

  • npx tsc --noEmit is clean.
  • CI=true npx vitest run: 344 files and 4,463 tests pass locally. On GitHub, Test, Build and Analyze are green on 25a872d.
  • Real-browser end-to-end tests: shop and vault (Sentinel blocks "Place order"; the vault refuses a phishing clone), missions, desktop under Xvfb.
  • Real local SMTP server: a reply covered by a grant is sent, anything out of scope asks, and with no human it is refused.
  • Real Chromium: CAPTCHA hand-off (detection, self-clearing wait, the human-only guard, press-and-hold relay, auto-resume) and lean mode (skips images/fonts/media, keeps the cache, turns off for a screenshot, PDF, takeover or live view).

🤖 Generated with Claude Code

https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen

claude added 5 commits October 3, 2026 00:46
- The agent can no longer reach its own approvals through the control center:
  requests are refused (403 [AGENT_BROWSER]) while the agent's browser has a tab
  on this control center (matched by Host / any local or tunnel address of the
  port); such tabs are sent back to about:blank in a QodeX-launched browser, and
  control-center login cookies (qx_ctl*, any port) are swept out of the agent's
  profile after each ?k= login, on eviction and every 20s.
- Live-view navigate refuses this control center and any QodeX control link
  (local/LAN/tunnel URL with a k= token): opening one in the agent's browser
  planted the login cookie there ([CONTROL_CENTER_URL]).
- `qodex control` ignored Ctrl+C in the real CLI: src/index.ts imports the tool
  registry, whose process-registry adds a non-exiting SIGINT listener. The
  command now installs (from the start) an always-exiting SIGINT handler with a
  bounded graceful stop; a second Ctrl+C exits at once.
- ?k= login with a percent-encoded key (%6B=) kept the token in the bounce target
  and bounced forever; stripTokenFromUrl now compares decoded keys.
- maskSecrets masks ?k=/&k= control tokens (e.g. mission live URLs) in text
  published to viewers/channels.
- A control-center takeover is handed back to the agent after 5 min with no
  dashboard connected (takeoverReleaseMs; 0 disables), with a bus notice —
  unattended runs no longer wait forever behind a closed tab.
- Control action errors map to proper statuses: *_NOT_FOUND 404,
  BAD_REQUEST/INVALID_* /*_BAD_* 400, *_NOT_PENDING/CONFLICT 409 (was always 500).
- missions-bridge: concurrent registerMissionControl calls raced into two event
  bridges (every mission event published twice) — now one shared setup; disposers
  are idempotent (a double release dropped other holders' actions);
  missions.resolveApproval threw nothing on failure (HTTP 200 + {ok:false}, the
  dashboard dropped the card) — now throws [APPROVAL_NOT_FOUND]/[APPROVAL_BAD_ANSWER]/
  [APPROVAL_NOT_PENDING]; a missions DB that cannot open no longer fails
  registration (and with it /control).
- LAN rebind: explicitly-undefined options no longer wipe kept settings.
- Dashboard: emoji, AltGr and macOS Option characters are typed as text (were sent
  as unknown key combos); 404/409 on a mission approval removes the card from the
  right list; the URL bar follows only the ACTIVE tab (Module A publishes
  'navigated' for every tab).
- Tests: real-Chromium end-to-end with the real QodexBrowserManager (takeover-only
  launch via navigate, CDP screencast frames, frame-coordinate clicks, typing and
  key names, frame.jpg, hand-back resumes waiters, dashboard driving the agent
  browser, close/relaunch, agent self-approval attempt), subprocess Ctrl+C test,
  missions-bridge tests, and server regressions for each fix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Desktop control (src/tools/computer/**), found by review and verified
against real Chromium under Xvfb where possible:

- Clipboard paste is now abort-safe. pasteText restores the previous
  clipboard in a finally, outside the aborted signal, and clears it when
  the old contents couldn't be read. Before, cancelling mid-paste left the
  typed text (possibly a password) on the clipboard (reproduced under Xvfb
  with xclip). On Windows, a killed paste script now clears the clipboard
  if it still holds our text.
- X11 'auto' typing pastes non-ASCII text. `xdotool type` of Persian
  types wrong letters but still exits 0 ("سلسم سنیا" for "سلام دنیا",
  3 of 4 runs, at any --delay), so the old fallback-on-failure never
  fired. An aborted `xdotool type` no longer falls through to the
  clipboard.
- X11/Wayland typed text now goes over stdin (`--file -`), not argv, so
  it no longer shows in ps or /proc/*/cmdline. Older tools fall back to
  argv.
- CRLF is normalized before typing (xdotool pressed Return twice per
  "\r\n").
- Window titles are attacker-controlled (a browser window's title is the
  page <title>). Output Sentinel does not fence (screenshot result and
  notes, out-of-bounds coordinate errors, screen_info, WINDOW_NOT_FOUND
  lists, backend errors) now only shows safeTitle(): truncated, invisible
  and bidi characters stripped, and withheld if scanInjection flags it.
- computer_use_agent passes ctx.askUser to the sub-agent runner, like
  `task` does, so its approvals reach the caller's human or channel.
- computer_use_open: coerceArgs turns host:port targets
  ("evil.example:8443/x", "localhost:3000") into the http:// URLs the
  backends open, so Sentinel's blocked, allowed and private-network
  domain policy applies. Before, they were reviewed as a local app. The
  tool also refuses QodeX's own secrets (vault, vault key, ~/.qodex/.env,
  browser profiles; file: URLs and symlinks too), and screenshot paths
  can't point into them.
- computer_use_click takes an optional `element`, auto-filled in
  coerceArgs from the last computer_use_locate hit at that point, so the
  review and the activity log can say what is clicked.
- computer_use_locate stops waiting on the vision model when the call is
  cancelled (vision_analyze ignores the signal).
- The locate parser accepts Qwen2.5-VL bbox_2d / point_2d and Qwen2-VL
  <|box_start|> tokens (qwen2.5vl is the recommended local vision model).
- computer_use_scroll with only one of x/y is an error instead of
  scrolling under the pointer.
- The default screenshots dir is created and tightened to 0700.

Tests: new test/desktop-hardening.test.ts (23 cases); desktop-backends
updated for stdin typing and auto-paste. tsc clean; full suite
158 files / 1916 tests pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Approval bypasses
- Foreground `mission start/resume --foreground --yes` on a TTY marked the
  worker as interactive, so Sentinel routed CRITICAL actions (purchase,
  payment, send, credential) through the step's askUser, which auto-answers
  "yes" in --yes mode: purchases were auto-approved. 'auto' mode no longer
  counts as a human at askUser; the terminal becomes an approval channel
  (TerminalApprovalChannel) racing the mission queue/control/Telegram.
- An agent inside a mission could approve its own requests or inject "user"
  steering notes via its shell (`qodex mission approve|steer`): refused when
  QODEX_MISSION_ID is inherited ([MISSION_APPROVAL_FORBIDDEN] /
  [MISSION_STEER_FORBIDDEN]).
- The live control-center URL (token = approval authority) was shown to the
  model by mission_status/inline mission_start and printed into the
  agent-readable worker log: token redacted there (CLI status for the human
  still shows it, except from inside a mission).
- mission_start ran no permission check at all (a prompt injection could start
  detached autonomous work): now goes through PermissionEngine/askUser like
  the equivalent shell command (allow / autoReject / ask).

Worker pid ownership
- runMission never released its pid: after an inline (TUI) mission paused or
  failed, `mission cancel` from another terminal/Telegram SIGTERMed the TUI,
  and a cancel from the same process left it "Cancelling…" forever. Pids are
  now released at finalize, only ACTIVE missions are signalled, and shared
  inline hosts (worker_shared) are never signalled.
- PID reuse after a crash/reboot made a stranger process look like the worker
  (stuck "running", SIGTERM to an unrelated process). The worker's start
  identity (/proc/<pid>/stat starttime, column pid_start) is recorded and
  checked (isWorkerAlive).
- Racing resumes could start two workers (spawn overwrote the pid; the worker
  check-then-set was not atomic): claimWorker() is a compare-and-set; a losing
  spawn reports [MISSION_BUSY] and a losing worker exits 3 without failing the
  mission.

Untrusted output / secrets
- Step results were pasted raw into the next step's prompt and the report
  prompt (injection propagation with user authority): fenced as data with an
  injection banner. mission_status is untrustedOutput; inline mission_start
  fences the mission text.
- Tool-result excerpts, notices, step/report excerpts, milestones, approval
  prompts (incl. Telegram/control rows), the --yes audit trail, worklog and
  notifications are secret-masked (maskSecrets); step results/report keep the
  real values.

Correctness
- `qodex mission start/status/list/resume` lost --yes, --model and --json to
  the root program's identically named global options (commander), so
  `--yes` and scheduled routines silently ran in 'ask' mode on the default
  model and --json printed text. Subcommands now read optsWithGlobals.
- A deleted/expired DB approval whose options have no 'n*' answer left the
  tool waiting forever: the poller withdraws it with broker.cancel().
- Approving a pending approval of a dead worker reported success into the
  void: answer paths reconcile first (approval expires, mission paused).
- Planner capping folded overflow steps into the last step but dropped their
  dependencies (folded work could run before what it needed).
- Schedule id prefix lookup used unescaped LIKE (`schedule rm %%%%` matched an
  arbitrary entry).
- Mission worker and schedule run logs (goal, output, report) are 0600 in a
  0700 dir.

Tests: test/missions-review.test.ts (22 tests, all failing before the fix),
plus a real-process E2E (detached worker against a hanging local model,
duplicate worker refused, SIGTERM cancel, resume, schedule tick).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
- '# Your Computer' is for the top-level agent only; sub-agents get a
  focused role brief (subagent-compressed prompt 7.1k -> 6.4k of 6.5k)
- shorter shared browser param descriptions (ref/selector/element/
  timeout/snapshot are repeated across ~20 tools)
- toolTokensNormal covers the five new relevance-gated tool families
  (measured ~44k over 163 tools, same ~13% headroom as before); the
  per-turn gated budget is unchanged and still passes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Guard bypasses and integrity gaps found after integration, each with a
failing test first (test/sentinel-review, test/sentinel-browser on real
Chromium via ToolRegistry, test/vault-frames on real Chromium):

- afterTool trusted a text prefix to decide "already fenced": page text
  starting with "<untrusted_content ...>" (browser_get_text, evaluate,
  clipboard) skipped both the injection scan and the fence. Now decided by
  result metadata only; banner excerpts can no longer carry markup.
- web_fetch / web_search / remote http_request and MCP tool output is now
  scanned and fenced too (localhost / LAN dev servers stay unfenced).
- browser_fill_secret judged a field's frame by its URL: an about:blank
  child of a cross-origin ad iframe (or a sandboxed srcdoc frame) "was" the
  page and received the password (reproduced in Chromium). Now uses the
  document's origin (window.origin); opaque/unknown about: frames refused.
- Vault: tenants on shared hosting (me.github.io, bucket.s3.amazonaws.com)
  now match exactly; "evil.bucket.s3.amazonaws.com" is someone else's.
- A scripted click routed around the purchase guard: browser_evaluate
  ".click()/requestSubmit()/dispatchEvent" and javascript: URLs (which run
  in the page) are now judged like a click on what they select (selector
  words + the described target elements; checkout/gateway pages). Space on
  a focused button is an activation (press Space on "Place order").
- browser_fill_form fields addressed by selector were never described or
  classified (password by selector slipped through); target-less
  browser_type now describes the focused element; selector words are used
  when an element can't be described.
- Self-approval: shell `qodex mission approve|deny`, `qodex control`,
  `qodex telegram pair|setup|unpair`, `mission start --yes` (also typed
  via computer_use_type / clipboard) need a human; any use of the
  in-process control center (port, token, tunnel) is hard-blocked; ?k=
  control tokens are masked in every tool result and in the audit log
  (mission_status exposed the worker's live URL with its token).
- Approval trust stores are write-protected like config.yaml: ~/.qodex/.env,
  sessions.db (SQL writes to mission_approvals), channels/ (Telegram
  pairing) and sentinel/ (audit). Integrity rules ignore autoApprove.
- tar / zip / rsync / cp -r / grep -r / s3_sync of ~/.qodex as a whole
  are blocked like the vault key and profiles themselves.
- workflow_run review: matched the real workflow format (upload `files`,
  tab-new URLs, param defaults, vault-backed secret params = fixed high)
  and found workflows by normalizeWorkflowName (Persian/spaced names were
  "contents unknown").
- browser_dialog accept reads the pending dialog text (confirm "Delete
  your account?" / "Confirm purchase?").
- Unreadable Sentinel config falls back to the defaults (was: allow all);
  the audit trail recreates its directory if it is removed mid-session.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Comment thread src/agent/recovery.ts Fixed
Comment thread src/channels/telegram/command.ts Fixed
Comment thread src/channels/telegram/command.ts Fixed
Comment thread src/channels/telegram/format.ts Fixed
Comment thread src/control/command.ts Fixed
Comment thread src/tools/browser/command.ts Fixed
Comment thread src/control/server.ts Fixed
Comment thread src/control/server.ts
Comment thread src/sentinel/guard.ts Fixed
Comment thread src/sentinel/guard.ts Fixed
claude added 9 commits October 3, 2026 00:50
Replay (src/workflows/replay.ts)
- Sentinel/vault saw a different element than replay acted on: text/label
  matches (and the recorded selector) pick the first VISIBLE match, but the
  guard got the bare selector and describeSelector() resolves .first() -- a
  hidden look-alike first laundered a payment click (reproduced on real
  Chromium with the real Sentinel). The selector handed to the guard and the
  vault is now pinned with `>> nth=i`; records keep the reusable selector.
- Recorded selectors are resolved with page.locator (the manager's locator
  is .first(), hiding extra / hidden matches); a precise candidate that now
  matches several visible elements yields to one that matches exactly one.
- Fail closed: without a ToolContext the guard was skipped entirely; now a
  deny-by-default context is used. The default guard is Sentinel (static
  import); if it can't be obtained every consequential step is refused.
- Never act after cancel / during a human takeover: abort + takeover are
  re-checked right before every page action (an approval can take minutes).
- Secret params no longer leak: composite values ("{{user}}:{{pass}}") are
  treated as secret, secrets in names/URLs are scrubbed from guard args,
  action records and reports (incl. URL-encoded forms), and a secret param in
  a navigation URL is refused.
- start_step past the last step replays nothing instead of being clamped back
  onto the last step (the resume hint for a hand-finished last step would
  repeat e.g. "Place order").
- Uploads: QodeX-home check also against the symlink-resolved home.
- Final URL moved inside the untrusted-data fence; title timer cleared.

Recorder (src/workflows/recorder.ts)
- The capture nonce leaked through window property names
  (__qxRecInstalled_<nonce>): any page could list them and forge steps
  (reproduced: a page injected a "Delete account" click). Names now use a
  one-way tag; the binding is never re-read from the page global.
- Only trusted input/change events are captured (pages can't author fills by
  dispatching synthetic events); input from cross-origin iframes is ignored.
- Enter in a text field recorded fill+press+fill+click (change event and the
  implicit-submission click); now one fill + one press.
- Mixed recordings: echo matching rewritten (value-aware, echoes of instant
  actions must precede the agent's record, submit:true also echoes Enter,
  Enter on the focused field matched by key) and navigations are attributed
  to the agent action by its echo time -- previously the wrong capture was
  dropped (the human's own input lost) and Enter was duplicated.
- Typed values are parameterized into later URLs only as whole query values /
  path segments, never into the host ("shop" on shop.example made the host a
  param, so replay navigated to another site).
- A slow capture install that outlived its recording left the binding/init
  script in the context forever; now disposed (generation-checked).
- Auto-detected vault fills (no ref/selector) keep a replayable step.

Skills / tools
- Generated SKILL.md (trusted, auto-injected instructions) carried
  page-derived labels and the model's description unscanned: workflows whose
  text trips Sentinel's / the skill installer's injection scanners get no
  skill; step labels are fenced as data with backticks neutralized.
- workflow_show output is untrustedOutput; workflow_run accepts params: null;
  skill allowed-tools cover every step kind.

Tests: test/workflows-hardening.test.ts (unit regressions),
test/workflows-e2e.test.ts (real QodexBrowserManager + browser_* tools +
real Sentinel on Chromium), chromium forge/synthetic-event test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sub-agent isolation (fresh child loop per run), cancellation-bound askUser,
run-mode enforcement at execution, wall budget that only fires on stalls,
real-browser core tests. Tests aligned with the headless fail-safe policy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
runSubagent now runs each child on a fresh AgentLoop, so the parent-instance
stub of buildInitialMessages no longer reached the child.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Comment thread src/sentinel/policy.ts Fixed
Comment thread src/missions/tools.ts Fixed
Comment thread src/missions/tools.ts Fixed
claude added 4 commits October 3, 2026 00:59
…ping)

- control center: an unexpected request error no longer echoes its text to
  the client (logged only)
- mission log tail reads size and data through one descriptor
- scheduler lock steal renames atomically and puts back a fresh lock it
  moved by mistake (two ticks stealing at once could both run)
- Windows browser-profile probe opens the lockfile directly
- Telegram plain-text fallback strips tags until stable

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
… initial values)

- scheduler lock: age, inode and holder pid read through one descriptor
- mission log: no exists-check before the open
- Telegram plain-text fallback drops stray angle brackets before decoding
- control center never stringifies a thrown non-Error into a response
- mission_start confirmation: no dead initial values

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Comment thread src/schedule/runner.ts
let stat: fsSync.Stats | null = null;
let holder: number | null = null;
const existing = await fs.open(lockPath, 'r').catch(() => null);
if (!existing) return fs.open(lockPath, 'wx').catch(() => null); // released meanwhile

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the same lock path as the thread above. The open flagged here is open(lockPath, 'wx'), an exclusive create, used when the lock disappears between the first EEXIST and the inspection. It can't be raced into a double hold, so I'm leaving it as is.


Generated by Claude Code

claude added 9 commits October 3, 2026 01:07
…fixes)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ept)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rigin

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- R1: process-registry / live-registry no longer install module-load SIGINT
  listeners that never exit (importing the tool registry disabled Node's
  default Ctrl+C exit for every subcommand); cleanup stays on 'exit'.
  qodex control keeps its always-exiting graceful-stop handler (comment
  fixed); qodex telegram start now always installs its own exiting handler
  (graceful stop confirms the update offset) instead of only when another
  listener existed. telegram-command test updated for that.
- R2: dev_server_* schemas put .describe() before .optional() so the
  descriptions reach the JSON schema; dev_server_start.env is a
  [{key, value}] array (z.record was advertised as a string), with
  coerceArgs still accepting a {NAME: value} object or its JSON string.
- R3: ToolRegistry redacts args (redactForAudit, hideTyped) in the
  'Executing tool' debug log and in the ARGUMENT_VALIDATION_ERROR echo, so
  typed text / password values / bodies never reach qodex.log or the
  transcript verbatim.
- R4: remember refuses facts with high-severity Sentinel injection findings
  ([MEMORY_REFUSED], nothing stored) — same bar as the loop's load filter.
- R5: gather / fanout / orchestrate declare timeoutSeconds = 2400 like task,
  so they aren't killed at the global tool timeout and their run time is
  excused from the parent's wall budget.
- R6: the agent loop's cancel listener settles the tool race for every
  tool, not only no-timeout ones: Ctrl+C no longer waits on a tool that
  ignores ctx.signal; the timeout rejects before aborting.
- R7: mission_status / mission_list and the background_job / dev_server
  status pollers are state-dependent (repeat guard keyed by result hash).
- R8: TUI tool summaries lift Sentinel's injection banner off the top of a
  fenced result (shown clipped after the tool's own lines); the fence
  itself was already stripped for display.
- R9: telegram.apiBase, telegram.botTokenEnv and control.host are honored
  only from ~/.qodex/config.yaml; a project .qodex/config.yaml setting them
  is ignored with a warning.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
…6), no unsafe-inline

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
claude added 6 commits October 3, 2026 12:05
…omium; stale login error fix

Fake manager: http (non-localhost) page and wrong-site rotation refused, failed
typing rolls back a new entry / restores the previous password, value scrubbed.
Real Chromium on local pages: sign-up fills password + confirm and saves the
entry, change-password rotates into new + confirm and keeps the old one,
maxlength respected, login-only form refused; browser_login classic,
identifier-first and TOTP flows, wrong url / redirect to another site refused
with nothing filled, one failed attempt then LOGIN_HALTED until the entry
changes, a payment-like submit button refused, submit:false. Every test checks
that no password, seed or code reaches a result, progress event, action record
or the console.

Fix: an error message already on the page before submitting (a script-driven
form after an earlier attempt) was read as this attempt failing; it now only
counts when the form is still there at the deadline.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Mount /api/secrets and /api/vault (full token only; loopback or https; sealed
ECDH+AES-GCM over tunnels; plain-http LAN refused) and register the control
center as a secret surface while it runs. Dashboard: a request card that seals
the typed login in the page, and a vault panel (masked list, add, change
password, edit sites/login URL, confirmed delete). Real-Chromium tests cover
local and tunnel sealing and assert no secret reaches responses, bus or approvals.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
… target field

The guard only gathered the page URL and the target element for browser_*
tools, so the approval prompt for vault_generate_and_fill named neither the
site nor the field. It now gets both (ref / selector described like
browser_fill_secret).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
… the running shell, unattended auto-ask timeout, project instructions first|all

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen

# Conflicts:
#	src/cli/ui.tsx
#	src/sentinel/guard.ts
….1, model aliases, /checkup prompt-audit, build-eval / hillclimb skills

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen

# Conflicts:
#	src/index.ts
…ault_generate_and_fill

A successful vault_generate_and_fill (and browser_login, explicitly) is
evidence of a real-world action for the completion gate; a failed login is
not. The recorder turns browser_login's auto-detected vault fills into steps
on the field's autocomplete token (username / email / current-password /
one-time-code) instead of skipping the username step.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Comment thread test/workflows-store.test.ts Fixed
Comment thread test/salvage-desktop.test.ts
Comment thread test/browser-real.test.ts Fixed
Comment thread test/automode-core.test.ts Fixed
Comment thread test/automode-core.test.ts Fixed
Comment thread src/workflows/command.ts Fixed
Comment thread src/tools/computer/backends/types.ts Fixed
Comment thread src/sentinel/guard.ts Fixed
Comment thread src/workflows/recorder.ts Fixed
Comment thread src/tools/browser/types.ts
claude added 7 commits October 3, 2026 12:11
The chat got 'unpaired' while the approval channel was still registered, so an approval
raised in that window could still be routed to it (the pairing test caught the gap under
load). Same order as /pair: channel state first, then the confirmation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
…ogin and vault_request_login

Tool relevance: a vault/login family (vault_*, browser_login,
browser_fill_secret) keyed on the user's own credentials and 2FA wording,
EN + FA ('log me in with my saved password', 'رمز عبورم', 'گاوصندوق', 'کد دو
مرحله') — never on a bare 'password', so coding tasks about password hashing
stay lean. Budgets unchanged: tool definitions 48143 / 50000 (173 tools),
compressed prompt 6326 / 6500.

Prompts: the Credentials line, the web-task addendum, the browser role and
browser_agent now lead with browser_login (one field: browser_fill_secret;
sign-up: vault_generate_and_fill). vault_request_login is named only when the
tool is registered (system prompt) or as 'if you have it'. The browser
sub-agent may use vault_generate_and_fill / vault_request_login. /vault lists
the edit / rotate / import commands.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
session.ts gains additive human-input observers (before a human click / Enter
under takeover, main-frame navigation, takeover switch). src/vault/capture.ts
reads the submitted login host-side from the owning frame (iframes followed by
focus / hit point, origin checked against the frame URL), keeps it in memory
only, and after the next navigation leaves the login form (or at takeover end)
asks 'Save the login for <host> (user <masked>)?' via ApprovalBroker — terminal
and control center — then adds or merge-patches the vault entry. Two-step logins
keep the step-1 username; an identical saved login is not asked again. Installed
by the control center and the TUI (local asker). Real-Chromium tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
…d full run

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
… wait is not missed

IDLE only reports EXISTS for mail that arrives after the SELECT, so a message landing between
the watcher's last check and the start of the wait sat until the next wake-up (up to the IDLE
timeout, 25 min by default). The watcher now passes the next UID it has not seen; the IMAP
transport compares it with UIDNEXT right after the SELECT (the fake does the same) and returns
at once. Regression test delivers a message exactly in that window.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
…-gated project mods, loop integration, qodex mod CLI

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen

# Conflicts:
#	src/cli/modes/headless.ts
#	src/cli/slash-catalog.ts
#	src/index.ts
…UI, built-in context-bar, you-should-know, sample-hello, /mod new, docs/MODS.md

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen

# Conflicts:
#	src/cli/slash-catalog.ts
#	src/cli/ui.tsx
Comment thread src/checkup/prompt-audit.ts
Comment thread src/checkup/prompt-audit.ts
Comment thread src/index.ts Fixed
claude added 2 commits October 3, 2026 12:22
…ion, CSV import, vault_generate_and_fill, browser_login

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen

# Conflicts:
#	src/sentinel/policy.ts
…rompt and a sealed control-center form, save-login capture, vault panel

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen

# Conflicts:
#	src/agent/tool-relevance.ts
#	src/control/dashboard.ts
#	src/control/server.ts
#	src/sentinel/policy.ts
#	src/vault/tools.ts
Comment thread test/cc-modsui-builtin.test.ts Fixed
Comment thread src/mods/command.ts Fixed
Comment thread src/mods/command.ts Fixed
Comment thread src/mods/api.ts Fixed
Comment thread src/cli/slash-commands.ts Fixed
Comment thread src/cli/ui.tsx Fixed
claude added 3 commits October 3, 2026 12:27
…watcher loop tests wait long enough under load

A page opened from a scoped hand-off link loaded the password panel script, which polled
/api/secrets (refused for scoped links, but the page should not ask). In hand-off mode the
panels stay hidden and nothing polls. The watcher loop tests now wait up to 10 s (still far
below the 60 s IDLE timeout they prove is not hit).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
…ringifying the promise

String(page.title()) gave '[object Promise]' — the title signal ('Just a moment…') was never
read when the in-page probe failed — and left the promise's rejection unhandled when the page
closed (seen as an unhandled rejection in the full suite).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
…eady aborted signal

A stop() landing while the watcher updated its state (or while the IMAP IDLE connection was
being set up) aborted the signal before the wait added its abort listener; that listener never
fires, so the wait ran to its timeout (60 s in tests, 25 min against a real server) and stop()
hung. The loop now re-checks the signal right before waiting, and the IMAP transport checks it
after connecting. Deterministic regression test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Comment thread src/vault/command.ts Fixed
Comment thread src/vault/vault.ts
Comment thread src/vault/vault.ts
Comment thread src/vault/vault.ts
Comment thread test/vault-core-keystore.test.ts
Comment thread src/control/secret-routes.ts Fixed
Comment thread src/vault/capture.ts
Comment thread src/vault/tools.ts Fixed
Comment thread src/vault/tools.ts Fixed
Comment thread src/vault/tools.ts Fixed
…arity (README, CHANGELOG, guides, Persian summary)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
@QodeXcli QodeXcli changed the title feat: QodeX agent platform — dedicated browser, desktop control, missions, Sentinel, control center, Telegram feat: QodeX agent platform — browser, desktop, missions, Sentinel, auto mode, mail, password vault, CAPTCHA hand-off, mods Oct 3, 2026
claude added 3 commits October 3, 2026 13:05
…, vault, CAPTCHA hand-off and mods

The previous docs commit only carried the MAIL.md move and the vault/CAPTCHA
guide; this adds the README sections and slash-command list, the CHANGELOG
entry, the standing-grants/watcher/rules section of docs/MAIL.md and the
Persian bullets in docs/AGENT_PLATFORM.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
…eeds the pixels

browser.lean (auto | on | off, default auto = only a headless browser QodeX
launched; QODEX_BROWSER_LEAN=1|0 forces it). One CDP session per tab enables
the Fetch domain for the Image / Font / Media resource types only and fails
them with BlockedByClient: the DOM, scripts, styles, XHR, forms, cookies and
the HTTP cache are untouched (context.route() turns the cache off for every
request: a cacheable script was fetched 3x over 3 loads with it, once with this).

Never on the user's own Chrome (cdpUrl) or on loopback / LAN / non-http pages
(dev servers). It switches itself off for the rest of the session once pixels
matter: browser_screenshot, browser_pdf (both say what to reload), the live
view, a takeover or a bot check. browser_status shows it; the request log
names it and the console drops the per-image "Failed to load resource" noise.

12-photo page, headless: ~0 MB vs 5 MB downloaded, ~270 ms vs ~720 ms for
launch + load, ~100-180 MB less Chromium RSS.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
…temp file, file races, dead /mod case)

CI "Test" was red on the last three heads: test/vault-entry-tui.test.ts read
Ink's output, which Ink keeps to itself until unmount when CI is set
(is-in-ci). The test now renders with debug: true like the other Ink tests.

That exposed a real race in the prompt itself: Ink attaches key handlers in
a deferred effect, so a key typed right after Enter could reach the previous
field's handler — a fast "pw⏎pw⏎" became a false "passwords do not match".
The form now routes every key through refs for the current field and text,
and a pasted value ending in Enter moves on. Two new tests fail on the old
component (type-ahead, paste) and pass now; 8/8 runs under CI=true.

CodeQL:
- wayland: measure the screen through a private mkdtemp dir, not a
  predictable name in the shared temp dir (symlink pre-creation).
- mods $.fs.read and `qodex vault import`: size check and read through one
  file handle (the file checked is the file read).
- recorder: the flush expression uses the hex-checked tag as a plain
  identifier instead of JSON-in-code.
- remove the unreachable second `case 'mod'` (the tested modsmith one stays),
  unused vault/tools imports, a dead initial value in the mail-grant check,
  a redundant `!pendingPrompt`, and clear typed secrets in place.
- tests: escape the reflected fixture value, one-handle reads, plain string
  matches instead of unanchored host regexes, explicit regex grouping.

Full suite under CI=true: 344 files, 4463 tests pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
Comment thread test/vault-entry-control.test.ts Fixed
claude added 2 commits October 3, 2026 13:37
…bstring check in page code)

CodeQL read `textContent.includes('shop.example.com')` as URL sanitization.
The test now waits for a non-empty row and asserts the host and the masked
username with toContain.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen
…identity (opt-in)

An honest alternative to detection evasion: with browser.botAuth on, QodeX
signs its own requests with an Ed25519 key so a site can recognise the agent
("this is QodeX, acting for its user") and choose to let it through. It never
hides that the browser is automated — no fingerprint spoofing, no hiding
navigator.webdriver, no synthetic human input.

- src/tools/browser/bot-auth.ts (pure): Ed25519 key management, RFC 7638 JWK
  thumbprint key id, RFC 9421 HTTP Message Signatures with the web-bot-auth tag
  over @authority + signature-agent, cached per authority, and the public-key
  directory (JWK Set) to publish. Unit tests verify the signature with the
  public key and reject a tampered authority.
- session.ts setupBotAuth(): a context.route signs same-site document / xhr /
  fetch on public hosts only; launched browser only (never the user's own
  Chrome over CDP); loopback / LAN pages are exempt (reuses the lean predicate).
  status().botAuth + botAuthSigner().
- config browser.botAuth (off by default; true/false/on/off/object;
  QODEX_BROWSER_BOT_AUTH=1|0); `qodex browser bot-auth --init|--directory` and a
  line in `qodex browser status` / browser_status.
- The private key lives in ~/.qodex/browser/bot-auth/ (0600); Sentinel adds it
  to the protected paths so the agent can never read or change it.
- Real-Chromium test: the server receives a valid signature on the document and
  the fetch (not the image), the public key verifies it, and loopback stays
  unsigned by default.

Docs: VAULT_AND_CAPTCHA.md (EN + FA), README, CHANGELOG, AGENT_PLATFORM.md (FA).
Full suite under CI=true: 345 files, 4477 tests pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ueof9NteyRfBNdeRpBxJen

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants