feat(computer-use): native desktop automation for macOS - #644
Conversation
GPT 5.6 Review — ✅ human override acceptedHuman judgment by @bolichen97 overrides the GPT 5.6 finding for This comment is updated in place on each push. The model was not re-run because an authorized human decision supersedes it. False positive or not applicable? A repository writer can comment: |
Arbiter — ✅ no blocking findingsArbiter found no unresolved long-term items that require action before merging Second-order review for Review detailsBoth files read. The diff file is empty (fully truncated), so I'm judging solely on the findings listed in
Nothing clears the deliberately narrow bar. Arbiter-Verdict: PASS No sub-threshold finding meets the long-term-impact bar. Suggested follow-ups (open as issues — non-blocking)
[ARBITER-REVIEWED] 99a2596 False positive or not applicable? A repository writer can comment: For a broader accepted-risk deferral, apply |
Design Review (Fable 5) — 🟡 CONCERNSAdvisory design-level review of Design-Verdict: CONCERNS AGENTS.md — the repo's own single-source-of-truth — ships stating the opposite of what this PR ships: it says Watch
Suggestions
[DESIGN-REVIEWED] 99a2596 |
de8c887 to
41edb60
Compare
027f8c7 to
ad834cc
Compare
cfee151 to
fcf2c9b
Compare
fcf2c9b to
83c2ef7
Compare
|
/ai-review override gpt 33ed3e2: Human override — scope decision, not a code defect. |
Human judgment recorded@bolichen97 marked the gpt AI finding as false positive, not applicable, or explicitly accepted for
This decision applies only to this commit. A new push requires a new judgment. |
Round 24 — reviewer findings addressed, plus one overrideSquashed to a single commit on current UX Review — both blockers were real. Fixed.1. The safety copy was false.
2. The screenshots were stale. Both showed an "Allow moving the mouse pointer" Also took the two cheap Not taken: the 3-minute permission-poll bound. It is deliberate (each tick spawns a Root cause worth naming: stale prose outlived the codeWhile verifying I found ~30 docstrings and comments across 13 files still Most importantly, the bundled All of it removed or rewritten, and the specs updated in the same commit. GPT 5.6 — overriding the prescription, not a defectThe finding's prescribed fix ("restore fail-closed governance checks and require "denying policy → PreToolUse is skipped → denied actions execute." The denylist is A pre-authorized MCP tool skips the PreToolUse gate and still hits this. " I suspect the stale docstrings above fed this finding: they promised a two-permit Gatesflake8 · isort · mypy (550 files) clean · backend 20,846 passed · frontend |
|
/ai-review override gpt 4423ea9: Prescription is a scope decision; the stated bypass does not reproduce. The finding asks to "restore fail-closed governance checks and require explicit keystone and policy For the record, the stated failure chain does not hold on this commit:
The inconsistency that likely produced this finding was real and is fixed in this push: ~30 docstrings |
Human judgment recorded@bolichen97 marked the gpt AI finding as false positive, not applicable, or explicitly accepted for
This decision applies only to this commit. A new push requires a new judgment. |
Round 25 — rebased onto
|
|
/ai-review override gpt 6995a9a: Both findings prescribe reinstating the governance model that was removed by product decision. Same two findings as the previous commit, now split in two. Both fixes ask for the thing this PR Two corrections for the record, since both chains name a mechanism that is not the one doing the work:
The prose/code inconsistency that plausibly produced these findings was real and is fixed in this push: |
Human judgment recorded@bolichen97 marked the gpt AI finding as false positive, not applicable, or explicitly accepted for
This decision applies only to this commit. A new push requires a new judgment. |
Round 26 — Design Review addressed (docs + description, not the security model)All three blockers were correct: the description, 1. Phantom security model — fixedPR description: the "Two enforcement planes" and "Governance: off completely, or
2. The
|
Round 27 — GPT found a real security regression. Fixed, not overridden.This one was mine, and it was introduced by this scope change, so I want to be BLOCKING —
|
Opus 5 Review — timed out, twice; not a finding
No findings were emitted at either attempt — there is nothing to answer. The job's own Every other gate is green on this commit, including the three reviewers that DID Flagging rather than re-running a third time, since two attempts hitting the same wall |
Round 28 — two blockers found by actually using it, plus a real GPT findingThe feature is now working end to end on macOS. Getting there surfaced two bugs that 1.
|
|
/ai-review override gpt b4051b1: design decision |
Human judgment recorded@bolichen97 marked the gpt AI finding as false positive, not applicable, or explicitly accepted for
This decision applies only to this commit. A new push requires a new judgment. |
Round 29 — the remaining reference ports, and two real GPT blockersTwo things in this round: the accessibility reads that were still missing, and GPT's What was still missing from the walkEvery field the tree omitted was a turn the model spent guessing a coordinate and
A secure element still discloses only its existence — no traits, no frame. The click ladder gains a last rung. When the addressed element refuses every verb, Verified live on a real background window (Zoom, arm64): bounds read, frames relative GPT 5.6 BLOCKING ×2 — both real1. Refused, not implemented: the private sequence was reverse-engineered for a left click Gated at the dispatch chokepoint on the resolved method (so a future 2. A malformed keystone turned the Settings GET into an HTTP 500. It falls back to an empty Reproduced both before fixing, and confirmed the tests fail against the reverted code. FINDING — the CHANGELOG. Correct, fixed.The entries still advertised the governance model, per-app narrowing, per-use approval FINDING — function-local imports. Not taken.
A latent bug found while wiring this up
The Windows shardThe two failures two SHAs ago were my own new test asserting POSIX-only semantics:
Gatesflake8 · isort · mypy (557 files) clean · backend 21,236 passed · frontend |
Round 30 — GPT's third blocker: the action header bypassed redaction. Real, fixed.
Correct, and worth spelling out because the shape is the interesting part. Every mutating tool returns Reproduced before fixing — the literal survived end-to-end, and Fixed by redacting Both halves are pinned. The first behaviourally — a credential-named app dispatched Confirmed both tests fail against the reverted code. Note on the two local test failures I checked
Gatesflake8 · isort · mypy (557 files) clean · backend 21,237 passed · frontend |
Round 31 — GPT is green. One advisory finding taken the other way round.
The one remaining advisory finding was real, but its prescription was backwards:
The inconsistency is real and I verified both ends: But the runtime is the part that is right, and it is a security control, not an So the doc was the bug. A test pins the doc against the runtime in both directions, so they cannot drift again: Gatesflake8 · isort · mypy (557 files) clean · backend 21,242 passed · frontend The single local backend failure is my gitignored |
Round 32 — the stale indexless claim was in the spec too, in the worse directionGPT repeated its advisory finding on
That was written during the scope change and never implemented — the runtime has always It is now a row in What still refuses, with the reasoning (an unnamed target has no A test now pins both documents against the runtime, so a third copy of this claim Worth noting what this cost: GPT's finding was correct twice and my first fix was Gatesflake8 · isort · mypy (557 files) clean · backend 21,243 passed · frontend |
Round 33 — GPT's fourth blocker: the plist read was check-then-open. Real, fixed.
Correct, and it was mine from an earlier round — a reviewer finding I fixed by adding
Two things came with it:
Also added a structural assertion that no bare I also swept the rest of the package for the same shape: the only other file reads in Gatesflake8 · isort · mypy (557 files) clean · backend 21,246 passed · frontend |
Read and drive native desktop apps through the macOS accessibility layer:
list on-screen windows, snapshot one as a numbered element tree, then click,
type, set values, scroll, drag or run a named action by element index or
screen coordinate. Pure ctypes over AXUIElement / CoreGraphics / ImageIO —
no pyobjc. Windows and Linux report unsupported rather than degrading.
Computer use is ONE operator opt-in. The enable lives on the keystone
`computer_use.json`, which `security._SENSITIVE_HOME_DIRS` fences the agent
away from, so a prompt-injected agent can neither read nor flip it. Past that
point the agent drives the desktop the way the operator would: there is no
governance model, no per-app allow-list ceiling, no unattended-surface
refusal, and no interactive-approval floor. That is a product decision, and
the residual risk is documented rather than implied — see
`docs/system-specs/modules/computer-use.md` -> "What is enforced, and what is
not" and "The keystone is the whole security boundary".
What still refuses, all enforced in band on the dispatch path:
- KiroCrew's own window, because driving our Settings UI would route around
the keystone that holds the enable;
- password fields — never read, and a window holding one is never captured;
- sensitive-text and secure-target checks on every input verb;
- credential redaction on the way out.
The real-pointer path (`click_method: "global"`) needs no second opt-in, but
the model must NAME it: `auto` never resolves onto a pointer-moving method, so
the operator's cursor is never warped by accident, and every such gesture gets
its own SEL `tool_kind` so "did the agent take my mouse?" is one log filter.
With the ceiling gone the audit trail is the accountability: every call is
recorded, allowed or refused.
One accessibility walk now reads what a model needs to act on the first try,
rather than leaving it to guess a coordinate and read back a screenshot:
- element frames (`AXPosition`/`AXSize`, unboxed from `AXValue` with the type
checked so a CGPoint can never be transposed into a CGSize), reported
WINDOW-LOCAL with the window origin published alongside them — the
screenshot is a crop of the window, so a screen-absolute rect could not be
related to any pixel the model can see;
- `editable` / `selected` / `expanded` traits, read as a tri-state so absent
never renders as false. `editable` comes from AXUIElementIsAttributeSettable,
not AXEnabled: a read-only field reports enabled with a readable value, so
the model used to type into it, get an ok result, and lose the text;
- the focused element and the user's text selection, read once per walk off
the application element (never system-wide, which would follow the operator
instead of the target app) and compared with CFEqual, since AXFocusedUIElement
returns a fresh reference that no address comparison would ever match;
- `AXRows` / `AXVisibleChildren` merged with `AXChildren`, deduplicated by
element identity. A table, outline or list often exposes its rows ONLY there,
so a children-only walk rendered a spreadsheet, a Finder list or a mail inbox
as an empty container.
A secure element still discloses only its existence — no traits and no frame,
for the same reason its value is withheld.
The click ladder gains a last rung: when the addressed element refuses every
verb, press the enclosing control. Web content renders a clickable row as a
plain AXStaticText inside a pressable wrapper, which left the whole row dead to
an element click. Bounded by hops AND by area ratio, declining outright when
the target has no frame to compare against, never for a right click, and the
result says it pressed the container so the model can tell "my click worked"
from "something near my click worked".
Also fixes a latent bug found while wiring this up: capture_snapshot_image
rebuilt the frozen Snapshot field by field, so every field added afterwards was
dropped whenever a screenshot was attached — and only then, which is why it had
gone unnoticed. It uses dataclasses.replace now, pinned by a test that walks
the whole dataclass so the next field is covered without anyone remembering.
The Windows shard's two failures were the new data-home pin test asserting
POSIX-only semantics: Path("/usr").resolve() is <drive>:\usr there, which the
resolver rightly accepts. Split into a portable root case and a POSIX-gated
system-directory case.
GPT 5.6 found two real blockers on the previous SHA; both are fixed rather than
overridden.
`sky_click` silently downgraded right and middle clicks to left. The recipe took
no button argument at all and built the left-button event codes unconditionally,
and nothing upstream refused the pair — so a right-click request through this
method activated the control instead of opening its context menu, on a background
window the operator cannot see. Refused now, not implemented: the private
sequence was reverse-engineered for a left click and the button number is one
field among nine, so a right-click variant would be invented rather than
observed. Gated at the dispatch chokepoint (policy.check_method_button, on the
RESOLVED method) and re-checked inside macos_skylight; the driver now passes the
button rather than assuming it, pinned by a test that reads the call site because
a behavioural test alone would keep passing if the argument were dropped again.
A malformed keystone turned the Settings GET into an HTTP 500. load_policy_config
raises on a present-but-malformed allowed_apps by design — coercing it to empty
would convert an operator's restriction into no restriction — but on the READ
path that escaped _snapshot() and made the only UI that can repair the file
unreachable. The page has to render precisely because the file is broken. It now
falls back to an empty PolicyConfig and publishes policy_error, which the panel
renders as a warning naming the file, since an empty allow-list otherwise reads
as "no restriction configured". The ceiling is unchanged: every dispatch still
loads the policy itself and still refuses, and a test asserts both halves.
Also rewrites the CHANGELOG entries, which still advertised the governance model,
per-app narrowing and per-use approval that the scope change deleted — GPT's
third finding, and correct.
A third GPT blocker, also real: the mutating tools' action header bypassed
credential redaction. Every mutator returns "<detail>\n\nRefreshed state:\n<tree>";
the tree half is redacted inside render_tree, but the header was concatenated
after that pass — and detail is not our prose. Every driver confirmation
interpolates app-supplied text (_click_text embeds app.name, the process name
macOS reports), so a process named "Notes key=AKIA…" put a raw credential
directly in front of a fully redacted tree.
detail is now redacted at the interpolation. Deliberately NOT by redacting the
joined string: render_tree appends its screenshot note after its own pass because
the per-user temp path contains a long random segment the bare-secret-key
heuristic masks, so a second pass would replace every screenshot path with a
placeholder and that channel would silently stop working. Header redacted alone,
already-redacted body untouched — both halves pinned by tests, the second one
structurally, since a mutator's refresh walk carries no image and so no
behavioural test in that file would catch the regression.
Also fixes a stale skill contract GPT flagged as advisory: SKILL.md advertised
computer_type_text(app, text, element_index?) with an "else the focused control"
fallback, which the runtime has never allowed. element_index is REQUIRED on both
keyboard tools because an unnamed target has no role or subrole for the
secure-field check to inspect, and press_key("tab") can move focus onto a password
box. GPT prescribed changing the schema to match the doc; taken the other way
round, since the runtime behaviour is the security control. A test now pins the
doc against the runtime so the two cannot drift again.
The same stale claim lived in the spec too, in the more dangerous direction:
"What no longer refuses" listed indexless keyboard input as working again, which
was written during the scope change and never implemented — so a reader auditing
the security posture from that document would conclude a control was gone that is
still enforced. Moved to "What still refuses" alongside two rows that were also
missing (the action header's redaction, and the non-left sky_click refusal), and
pinned by a test so neither document can drift from the runtime again.
A fourth GPT blocker, also real: the bundle Info.plist read was check-then-open.
It called is_sensitive_path and then opened the path in a separate step, so a
final-component symlink swapped in between would read a protected file's bytes on
a path that never touched the hardened gate — and the agent chooses which process
to target, so it can arrange the bundle. It now reads through
hooks.safe_read_prefix, which canonicalizes with realpath, re-checks the RESOLVED
target and opens with O_NOFOLLOW; that helper is the repo's stated requirement for
any read of an agent-influenced path.
The size cap moved off a getsize stat and onto the bytes actually read, since
statting a path and then opening it is the same raceable shape. Reading
MAX_INFO_PLIST_BYTES + 1 is what distinguishes "at the limit" from "over it"
without a second stat.
The existing floor test patched kiro_crew.security.is_sensitive_path, which the
helper resolves through its own import — so it would have passed against a
bypassed floor. Rewritten to stage a plist genuinely under ~/.ssh with $HOME
redirected, plus a symlinked-plist case, an oversized case, and a structural
assertion that no bare open() returns here (the original bug passed every
behavioural test in that file).
Fixes the Windows shard failure my own plist test introduced: it redirected only
$HOME, but os.path.expanduser reads USERPROFILE there, so the planted bundle was
not under a sensitive dir and the test asserted the opposite of what it meant. Both
vars are set now, and the symlink case is POSIX-gated (creating a symlink needs
elevation on Windows; the resolved-target check it exercises is platform-independent
and covered by the sensitive-dir case).
A fifth GPT blocker was half right, and the half that was right is fixed without
taking its prescription. An unresolved session key was forwarded as the empty
string, and SnapshotIndex namespaces by (session_key, window_key) — so every
unresolved session shared one ("", window) slot. Unresolved is the NORMAL case on
macOS, so two concurrent sessions observing the same window overwrote each other's
element indices, and each one's own verify_fingerprint still passed because both
trees describe the same window: a wrong-target action with nothing reporting it.
GPT prescribed refusing an empty key. Not taken — that is the refusal removed by
product decision, and it is what made the feature unusable on its only supported
platform. Fixed by namespacing instead: an unresolved key becomes
unresolved:<shim pid>, and kiro-cli spawns one shim per session, so the pid
separates the namespaces exactly as far as the sessions really are separate. Read at
call time so a forked child cannot inherit its parent's string and re-alias. The
prefix keeps it legible as a namespace separator rather than attribution. Nothing
is refused; only the cache key changed.
Round 34 — GPT's fifth blocker was half right. Fixed the real half, declined the prescription.
The diagnosis is correct and I had missed it. The prescribed fix is the one thing I cannot do. "Refuse an empty session key" is So the aliasing is fixed by namespacing instead: an unresolved key becomes Details that matter:
The Windows shard failure was mine, and instructive
Both vars are set now via a helper, and the symlink case is POSIX-gated — creating a Gatesflake8 · isort · mypy (557 files) clean · backend 21,252 passed · frontend |
|
/ai-review override gpt 99a2596: The bypass is a pre-existing repo-wide gap (arbitrary code execution defeats every keystone identically, verified on main) and the prescribed revert would strictly weaken the posture. |
Human judgment recorded@bolichen97 marked the gpt AI finding as false positive, not applicable, or explicitly accepted for
This decision applies only to this commit. A new push requires a new judgment. |
Round 35 — GPT's sixth blocker: pushing back, with the reproduction
The mechanism is real. I reproduced it. A But it is not this PR's bug, and the fix would make things worse. Three findings: 1. The identical technique defeats every keystone in the repoRan the same shape against keystones this PR never touched, all present on
Root cause: every matcher in 2. The prescribed revert strictly weakens the postureThe keystone is the security model here. Reverting it puts the enable in So the fix would open a read path while leaving the 3. The premise is already documented as an accepted residualThe gap requires arbitrary code execution, which is ungated by design in this repo What I am NOT claimingThe gap is real and it is not closed. Closing it needs OS-level enforcement (a I also added a scope-limit note in Overridden per-commit with that reasoning. |
|
/ai-review override fable 99a2596: Opus review timed out on oversized diff; deterministic CI and manual review completed. |
Human judgment recorded@bolichen97 marked the fable AI finding as false positive, not applicable, or explicitly accepted for
This decision applies only to this commit. A new push requires a new judgment. |
Problem
KiroCrew could drive a browser but nothing else on the machine. Anything that
lives outside a browser tab — a spreadsheet, a native internal tool, a desktop
app's export dialog, an error window — was simply unreachable, so a task that
touched one of them stopped at "I can't see that."
Why it matters
A large share of real work is in native apps. Without this the agent can research
and write code but cannot read a figure out of Excel, walk a desktop-only admin
tool, or tell you what an error dialog says. It also means the agent is blind
exactly where a user is most likely to be stuck.
Fix (symptoms → root cause → change)
Symptom: no capability existed for native windows.
Root cause: reaching them needs three OS primitives KiroCrew had never bound —
the accessibility tree (structure), window capture (pixels), and synthesized
input (action) — plus an authorization story, because a tool that can click
anything on your desktop is the highest-blast-radius thing the product ships.
Change: a new
computer_usepackage and a managedkirocrew-computerMCPserver exposing 10 tools. Written as our own Python rather than by vendoring
the reference runtime: that runtime's macOS bundle is signed by a third-party
Developer ID and is not notarized, and ad-hoc re-signing it — which KiroCrew's own
pipeline does to everything it ships — permanently breaks Accessibility
(verified 3/3 runs). The native layer is pure ctypes against system frameworks:
no pyobjc, no new dependency, no third-party binary.
Everything runs in the gateway, not the MCP sidecar
The stdio MCP process is a thin shim that resolves its identity strictly and
forwards over loopback; all native work, all refusals and all auditing happen in the
gateway, where the OS-resolved app identity and the addressed element's role are
known. Nothing relies on
hooks._governance_denial(the PreToolUse gate), which isfail-OPEN by deliberate repo policy and can be skipped entirely by a
pre-authorized tool — so the checks that matter are enforced in band on the
tools._dispatchpath that every caller goes through.One opt-in — computer use is deliberately NOT governed
This is a scope reversal from the first 23 review rounds, and it is intentional.
An earlier revision of this PR carried eight
SCOPE_CATALOGrows (capability,actions, apps, app_names, observations, targets, approval, pointer), an
unattended-surface refusal, an interactive-approval floor and a per-app disclosure
filter. All of it is gone.
gate.pyis audit-only and always permits;SCOPE_CATALOGgains zero computer-use rows; there is noallow_pointer_moveand nocapabilities.computer_use_pointer.CONTRACT_VERSIONstays 1 and the evaluator isuntouched — the removal is data rows and dead code, not an evaluator edit.
The product decision: computer use is one operator opt-in, and past that point the
agent drives the desktop the way the operator would. The residual risk is documented
rather than implied — see
docs/system-specs/modules/computer-use.md→ "What isenforced, and what is not" and "The keystone is the whole security boundary".
What still refuses, all in band on the dispatch path:
policy.check_app) — driving our Settings UI would routearound the keystone that holds the enable. Matched on bundle id, process name and
window title: the dashboard is also reachable as a browser tab, where the app
identity is Chrome's and an identity-only rule cannot fire. This is the one retained
denylist entry;
terminals, password managers and System Settings are no longer refused, because
that list was incomplete by construction (an IDE's embedded terminal never matched)
and got in the operator's way on their own machine.
What no longer refuses: unattended surfaces (a cron job driving the desktop is now
a supported flow), paste (
cmd+vis allowed), indexless keyboard input, observationchannels, and interactive approval.
Default posture
Off. Enabling it restarts your chat sessions and pins the data home into the MCP
spec's
env, so the feature works in the session you are sitting in — see "Two bugsfound by using it" below.
The enable lives on the keystone (
~/.kiro/crew/computer_use.json), notconfig.json, because an auto-approved agent shell can writeconfig.json(
is_sensitive_bash_commandreturnsNonefor it) — so the agent cannot enable itsown desktop automation. Flipping it restarts chat sessions, because kiro-cli
caches
tools/listper session and ACP has notools/list_changed; without that, anopen chat reports "0 tools" and the feature looks broken.
The real-pointer path (
click_method: "global") needs no second opt-in, but the modelmust name it —
autonever resolves onto a pointer-moving method, which is now theonly thing between an ordinary click and the operator's cursor. Every such gesture is
SEL-audited under its own
tool_kind, so "did the agent take my mouse?" is one logfilter.
Feature parity + PiP
Coordinate click, drag,
click_count,mouse_button,click_method(
auto/accessibility/app_post/sky_click/global), and Cursor Motion — areal-desktop fake cursor animated along a Bézier arc by a progress spring.
autonever resolves to a pointer-moving path, so the operator's cursor cannot be
warped by accident — the model has to name that method explicitly. Also adds a
picture-in-picture live view.
sky_clickis ported (an earlier revision of this PR deliberately did not).It clicks a window that is behind another one without raising it or moving the
pointer, which no public method can do:
accessibilityneeds an addressableelement and
app_postis ignored by Chromium/Catalyst renderers that hit-testagainst the window server. Hit in practice on a canvas behind an overlay. It is the
only path built on undocumented Apple ABI, so it is contained rather than accepted
wholesale — quarantined in
macos_skylight.py(a test fails if a private symbolname appears anywhere else), never reachable from
auto, and fully degrading to aclear refusal when a symbol is missing.
NOTICEcarries the attribution; nothird-party code is copied.
What one accessibility walk reads
The tree is the only channel the model reasons over, so every field it omitted was
a turn spent guessing a coordinate and reading back a screenshot to check. Added:
them and a line stating the conversion, because
computer_click(x, y)takesscreen coordinates. Window-local because the screenshot is a crop of the
window, so a screen-absolute rect could not be related to any pixel the model can
see (and it survives the user dragging the window).
AXPosition/AXSizearriveas
AXValueboxes, so each is type-checked before unboxing — a CGPoint read intoa CGSize would transpose
yintowidthand yield a plausible-looking rectpointing somewhere else, which is worse than no rect. A half-read yields
None.editable/selected/expandedtraits, tri-state so absent neverrenders as false.
editableis the load-bearing one and comes fromAXUIElementIsAttributeSettable, notAXEnabled: a read-only field (adisabled input, a log pane) reports enabled with a readable value, so the model
typed into it, got an
ok, and the text went nowhere.element. Not system-wide: that follows whatever the operator is in, so it would
mark a background app's element only when the target happened to be frontmost.
Compared with
CFEqual, sinceAXFocusedUIElementreturns a fresh reference noaddress comparison would ever match (that would be a silent no-op, not a visible
failure).
AXRows/AXVisibleChildrenmerged withAXChildren, deduplicated byelement identity. A table, outline or list often exposes its rows only
there, so a children-only walk rendered a spreadsheet, a Finder list or a mail
inbox as an empty container — which reads as "this app has no content". The
per-node child cap bounds the merged list, so three collections cannot together
exceed what one was allowed.
A secure element still discloses only its existence — no traits and no frame,
for the same reason its value is withheld (
editablewould confirm it acceptsinput; a rect would locate it for a coordinate click).
The click ladder gains a last rung. When the addressed element refuses every
verb, press the enclosing control: web content renders a clickable row as a plain
AXStaticTextinside a pressable wrapper, which left the whole row dead to anelement click while a coordinate click worked. Bounded twice — by hops and by
area ratio, since the real signal that an ancestor "is" the row is that it is
roughly the same size; a container orders of magnitude larger is the page. It
declines outright (before spending any AX round-trip) when the target has no frame
to judge against, never applies to a right click, and the result says it pressed
the container so the model can tell "my click worked" from "something near my click
worked".
Tests
29 new test files / ~1,150 cases. The load-bearing ones:
AXRole=AXTextField+AXSubrole=AXSecureTextFieldwith a readable value, soa role-only check misses every one. Also pins that a secure field past the node
budget still suppresses the window's pixels — the model picks
max_tree_nodes,so it could otherwise choose a budget that hid the field.
tree lines.
path listing still falls through to cookie auth when the secret header is absent).
(that path was an uncatchable
exit 134)._FN_SPECSrow has bothrestypeandargtypes;CGEventPostisconfined to the two
*_globalfunctions; CFRelease balance shows zero leaks.AXValuebox is refusedrather than transposed; a half-read frame is
None, not(x, y, 0, 0); anexplicit
AXSelected=falseand an absent one both render no trait; the same rowin two child collections appears once (identity, not handle — the fake models
the Create Rule by minting a fresh handle per read, so a pointer comparison
genuinely fails there); the ancestor fallback declines on a too-large container,
stops exactly at the hop bound, and never fires for a right click.
and a test fails if any private symbol name appears outside
macos_skylight.py.capture_snapshot_imageis purely additive — asserted field-by-field over thewhole
Snapshotdataclass, so the next field added is covered without anyoneremembering to (see the bug it caught, below).
harness, so everything runs on the Linux and Windows shards. The shipped fake
carries the new fields too (frames, traits, focus, selection, a non-zero window
origin so a window-local/screen mix-up is detectable), since it is the surface a
downstream suite sees.
Manual verification
Driven live on macOS 15 (arm64), because no CI job can:
helper pids
pgrepreturns).AXManualAccessibilityretry.0700dir with0600files; both ImageIOoption keys honored across a quality/size sweep.
screencaptureby design, click-through, and exits rc=0 on stdin EOF so agateway crash cannot strand a fake cursor on screen.
below).
~400KB-per-walk leak (measured, now 0.0KB) and synthesized keys inheriting live
modifier state (typing
abcproducedI Abc).Screenshots
Default state — off, with the advisory permission rows:
Enabled — Cursor Motion appears, and the card discloses what the agent may do:
Reviewer notes
The governance change is data rows only;
CONTRACT_VERSIONunchanged and noevaluator function touched.
The master enable is on the keystone, citing the
denied_commands.jsonprecedent (
security.py).scripts/scrub-allowlist.txtgains one anchored per-file entry forcomputer_use/policy.py. The same bundle id is already allowlisted for fourother files; the pattern itself is unchanged. It is there because KiroCrew
refuses to automate its own window — the dashboard can flip this feature's
master enable, so driving it would route around the keystone.
config-baseline.jsonis regenerated (the flagged config/infra signal).Known limitation, documented not papered over: an agent that can write
~/.kiro/agents/*.jsoncan add itself toallowedTools, which stops kiro-clisending permission requests and skips the PreToolUse gate entirely. That is a
pre-existing gap affecting every governed capability; hardening those paths is
deferred to its own PR so it is reviewed on its own merits.
Being explicit about what that now means, since an earlier draft of this
description claimed a second plane would still hold and that is no longer true:
with the governance model removed, an agent on that path can drive any application
except KiroCrew's own window — including a terminal, which reaches a shell without
passing the 137-rule command deny floor. The in-band refusals that survive
(KiroCrew's own window, password fields, the sensitive-text scan, redaction) still
apply, but they are not a substitute for the command floor. This is the accepted
consequence of the one-opt-in posture on a single-user machine where the operator is
trusted with their own desktop; it is written up in
docs/system-specs/modules/computer-use.md→ "The keystone is the whole securityboundary" rather than left for a reader to infer.
Two GPT 5.6 blockers on the previous SHA, both fixed rather than overridden —
and both are the kind that only show up when someone actually uses the thing:
sky_clicksilently downgraded right and middle clicks to left. The recipetook no button argument and built left-button codes unconditionally, and nothing
upstream refused the pair — so "open the context menu" became "activate the
control", on a background window the operator cannot see. Refused, not
implemented: the private sequence was reverse-engineered for a left click and
the button number is one field among nine, so a right-click variant would be
invented rather than observed. Same reasoning as
AX_MENU_LADDERnever fallingback to
AXPress— performing a different gesture than the one asked for isworse than performing none. Gated at the chokepoint on the resolved method, and
re-checked inside
macos_skylight; the driver passes the button now, pinned by atest that reads the call site, because a behavioural test alone would keep passing
if the argument were dropped again and the default took over.
load_policy_configraises on a malformedallowed_appsby design — coercingit to empty would turn an operator's restriction into no restriction — but on the
read path that escaped and made the only UI that can repair the file
unreachable. The page has to render precisely because the file is broken. It falls
back to an empty
PolicyConfigand publishespolicy_error, which the panelshows as a warning naming the file (an empty allow-list otherwise reads as "no
restriction configured" — the opposite of what the operator wrote). The ceiling
is unchanged: every dispatch still loads the policy itself and still refuses,
and a test asserts both halves.
The CHANGELOG entries are rewritten — GPT's third finding, and correct: they
still advertised the governance model, per-app narrowing and per-use approval that
the scope change deleted.
A latent bug fixed on the way, worth a look because of its shape:
capture_snapshot_imagerebuilt the frozenSnapshotfield by field, so everyfield added to the dataclass afterwards was silently dropped whenever a
screenshot was attached — and only then, which is exactly why it had gone
unnoticed. It uses
dataclasses.replacenow, pinned by a test that walks the wholedataclass. Verified by reverting the fix and watching the test fail.
The two Windows-shard failures on the previous SHA were my own new test
asserting POSIX-only semantics:
Path("/usr").resolve()is<drive>:\usrthere, sothe data-home resolver rightly accepts it. Split into a portable filesystem-root
case and a POSIX-gated system-directory case; the cross-platform invariant that
actually matters (the pin always agrees with the resolver) is still asserted on
every OS.
macOS-only; Windows and Linux backends degrade with a clear refusal.