Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 11 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -150,6 +150,7 @@ Also looked at: [mediar-ai/mcp-server-macos-use](https://github.com/mediar-ai/mc
- **Coordinates that land.** Every screenshot and zoom carries its screen origin and pixels-per-point. `zoom` captures any region at full physical resolution.
- **Your browser, your logins.** The Chrome extension gives the agent its own labelled tab group in your signed-in Chrome, and leaves your tabs alone unless you point it at one.
- **Or the tab you already have open.** `browser_list_tabs all=true` shows every tab in the browser and `browser_use_tab` takes one over in place — useful when the page is already signed in or mid-flow and re-opening the URL would throw that away. An adopted tab is not moved into the agent's group, not activated and not reloaded; cleanup releases it rather than closing it, and `browser_release_tab` hands it back early.
- **Tabs the agent's clicks open are the agent's too.** A `target=_blank` link or popup opened by an agent click or key press within 10 seconds joins that task's tabs and group, and its cleanup closes it. A tab opened from an agent tab without a recent agent action stays yours.
- **Parallel tasks, one extension.** Every MCP process gets its own tab group, and one process can run several tasks by passing a stable `session_id` on its browser calls. A task cannot drive or adopt another task's tabs, and cleanup (including a process exiting) closes only its own. Any number of MCP processes share the one extension: the first owns it and the rest go through it, and if the owner exits another takes over without closing anyone's tabs. Tasks share Chrome's cookies and logins, and desktop apps and the clipboard are not isolated.
- **Model-agnostic.** No vision model is required for interaction; local models work too.
- **Look → act → verify.** `hover` for mouse-over menus, `wait` for loads, `query` to find a control by label without reading a whole tree.
Expand Down Expand Up @@ -220,7 +221,9 @@ This is not a guarantee of clickability or
lack of occlusion. Frames and shadow
roots are not searched. Choose a selector specific to the **new expected UI**;
an element already present can satisfy it immediately. No arbitrary page script
is accepted.
is accepted, and a selector may not test the `value` attribute: React mirrors a
controlled input's value there, password fields included, so a wait could
otherwise confirm a password one guessed character at a time.

- `wait_timeout_ms`: defaults to `8000`, accepts integers `0`–`10000`; `0` checks
once. Requires `wait_for_selector`. The budget covers polling, not browser/API
Expand Down Expand Up @@ -248,8 +251,9 @@ Send this to `browser_snapshot`, inspect its readiness and fresh indices, then
choose the next action. A timeout or tool error does not roll back a click,
submission, typing, or navigation. Use current state to decide whether a retry is
safe. Invalid wait types/bounds or an action wait without `return_state: true` are
rejected before acting. These options require an updated extension as well as the
native server; older extensions may omit readiness metadata.
rejected before acting. `null` for an optional argument means the same as leaving
it out. These options require an updated extension as well as the native server;
when an older extension ignores them, the result says so and asks for a reload.

Regression checks: `node --test chrome-extension/background.test.mjs` covers the
wait contract and site-policy guards. The real-browser CI harness
Expand All @@ -267,7 +271,10 @@ node scripts/snapshot-dom.test.mjs --chrome /path/to/chrome --source original-ba
```

`browser_snapshot` automatically scopes to the topmost visible dialog when one is
open, so covered background controls cannot consume the dialog's budget. Hidden
open, so covered background controls cannot consume the dialog's budget. A
dialog's portalled popups stay in scope: elements its controls point at through
`aria-controls`/`aria-owns`, open popovers, and controls painted on top of the
dialog, such as a select's options rendered at the end of `<body>`. Hidden
and inert subtrees are excluded; reachable offscreen controls remain available.
Native modeless dialogs and explicit `aria-modal="false"` do not scope the page.
Native modal dialogs and explicit `aria-modal="true"` are definitive candidates.
Expand Down
235 changes: 218 additions & 17 deletions chrome-extension/background.js
Original file line number Diff line number Diff line change
Expand Up @@ -321,15 +321,65 @@ async function openTab(clientId, url) {
tabOwner.set(tab.id, clientId);
await ensureGroup(clientId, tab.id);
await persistOwnedState();
// Pages replace their favicon on load (Spotify, YouTube, …). Re-apply the
// pointer whenever the document finishes, and also when the tab's own icon
// changes, so the strip stays on the agent cursor rather than the site logo.
keepBadged(tab.id);
return { tabId: tab.id, url: tab.url, title: tab.title, clientId };
}

/// Pages replace their favicon on load (Spotify, YouTube, …). Re-apply the
/// pointer whenever the document finishes, and also when the tab's own icon
/// changes, so the strip stays on the agent cursor rather than the site logo.
function keepBadged(tabId) {
chrome.tabs.onUpdated.addListener(function badge(id, info) {
if (id !== tab.id) return;
if (info.status === "complete" || info.favIconUrl) markTab(tab.id);
if (!tabOwner.has(tab.id)) chrome.tabs.onUpdated.removeListener(badge);
if (id !== tabId) return;
if (info.status === "complete" || info.favIconUrl) markTab(tabId);
if (!tabOwner.has(tabId)) chrome.tabs.onUpdated.removeListener(badge);
});
return { tabId: tab.id, url: tab.url, title: tab.title, clientId };
}

/**
* When the agent last clicked or pressed a key in a tab, by tab id. A tab that
* opens from one of the agent's tabs shortly after such an action was opened by
* the agent (a target=_blank link, a window.open popup), so it belongs to that
* task. Without the time bound, a link the user clicks in a tab the agent holds
* would be claimed too — and then closed by the agent's cleanup.
*
* @type {Map<number, { clientId: string, at: number }>}
*/
const agentInputAt = new Map();
const CHILD_TAB_WINDOW_MS = 10_000;

function noteAgentInput(clientId, tabId) {
agentInputAt.set(tabId, { clientId, at: Date.now() });
}

/// Take ownership of a tab the agent's own action opened. It is agent-created,
/// not adopted, so the task's cleanup closes it like any tab it opened itself.
function claimChildTab(tab) {
const openerId = tab.openerTabId;
if (typeof tab.id !== "number" || typeof openerId !== "number") return;
if (tabOwner.has(tab.id)) return;
const input = agentInputAt.get(openerId);
if (!input || Date.now() - input.at > CHILD_TAB_WINDOW_MS) return;
const clientId = tabOwner.get(openerId);
// The opener changed hands (released, closed) since the action.
if (clientId === undefined || clientId !== input.clientId) return;
// Ownership is recorded now, synchronously, so a list_tabs or cleanup that is
// already queued sees the child; grouping waits its turn in the task's queue.
const state = clientState(clientId);
state.tabs.add(tab.id);
tabOwner.set(tab.id, clientId);
void persistOwnedState();
keepBadged(tab.id);
void enqueue(clientId, async () => {
if (tabOwner.get(tab.id) !== clientId) return;
try {
await ensureGroup(clientId, tab.id);
} catch {
// A popup window cannot hold a tab group. The tab is still the task's,
// and cleanup closes it wherever it is.
await markTab(tab.id).catch(() => {});
}
}).catch(() => {});
}

async function listTabs(clientId, all = false) {
Expand Down Expand Up @@ -627,9 +677,39 @@ const SNAPSHOT_JS = `((options = {}) => {
const r = dialog.getBoundingClientRect();
return r.right > 0 && r.bottom > 0 && r.left < innerWidth && r.top < innerHeight;
}) || null;
const scope = modal || document;
for (const el of scope.querySelectorAll(sel)) {
if (!visible(el)) continue;
// A select, combobox or menu inside the modal often renders its options in a
// portal at the end of <body>, outside the dialog. Those stay in scope: roots
// the modal's controls point at (aria-controls/aria-owns), open popovers, and
// any control painted on top of the modal: inside the modal's rectangle, and
// the topmost element at its own centre.
const floating = [];
if (modal) {
for (const owner of modal.querySelectorAll('[aria-controls],[aria-owns]')) {
const ids = ((owner.getAttribute('aria-controls') || '') + ' ' + (owner.getAttribute('aria-owns') || '')).split(/\s+/);
for (const id of ids) {
const root = id && document.getElementById(id);
if (root && !modal.contains(root)) floating.push(root);
}
}
try {
for (const pop of document.querySelectorAll(':popover-open')) {
if (!modal.contains(pop) && !pop.contains(modal)) floating.push(pop);
}
} catch {}
}
const bounds = modal ? modal.getBoundingClientRect() : null;
const inScope = (el) => {
if (!modal || modal.contains(el)) return true;
if (floating.some((root) => root.contains(el))) return true;
const r = el.getBoundingClientRect();
const x = r.left + r.width / 2, y = r.top + r.height / 2;
if (x < Math.max(0, bounds.left) || y < Math.max(0, bounds.top) ||
x >= Math.min(innerWidth, bounds.right) || y >= Math.min(innerHeight, bounds.bottom)) return false;
const hit = document.elementFromPoint(x, y);
return !!hit && (hit === el || el.contains(hit));
};
for (const el of document.querySelectorAll(sel)) {
if (!inScope(el) || !visible(el)) continue;
const index = i++;
// Tag the full selected scope so indices stay global and earlier pages remain
// actionable when paging an unchanged UI. Only label/serialize this page.
Expand Down Expand Up @@ -990,18 +1070,106 @@ async function typeText(tabId, text) {
return { typed: text.length };
}

/// Watch the next Tab keydown, so the follow-up can tell a page that handled
/// the key itself (an editor indenting, a widget trapping focus) from one where
/// nothing happened.
const TAB_PROBE_JS = `(() => {
const probe = { before: document.activeElement };
probe.listener = (e) => {
if (e.key !== 'Tab') return;
probe.event = e;
};
addEventListener('keydown', probe.listener, true);
window.__cuTabProbe = probe;
return true;
})()`;

/// Finish a Tab press. Chrome runs its own focus traversal only in a page that
/// has focus, and an agent tab in the background never does: the keystroke
/// reaches the page's handlers and then focus stays put, so the next type lands
/// in the field the agent meant to leave. When that happens, move focus to the
/// next element in sequential focus order the way the browser would have.
const TAB_FINISH_JS = `(() => {
const probe = window.__cuTabProbe;
delete window.__cuTabProbe;
if (probe) removeEventListener('keydown', probe.listener, true);
const name = (el) => {
if (!el || el === document.body || el === document.documentElement) return null;
const label = el.getAttribute('aria-label') || el.labels?.[0]?.innerText ||
el.getAttribute('placeholder') || el.getAttribute('name') || el.getAttribute('title') ||
(el.tagName === 'INPUT' || el.tagName === 'TEXTAREA' ? '' : el.innerText) || el.id || '';
const index = el.getAttribute('data-cu-idx');
return (index !== null ? '[' + index + '] ' : '') + el.tagName.toLowerCase() +
(label ? ' "' + String(label).trim().replace(/\\s+/g, ' ').slice(0, 60) + '"' : '');
};
const before = probe ? probe.before : null;
const active = document.activeElement;
if (probe && active !== before) return { moved: true, focused: name(active) };
if (probe?.event?.defaultPrevented) return { moved: false, handled: true, focused: name(active) };
if (active && (active.tagName === 'IFRAME' || active.tagName === 'FRAME' || active.shadowRoot)) {
return { moved: false, reason: 'focus is inside a frame or component the extension cannot step through' };
}
const shown = (el) => {
if (el.disabled || el.closest('[inert]') || !el.getClientRects().length) return false;
const style = getComputedStyle(el);
return style.visibility !== 'hidden' && style.visibility !== 'collapse';
};
// A radio group is one stop: its checked button, or its first if none is.
const radioStop = (el) => {
if (el.type !== 'radio' || !el.name) return true;
const scope = el.form || document;
const group = Array.from(scope.querySelectorAll('input[type=radio]'))
.filter((r) => r.name === el.name && shown(r));
const checked = group.find((r) => r.checked);
return checked ? checked === el : group[0] === el;
};
const sel = 'a[href], area[href], button, input:not([type=hidden]), select, textarea, iframe, ' +
'summary, [contenteditable]:not([contenteditable=false]), [tabindex]';
const stops = Array.from(document.querySelectorAll(sel))
.filter((el) => el.tabIndex >= 0 && shown(el) && radioStop(el));
const order = stops.filter((el) => el.tabIndex > 0).sort((a, b) => a.tabIndex - b.tabIndex)
.concat(stops.filter((el) => el.tabIndex === 0));
const at = order.indexOf(active);
let next;
if (at >= 0) next = order[at + 1];
else if (active && active !== document.body && active !== document.documentElement) {
next = order.find((el) => active.compareDocumentPosition(el) & Node.DOCUMENT_POSITION_FOLLOWING);
} else next = order[0];
if (!next) return { moved: false, reason: 'focus is already on the last element of the page' };
next.focus();
if (document.activeElement !== next) return { moved: false, reason: 'the next element refused focus' };
// Tabbing into a text field selects its contents, as it does for a person.
if (next.tagName === 'TEXTAREA' || (next.tagName === 'INPUT' && typeof next.select === 'function' &&
/^(text|search|url|tel|email|password|number)$/.test(next.type))) {
try { next.select(); } catch {}
}
return { moved: true, focused: name(next) };
})()`;

async function pressKey(tabId, key) {
const map = {
Enter: { windowsVirtualKeyCode: 13, key: "Enter", text: "\r" },
Tab: { windowsVirtualKeyCode: 9, key: "Tab" },
Escape: { windowsVirtualKeyCode: 27, key: "Escape" },
Backspace: { windowsVirtualKeyCode: 8, key: "Backspace" },
Enter: { windowsVirtualKeyCode: 13, code: "Enter", key: "Enter", text: "\r" },
Tab: { windowsVirtualKeyCode: 9, code: "Tab", key: "Tab" },
Escape: { windowsVirtualKeyCode: 27, code: "Escape", key: "Escape" },
Backspace: { windowsVirtualKeyCode: 8, code: "Backspace", key: "Backspace" },
};
const spec = map[key];
if (!spec) throw new Error(`unsupported key: ${key}`);
if (key === "Tab") await send(tabId, "Runtime.evaluate", { expression: TAB_PROBE_JS, returnByValue: true });
await send(tabId, "Input.dispatchKeyEvent", { type: "keyDown", ...spec });
await send(tabId, "Input.dispatchKeyEvent", { type: "keyUp", ...spec });
return { pressed: key };
if (key !== "Tab") return { pressed: key };
const res = await send(tabId, "Runtime.evaluate", { expression: TAB_FINISH_JS, returnByValue: true });
const focus = res?.result?.value || {};
// Reporting success when focus stayed put sent the next type into the wrong
// field, so an unmoved Tab is an error the agent can act on.
if (!focus.moved && !focus.handled) {
throw new Error(
`Tab did not move focus (${focus.reason || "the page did not respond"}) — ` +
"click the field you want with browser_click instead",
);
}
return { pressed: key, focused: focus.focused ?? null, handledByPage: focus.handled === true };
}

async function screenshot(tabId) {
Expand Down Expand Up @@ -1184,12 +1352,33 @@ function validateReadiness(p, action = false) {
if (typeof p.waitForSelector !== "string" || !p.waitForSelector.trim() || p.waitForSelector.length > 2000) {
throw new Error("wait_for_selector must be a nonempty CSS selector (at most 2000 characters)");
}
if (testsValueAttribute(p.waitForSelector)) {
throw new Error("wait_for_selector cannot test an element's value attribute");
}
if (action && p.returnState !== true) throw new Error("wait_for_selector requires return_state=true on actions");
if (p.waitTimeoutMs !== undefined && (!Number.isInteger(p.waitTimeoutMs) || p.waitTimeoutMs < 0 || p.waitTimeoutMs > 10000)) {
throw new Error("wait_timeout_ms must be an integer from 0 to 10000");
}
}

/**
* Whether a selector reads the `value` attribute. React mirrors a controlled
* input's value into it, password fields included, so a wait on
* `input[type=password][value^="a"]` would answer met/timeout one guessed
* character at a time — the password the agent is never meant to see.
* Escapes and comments are undone first, so `[\76 alue]` is caught too.
*/
function testsValueAttribute(selector) {
const plain = selector
.replace(/\/\*[\s\S]*?\*\//g, "")
.replace(/\\([0-9a-fA-F]{1,6})[ \t\n\r\f]?/g, (_, hex) => {
const code = parseInt(hex, 16);
return code > 0 && code <= 0x10ffff ? String.fromCodePoint(code) : "\ufffd";
})
.replace(/\\([\s\S])/g, "$1");
return /\[\s*(?:[\w*-]*\|)?value\s*(?:[~|^$*]?=|\])/i.test(plain);
}

/** A visible match in the top-level document, not arbitrary application readiness. */
function visibleSelectorInPage(selector, visible) {
let matches;
Expand Down Expand Up @@ -1225,7 +1414,11 @@ async function waitForReadiness(p) {
returnByValue: true,
});
} catch (error) {
return outcome("error", { error: error.message || "evaluate failed" });
// A navigation committing mid-wait destroys the execution context; the
// new document is the one to observe, so poll again until the deadline.
if (Date.now() >= deadline) return outcome("error", { error: error.message || "evaluate failed" });
await sleep(Math.min(100, deadline - Date.now()));
continue;
}
if (res?.exceptionDetails) return outcome("error", { error: res.exceptionDetails.text || "evaluate failed" });
const value = res?.result?.value;
Expand Down Expand Up @@ -1690,6 +1883,7 @@ const handlers = {
const clientId = requireClientId(p);
assertOwned(clientId, p.tabId);
await checkTab(clientId, p.tabId, p.sites, "click in");
noteAgentInput(clientId, p.tabId);
const result = p.index !== undefined ? await clickElement(p.tabId, p.index) : await clickAt(p.tabId, p.x, p.y);
return withState(p, result);
},
Expand All @@ -1705,6 +1899,7 @@ const handlers = {
const clientId = requireClientId(p);
assertOwned(clientId, p.tabId);
await checkTab(clientId, p.tabId, p.sites, "press keys in");
noteAgentInput(clientId, p.tabId);
return withState(p, await pressKey(p.tabId, p.key));
},
screenshot: async (p) => {
Expand All @@ -1723,7 +1918,10 @@ const handlers = {

async function handleCommand(msg, replyPort) {
await ensureStateReady();
const { id, command, params = {} } = msg || {};
const { id, command, params: raw = {} } = msg || {};
// MCP clients may send null for an optional argument they did not set, and
// the hosts pass it through; null means absent everywhere below.
const params = Object.fromEntries(Object.entries(raw ?? {}).filter(([, value]) => value !== null));
try {
const processId = requireClientId(params);
if (command === "close_client_tabs") {
Expand Down Expand Up @@ -1758,8 +1956,11 @@ async function handleCommand(msg, replyPort) {
}
}

chrome.tabs.onCreated.addListener(claimChildTab);

chrome.tabs.onRemoved.addListener((tabId) => {
siteFavicons.delete(tabId);
agentInputAt.delete(tabId);
if (!tabOwner.has(tabId) && !attached.has(tabId)) return;
const clientId = tabOwner.get(tabId);
if (clientId) {
Expand Down
Loading
Loading