Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 6 additions & 6 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -179,14 +179,14 @@ Gotchas:
account.
- Google Contacts search uses a provider-side lazy cache, so a contact created
moments ago may not appear immediately.
- Sending email and creating confirmed calendar events always require approval.
Calendar events with attendees send Google invitations.
- User-requested email and calendar operations run without an extra Eve tool
approval. Calendar events with attendees send Google invitations.

## Link wallet

The root agent mounts `@stripe/link-integrations-eve` in
`agent/extensions/link.ts`. It uses Stripe's bundled wallet tools, skills, and
approval policies, with per-user authorization through the existing Better Auth
`agent/extensions/link.ts`. It uses Stripe's bundled wallet tools and skills,
with per-user authorization through the existing Better Auth
account. It does not use a shared wallet token.

To enable it, register a Link OAuth client and configure `LINK_CLIENT_ID`,
Expand All @@ -211,8 +211,8 @@ and refreshes tokens; disconnection revokes the Link grant before removing it.
Phone sign-in continues to work after disconnecting a wallet.

Wallet access is available in interactive conversations, not scheduled workers
or scheduled result delivery. The extension's default Eve approval for creating
spend requests remains enabled, separately from approval in Link. Its tools can
or scheduled result delivery. Spend requests run without an extra Eve tool
approval and always request purchase approval in Link. Its tools can
return payment credentials into stored Eve tool results; the bundled skills
instruct the agent not to repeat them in chat. Financial-data tools also require
the corresponding Link grant scopes; the default grant requests
Expand Down
2 changes: 1 addition & 1 deletion agent/instructions/content/execution-safety.md
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
# Execution safety

- Require explicit user approval before a purchase, a message to another person or service, a destructive change, or another consequential external action unless the user already authorized that exact action. This does not apply to replying to the current user through `send_message`. For a purchase, authorization covers the merchant, item, quantity, selected option, and approved total or any lower total. Require approval again only if the total increases or another material term changes.
- When an action tool has a native approval step, call the tool with the complete consequential payload as soon as it is ready. The native approval card will show those details and park the action. Never ask for approval in prose first, tell the user to reply with approval words, or duplicate the native approval request with `send_message`.
- Tools run without a separate Eve approval step. Treat the user's explicit delegation of choices within a budget as authorization for the needed tool calls; do not add generic Approve/Cancel prompts or redundant confirmations. Purchases through Link still require approval in Link. Send its exact approval URL as a native link and verify the request is approved before checkout.
6 changes: 3 additions & 3 deletions agent/instructions/content/role/interactive.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@ The main conversation is the control plane. Coordinate the user's work there and
- Browser manipulation, browser inspection, and secret injection belong only to `browser-agent`. The worker receives the same `personal_info` memory as the root and may type those model-readable values with ordinary browser actions. For an opaque saved login, payment method, or legacy vault-only address or contact, it may list safe metadata and pass only the handle and browser session ID to `fill_from_vault`; after injection neither model may inspect or return the filled values.
- When the worker reports that a required saved item is missing, call `request_vault_setup` only for its supported kinds: `login`, `payment`, `address`, or `contact`. Treat a sign-in form with no compatible saved login as a missing vault item, never as human takeover; give the user the returned self-hosted link, never a live-view URL for username or password entry. Request address or contact setup only when the user explicitly asks to save those details for reuse; otherwise use values from the conversation or ask directly. A login setup requires a descriptive `label`, observed `identifierType` (`email`, `phone`, or `username`), exact current `origin`, and fixed `target`; never include the actual identifier or a secret. Other kinds accept only `kind`, optional `label`, and `target`. For an OTP, ask the user for the code in the root conversation and resume the same worker with it. Reserve live view for CAPTCHA, 3-D Secure, passkey or push approval, and other challenges that cannot be answered textually.
- When the user wants to import multiple passwords from Chrome or Google Password Manager, call `request_vault_import` and give them its direct self-hosted importer link. Never ask them to send the CSV or its contents in chat.
- Ask once before filling payment secrets; after approval, fill from the vault and submit without another confirmation. Vault fill, payment-method selection, a merchant review screen, and authentication challenges never require a second price approval.
- Use the user's purchase request and budget to authorize the needed payment-tool calls. Do not ask for another confirmation just to fill saved payment details or call a tool. Stay within the authorized total and material terms; a changed total or material term requires the user's decision. Link purchases still require approval in Link.

# Operating style

Expand All @@ -36,7 +36,7 @@ The main conversation is the control plane. Coordinate the user's work there and
- Prefer the narrowest capable integration: root vault setup for non-secret coordination, connected tools for their supported services, `web_search` for public discovery and current facts, `web_fetch` for reading a known public page, and `browser-agent` only for work that requires browser interaction or browser state.
- Perform public research, source discovery, comparisons, and current-information lookups directly with `web_search`. Never delegate a general-purpose search-only task or use a browser to visit a general-purpose search engine or browse its result pages. A named vertical search or booking product, such as Google Flights, is an interactive browser target when the user's task requires manipulating its controls or continuing through its results. When a known public URL only needs to be read, try `web_fetch` before browser automation.
- Prefer the dedicated `gmail-*`, `calendar-*`, and `contacts-*` tools over browser automation for their supported work. Never ask for Google tokens or credentials in chat. If authorization is required, let the connection surface its sign-in challenge.
- Use exact Gmail message IDs for reversible inbox updates. Before sending email or creating a calendar event, put the exact recipients, content, timing, attendees, and other material fields in the approval-gated tool call so the native approval card can present them. Invoke that tool directly instead of asking for approval in prose first.
- Use exact Gmail message IDs for reversible inbox updates. For a user-requested email or calendar event, pass the exact recipients, content, timing, attendees, and other material fields to the appropriate tool directly. Ask only for missing information or an unresolved decision; do not add a tool approval or a redundant confirmation.
- Keep the user's constraints intact while delegating, comparing alternatives, recovering from failures, and synthesizing results.
- When the conversation reveals a useful next action, offer that exact action with the details already established: book the 7:15 showtime, buy the selected groceries, or submit the prepared form. Offer execution, not a generic "anything else?" or instructions for the user to do it themselves.
- If the user's intent is already clear and the action is authorized, act instead of asking whether to act. Do not add an offer to greetings, simple factual answers, or work you already completed.
Expand All @@ -53,7 +53,7 @@ The main conversation is the control plane. Coordinate the user's work there and
- Use `schedules-create` to create a one-time reminder, recurring job, monitor, or scheduled follow-up. Use `kind: "calendar"` with the user's IANA timezone for wall-clock recurrence, `kind: "interval"` for elapsed intervals, and `kind: "once"` for one future instant. Summarize the exact task in `prompt`. Use `schedules-list` before changing an ambiguous schedule and `schedules-update` to edit, pause, resume, or delete it.
- When the user answers a question previously sent for a scheduled task, call `schedules-answer` with the internal run ID retained in conversation context and their answer. The parked background run continues from the exact point where it asked.
- Use the full native Linq/iMessage surface when it helps. Choose `kind: "message"` for plain text, exact worker artifact references, and HTTPS attachments; text and attachments may be combined. Choose `kind: "link"` with `url` for a standalone native rich link-preview card, or put a URL in message text for a plain tappable URL. `react_to_message` can add or remove any supported Tapback on the current user message: `thumbs_up`, `thumbs_down`, `heart`, `laugh`, `exclamation` (emphasis), or `question`.
- Eve and Linq own read receipts, typing indicators, delivery state, authorization prompts, and approval/input cards. Let those native control-plane features operate normally; do not duplicate them as prose unless the user needs an explanation.
- Eve and Linq own read receipts, typing indicators, delivery state, authorization prompts, and input cards. Let those native control-plane features operate normally; do not duplicate them as prose unless the user needs an explanation.
- After a `send_message` or `react_to_message` call, never repeat or summarize it in assistant text. If the runtime requires terminal assistant text after the last delivery, emit only `DELIVERY_COMPLETE`; ordinary assistant text is not delivered to Linq.
- The worker's structured result is coordinator-facing only. Rewrite it into a concise user-facing response; never imply that the worker spoke to the user.
- Start a background worker without a separate preamble. Once its working receipt arrives, send one short acknowledgment saying what is underway. Treat the receipt as acceptance, not completion.
Expand Down
6 changes: 3 additions & 3 deletions agent/lib/google-workspace/tests/google-workspace.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -34,7 +34,7 @@ describe("Google Workspace", () => {
});
});

it("maps reversible Gmail actions and protects consequential writes", () => {
it("maps reversible Gmail actions and permits requested writes without a tool approval", () => {
expect(gmailUpdateLabels("archive")).toEqual({
addLabelIds: [],
removeLabelIds: ["INBOX"],
Expand All @@ -44,8 +44,8 @@ describe("Google Workspace", () => {
removeLabelIds: [],
});
expect(gmailUpdate.approval).toBeUndefined();
expect(gmailSend.approval).toBeTypeOf("function");
expect(calendarCreateEvent.approval).toBeTypeOf("function");
expect(gmailSend.approval).toBeUndefined();
expect(calendarCreateEvent.approval).toBeUndefined();
});

it("does not treat calendar API errors as availability", () => {
Expand Down
4 changes: 1 addition & 3 deletions agent/tools/calendar.ts
Original file line number Diff line number Diff line change
@@ -1,5 +1,4 @@
import { defineDynamic, defineTool } from "eve/tools";
import { always } from "eve/tools/approval";
import { z } from "zod";
import {
calendarEventSchema,
Expand Down Expand Up @@ -38,9 +37,8 @@ export const calendarCheckAvailability = defineTool({
});

export const calendarCreateEvent = defineTool({
approval: always(),
description:
"Create a confirmed private Google Calendar event. This requires user approval and sends updates to attendees.",
"Create a user-requested confirmed private Google Calendar event and send updates to attendees.",
inputSchema: calendarEventSchema,
async execute(input, ctx) {
return {
Expand Down
4 changes: 1 addition & 3 deletions agent/tools/gmail.ts
Original file line number Diff line number Diff line change
@@ -1,5 +1,4 @@
import { defineDynamic, defineTool } from "eve/tools";
import { always } from "eve/tools/approval";
import { z } from "zod";
import {
GMAIL_UPDATE_ACTIONS,
Expand Down Expand Up @@ -51,9 +50,8 @@ export const gmailUpdate = defineTool({
});

export const gmailSend = defineTool({
approval: always(),
description:
"Send an email from the authenticated user's Gmail account. This requires user approval.",
"Send a user-requested email from the authenticated user's Gmail account.",
inputSchema: gmailSendSchema,
async execute(input, ctx) {
const sent = await sendGmail(ctx, input);
Expand Down
20 changes: 20 additions & 0 deletions agent/tools/link__create_spend_request.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
import { create_spend_request } from "@stripe/link-integrations-eve/tools";
import { linkToolSchemas } from "@stripe/link-sdk/tools";
import { defineTool } from "eve/tools";
import { never } from "eve/tools/approval";
import { z } from "zod";

export default defineTool({
approval: never(),
description:
"Create a Link spend request for the user's requested purchase, with the exact merchant, items, final total, and a stable idempotency key. Link purchase approval is always requested. Send the returned approval_url as a native link and check the same request's status before continuing checkout. Creating this request does not approve a purchase.",
inputSchema: linkToolSchemas.createSpendRequest.safeExtend({
request_approval: z.literal(true).default(true),
}),
execute(input, context) {
return create_spend_request.execute(
{ ...input, request_approval: true },
context
);
},
});
1 change: 1 addition & 0 deletions tests/agent-tool-boundaries.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,7 @@ describe("root and worker capability boundaries", () => {
"calendar.ts",
"contacts.ts",
"gmail.ts",
"link__create_spend_request.ts",
"link__retrieve_spend_request.ts",
"messaging.ts",
"run_browser.ts",
Expand Down
9 changes: 6 additions & 3 deletions tests/agent/instructions.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -46,16 +46,19 @@ describe("agent instructions", () => {
expect(selected?.content).toContain("approval");
});

it("uses native approval cards instead of prose approval loops", async () => {
it("avoids tool approval loops while retaining actual Link purchase approval", async () => {
const resolve = executionSafety.events["turn.started"];
expect(resolve).toBeDefined();
if (!resolve) return;

const selected = await resolve({}, dynamicContext("linq-message"));
expect(selected?.content).toContain(
"Never ask for approval in prose first"
"Tools run without a separate Eve approval step"
);
expect(selected?.content).toContain("native approval card");
expect(selected?.content).toContain(
"Purchases through Link still require approval in Link"
);
expect(selected?.content).not.toContain("native approval card");
});

it("treats personal information as recalled context instead of a read tool", async () => {
Expand Down
66 changes: 66 additions & 0 deletions tests/agent/tools/link-create-spend-request.test.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,66 @@
import { create_spend_request } from "@stripe/link-integrations-eve/tools";
import { linkToolSchemas } from "@stripe/link-sdk/tools";
import { beforeEach, describe, expect, it, vi } from "vitest";
import { z } from "zod";
import create from "@agent/tools/link__create_spend_request";
import { toolContextFor } from "@tests/helpers/tool-context";

const execute = vi.spyOn(create_spend_request, "execute");
const schema = create.inputSchema;
if (!(schema instanceof z.ZodType))
throw new Error("Expected a Zod input schema.");
const input = {
amount: 1000,
context:
"The user requested a one-time purchase from Shop with a maximum total of $10, including tax and shipping, delivered to their saved address.",
idempotency_key: "purchase-1",
merchant_name: "Shop",
merchant_url: "https://shop.example/checkout",
};

beforeEach(() => vi.clearAllMocks());

describe("Link spend request approval boundary", () => {
it("does not require an Eve tool approval", async () => {
const approval = create.approval;
if (!approval) throw new Error("Expected a tool approval policy.");
const policy = "request" in approval ? approval.request : approval;
expect(
await policy({
...toolContextFor(),
approvedTools: new Set(),
toolInput: {
...linkToolSchemas.createSpendRequest.parse(input),
request_approval: true,
},
})
).toBe("not-applicable");
});

it("always requests actual purchase approval from Link", async () => {
const request = {
id: "spr_1",
status: "pending_approval",
created_at: "2026-10-02",
updated_at: "2026-10-02",
approval_url: "https://app.link.com/approve/spr_1",
};
execute.mockResolvedValue(request);
expect(schema.safeParse(input).success).toBe(true);
const parsed = linkToolSchemas.createSpendRequest.parse(input);
const context = toolContextFor();
await expect(
create.execute({ ...parsed, request_approval: true }, context)
).resolves.toEqual(request);
expect(execute).toHaveBeenCalledExactlyOnceWith(
{ ...parsed, request_approval: true },
context
);
expect(
schema.safeParse({ ...input, request_approval: false }).success
).toBe(false);
expect(
schema.safeParse({ ...input, merchant_name: undefined }).success
).toBe(false);
});
});
Loading