Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
148 changes: 131 additions & 17 deletions build/api/index.js

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion build/api/src/data/model/application_error.d.ts
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
export type ApplicationErrorKind = 'configuration' | 'authorization' | 'provider' | 'agent' | 'validation' | 'workflow' | 'unknown';
export type ApplicationErrorCode = 'configuration.invalid' | 'configuration.unsupported' | 'authorization.denied' | 'authorization.credential-invalid' | 'provider.not-found' | 'provider.conflict' | 'provider.rate-limited' | 'provider.unavailable' | 'provider.contract-invalid' | 'agent.policy-rejected' | 'agent.failed' | 'locale.translation-failed' | 'validation.invalid-input' | 'workflow.invalid-event' | 'workflow.stale' | 'workflow.cancelled' | 'workflow.failed' | 'timeout' | 'unexpected';
export type ApplicationErrorCode = 'configuration.invalid' | 'configuration.unsupported' | 'authorization.denied' | 'authorization.credential-invalid' | 'provider.not-found' | 'provider.conflict' | 'provider.rate-limited' | 'provider.unavailable' | 'provider.contract-invalid' | 'agent.policy-rejected' | 'agent.failed' | 'locale.output-invalid' | 'locale.translation-failed' | 'validation.invalid-input' | 'workflow.invalid-event' | 'workflow.stale' | 'workflow.cancelled' | 'workflow.failed' | 'timeout' | 'unexpected';
interface ApplicationErrorMetadata {
readonly kind: ApplicationErrorKind;
readonly retryable: boolean;
Expand Down
319 changes: 246 additions & 73 deletions build/cli/index.js

Large diffs are not rendered by default.

319 changes: 246 additions & 73 deletions build/github_action/index.js

Large diffs are not rendered by default.

9 changes: 9 additions & 0 deletions docs/agents/execution-contract.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,15 @@ All roles deny agent-tool network, approval escalation, MCP, plugins, subagents,

The hard maxima are 15 minutes, 512 KiB of prompt bytes, and 4 MiB of combined process output. Structured tasks declare a JSON Schema. Before Codex starts, Copilot requires a closed object schema in which every property is required, optional values are nullable, nested objects reject additional properties, and arrays define their items. Copilot then validates the returned response locally as well. An incompatible contract, invalid or oversized output, and prose-wrapped JSON are rejected; removed response shapes have no legacy conversion path.

Every product-facing read-only task also receives the canonical effective
`targetLocale` and must return the exact same tag in required `outputLocale`
metadata. This covers Think and help answers, recommendations, progress,
pull-request descriptions, and Bugbot findings. A missing, malformed,
non-canonical, or different output locale is rejected before a comment, review,
description, label, or recommendation state can change. There is no
best-effort mixed-language publication. Technical paths, refs, commands, links,
identifiers, and structured facts remain locale-independent.

Each invocation emits bounded lifecycle observations for plan start, admitted
preflight, and terminal execution. They include only provider, capability,
manifest/version, workspace and output modes, artifact hashes, duration, byte
Expand Down
1 change: 1 addition & 0 deletions docs/agents/failure-policy.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,7 @@ The Action fails closed. It MUST stop rather than silently selecting another run
| CLI timeout | Execution | Stop; do not switch provider |
| Non-zero CLI exit | Execution | Stop; preserve sanitized diagnostics only |
| Invalid agent response | Parsing | Stop; do not publish unsafe findings |
| Missing or mismatched `outputLocale` | Parsing | Stop before product publication or state mutation; retain existing state |
| Verification command fails | Autofix | Do not commit or push |
| Sensitive workspace change | Commit policy | Reject automated commit |

Expand Down
6 changes: 6 additions & 0 deletions docs/bugbot/detection.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -93,6 +93,12 @@ Each child comment is attached to an exact line/range confirmed from the current

The action checks the PR head before invoking the reviewer, before mutation, after finding mutations, and before projecting current state. If a newer revision owns the PR, the old run is recorded as superseded and does not overwrite current-state surfaces.

Finding titles, descriptions, evidence, and suggestions use the effective PR
locale. Paths, line numbers, symbols, code, severities, IDs, commands, and links
remain structured and unchanged. The findings response must echo the exact
canonical locale in `outputLocale`; a missing or different tag rejects the
whole analysis before any finding or review is published.

When the configured agent later reports a finding as **resolved** (e.g. after you or the bot fixed the code), the action:

- **Updates** every existing stored destination so the marker shows `resolved: true`.
Expand Down
22 changes: 14 additions & 8 deletions docs/development/agent-functionality-audit.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -11,14 +11,14 @@ This page is the maintained acceptance contract for Copilot's agent surface. An

| Capability | Trigger | Role | Agent output | Authorized side effect |
| --- | --- | --- | --- | --- |
| Initial help and assistance | Opened question/help issue | `planner` | Schema-validated answer | Sanitized issue comment |
| Recommendations | Issue lifecycle, single action, or CLI | `planner` | Markdown recommendation | Managed issue comment/state |
| Progress analysis | Push, single action, or CLI | `findings` | Schema-validated progress | Issue/PR progress state |
| PR description | PR lifecycle or `/copilot description` | `planner` | Sanitized Markdown | Managed append block by default; explicit replace mode only |
| Think, explain, diagnose, analyze, and conversational commands | Authorized issue/PR comment or CLI | `planner`, `findings`, or `reviewer` by command | Schema-validated answer | Sanitized comment |
| Test plan | `/copilot test-plan` | `tester` | Schema-validated answer | Sanitized comment |
| Translation | Explicitly addressed issue or PR comment | `findings` | Schema-validated language decision and translation | Idempotent comment update |
| Bugbot detection | Push, PR, `/copilot review`, `/copilot recheck`, single action, or CLI | `findings` or `reviewer` | Bounded findings schema | Review/comments plus persisted finding lifecycle |
| Initial help and assistance | Opened question/help issue | `planner` | Locale-tagged schema-validated answer | One sanitized issue reply |
| Recommendations | Issue lifecycle, single action, or CLI | `planner` | Locale-tagged Markdown recommendation or typed unchanged state | One durable plan card/state |
| Progress analysis | Push, single action, or CLI | `findings` | Locale-tagged progress facts and prose | One durable progress card plus issue/PR labels |
| PR description | PR lifecycle or `/copilot description` | `planner` | Locale-tagged sanitized Markdown | Managed append block or configured replace mode |
| Think, explain, diagnose, analyze, and conversational commands | Authorized issue/PR comment or CLI | `planner`, `findings`, or `reviewer` by command | Locale-tagged schema-validated answer | One sanitized reply |
| Test plan | `/copilot test-plan` | `tester` | Locale-tagged schema-validated answer | One sanitized reply |
| Translation | Explicitly addressed issue or PR comment | `findings` | One schema-validated adaptation decision | Internal localized request; source comment remains unchanged |
| Bugbot detection | Push, PR, `/copilot review`, `/copilot recheck`, single action, or CLI | `findings` or `reviewer` | Locale-tagged bounded findings schema | Canonical status/review plus persisted finding lifecycle |
| Bugbot autofix | Explicit `/copilot fix` or authorized intent | `fixer` | Workspace edits | Exact-path commit and token-scoped push after verification |
| General implementation | Explicit `/copilot implement`, authorized intent, or CLI | `fixer` | Workspace edits | Exact-path commit and token-scoped push after verification |

Expand All @@ -30,6 +30,12 @@ available. File mutation requires an authorized actor, a clean workspace, the
expected branch, a bounded exact-path change set, sensitive-path rejection,
isolated verification commands, and a second postflight check before commit.

All six product-facing query families receive the effective surface locale and
must echo it exactly as `outputLocale`. The common parser rejects absent,
malformed, non-canonical, or mismatched metadata before a product-facing side
effect. Fixer responses and intent classifiers are not public prose and remain
separate contracts; trusted application code owns their eventual status UI.

## Runtime matrix

| Runtime | Authentication | Analysis boundary | Fixer boundary | Structured output |
Expand Down
10 changes: 9 additions & 1 deletion docs/development/architecture.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,14 @@ cached language-agent request. A missing or invalid dynamic slice falls back
atomically to English and records only locale, source, descriptor count, and a
stable reason in the Job Summary.

Product-facing agent calls cross a second closed locale boundary. The fixed
inventory is Think, initial help, recommendations, progress, pull-request
descriptions, and Bugbot review. Each call uses the common structured-query
options builder, receives `targetLocale`, requires `outputLocale`, and validates
an exact canonical match before any product mutation. An architecture test
compares the call sites with that inventory, while strict-schema tests require
the locale field in every response contract.

Language adaptation is separate from catalog resolution. Comment admission,
command parsing, and authorization establish trusted command identity before at
most one prose adaptation. The language port has no comment-update capability:
Expand Down Expand Up @@ -101,7 +109,7 @@ reconciliation use cases receive narrow context contracts containing only the
facts they need; they do not depend on the complete `Execution` aggregate.
This keeps pure decisions independently testable and makes new event sources
less likely to couple unrelated capabilities. Application failures use the
closed `ApplicationError` contract: one of 18 semantic codes, a derived
closed `ApplicationError` contract: one of 20 semantic codes, a derived
category and retry decision, impact, recommended action, retained-state
summary, and request correlation UUID. A raw cause is ECMAScript-private,
never serialized or logged, and cannot enter `Result.errors`. CLI output,
Expand Down
7 changes: 6 additions & 1 deletion docs/development/testing.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -85,7 +85,12 @@ all-branch observer workflow is also contract-checked to remain free of agent in

Architecture tests additionally verify that production imports remain acyclic,
application code does not depend on concrete adapters, and pure model/policy
code remains free of runtime, provider, and logging dependencies. Refresh the
code remains free of runtime, provider, and logging dependencies. The
product-facing agent inventory test also requires exactly one guarded query for
Think, help, recommendations, progress, PR descriptions, and Bugbot, and prompt
tests require both `targetLocale` and `outputLocale`. Workflow tests prove that
a mismatched response cannot publish or mutate labels, descriptions,
recommendation state, or findings. Refresh the
local topology before an architecture review with Graphify, and use the
RepoWise index-only health/dead-code reports to distinguish current complexity
from history-based churn signals:
Expand Down
3 changes: 2 additions & 1 deletion docs/pull-requests/ai-description.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,8 @@ description: How the configured agent fills your pull request template from the
- Computes the **diff** between base and head (e.g. `git diff base..head`) to understand what changed.
- Uses the **issue description** (when the PR branch is linked to an issue) as context; otherwise it infers intent from the PR metadata and diff.
3. The agent **fills the template** with a structured description: summary, scope of changes, technical details, how to test, breaking changes, deployment notes, etc., following the same sections and format as your template.
4. The action writes the result to the PR body (prefixed with the issue number for reference).
4. The request carries the effective PR locale (`pull-requests-locale`, otherwise `repository-locale`, otherwise `en-US`). The response must echo that exact canonical tag in `outputLocale`; a mismatch stops before the body changes.
5. The action writes the validated result to the PR body according to the configured ownership mode.

No pre-computed file list or patches are sent from the action; the agent has access to the workspace and computes the diff itself, similar to the [check progress](/single-actions) flow.

Expand Down
2 changes: 1 addition & 1 deletion docs/security-operations/security/prompt-injection.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -40,7 +40,7 @@ only for a structured language decision or translation. The workflow:
- preserves the original comment as escaped inert text;
- neutralizes mentions, slash commands, and HTML-comment markers in the
translated publication;
- writes a versioned marker so a translated comment is not translated again;
- writes a versioned marker only in the bot-owned response so the adaptation context cannot re-enter command routing;
- never treats translated text as an action request.

### Findings and recommendations
Expand Down
2 changes: 1 addition & 1 deletion src/application/errors/__tests__/application_error.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ describe('ApplicationError', () => {
it('exposes the closed semantic contract for every error code', () => {
const codes = Object.keys(APPLICATION_ERROR_METADATA) as ApplicationErrorCode[];

expect(codes).toHaveLength(19);
expect(codes).toHaveLength(20);
for (const code of codes) {
const error = new ApplicationError(code, 'Safe public message.', { correlationId: CORRELATION_ID });
expect(error).toMatchObject({
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,68 @@
import {
PRODUCT_FACING_AGENT_TASKS,
agentOutputLocaleFailureMessage,
productFacingAgentQueryOptions,
validateAgentOutputLocale,
} from '../agent_output_locale_policy';

describe('agent output locale policy', () => {
const schema = {
type: 'object',
properties: { outputLocale: { type: 'string' }, prose: { type: 'string' } },
required: ['outputLocale', 'prose'],
additionalProperties: false,
} as const;

it('keeps the complete product-facing task inventory closed and builds strict query options', () => {
expect(PRODUCT_FACING_AGENT_TASKS).toEqual([
'think',
'answer-issue-help',
'progress',
'recommend-steps',
'pull-request-description',
'bugbot-review',
]);
expect(PRODUCT_FACING_AGENT_TASKS.map(task => productFacingAgentQueryOptions(task, schema).schemaName))
.toEqual([
'think_response',
'answer_issue_help_response',
'progress_response',
'recommend_steps_response',
'pull_request_description_response',
'bugbot_findings',
]);
});

it('rejects a schema that does not require outputLocale before provider use', () => {
expect(() => productFacingAgentQueryOptions('think', { properties: {}, required: [] }))
.toThrow('must require outputLocale');
});

it('accepts exact canonical metadata and returns an immutable copy', () => {
const input = { outputLocale: 'pt-BR', prose: 'ConcluΓ­do' };
const result = validateAgentOutputLocale(input, 'pt-BR');
expect(result).toMatchObject({ kind: 'valid', expectedLocale: 'pt-BR', payload: input });
if (result.kind === 'valid') {
expect(result.payload).not.toBe(input);
expect(Object.isFrozen(result.payload)).toBe(true);
}
});

it.each([
[undefined, 'response-not-object'],
['text', 'response-not-object'],
[{}, 'output-locale-missing'],
[{ outputLocale: '@@' }, 'output-locale-invalid'],
[{ outputLocale: 'en_us' }, 'output-locale-mismatch'],
[{ outputLocale: ' en-US ' }, 'output-locale-mismatch'],
[{ outputLocale: 'fr-FR' }, 'output-locale-mismatch'],
])('rejects invalid locale-tagged output %#', (response, reason) => {
const result = validateAgentOutputLocale(response, 'en-US');
expect(result).toMatchObject({ kind: 'invalid', reason, expectedLocale: 'en-US' });
if (result.kind === 'invalid') {
expect(agentOutputLocaleFailureMessage(result)).toContain(String(reason));
const serialized = JSON.stringify(response);
if (serialized) expect(agentOutputLocaleFailureMessage(result)).not.toContain(serialized);
}
});
});
17 changes: 17 additions & 0 deletions src/application/policies/__tests__/agent_response_schemas.test.ts
Original file line number Diff line number Diff line change
@@ -1,18 +1,35 @@
import {
LANGUAGE_CHECK_RESPONSE_SCHEMA,
PULL_REQUEST_DESCRIPTION_RESPONSE_SCHEMA,
RECOMMEND_STEPS_RESPONSE_SCHEMA,
THINK_RESPONSE_SCHEMA,
TRANSLATION_RESPONSE_SCHEMA,
} from '../agent_response_schemas';
import { assertStrictOutputSchema } from '../agent_execution/strict_output_schema_policy';
import { PROGRESS_RESPONSE_SCHEMA } from '../../usecases/actions/progress_response';
import { BUGBOT_RESPONSE_SCHEMA } from '../../usecases/steps/commit/bugbot/schema';

describe('production agent response schemas', () => {
it.each([
['translation', TRANSLATION_RESPONSE_SCHEMA],
['language check', LANGUAGE_CHECK_RESPONSE_SCHEMA],
['think', THINK_RESPONSE_SCHEMA],
['progress', PROGRESS_RESPONSE_SCHEMA],
['recommend steps', RECOMMEND_STEPS_RESPONSE_SCHEMA],
['pull-request description', PULL_REQUEST_DESCRIPTION_RESPONSE_SCHEMA],
['bugbot review', BUGBOT_RESPONSE_SCHEMA],
])('%s uses the strict native structured-output contract', (_name, schema) => {
expect(() => assertStrictOutputSchema(schema)).not.toThrow();
});

it.each([
THINK_RESPONSE_SCHEMA,
PROGRESS_RESPONSE_SCHEMA,
RECOMMEND_STEPS_RESPONSE_SCHEMA,
PULL_REQUEST_DESCRIPTION_RESPONSE_SCHEMA,
BUGBOT_RESPONSE_SCHEMA,
])('requires exact output-locale metadata for product-facing prose', (schema) => {
expect(schema.required).toContain('outputLocale');
expect(schema.properties.outputLocale).toMatchObject({ type: 'string', maxLength: 255 });
});
});
Loading