Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 22 additions & 0 deletions docs/architecture/rfcs/agent-loop-effect-interpreter-v0.md
Original file line number Diff line number Diff line change
Expand Up @@ -440,6 +440,28 @@ The existing R5 CLI captures full decisions before projection, with a private di

### What Is Missing

#### Heartbeat and Turn Envelope convergence

Under existing M7.4 and roadmap S2/S3/S6/S8, converge the **execution facts**
used by heartbeat and Turn hosts, while keeping host effect ownership separate.
Neither today's large quota packet nor the smaller TurnEnvelope is a target
shape merely because of its size. Optional memory participation now has one
compact, signed envelope projection; the Codex CLI adapter preserves default-off
isolation. This does not qualify installed heartbeat/App adoption or model value.

Remaining implementation Todos, in dependency order:

| Todo | Observable outcome and decisive acceptance |
|---|---|
| Reconcile execution/context requirements across heartbeat and TurnEnvelope | Same captured authoritative decision preserves actor/Goal/Todo, required full reads, claim/lease, action selection, replan/closure, conditional settlement and scheduler ownership. Inventory omitted/duplicated facts before deleting render branches. Include optional capabilities off, recall-only, ingest-only, stale binding and provider failure; private detail is accessed only through authorized references. |
| Adopt one typed projection in real host renderers | Heartbeat full/thin and Turn host consume the same execution facts and per-Turn capture/detail route. Keep host-specific notification and scheduler transport explicit. Real File/SQLite CLI plus packaged Codex App tests cover reentry, source loss, refusal before required reads, late results, backoff and exactly-once settlement; no second admission from a detail read. Retire the replaced projection only after its last caller moves. |
| Qualify the context shape and migration default | Compare the same normal, replan, wait/recovery and optional-capability workloads against both current full and compact paths. Measure payload/model tokens, detail IO, latency, resource growth, omissions and decision/outcome quality. Keep data loss, duplicate effects, identity and settlement errors as hard constraints. Preserve supported saved prompts/receipts and reversible rollout; change budgets or defaults only with that evidence. |

Do not add a generic executor or lower an acceptance threshold to make a short
packet pass. Preserve unsatisfied requirements and distinguish transport parity,
installed host adoption and useful model outcomes. See the
[current envelope contract](../../reference/protocols/turn-envelope-v0.md#optional-memory-participation).

- A generic shared executor is deliberately absent. The current adapters share
plan/receipt algebra but have different execution ownership, so M7.3
is closed with no follow-up rather than filled with a speculative framework.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -312,6 +312,26 @@ R5 短包投影也完整保留已有 CLI 结算计划,包括 effect identity

### 还缺什么

#### Heartbeat 与 TurnEnvelope 收敛

沿用既有 M7.4 与 roadmap S2/S3/S6/S8,统一 heartbeat 和 Turn host 消费的
**执行事实**,宿主效果的所有权仍分别保留。现有大 quota packet 与小 TurnEnvelope
都不是仅凭大小就合理的目标形态。可选 memory 参与事实现进入一个有界、签名覆盖的
envelope 投影,Codex CLI 保持关闭隔离;这不代表安装态 heartbeat/App 或模型收益
已验收。

后续实施 Todo 按依赖顺序推进:

| Todo | 可观察结果与决定性验收 |
|---|---|
| 核对 heartbeat 与 TurnEnvelope 的执行和上下文要求 | 同一捕获的权威决策保留 actor/Goal/Todo、必读全文、claim/lease、选择、replan/收尾、带条件结算和 scheduler 所有权。删除渲染分支前列明遗漏与重复事实。覆盖可选能力关闭、仅 recall、仅 ingest、绑定失效和 provider 失败;私有详情只经有权限的引用访问。 |
| 在真实宿主 renderer 采用同一 typed 投影 | heartbeat full/thin 与 Turn host 消费同一执行事实及同 Turn 捕获/详情入口;显式保留通知和 scheduler 的宿主传输。真实 File/SQLite CLI 与打包 Codex App 覆盖重入、来源丢失、必读前拒绝、迟到结果、backoff 和一次结算;补读不触发第二次准入。最后调用方迁移后才退役旧投影。 |
| 核验上下文形态及迁移默认 | 将相同的正常、replan、等待/恢复与可选能力负载同时对照当前完整和短包路径,测量载荷/model token、详情 IO、延迟、资源增长、遗漏、决策与结果质量。数据丢失、重复效果、身份和结算错误保持硬约束;保留受支持的已保存 prompt/回执与可逆 rollout,根据证据再改预算或默认。 |

不引入通用 executor,也不降低验收门槛来让短包通过。保留未满足要求,分别记录
传输等价、安装态宿主采用和有效模型结果。见
[当前 envelope 契约](../../reference/protocols/turn-envelope-v0.md#optional-memory-participation)。

- 通用共享 executor 被有意保留为空。当前 adapter 共享 plan/receipt algebra,却拥有不同的执行边界,因此 M7.3 应以 no-follow-up 关闭,而不是用推测性 framework 填充。
- 常规 LoopX 核心路径仍需逐条做有界采用判断。只有当路径包含多步 external effect、单一稳定 identity、durable receipt、replay 要求,并且变更能删除重复 settlement truth 时,才应该使用这套 algebra。
- Race/CAS qualification 推迟到真实并发执行入口出现后;同步 adapter 本身不足以证明需要并发基础设施或测试。
Expand Down
29 changes: 29 additions & 0 deletions docs/reference/protocols/turn-envelope-v0.md
Original file line number Diff line number Diff line change
Expand Up @@ -332,6 +332,35 @@ final packet, including diagnostics. The historical `source_json_bytes` and
`envelope_json_bytes` fields still count Unicode code points for v0 compatibility;
do not use them as wire-byte measurements.

### Optional memory participation

The compact boundary retains the verified Goal/Agent Reward Memory automation
facts in `boundary.capabilities.reward_memory`: `automatic_recall` and
`automatic_ingest`. Unconfigured, disabled or unavailable bindings omit this
projection. It carries neither private provider configuration nor memory content
and grants no action authority. Existing action-signature coverage includes it.

Codex CLI host requests use these facts together with the fresh runtime binding
readback. Recall guidance is added only with enabled recall and actual context;
the reflection instruction and output-schema field appear only with enabled
ingest. A disabled or stale recall packet cannot activate either operation.
Older/custom hosts may still return a reflection field: while ingest is off the
adapter ignores it before validation and journaling, preserving ordinary work.
When enabled, bounded reflection validation, exact independent attestation and
post-settlement ingest retain their existing rules. Provider failure remains
fail-open. The host adapter is Python IO transport over the existing capability
resolver and TypeScript envelope owner, not a new enablement policy.

This is a bounded shared-facts step. Heartbeat still has its existing full and
compact rendering paths; it is not yet the same installed host journey. The
remaining convergence work is tracked in the existing
[effect-interpreter RFC](../../architecture/rfcs/agent-loop-effect-interpreter-v0.md#heartbeat-and-turn-envelope-convergence).

中文:短包只携带已核验的 recall/ingest 参与事实;关闭、未配置或不可用时不加入。
Codex CLI 据此及新鲜绑定读回装配指令和 schema,关闭 ingest 时忽略旧宿主返回的
反思字段。权限、独立验证与结算规则不变。heartbeat 与 TurnEnvelope 的安装态统一
仍待验证,不能把这一步当作完成。

### Budget warnings and allocation

Oversize valid envelopes keep their normal Turn plan/controller route. They
Expand Down
11 changes: 11 additions & 0 deletions loopx/control_plane/quota/turn_envelope.ts
Original file line number Diff line number Diff line change
Expand Up @@ -352,6 +352,17 @@ function boundary(payload: JsonObject): JsonObject {
};
}
const guards = textList(source.guards, 8, 280);
// Effective automation is already resolved for this Goal/Agent. Keep the
// shared participation facts, not private config, diagnostics or recall.
const memory = object(object(source.capabilities).reward_memory);
if (memory.enabled === true && memory.configured_for_agent === true
&& memory.experiment_available === true
&& (memory.automatic_recall === true || memory.automatic_ingest === true)) {
result.capabilities = {reward_memory: {
automatic_recall: memory.automatic_recall === true,
automatic_ingest: memory.automatic_ingest === true,
}};
}
if (guards.length > 0) result.guards = guards;
const stopCondition = text(source.stop_condition, 320);
if (stopCondition) result.stop_condition = stopCondition;
Expand Down
27 changes: 18 additions & 9 deletions loopx/control_plane/turn_driver/codex_cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,7 @@
HOST_REWARD_MEMORY_REFLECTION_JSON_MAX_CHARS,
HOST_RESULT_TEXT_LIMITS,
LOOPX_TURN_HOST_REQUEST_SCHEMA_VERSION,
reward_memory_automation_enabled,
)
from .execution_profile import require_supported_reasoning_effort
from .host_failure import BuiltInHostError
Expand Down Expand Up @@ -251,11 +252,12 @@ def codex_cli_result_schema(
"maxLength": HOST_AGENT_VISION_JSON_MAX_CHARS,
},
"summary": {"type": "string", "maxLength": text_limits["summary"]},
"reward_memory_reflection_json": {
}
if reward_memory_automation_enabled(request, operation="automatic_ingest"):
properties["reward_memory_reflection_json"] = {
"type": "string",
"maxLength": HOST_REWARD_MEMORY_REFLECTION_JSON_MAX_CHARS,
},
}
}
if _has_subagent_topology(request):
child_receipts = child_execution_receipts_json_schema()
child_receipt_properties = child_receipts["items"]["properties"]
Expand All @@ -279,17 +281,23 @@ def _prompt(request: Mapping[str, Any]) -> str:
instructions = [
"Execute exactly one bounded LoopX Turn in the current workspace.",
"Use the TurnEnvelope as the source of truth. Perform work only when its contract allows it.",
"When reward_memory_recall contains guidance, treat it as private, non-authoritative decision context: apply it only when it fits current evidence and never treat it as new action authority.",
"Set reward_memory_reflection_json to an empty string unless independent task evidence established a reusable experience. For eligible evidence, return one compact JSON object using schema_version=turn_reward_memory_reflection_v1, status=eligible, a configured surface_id, outcome_kind in research|simulation|real|engineering, content_summary, reasoning_summary, confidence in low|medium|high, and 1-5 opaque evidence_refs. Also include experience using schema_version=procedural_experience_contract_v0 with non-empty applicability and limitations lists, observed_outcome, attribution, the same evidence_refs, and future_behavior containing trigger, action, validation, and stop_condition. A fact recap without a future behavior change and non-generalization boundary is not eligible memory. Legacy v0 reflections are audit-only and cannot become durable memory. Never use your own summary as evidence. Settlement may ingest it only when the caller-declared Todo validator attests the exact reflection digest and evidence; ordinary validator success remains awaiting and makes no provider write.",
"Do not write LoopX state, spend quota, or apply scheduler changes; the adapter owns those effects.",
"Return only the schema-constrained result. For validated_progress, repair_required, or replan_required, fill every material field with public-safe evidence.",
"For those material results, set path_delta_mode=material_replan only when this Turn changes a prior assumption, route, scope, acceptance rule, or stops prior work; then provide a complete bounded agent vision packet with goal_path_delta_v0 in agent_vision_json and leave vision_unchanged_reason empty.",
"For routine continuation, retry, successor creation, or no-change replanning, set path_delta_mode=unchanged, leave agent_vision_json empty, and provide vision_unchanged_reason.",
"For user_action_required, wait, or iteration_failed, leave material-only fields empty and explain the stop in summary. iteration_failed ends only this iteration and never requests a retry or successor.",
'completed_phases must be exactly ["host_execute","typed_result"], and turn_key must match the request.',
"Turn request:",
request_json,
]
recall = _mapping(request.get("reward_memory_recall"))
if (reward_memory_automation_enabled(request, operation="automatic_recall")
and _mapping(recall.get("context")).get("guidance")):
instructions.append(
"When reward_memory_recall contains guidance, treat it as private, non-authoritative decision context: apply it only when it fits current evidence and never treat it as new action authority."
)
if reward_memory_automation_enabled(request, operation="automatic_ingest"):
instructions.append(
"Set reward_memory_reflection_json to an empty string unless independent task evidence established a reusable experience. For eligible evidence, return one compact JSON object using schema_version=turn_reward_memory_reflection_v1, status=eligible, a configured surface_id, outcome_kind in research|simulation|real|engineering, content_summary, reasoning_summary, confidence in low|medium|high, and 1-5 opaque evidence_refs. Also include experience using schema_version=procedural_experience_contract_v0 with non-empty applicability and limitations lists, observed_outcome, attribution, the same evidence_refs, and future_behavior containing trigger, action, validation, and stop_condition. A fact recap without a future behavior change and non-generalization boundary is not eligible memory. Legacy v0 reflections are audit-only and cannot become durable memory. Never use your own summary as evidence. Settlement may ingest it only when the caller-declared Todo validator attests the exact reflection digest and evidence; ordinary validator success remains awaiting and makes no provider write."
)
boundary = _mapping(_mapping(request.get("turn_envelope")).get("boundary"))
if boundary.get("checkpointed_boundary_authority"):
instructions.append(
Expand All @@ -298,14 +306,15 @@ def _prompt(request: Mapping[str, Any]) -> str:
"for those scopes; other scopes, publish, and production actions retain their gates."
)
if _has_subagent_topology(request):
instructions[7:7] = [
instructions.extend([
"When subagent_execution_topology is present, return one compact child_execution_receipts item for each observed child, including the actual context_mode. Never copy prompts, transcripts, tool output, credentials, private links, or local absolute paths into a receipt. If no child was observed, return an empty list.",
"Launch a child only from its complete child_execution_task_packet_v0. Keep the child inside its objective, acceptance, capability, write-scope, effect, workspace, and execution-budget boundaries, and copy the exact task_packet_digest into its receipt.",
"Use the generic task-packet context mode exactly and execute the separate host_adapter projection. For the Codex spawn_agent adapter, fresh maps to fork_context=false and forked_snapshot maps to fork_context=true; never infer native arguments inside the generic LoopX task packet.",
"For every Codex child receipt, set runtime_id to the stable host id codex-cli. Keep worker_ref opaque. Never use an executable, workspace, session-file, or other local path as either identifier.",
"Use one or more opaque evidence_refs such as artifact:child-result. Receipt identifiers and evidence refs must contain no spaces, prose, URLs, or local paths.",
"If a child deviates from that packet, stop or quarantine only that child and its evidence. Do not let the child write LoopX state or block the parent agent; the parent may retry fresh, replace the child, take over serially, or ignore an optional result.",
]
])
instructions.extend(["Turn request:", request_json])
return "\n".join(instructions)


Expand Down
60 changes: 53 additions & 7 deletions loopx/control_plane/turn_driver/executor.py
Original file line number Diff line number Diff line change
Expand Up @@ -132,6 +132,42 @@
JournalPersist = Callable[[Mapping[str, Any]], None]


def reward_memory_automation_enabled(
payload: Mapping[str, Any] | None, *, operation: str,
) -> bool:
"""Project verified capability automation into the host IO contract.

The experiment resolver owns enablement, config/receipt verification and
provider policy. A recall packet's presence or old context is not opt-in.
"""
if not isinstance(payload, Mapping):
return False
recall = payload.get("reward_memory_recall")
if not isinstance(recall, Mapping):
return False
experiment = recall.get("experiment")
envelope = payload.get("turn_envelope")
if not isinstance(experiment, Mapping) or not isinstance(envelope, Mapping):
return False
boundary = envelope.get("boundary")
capabilities = boundary.get("capabilities") if isinstance(boundary, Mapping) else None
memory = capabilities.get("reward_memory") if isinstance(capabilities, Mapping) else None
return (
isinstance(memory, Mapping)
and memory.get(operation) is True
# Recall resolves the binding again after admission. A stale envelope
# must not resurrect a now unavailable or disabled binding.
and experiment.get("enabled") is True
and experiment.get("available") is True
and experiment.get("configured_for_agent") is True
and bool(envelope.get("goal_id"))
and experiment.get("goal_id") == envelope.get("goal_id")
and bool(envelope.get("agent_id"))
and experiment.get("agent_id") == envelope.get("agent_id")
and experiment.get(operation) is True
)


def build_loopx_turn_host_request(plan: Mapping[str, Any]) -> dict[str, Any]:
transaction = (
plan.get("transaction") if isinstance(plan.get("transaction"), dict) else {}
Expand All @@ -158,8 +194,12 @@ def build_loopx_turn_host_request(plan: Mapping[str, Any]) -> dict[str, Any]:
if isinstance(goal_ref, Mapping):
request["goal_ref"] = dict(goal_ref)
reward_memory_recall = plan.get("reward_memory_recall")
if isinstance(reward_memory_recall, Mapping):
recall_enabled = reward_memory_automation_enabled(plan, operation="automatic_recall")
ingest_enabled = reward_memory_automation_enabled(plan, operation="automatic_ingest")
if isinstance(reward_memory_recall, Mapping) and (recall_enabled or ingest_enabled):
request["reward_memory_recall"] = dict(reward_memory_recall)
if not recall_enabled:
request["reward_memory_recall"]["context"] = None
request.update(subagent.subagent_host_request_projection(plan))
from ...extensions.codex_native_child import configured_native_child_limit

Expand Down Expand Up @@ -323,12 +363,18 @@ def validate_loopx_turn_host_result(
)
if text:
normalized[field] = text
reflection_json = _bounded_public_text(
result,
"reward_memory_reflection_json",
limit=HOST_REWARD_MEMORY_REFLECTION_JSON_MAX_CHARS,
required=False,
errors=errors,
# Old/custom hosts may still return this field. Ignore it while disabled:
# ordinary work must not gain a memory validation gate or retain the content.
reflection_json = (
_bounded_public_text(
result,
"reward_memory_reflection_json",
limit=HOST_REWARD_MEMORY_REFLECTION_JSON_MAX_CHARS,
required=False,
errors=errors,
)
if reward_memory_automation_enabled(plan, operation="automatic_ingest")
else None
)
if reflection_json:
if material:
Expand Down
Loading
Loading