diff --git a/docs/architecture/rfcs/agent-loop-effect-interpreter-v0.md b/docs/architecture/rfcs/agent-loop-effect-interpreter-v0.md index afc9d92444..d240053cf9 100644 --- a/docs/architecture/rfcs/agent-loop-effect-interpreter-v0.md +++ b/docs/architecture/rfcs/agent-loop-effect-interpreter-v0.md @@ -440,6 +440,28 @@ The existing R5 CLI captures full decisions before projection, with a private di ### What Is Missing +#### Heartbeat and Turn Envelope convergence + +Under existing M7.4 and roadmap S2/S3/S6/S8, converge the **execution facts** +used by heartbeat and Turn hosts, while keeping host effect ownership separate. +Neither today's large quota packet nor the smaller TurnEnvelope is a target +shape merely because of its size. Optional memory participation now has one +compact, signed envelope projection; the Codex CLI adapter preserves default-off +isolation. This does not qualify installed heartbeat/App adoption or model value. + +Remaining implementation Todos, in dependency order: + +| Todo | Observable outcome and decisive acceptance | +|---|---| +| Reconcile execution/context requirements across heartbeat and TurnEnvelope | Same captured authoritative decision preserves actor/Goal/Todo, required full reads, claim/lease, action selection, replan/closure, conditional settlement and scheduler ownership. Inventory omitted/duplicated facts before deleting render branches. Include optional capabilities off, recall-only, ingest-only, stale binding and provider failure; private detail is accessed only through authorized references. | +| Adopt one typed projection in real host renderers | Heartbeat full/thin and Turn host consume the same execution facts and per-Turn capture/detail route. Keep host-specific notification and scheduler transport explicit. Real File/SQLite CLI plus packaged Codex App tests cover reentry, source loss, refusal before required reads, late results, backoff and exactly-once settlement; no second admission from a detail read. Retire the replaced projection only after its last caller moves. | +| Qualify the context shape and migration default | Compare the same normal, replan, wait/recovery and optional-capability workloads against both current full and compact paths. Measure payload/model tokens, detail IO, latency, resource growth, omissions and decision/outcome quality. Keep data loss, duplicate effects, identity and settlement errors as hard constraints. Preserve supported saved prompts/receipts and reversible rollout; change budgets or defaults only with that evidence. | + +Do not add a generic executor or lower an acceptance threshold to make a short +packet pass. Preserve unsatisfied requirements and distinguish transport parity, +installed host adoption and useful model outcomes. See the +[current envelope contract](../../reference/protocols/turn-envelope-v0.md#optional-memory-participation). + - A generic shared executor is deliberately absent. The current adapters share plan/receipt algebra but have different execution ownership, so M7.3 is closed with no follow-up rather than filled with a speculative framework. diff --git a/docs/architecture/rfcs/agent-loop-effect-interpreter-v0.zh-CN.md b/docs/architecture/rfcs/agent-loop-effect-interpreter-v0.zh-CN.md index 49b5df1298..6d5cc94436 100644 --- a/docs/architecture/rfcs/agent-loop-effect-interpreter-v0.zh-CN.md +++ b/docs/architecture/rfcs/agent-loop-effect-interpreter-v0.zh-CN.md @@ -312,6 +312,26 @@ R5 短包投影也完整保留已有 CLI 结算计划,包括 effect identity ### 还缺什么 +#### Heartbeat 与 TurnEnvelope 收敛 + +沿用既有 M7.4 与 roadmap S2/S3/S6/S8,统一 heartbeat 和 Turn host 消费的 +**执行事实**,宿主效果的所有权仍分别保留。现有大 quota packet 与小 TurnEnvelope +都不是仅凭大小就合理的目标形态。可选 memory 参与事实现进入一个有界、签名覆盖的 +envelope 投影,Codex CLI 保持关闭隔离;这不代表安装态 heartbeat/App 或模型收益 +已验收。 + +后续实施 Todo 按依赖顺序推进: + +| Todo | 可观察结果与决定性验收 | +|---|---| +| 核对 heartbeat 与 TurnEnvelope 的执行和上下文要求 | 同一捕获的权威决策保留 actor/Goal/Todo、必读全文、claim/lease、选择、replan/收尾、带条件结算和 scheduler 所有权。删除渲染分支前列明遗漏与重复事实。覆盖可选能力关闭、仅 recall、仅 ingest、绑定失效和 provider 失败;私有详情只经有权限的引用访问。 | +| 在真实宿主 renderer 采用同一 typed 投影 | heartbeat full/thin 与 Turn host 消费同一执行事实及同 Turn 捕获/详情入口;显式保留通知和 scheduler 的宿主传输。真实 File/SQLite CLI 与打包 Codex App 覆盖重入、来源丢失、必读前拒绝、迟到结果、backoff 和一次结算;补读不触发第二次准入。最后调用方迁移后才退役旧投影。 | +| 核验上下文形态及迁移默认 | 将相同的正常、replan、等待/恢复与可选能力负载同时对照当前完整和短包路径,测量载荷/model token、详情 IO、延迟、资源增长、遗漏、决策与结果质量。数据丢失、重复效果、身份和结算错误保持硬约束;保留受支持的已保存 prompt/回执与可逆 rollout,根据证据再改预算或默认。 | + +不引入通用 executor,也不降低验收门槛来让短包通过。保留未满足要求,分别记录 +传输等价、安装态宿主采用和有效模型结果。见 +[当前 envelope 契约](../../reference/protocols/turn-envelope-v0.md#optional-memory-participation)。 + - 通用共享 executor 被有意保留为空。当前 adapter 共享 plan/receipt algebra,却拥有不同的执行边界,因此 M7.3 应以 no-follow-up 关闭,而不是用推测性 framework 填充。 - 常规 LoopX 核心路径仍需逐条做有界采用判断。只有当路径包含多步 external effect、单一稳定 identity、durable receipt、replay 要求,并且变更能删除重复 settlement truth 时,才应该使用这套 algebra。 - Race/CAS qualification 推迟到真实并发执行入口出现后;同步 adapter 本身不足以证明需要并发基础设施或测试。 diff --git a/docs/reference/protocols/turn-envelope-v0.md b/docs/reference/protocols/turn-envelope-v0.md index fde412f43b..42b78a5920 100644 --- a/docs/reference/protocols/turn-envelope-v0.md +++ b/docs/reference/protocols/turn-envelope-v0.md @@ -332,6 +332,35 @@ final packet, including diagnostics. The historical `source_json_bytes` and `envelope_json_bytes` fields still count Unicode code points for v0 compatibility; do not use them as wire-byte measurements. +### Optional memory participation + +The compact boundary retains the verified Goal/Agent Reward Memory automation +facts in `boundary.capabilities.reward_memory`: `automatic_recall` and +`automatic_ingest`. Unconfigured, disabled or unavailable bindings omit this +projection. It carries neither private provider configuration nor memory content +and grants no action authority. Existing action-signature coverage includes it. + +Codex CLI host requests use these facts together with the fresh runtime binding +readback. Recall guidance is added only with enabled recall and actual context; +the reflection instruction and output-schema field appear only with enabled +ingest. A disabled or stale recall packet cannot activate either operation. +Older/custom hosts may still return a reflection field: while ingest is off the +adapter ignores it before validation and journaling, preserving ordinary work. +When enabled, bounded reflection validation, exact independent attestation and +post-settlement ingest retain their existing rules. Provider failure remains +fail-open. The host adapter is Python IO transport over the existing capability +resolver and TypeScript envelope owner, not a new enablement policy. + +This is a bounded shared-facts step. Heartbeat still has its existing full and +compact rendering paths; it is not yet the same installed host journey. The +remaining convergence work is tracked in the existing +[effect-interpreter RFC](../../architecture/rfcs/agent-loop-effect-interpreter-v0.md#heartbeat-and-turn-envelope-convergence). + +中文:短包只携带已核验的 recall/ingest 参与事实;关闭、未配置或不可用时不加入。 +Codex CLI 据此及新鲜绑定读回装配指令和 schema,关闭 ingest 时忽略旧宿主返回的 +反思字段。权限、独立验证与结算规则不变。heartbeat 与 TurnEnvelope 的安装态统一 +仍待验证,不能把这一步当作完成。 + ### Budget warnings and allocation Oversize valid envelopes keep their normal Turn plan/controller route. They diff --git a/loopx/control_plane/quota/turn_envelope.ts b/loopx/control_plane/quota/turn_envelope.ts index 96bb3d7242..cfdbc752fa 100644 --- a/loopx/control_plane/quota/turn_envelope.ts +++ b/loopx/control_plane/quota/turn_envelope.ts @@ -352,6 +352,17 @@ function boundary(payload: JsonObject): JsonObject { }; } const guards = textList(source.guards, 8, 280); + // Effective automation is already resolved for this Goal/Agent. Keep the + // shared participation facts, not private config, diagnostics or recall. + const memory = object(object(source.capabilities).reward_memory); + if (memory.enabled === true && memory.configured_for_agent === true + && memory.experiment_available === true + && (memory.automatic_recall === true || memory.automatic_ingest === true)) { + result.capabilities = {reward_memory: { + automatic_recall: memory.automatic_recall === true, + automatic_ingest: memory.automatic_ingest === true, + }}; + } if (guards.length > 0) result.guards = guards; const stopCondition = text(source.stop_condition, 320); if (stopCondition) result.stop_condition = stopCondition; diff --git a/loopx/control_plane/turn_driver/codex_cli.py b/loopx/control_plane/turn_driver/codex_cli.py index 5442e89dd5..0e6302ee24 100644 --- a/loopx/control_plane/turn_driver/codex_cli.py +++ b/loopx/control_plane/turn_driver/codex_cli.py @@ -35,6 +35,7 @@ HOST_REWARD_MEMORY_REFLECTION_JSON_MAX_CHARS, HOST_RESULT_TEXT_LIMITS, LOOPX_TURN_HOST_REQUEST_SCHEMA_VERSION, + reward_memory_automation_enabled, ) from .execution_profile import require_supported_reasoning_effort from .host_failure import BuiltInHostError @@ -251,11 +252,12 @@ def codex_cli_result_schema( "maxLength": HOST_AGENT_VISION_JSON_MAX_CHARS, }, "summary": {"type": "string", "maxLength": text_limits["summary"]}, - "reward_memory_reflection_json": { + } + if reward_memory_automation_enabled(request, operation="automatic_ingest"): + properties["reward_memory_reflection_json"] = { "type": "string", "maxLength": HOST_REWARD_MEMORY_REFLECTION_JSON_MAX_CHARS, - }, - } + } if _has_subagent_topology(request): child_receipts = child_execution_receipts_json_schema() child_receipt_properties = child_receipts["items"]["properties"] @@ -279,17 +281,23 @@ def _prompt(request: Mapping[str, Any]) -> str: instructions = [ "Execute exactly one bounded LoopX Turn in the current workspace.", "Use the TurnEnvelope as the source of truth. Perform work only when its contract allows it.", - "When reward_memory_recall contains guidance, treat it as private, non-authoritative decision context: apply it only when it fits current evidence and never treat it as new action authority.", - "Set reward_memory_reflection_json to an empty string unless independent task evidence established a reusable experience. For eligible evidence, return one compact JSON object using schema_version=turn_reward_memory_reflection_v1, status=eligible, a configured surface_id, outcome_kind in research|simulation|real|engineering, content_summary, reasoning_summary, confidence in low|medium|high, and 1-5 opaque evidence_refs. Also include experience using schema_version=procedural_experience_contract_v0 with non-empty applicability and limitations lists, observed_outcome, attribution, the same evidence_refs, and future_behavior containing trigger, action, validation, and stop_condition. A fact recap without a future behavior change and non-generalization boundary is not eligible memory. Legacy v0 reflections are audit-only and cannot become durable memory. Never use your own summary as evidence. Settlement may ingest it only when the caller-declared Todo validator attests the exact reflection digest and evidence; ordinary validator success remains awaiting and makes no provider write.", "Do not write LoopX state, spend quota, or apply scheduler changes; the adapter owns those effects.", "Return only the schema-constrained result. For validated_progress, repair_required, or replan_required, fill every material field with public-safe evidence.", "For those material results, set path_delta_mode=material_replan only when this Turn changes a prior assumption, route, scope, acceptance rule, or stops prior work; then provide a complete bounded agent vision packet with goal_path_delta_v0 in agent_vision_json and leave vision_unchanged_reason empty.", "For routine continuation, retry, successor creation, or no-change replanning, set path_delta_mode=unchanged, leave agent_vision_json empty, and provide vision_unchanged_reason.", "For user_action_required, wait, or iteration_failed, leave material-only fields empty and explain the stop in summary. iteration_failed ends only this iteration and never requests a retry or successor.", 'completed_phases must be exactly ["host_execute","typed_result"], and turn_key must match the request.', - "Turn request:", - request_json, ] + recall = _mapping(request.get("reward_memory_recall")) + if (reward_memory_automation_enabled(request, operation="automatic_recall") + and _mapping(recall.get("context")).get("guidance")): + instructions.append( + "When reward_memory_recall contains guidance, treat it as private, non-authoritative decision context: apply it only when it fits current evidence and never treat it as new action authority." + ) + if reward_memory_automation_enabled(request, operation="automatic_ingest"): + instructions.append( + "Set reward_memory_reflection_json to an empty string unless independent task evidence established a reusable experience. For eligible evidence, return one compact JSON object using schema_version=turn_reward_memory_reflection_v1, status=eligible, a configured surface_id, outcome_kind in research|simulation|real|engineering, content_summary, reasoning_summary, confidence in low|medium|high, and 1-5 opaque evidence_refs. Also include experience using schema_version=procedural_experience_contract_v0 with non-empty applicability and limitations lists, observed_outcome, attribution, the same evidence_refs, and future_behavior containing trigger, action, validation, and stop_condition. A fact recap without a future behavior change and non-generalization boundary is not eligible memory. Legacy v0 reflections are audit-only and cannot become durable memory. Never use your own summary as evidence. Settlement may ingest it only when the caller-declared Todo validator attests the exact reflection digest and evidence; ordinary validator success remains awaiting and makes no provider write." + ) boundary = _mapping(_mapping(request.get("turn_envelope")).get("boundary")) if boundary.get("checkpointed_boundary_authority"): instructions.append( @@ -298,14 +306,15 @@ def _prompt(request: Mapping[str, Any]) -> str: "for those scopes; other scopes, publish, and production actions retain their gates." ) if _has_subagent_topology(request): - instructions[7:7] = [ + instructions.extend([ "When subagent_execution_topology is present, return one compact child_execution_receipts item for each observed child, including the actual context_mode. Never copy prompts, transcripts, tool output, credentials, private links, or local absolute paths into a receipt. If no child was observed, return an empty list.", "Launch a child only from its complete child_execution_task_packet_v0. Keep the child inside its objective, acceptance, capability, write-scope, effect, workspace, and execution-budget boundaries, and copy the exact task_packet_digest into its receipt.", "Use the generic task-packet context mode exactly and execute the separate host_adapter projection. For the Codex spawn_agent adapter, fresh maps to fork_context=false and forked_snapshot maps to fork_context=true; never infer native arguments inside the generic LoopX task packet.", "For every Codex child receipt, set runtime_id to the stable host id codex-cli. Keep worker_ref opaque. Never use an executable, workspace, session-file, or other local path as either identifier.", "Use one or more opaque evidence_refs such as artifact:child-result. Receipt identifiers and evidence refs must contain no spaces, prose, URLs, or local paths.", "If a child deviates from that packet, stop or quarantine only that child and its evidence. Do not let the child write LoopX state or block the parent agent; the parent may retry fresh, replace the child, take over serially, or ignore an optional result.", - ] + ]) + instructions.extend(["Turn request:", request_json]) return "\n".join(instructions) diff --git a/loopx/control_plane/turn_driver/executor.py b/loopx/control_plane/turn_driver/executor.py index f70f313efc..3b1649b50a 100644 --- a/loopx/control_plane/turn_driver/executor.py +++ b/loopx/control_plane/turn_driver/executor.py @@ -132,6 +132,42 @@ JournalPersist = Callable[[Mapping[str, Any]], None] +def reward_memory_automation_enabled( + payload: Mapping[str, Any] | None, *, operation: str, +) -> bool: + """Project verified capability automation into the host IO contract. + + The experiment resolver owns enablement, config/receipt verification and + provider policy. A recall packet's presence or old context is not opt-in. + """ + if not isinstance(payload, Mapping): + return False + recall = payload.get("reward_memory_recall") + if not isinstance(recall, Mapping): + return False + experiment = recall.get("experiment") + envelope = payload.get("turn_envelope") + if not isinstance(experiment, Mapping) or not isinstance(envelope, Mapping): + return False + boundary = envelope.get("boundary") + capabilities = boundary.get("capabilities") if isinstance(boundary, Mapping) else None + memory = capabilities.get("reward_memory") if isinstance(capabilities, Mapping) else None + return ( + isinstance(memory, Mapping) + and memory.get(operation) is True + # Recall resolves the binding again after admission. A stale envelope + # must not resurrect a now unavailable or disabled binding. + and experiment.get("enabled") is True + and experiment.get("available") is True + and experiment.get("configured_for_agent") is True + and bool(envelope.get("goal_id")) + and experiment.get("goal_id") == envelope.get("goal_id") + and bool(envelope.get("agent_id")) + and experiment.get("agent_id") == envelope.get("agent_id") + and experiment.get(operation) is True + ) + + def build_loopx_turn_host_request(plan: Mapping[str, Any]) -> dict[str, Any]: transaction = ( plan.get("transaction") if isinstance(plan.get("transaction"), dict) else {} @@ -158,8 +194,12 @@ def build_loopx_turn_host_request(plan: Mapping[str, Any]) -> dict[str, Any]: if isinstance(goal_ref, Mapping): request["goal_ref"] = dict(goal_ref) reward_memory_recall = plan.get("reward_memory_recall") - if isinstance(reward_memory_recall, Mapping): + recall_enabled = reward_memory_automation_enabled(plan, operation="automatic_recall") + ingest_enabled = reward_memory_automation_enabled(plan, operation="automatic_ingest") + if isinstance(reward_memory_recall, Mapping) and (recall_enabled or ingest_enabled): request["reward_memory_recall"] = dict(reward_memory_recall) + if not recall_enabled: + request["reward_memory_recall"]["context"] = None request.update(subagent.subagent_host_request_projection(plan)) from ...extensions.codex_native_child import configured_native_child_limit @@ -323,12 +363,18 @@ def validate_loopx_turn_host_result( ) if text: normalized[field] = text - reflection_json = _bounded_public_text( - result, - "reward_memory_reflection_json", - limit=HOST_REWARD_MEMORY_REFLECTION_JSON_MAX_CHARS, - required=False, - errors=errors, + # Old/custom hosts may still return this field. Ignore it while disabled: + # ordinary work must not gain a memory validation gate or retain the content. + reflection_json = ( + _bounded_public_text( + result, + "reward_memory_reflection_json", + limit=HOST_REWARD_MEMORY_REFLECTION_JSON_MAX_CHARS, + required=False, + errors=errors, + ) + if reward_memory_automation_enabled(plan, operation="automatic_ingest") + else None ) if reflection_json: if material: diff --git a/tests/control_plane/reward_memory_host_fixture.py b/tests/control_plane/reward_memory_host_fixture.py new file mode 100644 index 0000000000..9cfe6c4bfe --- /dev/null +++ b/tests/control_plane/reward_memory_host_fixture.py @@ -0,0 +1,49 @@ +"""Disposable verified memory bindings for host contract tests.""" + +from __future__ import annotations + +import hashlib +import json +from pathlib import Path + +def enable_plan_memory(plan: dict) -> None: + """Explicit verified participation for pre-existing host wiring fixtures.""" + envelope = plan["turn_envelope"] + plan["reward_memory_recall"] = {"experiment": { + "goal_id": envelope["goal_id"], "agent_id": envelope["agent_id"], + "enabled": True, "available": True, "configured_for_agent": True, + "automatic_recall": True, "automatic_ingest": True, + }} + envelope.setdefault("boundary", {})["capabilities"] = {"reward_memory": { + "automatic_recall": True, "automatic_ingest": True, + }} + + +def enable_live_memory(registry: Path, *, recall: bool = True, ingest: bool = True) -> Path: + # Synthetic pre-verified binding in a disposable project. Runtime resolution + # still checks the exact digest, actor scope and receipt; no live promotion. + from tests.capabilities.test_agent_turn_recall import raw_config + + payload = json.loads(registry.read_text()) + goal = payload["goals"][0] + agent = goal["coordination"]["registered_agents"][0] + config = raw_config(peer_ref=f"agent:{agent}", goal_id=goal["id"], agent_id=agent) + config["automation"].update(automatic_recall=recall, automatic_ingest=ingest) + config["project_provider_binding"]["provider_binary"] = "fixture-unavailable-memory-provider" + path = Path(goal["repo"]) / ".loopx/config/reward-memory/fixture.json" + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(json.dumps(config), encoding="utf-8") + digest = "sha256:" + hashlib.sha256(path.read_bytes()).hexdigest() + goal["control_plane"] = {"reward_memory": { + "enabled": True, "experimental": True, "enabled_agents": [agent], + "config_path": str(path.relative_to(goal["repo"])), "config_digest": digest, + "enablement_receipts": {agent: { + "schema_version": "reward_memory_enablement_receipt_v0", "status": "verified", + "goal_id": goal["id"], "agent_id": agent, "config_digest": digest, + "provider_id": "openviking", "isolation_mode": "goal_scoped_agent_private", + "actor_binding_verified": True, "writability_verified": True, + "exact_readback_verified": True, + }}, + }} + registry.write_text(json.dumps(payload), encoding="utf-8") + return path diff --git a/tests/control_plane/test_reward_memory_host_isolation.py b/tests/control_plane/test_reward_memory_host_isolation.py new file mode 100644 index 0000000000..58d19b1431 --- /dev/null +++ b/tests/control_plane/test_reward_memory_host_isolation.py @@ -0,0 +1,192 @@ +"""Optional memory must not change the ordinary Turn host contract.""" + +from __future__ import annotations + +import copy +import contextlib +import io +import json +import sys +from pathlib import Path + +import pytest + +from loopx.control_plane.turn_driver import build_loopx_turn_host_request, build_loopx_turn_plan +from loopx.control_plane.turn_driver.codex_cli import _prompt, codex_cli_result_schema +from loopx.control_plane.turn_driver.executor import validate_loopx_turn_host_result +from test_loopx_turn_driver import _envelope + + +def _plan(*, recall: bool = True, ingest: bool = True) -> dict: + plan = build_loopx_turn_plan(_envelope(), host="codex-cli", execution_mode="interactive-visible") + plan["reward_memory_recall"] = { + "schema_version": "agent_turn_recall_v0", + "status": "applied", + "experiment": { + "schema_version": "reward_memory_experiment_status_v1", + "goal_id": "fixture-goal", + "agent_id": "codex-fixture", + "enabled": True, + "configured_for_agent": True, + "available": True, + "automatic_recall": recall, + "automatic_ingest": ingest, + }, + "context": {"guidance": [{"content_summary": "Use the verified fixture sequence."}]}, + "grants_new_action_authority": False, + } + plan["turn_envelope"]["boundary"] = {"capabilities": {"reward_memory": { + "automatic_recall": recall, "automatic_ingest": ingest, + }}} + return plan + + +@pytest.mark.parametrize("disabled_by", ["missing", "enabled", "available", "configured_for_agent", "agent_id", "both_automations"]) +def test_disabled_memory_has_identical_host_request_prompt_and_schema(disabled_by: str) -> None: + plan = _plan() + plain = copy.deepcopy(plan) + plain.pop("reward_memory_recall") + plain["turn_envelope"].pop("boundary") + if disabled_by == "missing": + plan.pop("reward_memory_recall") + elif disabled_by == "both_automations": + plan["reward_memory_recall"]["experiment"].update(automatic_ingest=False, automatic_recall=False) + else: + plan["reward_memory_recall"]["experiment"][disabled_by] = "another-agent" if disabled_by == "agent_id" else False + plan["turn_envelope"].pop("boundary") + request = build_loopx_turn_host_request(plan) + ordinary = build_loopx_turn_host_request(plain) + assert request == ordinary + assert _prompt(request) == _prompt(ordinary) + assert "reward_memory" not in _prompt(request) + schema = codex_cli_result_schema(request) + assert schema == codex_cli_result_schema(ordinary) + assert "reward_memory_reflection_json" not in schema["properties"] + + +@pytest.mark.parametrize("invalidated", ["enabled", "available", "configured_for_agent", "goal_id", "agent_id", "automatic_ingest", "automatic_recall"]) +def test_admitted_memory_does_not_resurrect_invalidated_runtime_binding(invalidated: str) -> None: + plan = _plan() + experiment = plan["reward_memory_recall"]["experiment"] + experiment[invalidated] = "another-actor" if invalidated.endswith("_id") else False + request = build_loopx_turn_host_request(plan) + valid_binding = invalidated.startswith("automatic_") + assert ("reward_memory_recall" in request) is valid_binding + assert ("reward_memory_reflection_json" in codex_cli_result_schema(request)["properties"]) is (invalidated == "automatic_recall") + assert ("When reward_memory_recall contains guidance" in _prompt(request)) is (invalidated == "automatic_ingest") + + +@pytest.mark.parametrize(("recall", "ingest"), [(True, False), (False, True), (True, True)]) +def test_recall_and_reflection_are_independently_opted_in(recall: bool, ingest: bool) -> None: + request = build_loopx_turn_host_request(_plan(recall=recall, ingest=ingest)) + prompt = _prompt(request) + assert ("When reward_memory_recall contains guidance" in prompt) is recall + assert ("Set reward_memory_reflection_json" in prompt) is ingest + assert ("reward_memory_reflection_json" in codex_cli_result_schema(request)["properties"]) is ingest + assert bool(request["reward_memory_recall"].get("context")) is recall + if ingest: + assert "attests the exact reflection digest and evidence" in prompt + if recall: + assert "never treat it as new action authority" in prompt + + +def test_unavailable_provider_does_not_instruct_use_of_missing_recall_context() -> None: + plan = _plan() + plan["reward_memory_recall"].update(status="provider_unavailable", context=None) + request = build_loopx_turn_host_request(plan) + assert "When reward_memory_recall contains guidance" not in _prompt(request) + # Recall failure does not silently disable independently configured outcome ingest. + assert "reward_memory_reflection_json" in codex_cli_result_schema(request)["properties"] + + +@pytest.mark.parametrize("ingest", [False, True]) +def test_old_host_reflection_is_ignored_when_ingest_is_off(ingest: bool) -> None: + plan = _plan(recall=False, ingest=ingest) + result = { + "schema_version": "loopx_turn_result_v0", + "turn_key": plan["transaction"]["turn_key"], + "result_kind": "validated_progress", + "completed_phases": ["host_execute", "typed_result"], + "classification": "fixture_progress", + "recommended_action": "Continue the fixture", + "next_action": "Run the next fixture", + "delivery_batch_scale": "single_surface", + "delivery_outcome": "outcome_progress", + "vision_unchanged_reason": "The original fixture remains open.", + "reward_memory_reflection_json": '{"status":"no_evidence"}', + } + validated = validate_loopx_turn_host_result(plan, result) + assert validated["ok"], validated + assert ("reward_memory_reflection_json" in validated["result"]) is ingest + + +@pytest.mark.parametrize("provider", ["file", "sqlite"]) +@pytest.mark.parametrize("mode", ["unconfigured", "disabled", "stale", "both_off", "recall_only", "ingest_only", "both_on"]) +def test_real_turn_cli_memory_projection_and_settlement(tmp_path: Path, monkeypatch: pytest.MonkeyPatch, provider: str, mode: str) -> None: + from loopx.cli import main + from loopx.control_plane.coordination.runtime_shadow import build_todo_runtime_shadow_projection + from tests.control_plane.canonical_authority_fixture import initialize_canonical_authority, isolate_sqlite_runtime + from tests.control_plane.reward_memory_host_fixture import enable_live_memory + from test_loopx_turn_driver import _write_live_fixture + + isolate_sqlite_runtime(tmp_path, monkeypatch) + project, runtime, registry = _write_live_fixture(tmp_path) + state = project / ".codex/goals/loopx-turn-fixture/ACTIVE_GOAL_STATE.md" + projection = build_todo_runtime_shadow_projection(goal_id="loopx-turn-fixture", handoff_mode="soft_claim", todos=[{ + "schema_version": "todo_item_v0", "todo_id": "todo_fixture0001", "role": "agent", + "status": "open", "archive_state": "active", "source_section": "Agent Todo", + "task_class": "advancement_task", "action_kind": "fixture", "claimed_by": "codex-fixture", + "priority": "P0", "done": False, "text": "Advance one public fixture.", + }]) + initialize_canonical_authority(runtime, "loopx-turn-fixture", projection, state_path=state, provider=provider) + ingest = mode in {"ingest_only", "both_on"} + if mode != "unconfigured": + config = enable_live_memory(registry, recall=mode in {"recall_only", "both_on"}, ingest=ingest) + if mode == "disabled": + data = json.loads(registry.read_text()) + data["goals"][0]["control_plane"]["reward_memory"]["enabled"] = False + registry.write_text(json.dumps(data)) + elif mode == "stale": + config.write_text(config.read_text() + "\n") + capture = tmp_path / "host-capture.json" + host = tmp_path / "fixture-codex" + host.write_text(f"#!{sys.executable}\n" + '''import json, os, pathlib, sys +args = sys.argv[1:] +prompt = sys.stdin.read() +schema = json.loads(pathlib.Path(args[args.index('--output-schema') + 1]).read_text()) +request = json.loads(prompt.split('Turn request:\\n', 1)[1]) +pathlib.Path(os.environ['FIXTURE_HOST_CAPTURE']).write_text(json.dumps({'prompt': prompt, 'schema': schema, 'request': request})) +result = {'schema_version': 'loopx_turn_result_v0', 'turn_key': request['turn_key'], + 'result_kind': 'validated_progress', 'completed_phases': ['host_execute', 'typed_result'], + 'classification': 'fixture_progress', 'recommended_action': 'Continue the original fixture', + 'next_action': 'Run the next bounded fixture', 'delivery_batch_scale': 'single_surface', + 'delivery_outcome': 'outcome_progress', 'vision_unchanged_reason': 'The original fixture remains open.', + 'path_delta_mode': 'unchanged', 'agent_vision_json': '', 'summary': 'Independent fixture validation passed.'} +# Simulate an older/custom host returning a memory field even when the current +# result schema does not request it. The real adapter must isolate that input. +result['reward_memory_reflection_json'] = '{"status":"no_evidence"}' +pathlib.Path(args[args.index('--output-last-message') + 1]).write_text(json.dumps(result)) +print(json.dumps({'type': 'thread.started', 'thread_id': 'fixture-memory-session'})) +''', encoding="utf-8") + host.chmod(0o755) + monkeypatch.setenv("FIXTURE_HOST_CAPTURE", str(capture)) + output = io.StringIO() + with contextlib.redirect_stdout(output): + code = main(["--registry", str(registry), "--runtime-root", str(runtime), "--format", "json", + "turn", "run-once", "--goal-id", "loopx-turn-fixture", "--agent-id", "codex-fixture", + "--host", "codex-cli", "--project", str(project), "--codex-bin", str(host), + "--iteration-context", "fresh", "--validation-command-json", json.dumps([sys.executable, "-c", "pass"]), + "--scan-root", str(project), "--no-global-sync", "--execute"]) + result = json.loads(output.getvalue()) + assert code == 0, result + observed = json.loads(capture.read_text()) + assert ("reward_memory_reflection_json" in observed["schema"]["properties"]) is ingest + assert ("Set reward_memory_reflection_json" in observed["prompt"]) is ingest + assert "When reward_memory_recall contains guidance" not in observed["prompt"] + if mode in {"unconfigured", "disabled", "stale", "both_off"}: + assert "reward_memory" not in observed["prompt"] + assert result["effects"]["quota_spent"] is True + assert result["post_settlement"]["external_writes_performed"] is False + journals = list((runtime / "goals/loopx-turn-fixture/turns").glob("*.json")) + assert len(journals) == 1 + assert ("reward_memory_reflection_json" in json.loads(journals[0].read_text())["host_result"]) is ingest diff --git a/tests/control_plane_ts/turn_envelope.test.ts b/tests/control_plane_ts/turn_envelope.test.ts index c814a7a1e1..31f7293131 100644 --- a/tests/control_plane_ts/turn_envelope.test.ts +++ b/tests/control_plane_ts/turn_envelope.test.ts @@ -73,6 +73,38 @@ const protocolActionFields = { agent_action: "advance one bounded segment", }; +test("optional memory participation is compact, verified and signed", () => { + const plain = payload(); + const build = (source: JsonObject) => buildTurnEnvelope({payload: source, + protocol_action_fields: protocolActionFields, scheduler_execution_args: ""}); + const ordinary = build(plain); + const memory = {enabled: true, configured_for_agent: true, experiment_available: true, + automatic_recall: true, automatic_ingest: false, config_path: "private-config", + enablement_receipt: {diagnostic: "not-host-context"}}; + for (const flag of ["enabled", "configured_for_agent", "experiment_available"]) { + const source = payload(); + (source.goal_boundary as JsonObject).capabilities = {reward_memory: {...memory, [flag]: false}}; + assert.deepEqual(build(source).boundary, ordinary.boundary); + } + for (const [recall, ingest] of [[true, false], [false, true], [true, true], [false, false]]) { + const source = payload(); + (source.goal_boundary as JsonObject).capabilities = {reward_memory: {...memory, + automatic_recall: recall, automatic_ingest: ingest}}; + const envelope = build(source); + const projected = (envelope.boundary as JsonObject).capabilities; + assert.deepEqual(projected, recall || ingest ? {reward_memory: { + automatic_recall: recall, automatic_ingest: ingest, + }} : undefined); + assert.deepEqual(quotaActionSignatureDocument(source, protocolActionFields), + turnEnvelopeActionSignatureDocument(envelope)); + if (recall || ingest) { + const signed = JSON.parse(JSON.stringify(turnEnvelopeActionSignatureDocument(envelope))); + ((projected as JsonObject).reward_memory as JsonObject).automatic_ingest = !ingest; + assert.notDeepEqual(turnEnvelopeActionSignatureDocument(envelope), signed); + } + } +}); + test("settlement-only replans preserve owed commands and do not demand another outcome", () => { const prefix = "loopx --runtime-root /" + "long-path/".repeat(50); for (const commands of [ @@ -223,7 +255,7 @@ test("pending capability action outranks stale replan commands and remains signe const envelope = buildTurnEnvelope({payload: source, protocol_action_fields: protocolActionFields, scheduler_execution_args: ""}); assert.equal(envelope.replan_action_packet, null); assert.deepEqual((envelope.writeback as JsonObject).next_cli_actions, [command]); - const signed = turnEnvelopeActionSignatureDocument(envelope); + const signed = JSON.parse(JSON.stringify(turnEnvelopeActionSignatureDocument(envelope))); assert.deepEqual((signed.action as JsonObject).capability_intent, source.pending_capability_intent); (source.pending_capability_intent as JsonObject).command = "untrusted replacement"; assert.throws(() => buildTurnEnvelope({payload: source, protocol_action_fields: protocolActionFields, diff --git a/tests/test_loopx_turn_driver.py b/tests/test_loopx_turn_driver.py index b824080506..6f44e65068 100644 --- a/tests/test_loopx_turn_driver.py +++ b/tests/test_loopx_turn_driver.py @@ -839,12 +839,14 @@ def test_turn_host_request_carries_typed_child_operations() -> None: def test_turn_host_request_carries_reward_memory_decision_context() -> None: + from tests.control_plane.reward_memory_host_fixture import enable_plan_memory plan = build_loopx_turn_plan( _adaptive_envelope(), host="codex-cli", execution_mode="interactive-visible", ) - plan["reward_memory_recall"] = { + enable_plan_memory(plan) + plan["reward_memory_recall"].update({ "schema_version": "agent_turn_recall_v0", "status": "applied", "context": { @@ -858,7 +860,7 @@ def test_turn_host_request_carries_reward_memory_decision_context() -> None: ], }, "grants_new_action_authority": False, - } + }) request = build_loopx_turn_host_request(plan) @@ -3035,6 +3037,8 @@ def test_turn_run_once_commits_independently_validated_progress( host="generic-cli", execution_mode="isolated-headless", ) + from tests.control_plane.reward_memory_host_fixture import enable_plan_memory + enable_plan_memory(plan) def host_runner(request: dict[str, object]) -> dict[str, object]: return { @@ -3462,6 +3466,8 @@ def test_turn_run_once_codex_cli_wires_validated_reflection_post_settlement( monkeypatch: pytest.MonkeyPatch, ) -> None: project, runtime, registry = _write_live_fixture(tmp_path) + from tests.control_plane.reward_memory_host_fixture import enable_live_memory + enable_live_memory(registry) reflection = json.dumps( { "schema_version": "turn_reward_memory_reflection_v1",