Skip to content

chore(release): prepare 1.3.0 long-horizon semantic control quality - #5682

Draft
loopx-agent wants to merge 1 commit into
mainfrom
codex/release-1.3.0
Draft

loopx-agent wants to merge 1 commit into
mainfrom
codex/release-1.3.0

Conversation

@loopx-agent

@loopx-agent loopx-agent commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Prepare the 1.3.0 version identity, generated manual and Developer Book checkpoint around long-horizon benchmark diagnosis and semantic control-plane quality. The existing configuration-backup command is classified on the advanced manual surface, preserving bounded default help.

Publication is held: the release candidate does not certify final source qualification or authorize control-plane self-merge. The existing regression repair PR #5533 is being reconciled rather than duplicated. Final qualification and artifact/update readback will run on the frozen merged release source.

Changed surfaces: version metadata, Developer Book checkpoint, generated manual and advanced command classification. Placement: configuration-backup extends the existing local help-visibility owner; no new shared protocol or execution owner. Future-facing pass: no additional abstraction was needed.

Validation on the frozen candidate: 12/12 selected canary checks, five direct checks, 22 CLI output-budget tests, version/help/book contracts, required-scope Ruff and 19-file Mypy passed. One unchanged maintainability advisory remains. Preliminary real DS V4.1 Flash/high native Goal qualification passed with two Turns and two settlements. Baseline regression failures and unavailable final-source qualification are retained below. No benchmark jobs were launched.

Bilingual release-body candidate, usage and contribution attribution

#LoopX v1.3.0 — Long-horizon work, stronger semantic control

Release candidate — publication held until final exact-source qualification and maintainer review are complete.

Personal illustrated candidate guide — verified user owner, text, two image blocks and exported rendering; remains a candidate until release artifacts are verified.

Long-horizon benchmark diagnosis and Astra-assisted engineering shape this release's theme: retain the right work and evidence, make recovery actionable, and preserve the authority behind every next step. The latest research blog explains the failures that motivated the work. Astra describes the development and semantic-review effort; the reported EdgeBench worker experiment uses gpt-6.1-sol / xhigh.

Release Decision

Who should upgrade: Operators of continuing Codex work, evidence-driven replanning, and workspace Chat should consider this version once it is published. Users satisfied with v1.2.4 can remain there while the candidate is qualified.

What this release solves: Useful history could be crowded out by repeated observations; an already-known replan could demand a failed round trip; task steps, permission failures and settlement hints could point to the wrong recovery. This release collects bounded fixes to those paths, alongside clearer workspace Chat and explicit configuration recovery.

Breaking changes: No intentional breaking migration. Ordinary host-declared project Chat now defaults to workspace_write; choose workspace_read explicitly when needed. Existing read-only App bindings retain their grant. New generated settlement commands carry correctly placed global --format json; direct CLI defaults and stored receipts are unchanged. Managed execution names deepseek-flash; explicit historical model settings remain respected. Canonical new-Goal creation and Explore execution remain opt-in. Existing Goals are not migrated by a device preference.

How to verify: Expect loopx 1.3.0, a healthy owning installation, and a current scoped status/diagnostic readback after upgrading. Investigate unavailable or blocked results before starting work.

loopx --version
loopx doctor
loopx --format json status
loopx diagnose --goal-id "$GOAL_ID"

Contributors: Release maintainer @huangruiteng, with maintainer automation through @loopx-agent. The community contributions listed below are derived from v1.2.4 → candidate merged PRs.

State Kernel & Control Plane

  • Continue the correct task: Next Action writeback binds to the selected Todo, Agent and current basis; accountable Goal-level writes remain valid. Old or unrelated steps cannot silently take over the next turn. #5531, #5588

  • Use evidence before repeating a route: Dense typed replan context retains outcome and route diversity, with exact omitted-history recovery. Explicit successor selection can finish the known replan admission in one CLI call; genuinely fresh vision gaps can rearm it. #5536, #5624, #5628, #5646

  • Settle with the original authority: Checkpoint recovery, unchanged-artifact/negative-evidence guidance and vision closeout follow the canonical settlement plan. Full and compact packets preserve conditional steps and identities; generated JSON commands run as returned. #5573, #5629, #5633, #5640, #5636, #5667, #5672

  • Separate denial from failure: Host permission diagnostics retain actionable recovery; isolated edits preserve the complete evidence basis. Delegated stop acknowledges the request separately from proven native-process drain and lease settlement. #5644, #5587, #5308

Capabilities & Workflows

  • Explore Harness becomes usable in an ordinary turn: Evidence-only and planning modes share the existing owner and a turn-start context; finding titles survive projection. Configuration proves availability, while reading and changed decisions still need evidence. #5610, #5658

  • Work in an ordinary workspace conversation: The App Scope picker opens a granted workspace through the existing composer without synthesizing a Goal or borrowing portfolio authority. Writes stay inside the host grant; revoked or changed contexts reject new work. #5540, #5555

  • Recover configuration explicitly: Export and verify complete stored configuration, then restore to an isolated checkpoint. Device settings can opt future empty Goals into canonical File/SQLite authority and a frozen soft-claim/hard-lease policy; existing Goals retain their owner. #5557, #5569

  • Inspect public GitHub evidence: A SHA-pinned anonymous provider retrieves bounded sources; parent admission and downstream ledger readback remain separate. CLI/provider and conversation readback are shipped; complete frontend initiation/admission remains a staged boundary. #5459

Quality & Testing

  • Test semantics and retire duplicate owners: Native child receipts preserve resumed result correlation; stop fixtures retain real drain obligations. Redundant smokes, Python Todo transition/admission adapters and an unused shadow producer are retired; four document versions and the Lark visibility pair each use their defining owner. #5455, #5539, #5613, #5621, #5597, #5664, #5666, #5670

  • Reduce bounded preflight cost: The affected Turn/MCP path reuses typed projections; checkpoint context resolves in one TypeScript request. These measured engineering slices do not close sustained-operation or model-efficiency acceptance. #5283, #5585

Benchmarks & Integrations

  • Diagnose long-horizon work with controlled profiles: Native EdgeBench supports official, single-call, native Goal, heartbeat-resume and heartbeat-Explore workers. Trial-wide deadlines, terminal visualizer results and runnable successor admission use their intended lifecycle; feedback policy is distinguished from evaluator isolation. #5591, #5600, #5618, #5617, #5635

  • Private Lark Chat composes with the same workspace owner: Explicit App/source bindings support ordinary Chat, confirmed steward commissions, scoped status/help and authorized attached-Agent selection. One listener and durable Core requests preserve original audiences. Installed/live Lark and mobile qualification remain separate. #5541, #5542, #5544, #5546, #5550, #5637

  • More explicit provider observations: The managed model alias names DeepSeek V4.1 Flash, with operator-configurable Ark rows. Finance 0.8.3 distinguishes producer accuracy, encoded periods and parent-declared economic periods; the earlier position guard reads filled protection separately from trade authority. #5561, #5631, #5521

Documentation & Compatibility

  • Explain what the experiment establishes: The latest bilingual blog separates prompt confounds, task/judge accounting, genuine strategy failure, control-protocol recovery and context cost. It retains failed samples and open causal questions. No matched v1.2.4-versus-v1.3.0 outcome trial has been qualified; this release makes no benchmark-score, statistical superiority or long-horizon success-rate claim. #5649, #5635

  • Keep entrypoints current: Wheels ship project-scoped skill sources; current authority-format reads avoid migration locks; refresh authoring help explains the 1,200-character step limit and in-flight versus replan closeout. PR review supports explicit additional owner accounts and per-Agent direction without granting merge authority. #5560, #5661, #5653, #5672, #5608, #5656

Community Contributors

  • @Inference1 — first-time external contributor: four document-version owners and one Lark visibility owner (#5621, #5597).
  • @mikamikasuki — first-time external contributor: the refresh no-write mutation regression and retirement of completed contributor/RFC checkpoints (#5613, #5574, #5576, #5577, #5579).
  • @jackie-cqz — external contributor: native child result correlation, public GitHub evidence, shared qualification repair and long Windows workspace inputs (#5455, #5459, #5526, #5603).
  • @Duang777 — external contributor: Goal recreation fences, validated inbox read commits and accountable Goal-level writeback (#5338, #5563, #5588).
  • @hhyykk — external contributor: one-request TypeScript checkpoint context (#5585).
  • @catwithtudou — external contributor: repository identity normalization for leases and preserved diagnostic chronology (#5584, #5572).
  • @songoow — external contributor: governed delegation stop, native fixture cleanup, exact-target CI recovery and focused smoke governance (#5308, #5534, #5535, #5539, #5567).
  • @BigDataDZ — external contributor: the committed goal-direction F2 revision-drift fixture (#5549).
  • @maxliux5 — external contributor: personal follow-through documentation consolidated into the existing manager profile (#5384).

Optional Capability Activation & Use

Explore Harness

Activation: In Goal settings → Capability Center choose evidence or planning. CLI planning opt-in is shown below; the default is off.

Validation: Read turn context and summary for the same Goal/Agent; recorded nodes alone do not prove adoption.

Disable / rollback: Set --explore-mode off --execute; retained evidence stays readable.

Authority boundary: Analysis and planning grant no worker spawning, claims, quota spend, merge or external publication.

Docs: Versioned guide

loopx configure-goal --goal-id "$GOAL_ID" --explore-mode planning --explore-harness-profile adaptive-resilient --execute
loopx explore turn-context --goal-id "$GOAL_ID" --agent-id "$AGENT_ID"
loopx explore summary --goal-id "$GOAL_ID"
loopx configure-goal --goal-id "$GOAL_ID" --explore-mode off --execute

TurnEnvelope and captured decisions

Activation: Opt in per guard invocation with --turn-envelope; add --decision-output-dir only with an explicit Turn id and a new directory whose parent exists.

Validation: Read the returned capture and verify Goal/Agent/Turn, original source hash and ok; observe rejection and incomplete publication honestly.

Disable / rollback: Omit both options. Delete only no-longer-needed private captures through ordinary file management.

Authority boundary: Saved decisions are private observations, not fresh admission; selection, leases, cancellation and mutation-time checks remain mandatory. Frontend/Lark do not consume these files.

Docs: Versioned guide

loopx --format json quota should-run --goal-id "$GOAL_ID" --agent-id "$AGENT_ID" --turn-instance-id "$TURN_ID" --turn-envelope --decision-output-dir ./guard-001
cat ./guard-001/decision.json

Ordinary workspace Chat

Activation: Run loopx chat; in the App steward conversation choose an already granted workspace in Scope. Host-declared roots now default to workspace_write; use the explicit read-only command below when required.

Validation: Read the selected scope, effective grant and returned result in the same conversation. A changed grant creates a new context; old history remains.

Disable / rollback: Restart the service with --project-workspace-grant workspace_read, or revoke the host workspace grant; new work on the old context is rejected.

Authority boundary: Workspace writes follow AGENTS.md and the actual sandbox. No hidden Goal, portfolio grant, peer delegation or external-send authority is created.

Docs: Versioned guide

loopx chat --project-workspace-grant workspace_read --no-open

Owner private Lark conversations

Activation: In Settings → Lark explicitly verify the selected App and personal owner, choose a workspace, executor, grant and ordinary Chat or steward role, then connect. /delegate --tokens N objective requires an original-source confirmation. /agents, /agent TARGET_REF and /project use separately authorized attached-Agent targets.

Validation: In the bound private conversation send /status and /help; verify the App/source, role, workspace grant, original Session/Turn and queue. Inspect the same binding in Settings.

Disable / rollback: Disconnect that exact App binding in Settings → Lark; /stop targets the current ordinary Turn and /stop-commission targets the bound commission. Remove exact Agent target grants separately.

Authority boundary: Personal credentials, source/owner and listener identity remain explicit. Ordinary Chat creates no Goal; commission confirmation does not grant arbitrary writes, automatic heartbeat or canonical task acceptance. Live Lark/mobile qualification is separate.

Docs: Versioned guide

loopx chat --project-workspace-grant workspace_read --no-open
loopx chat --help

Configuration checkpoints

Activation: Settings → Capability Center → Configuration backup and recovery downloads a private checkpoint. CLI export uses a new destination, preview first, then --execute.

Validation: Run configuration-backup verify on the exact file; compare source scope and digest. Integrity verification does not certify privacy.

Disable / rollback: An isolated restored checkpoint can be removed without affecting live settings. Revert any adopted setting through its existing revision-checked editor or machine-config/configure-goal transaction.

Authority boundary: Capture and isolated restore copy no credential store, Host session, grant, provider selection, live registry, fence, lease or scheduler. Backups remain private, including secrets already embedded in configuration.

Docs: Versioned guide

loopx --format json configuration-backup export --goal-id "$GOAL_ID" --output "$NEW_CHECKPOINT_FILE"
loopx --format json configuration-backup export --goal-id "$GOAL_ID" --output "$NEW_CHECKPOINT_FILE" --execute
loopx --format json configuration-backup verify --input "$NEW_CHECKPOINT_FILE"

Canonical new-Goal creation

Activation: In Device defaults → New Goal authority opt in to canonical creation and choose File/SQLite plus soft_claim/hard_lease. The same v1 document uses the existing preview/apply transaction; default remains off.

Validation: Inspect machine settings, bootstrap a new empty project, then read its Todos and native authority receipt; existing Goals are not retargeted.

Disable / rollback: Preview/apply canonical_creation=false to disable future creation, or remove the goal_storage namespace with its exact removal-plan revision. Existing Goals retain storage and fences.

Authority boundary: Storage and execution policy are independent of tools, accounts, network, scheduling and migration authority. Missing authority cannot be recreated as empty by forced bootstrap.

Docs: Versioned guide

Save / 保存为 goal-storage.json:

{"schema_version":"loopx_goal_storage_defaults_v1","new_goal_provider":"sqlite","canonical_creation":true,"new_goal_handoff_mode":"hard_lease"}
loopx machine-config preview --namespace goal_storage --config-json goal-storage.json
loopx machine-config apply --namespace goal_storage --config-json goal-storage.json --expected-plan-revision "$PLAN_REVISION" --execute
loopx machine-config inspect
loopx machine-config remove --namespace goal_storage
#Review removal, then use its returned revision.
loopx machine-config remove --namespace goal_storage --expected-plan-revision "$REMOVAL_REVISION" --execute

Governed delegation stop

Activation: Use an already configured exact binding, registered requester and operation; delegation stop --execute explicitly requests stop. CLI/MCP and the App team surface share that owner.

Validation: delegation read distinguishes acknowledgement, proved native process drain and settlement. Missing supervision remains unknown and requires the returned recovery.

Disable / rollback: Remove the exact operator binding to revoke new execution; use stop on an active operation. A stopped/settled receipt is not a resume grant.

Authority boundary: Do not infer Host exit, lease release or non-execution from absent records. Windows/unsupported stopping refuses before launch-side cancellation effects.

Docs: Versioned guide

loopx delegation stop --goal-id "$GOAL_ID" --agent-id "$AGENT_ID" --execution-config "$DELEGATION_CONFIG" --operation-id "$OPERATION_ID" --execute
loopx delegation read --goal-id "$GOAL_ID" --agent-id "$AGENT_ID" --execution-config "$DELEGATION_CONFIG" --operation-id "$OPERATION_ID"

PR review queue ownership and direction

Activation: Capability Center configures additional owner logins and forward/reverse direction; Goal CLI can set exact per-Agent direction. Current-session request intake precedes generic queue discovery.

Validation: Inspect configure-goal and the read-only pr-review packet; compare effective accounts, Agent and direction. Queue ownership is separate from GitHub author identity.

Disable / rollback: Use --clear-pr-review-owner-logins --execute, --pr-review-agent-order AGENT=inherit, or --clear-pr-review-configuration --execute. Remove owner_logins from Goal/device settings before downgrading.

Authority boundary: These settings grant no GitHub review, comment, dismissal, merge, cross-Agent write or scheduler authority; review depth and CI policy retain their owner.

Docs: Versioned guide

loopx configure-goal --goal-id "$GOAL_ID" --pr-review-owner-login maintainer --pr-review-agent-order reviewer-a=forward --execute
loopx configure-goal --goal-id "$GOAL_ID"
loopx pr-review --goal-id "$GOAL_ID" --agent-id reviewer-a --repo "$REPOSITORY" --format json
loopx configure-goal --goal-id "$GOAL_ID" --clear-pr-review-configuration --execute

Public GitHub evidence

Activation: Opt in per plan with --public-github and a full commit SHA source; execute --execute authorizes anonymous source reads.

Validation: Read back the exact plan and execution receipt. Parent admit/reject and downstream research-ledger coverage are separate explicit operations.

Disable / rollback: Omit --public-github and --execute; no persistent provider switch is installed. Retire evidence only under the admitted downstream-coverage rules.

Authority boundary: No token, cookie, private repository, mutable branch source, raw-page persistence, automatic admission or outbound message permission. Complete frontend initiation remains open.

Docs: Versioned guide

loopx external-evidence plan --public-github --objective "Inspect public source" --user-activity "Choose a source" --decision "Whether a literal is present" --evidence-kind literal_match --source "$SHA_PINNED_PUBLIC_URL" --search-term LoopX --format json > plan.json
loopx external-evidence execute --plan-json plan.json --execute --format json > execution.json
loopx external-evidence readback --plan-json plan.json --receipt-json execution.json

Managed DeepSeek model selection

Activation: Managed execution now names deepseek-flash@high. Explicit LOOPX_TURN_MODEL/DSH_MODEL or --dsh-model overrides are still honored; configure credentials through the owning provider.

Validation: Read turn run-once --help and the returned runtime profile/actual provider identity; a model alias alone is not live qualification.

Disable / rollback: Set LOOPX_TURN_MODEL to the previous explicitly selected model and restart the owning host. Do not change an existing session silently; omitted configuration restores the shipped managed default.

Authority boundary: A model setting grants no API credential, paid invocation, workspace, permission, Goal or provider promotion authority. CPA routing remains separately installed/operator-owned.

Docs: Versioned guide

LOOPX_TURN_MODEL=deepseek-flash loopx turn run-once --help
#Explicit rollback selection for the next owning host invocation:
export LOOPX_TURN_MODEL="$PREVIOUS_MODEL"

Finance evidence assessment

Activation: Finance Value Discovery is a separately installed optional extension, version 0.8.3; use the pinned source package and register/enable its manifest. The additive assess-period input is finance_period_comparison_input_v1.

Validation: Run extension doctor, then assess-period on a reviewed local input. Producer accuracy, period eligibility, source authenticity and investment truth remain distinct.

Disable / rollback: Disable loopx-finance-value-discovery with the command below. Retain private account/position observations outside public research projections.

Authority boundary: The package performs deterministic evidence assessments; it grants no investment advice, order, transfer, signing or execution authority. Missing source/timezone/binding evidence stays explicit.

Docs: Versioned guide

#In a v1.3.0 source checkout:
python3 -m pip install ./packages/loopx-finance-value-discovery
loopx extension install --manifest packages/loopx-finance-value-discovery/extension.toml --execute --format json
loopx extension enable loopx-finance-value-discovery --execute
loopx extension doctor loopx-finance-value-discovery --execute
loopx-finance-value-discovery assess-period --input-json period.json
loopx extension disable loopx-finance-value-discovery --execute

EdgeBench native trials

Activation: Explicit research-only invocation from a v1.3.0 checkout; Linux, Docker, pinned EdgeBench/SForge/Harbor, selected Codex binary and authorized model access are prerequisites. Use the selected worker and feedback profile.

Validation: Run --help before the reviewed trial; read session/model, native terminal result, evaluator completion and isolation receipts. Registration and sampling counts do not prove countable outcomes.

Disable / rollback: Do not start another trial; stop the exact owned trial/controller and TLS relay through their existing cancellation lifecycle. Preserve private logs and incomplete results.

Authority boundary: A release launches no benchmark jobs. The native judge owns task/scoring; blind policy requires separately qualified credential/network/submission isolation. Raw evidence and secrets remain private.

Docs: Versioned guide

python -m benchmark.edgebench.run --help
#Paid execution only after the operator reviews task, pin, isolation and budget:
python -m benchmark.edgebench.run --task "$TASK_ID" --tasks-dir "$TASKS_DIR" --log-dir "$PRIVATE_LOG_DIR" --run-id "$UNIQUE_ATTEMPT" --worker heartbeat-resume --model "$MODEL" --effort xhigh --judge-url "$JUDGE_URL"

Install / Update

When published, the package version/tag will be 1.3.0 / v1.3.0. Python 3.11+ and Node.js 22.22.3+ remain required. New PyPI installs use:

python3 -m pip install --upgrade loopx
loopx workflow-skills --install
loopx slash-commands --install
loopx doctor

Existing installs preserve their acquisition owner:

loopx update check
loopx update plan
loopx update apply
loopx doctor
loopx extension doctor --all-enabled --execute

Do not install this candidate from a named stable channel yet. After publication, inspect the package/assets, source manifest and update feed independently; an uploaded source package does not prove a signed desktop build or installed runtime activation.

Qualification

Pending release gates: final exact-commit qualification, maintainer integration, published artifacts and final guide readback remain open. Preliminary real-model execution and personal guide ownership/rendering have been verified on the stated candidate source; no prior-source result is relabeled as final-release evidence.

  • Passed on the candidate lineage: release version/readiness contracts, generated help/manpage (after explicit advanced-command classification), bilingual Developer Book checkpoint, required-scope Ruff, Mypy (19 files), packaged Chat build, 42 installed-bundle dashboard command tests, and the eight changed public paths' boundary scan.
  • Regression repair is carried in the existing #5533, currently at 3d3beba. It updates stale current-contract fixtures and repairs the host-neutral scope prompt, preview retirement and unsent runtime-refusal recovery. Current repair checks include 19/19 risk canaries with 5 direct checks, Ruff/Mypy, real cleanup-failure enforcement, 68 source/revocation cases, and packaged configuration recovery with damaged-file rejection. The same 40 independent transport/projection cases pass on pinned main and the repair head; the unchanged TraeX oracle changes from two baseline failures to three candidate passes. Full pytest remains running and has reported a failure awaiting final detail. Earlier 174-pass/6-failure and 1-pass/22-failure observations remain historical evidence. This repair is not yet included in the release source; no full-suite or final release pass is claimed.
  • Final frozen risk canary: 12 selected / 12 executed, no blocking failures; 5 direct checks passed. CLI output-budget tests: 22 passed. One pre-existing maintainability advisory remains. Earlier failed attempts are retained separately and are superseded only by this unchanged-head rerun.
  • Broad scripts-inclusive Ruff reports six unchanged E402 import-placement findings; the actual CI lint scope passes. Initial no-clone smoke failed during disposable-home cleanup; fresh-clone and update smokes passed separately. Repository-hygiene reports symbolic Basic-auth construction and synthetic private-network fixture literals, not a leaked credential; the changed-path scan is clean.
  • Preliminary native Codex Goal on candidate 8056cc6 passed with DeepSeek V4.1 Flash / high: two real model Turns, two Todo settlements and independent acceptance. Final exact-source qualification will be repeated. One real Doubao evolving request failed authentication: the recovered key belongs to Ark Agent Plan and was not accepted by the public evolving endpoint. No benchmark trial was launched.
  • Verified personal user ownership and bounded guide history: the last substantive guide covers v1.2.4 and was updated on October 3, 2026. The new candidate guide has verified owner, complete text, two image blocks and nine-page exported rendering. Final released artifacts and guide status still need readback.

中文摘要

主题:长程 benchmark 与 Astra 辅助工程推动语义控制面质量提升。 重点是续接正确工作、让证据进入重规划,以及按原权限恢复和结算。最新 blog 保留实验混杂、评分修复、真实策略失败及开放问题。Astra 是开发与语义审查主线;本文 EdgeBench worker 实验使用的基础模型仍为 gpt-6.1-sol / xhigh。

候选版本,尚未发布。 个人图文候选指南已验证个人归属、正文、两张图与导出排版;正式发布后按实际产物原地更新。

升级决策

**谁需要升级:**需要持续 Codex 工作、证据驱动重规划或普通 workspace Chat 的用户,在正式发布后可考虑升级;当前满足需求的 v1.2.4 用户可等待候选验证完成。

**解决了什么:**早期有效结果被重复观察淹没、已知 replan 仍要求失败重入、任务步骤与权限/结算提示错位;本次集合这些有界修复,并改善 workspace 对话与配置恢复。

**是否有破坏性变更:**无主动 breaking migration。host-declared project Chat 新默认 workspace_write,需只读时显式选择 workspace_read,旧 App 只读 grant 保留。新生成结算命令自带位置正确的 JSON 参数,历史命令与直接 CLI 默认不变。managed 模型命名改为 deepseek-flash,明确旧配置保持优先。canonical 新 Goal 与 Explore 执行保持 opt-in,不迁移已有 Goal。

**如何验证:**正式升级后期望 loopx 1.3.0、健康安装及正确作用域状态。使用上方 loopx --version、loopx doctor、loopx --format json status 与 loopx diagnose --goal-id "$GOAL_ID",遇到 unavailable/blocked 先恢复再工作。

**贡献者:**维护者 @huangruiteng,自动化维护账号 @loopx-agent;以下社区贡献来自 v1.2.4 到候选的真实合并范围。

状态内核与控制面

  • **续接正确任务:**下一步写回绑定 Todo、Agent 与当前依据,同时保留真实的 Goal 级责任写回;旧步骤和无关任务不能覆盖当前工作。#5531, #5588

  • **让证据改变下一步:**重规划保留结果与路线差异,并可恢复被省略历史;显式后继选择可在一次 CLI 调用内完成已知重规划准入,新 vision gap 仍会重新触发。#5536, #5624, #5628, #5646

  • **按原权限结算:**checkpoint 恢复、未改产物/负证据与 vision 收口采用规范计划;完整包和短包保留条件、身份及可直接执行的 JSON 命令。#5573, #5629, #5633, #5640, #5636, #5667, #5672

  • **区分拒绝与故障:**宿主权限错误保留可操作恢复,隔离编辑保留完整依据;委派停止区分已接收、真实进程退出与 lease 结算。#5644, #5587, #5308

能力与工作流

  • **普通轮可使用 Explore Harness:**证据与规划模式复用已有 owner 和轮前入口,finding 标题保留;开启、读取与决策改变仍是独立证据。#5610, #5658

  • **直接处理工作区请求:**App 的 Scope 选择器沿用原 composer,无须创建 Goal;写入受宿主 grant 约束,撤权或上下文变化后拒绝新工作。#5540, #5555

  • **显式恢复配置:**导出并验证完整存储配置,只恢复到隔离 checkpoint;设备设置可为未来空 Goal 选择规范 File/SQLite 与冻结策略,已有 Goal 保持归属。#5557, #5569

  • **读取公开 GitHub 证据:**匿名 provider 按 SHA 取有限来源,父 Agent 采纳与下游账本回读分开;已交付 CLI/provider 和会话回读,完整前端发起/采纳仍为阶段边界。#5459

质量与测试

  • **验证语义并退役重复规则:**原生 child 回执保持续接结果归属,停止验证保留真实 drain 义务;删除冗余 smoke、Python Todo 适配与无用 shadow producer,文档版本和 Lark visibility 复用各自 owner。#5455, #5539, #5613, #5621, #5597, #5664, #5666, #5670

  • **减少有界预检开销:**Turn/MCP 复用类型化投影,checkpoint context 合并为一次 TS 请求;这不等于多日稳定性或模型效率验收完成。#5283, #5585

基准与集成

  • **用受控 profile 诊断长程工作:**EdgeBench 接入五种 worker,整场 deadline、终态显示与可运行后继各守生命周期;反馈政策与 evaluator 隔离明确区分。#5591, #5600, #5618, #5617, #5635

  • **私聊 Lark 复用工作区 owner:**明确 App/source 绑定支持普通对话、确认委托、作用域 status/help 与授权 Agent 选择;单 listener 与耐久请求保持原受众,安装/真实飞书和手机验收另行判断。#5541, #5542, #5544, #5546, #5550, #5637

  • **更明确的 provider 观察:**managed 模型使用 DeepSeek V4.1 Flash 名称,Ark 行可由 operator 配置;Finance 0.8.3 区分 producer 精度、编码期间与父调用者声明的经济期间,持仓守护的保护读回不授予交易权限。#5561, #5631, #5521

文档与兼容性

  • **说明实验到底证明了什么:**最新中英 blog 区分提示混杂、评分账本、策略失败、协议恢复和上下文成本,保留失败与因果问题。本次未完成 v1.2.4/v1.3.0 配对 outcome 验证,不声明 benchmark 涨分、统计优越性或长程成功率提升。#5649, #5635

  • **保持入口可用:**wheel 交付项目 skill,当前 authority 格式读取无需迁移锁;refresh help 解释 1,200 字符限制与 in-flight/replan 边界;PR review 支持 owner 列表与 Agent 方向,但不增加合并权限。#5560, #5661, #5653, #5672, #5608, #5656

社区贡献者

  • @Inference1 — 首次外部贡献者:四个文档版本 owner 与单一 Lark visibility owner(#5621, #5597)。
  • @mikamikasuki — 首次外部贡献者:refresh no-write mutation 回归及已完成贡献/RFC 检查点退役(#5613, #5574, #5576, #5577, #5579)。
  • @jackie-cqz — 外部贡献者:原生 child 结果归属、公开 GitHub 证据、共享验证修复与长 Windows 路径(#5455, #5459, #5526, #5603)。
  • @Duang777 — 外部贡献者:Goal 重建 fence、校验后的 inbox 读取提交与 Goal 级责任写回(#5338, #5563, #5588)。
  • @hhyykk — 外部贡献者:单次 TS 请求的 checkpoint context(#5585)。
  • @catwithtudou — 外部贡献者:lease 仓库身份归一及诊断时间顺序(#5584, #5572)。
  • @songoow — 外部贡献者:受治理委派停止、原生 fixture 清理、精确目标 CI 恢复与 smoke 治理(#5308, #5534, #5535, #5539, #5567)。
  • @BigDataDZ — 外部贡献者:Goal direction F2 revision-drift fixture(#5549)。
  • @maxliux5 — 外部贡献者:将个人 follow-through 文档归并到已有 manager profile(#5384)。

可选能力启用与使用

Explore Harness

**启用:**在 Goal 设置 → 能力中心选择 evidence 或 planning;下方命令启用 planning,默认关闭。

**验证:**读取同一 Goal/Agent 的 turn-context 与 summary;节点存在不等于已采用。

**停用 / 回退:**执行 --explore-mode off --execute;保留既有证据。

**权限边界:**分析与规划不授予 spawn、claim、扣额、合并或外发权限。

文档:版本固定指南

loopx configure-goal --goal-id "$GOAL_ID" --explore-mode planning --explore-harness-profile adaptive-resilient --execute
loopx explore turn-context --goal-id "$GOAL_ID" --agent-id "$AGENT_ID"
loopx explore summary --goal-id "$GOAL_ID"
loopx configure-goal --goal-id "$GOAL_ID" --explore-mode off --execute

TurnEnvelope and captured decisions

**启用:**每次 guard 显式加 --turn-envelope;保存完整 decision 还需明确 Turn id、已存在父目录与全新目标目录。

**验证:**读取 capture,核对 Goal/Agent/Turn、原始 source hash 与 ok;拒绝或不完整保存不可当成功。

**停用 / 回退:**省略两个参数;仅通过普通文件管理删除不再需要的私有 capture。

**权限边界:**旧观察不提供新准入;选择、lease、取消与写入时校验仍有效。frontend/Lark 不消费这些文件。

文档:版本固定指南

loopx --format json quota should-run --goal-id "$GOAL_ID" --agent-id "$AGENT_ID" --turn-instance-id "$TURN_ID" --turn-envelope --decision-output-dir ./guard-001
cat ./guard-001/decision.json

Ordinary workspace Chat

**启用:**运行 loopx chat,在管家对话的 Scope 选择已授权 workspace。宿主声明的目录默认 workspace_write;需要只读时用下方命令。

**验证:**核对 Scope、生效 grant 与同一对话的结果;grant 改变创建新上下文,旧历史保留。

**停用 / 回退:**以 --project-workspace-grant workspace_read 重启服务,或撤销宿主 workspace grant;旧上下文的新工作被拒绝。

**权限边界:**写入遵循 AGENTS.md 与实际 sandbox;不创建隐藏 Goal,不借 portfolio、peer 委派或外发权限。

文档:版本固定指南

loopx chat --project-workspace-grant workspace_read --no-open

Owner private Lark conversations

**启用:**在设置 → Lark 核验 App 与个人 owner,选择 workspace、executor、grant 和普通 Chat/管家角色后连接。/delegate --tokens N objective 需原来源确认;/agents、/agent TARGET_REF、/project 使用另行授权目标。

**验证:**在绑定私聊发送 /status、/help,核对 App/source、角色、workspace grant、原 Session/Turn 与队列;设置回读同一绑定。

**停用 / 回退:**在设置 → Lark 断开精确 App;/stop 停止普通 Turn,/stop-commission 停止绑定委托;另行撤销精确 Agent grant。

**权限边界:**个人凭据、source/owner、listener 身份仍明确。普通 Chat 不创建 Goal,委托确认不授予任意写入、自动 heartbeat 或规范任务验收。真实 Lark/手机资格另行验收。

文档:版本固定指南

loopx chat --project-workspace-grant workspace_read --no-open
loopx chat --help

Configuration checkpoints

**启用:**设置 → 能力中心 → 配置备份与恢复可下载私有 checkpoint;CLI export 使用新目标,先预览再 --execute。

**验证:**对精确文件运行 configuration-backup verify,核对来源范围与摘要;完整性不证明隐私安全。

**停用 / 回退:**未采用的隔离 checkpoint 可删除而不影响 live settings;已采用设置通过已有 revision 校验 editor 或 machine-config/configure-goal 回退。

**权限边界:**不复制 credential store、Host session、grant、live registry、fence、lease 或 scheduler;配置本来包含的秘密仍属私有。

文档:版本固定指南

loopx --format json configuration-backup export --goal-id "$GOAL_ID" --output "$NEW_CHECKPOINT_FILE"
loopx --format json configuration-backup export --goal-id "$GOAL_ID" --output "$NEW_CHECKPOINT_FILE" --execute
loopx --format json configuration-backup verify --input "$NEW_CHECKPOINT_FILE"

Canonical new-Goal creation

**启用:**设备默认 → 新 Goal 的权威存储显式启用 canonical creation,选择 File/SQLite 及 soft_claim/hard_lease;v1 document 采用已有 preview/apply,默认关闭。

**验证:**inspect 后 bootstrap 新空项目,再读 Todos 和原生 authority receipt;已有 Goal 不改目标。

**停用 / 回退:**预览/应用 canonical_creation=false 关闭未来创建,或按精确 removal-plan revision 删除 goal_storage namespace;已有存储与 fence 保持。

**权限边界:**存储/执行策略不授权工具、账号、网络、调度或迁移;丢失的 authority 不能以 forced bootstrap 当空数据重建。

文档:版本固定指南

Save / 保存为 goal-storage.json:

{"schema_version":"loopx_goal_storage_defaults_v1","new_goal_provider":"sqlite","canonical_creation":true,"new_goal_handoff_mode":"hard_lease"}
loopx machine-config preview --namespace goal_storage --config-json goal-storage.json
loopx machine-config apply --namespace goal_storage --config-json goal-storage.json --expected-plan-revision "$PLAN_REVISION" --execute
loopx machine-config inspect
loopx machine-config remove --namespace goal_storage
#Review removal, then use its returned revision.
loopx machine-config remove --namespace goal_storage --expected-plan-revision "$REMOVAL_REVISION" --execute

Governed delegation stop

**启用:**使用已配置的精确 binding、注册 requester 与 operation,delegation stop --execute 显式请求停止;CLI/MCP 与 App 团队表面复用 owner。

验证:delegation read 区分确认、已证明的原生进程 drain 与结算;监督缺失保持 unknown,执行返回的恢复路径。

**停用 / 回退:**移除精确 operator binding 撤销新执行,对 active operation 请求 stop;停止/结算回执不是 resume grant。

**权限边界:**缺少记录不证明 Host 退出、lease 释放或未执行;不支持的停止平台在 launch-side 取消效果之前拒绝。

文档:版本固定指南

loopx delegation stop --goal-id "$GOAL_ID" --agent-id "$AGENT_ID" --execution-config "$DELEGATION_CONFIG" --operation-id "$OPERATION_ID" --execute
loopx delegation read --goal-id "$GOAL_ID" --agent-id "$AGENT_ID" --execution-config "$DELEGATION_CONFIG" --operation-id "$OPERATION_ID"

PR review queue ownership and direction

**启用:**能力中心配置额外 owner 登录名与 forward/reverse;Goal CLI 可设置精确 Agent 方向。当前会话请求先于普通队列发现。

**验证:**读取 configure-goal 与只读 pr-review packet,核对实际账号、Agent 和方向;queue owner 不等于 GitHub author。

**停用 / 回退:**用 --clear-pr-review-owner-logins --execute、--pr-review-agent-order AGENT=inherit 或 --clear-pr-review-configuration --execute;降级前移除 Goal/设备 owner_logins。

**权限边界:**不授予 GitHub review、comment、dismiss、merge、跨 Agent 写入或调度权限;review 深度和 CI policy 沿用已有 owner。

文档:版本固定指南

loopx configure-goal --goal-id "$GOAL_ID" --pr-review-owner-login maintainer --pr-review-agent-order reviewer-a=forward --execute
loopx configure-goal --goal-id "$GOAL_ID"
loopx pr-review --goal-id "$GOAL_ID" --agent-id reviewer-a --repo "$REPOSITORY" --format json
loopx configure-goal --goal-id "$GOAL_ID" --clear-pr-review-configuration --execute

Public GitHub evidence

**启用:**每个 plan 使用 --public-github 和完整 SHA 来源;execute --execute 只授权匿名来源读取。

**验证:**回读精确 plan/execution receipt;父采纳/拒绝与研究账本覆盖另行显式操作。

**停用 / 回退:**省略 --public-github 与 --execute;没有持久 provider switch。证据退休仍遵循采纳与下游覆盖规则。

**权限边界:**不使用 token、cookie、私有仓库、可变分支、原始页面持久化、自动采纳或外发权限;完整 frontend 发起仍开放。

文档:版本固定指南

loopx external-evidence plan --public-github --objective "Inspect public source" --user-activity "Choose a source" --decision "Whether a literal is present" --evidence-kind literal_match --source "$SHA_PINNED_PUBLIC_URL" --search-term LoopX --format json > plan.json
loopx external-evidence execute --plan-json plan.json --execute --format json > execution.json
loopx external-evidence readback --plan-json plan.json --receipt-json execution.json

Managed DeepSeek model selection

**启用:**managed execution 使用 deepseek-flash@high;显式 LOOPX_TURN_MODEL/DSH_MODEL 或 --dsh-model 仍优先,凭据由 provider 配置。

**验证:**读 turn run-once --help 及返回的 runtime profile/实际 provider 身份;alias 本身不证明 live qualification。

**停用 / 回退:**将 LOOPX_TURN_MODEL 设置为之前明确选择的模型,重启 owning host;不要静默修改旧 session;省略配置恢复 shipped managed default。

**权限边界:**模型设置不提供 API 凭据、付费调用、workspace、Goal 或 provider 晋升权限;CPA route 仍为独立 operator-owned 部署。

文档:版本固定指南

LOOPX_TURN_MODEL=deepseek-flash loopx turn run-once --help
#Explicit rollback selection for the next owning host invocation:
export LOOPX_TURN_MODEL="$PREVIOUS_MODEL"

Finance evidence assessment

**启用:**Finance Value Discovery 是独立安装的可选 extension 0.8.3;从版本固定源码安装,再注册/启用 manifest。assess-period 输入为 finance_period_comparison_input_v1。

**验证:**运行 extension doctor,再对已审阅本地输入 assess-period;producer 精度、期间资格、来源真实性与投资真值分开。

**停用 / 回退:**用下方命令 disable loopx-finance-value-discovery;私有账户/持仓观察不放公开研究投影。

**权限边界:**只做确定性证据评估,不授予投资建议、下单、转账、签名或执行权限;来源/timezone/binding 缺失保持明确。

文档:版本固定指南

#In a v1.3.0 source checkout:
python3 -m pip install ./packages/loopx-finance-value-discovery
loopx extension install --manifest packages/loopx-finance-value-discovery/extension.toml --execute --format json
loopx extension enable loopx-finance-value-discovery --execute
loopx extension doctor loopx-finance-value-discovery --execute
loopx-finance-value-discovery assess-period --input-json period.json
loopx extension disable loopx-finance-value-discovery --execute

EdgeBench native trials

**启用:**仅在 v1.3.0 checkout 中显式启动研究:需要 Linux、Docker、固定 EdgeBench/SForge/Harbor、选定 Codex 与已授权模型。明确 worker 和反馈 profile。

**验证:**先 --help 再审阅试验命令;回读 session/model、终态、评测完成与隔离回执。注册/采样计数不证明有效 outcome。

**停用 / 回退:**不再启动新试验;按已有 cancellation lifecycle 停止精确 trial/controller 与 TLS relay,保留私有日志和未完成结果。

**权限边界:**发布不会启动 benchmark;judge 拥有任务/评分,blind 还需凭据、网络、提交资格隔离验收;原始证据与秘密保持私有。

文档:版本固定指南

python -m benchmark.edgebench.run --help
#Paid execution only after the operator reviews task, pin, isolation and budget:
python -m benchmark.edgebench.run --task "$TASK_ID" --tasks-dir "$TASKS_DIR" --log-dir "$PRIVATE_LOG_DIR" --run-id "$UNIQUE_ATTEMPT" --worker heartbeat-resume --model "$MODEL" --effort xhigh --judge-url "$JUDGE_URL"

发布验证

包/tag 为 1.3.0 / v1.3.0,维护者集成、最终冻结提交、产物及正式发布时间仍待门禁完成。候选已通过版本/readiness、help/manual、双语书籍检查点、必需范围 Ruff、19 文件 Mypy、打包 Chat、42 项 dashboard 命令及改动路径公开边界检查。先前 174 通过/6 失败和任务板 1 通过/22 失败保留为历史证据。

主线回归复用已有 #5533,当前 head 为 3d3beba538fd4c6732dc65a0dae047ebef1f3238。修复已通过 19/19 risk canary、5 项直接检查、Ruff/Mypy、真实进程清理故障、68 项 source/撤权测试,以及打包配置恢复与损坏文件拒绝。固定主线与修复 head 使用相同独立断言的 40 项 transport/projection 测试均通过;未改动的 TraeX 断言从基线 2 项失败变为候选 3 项通过。完整 pytest 仍在运行,并出现一项待最终详情的失败;该修复尚未纳入发布源码,不能称完整回归或最终发布通过。

DS V4.1 Flash/high 已在 8056cc61d 候选隔离原生 Goal 中真实通过:2 Turns、2 Todo 结算及独立验收。最终发布提交仍须重跑。Doubao evolving 的一次真实请求鉴权失败:取到的密钥属于 Ark Agent Plan,公共 evolving endpoint 未接受;这不是模型语义失败。没有启动新的 benchmark。

个人飞书账号权限及旧指南历史已核验;1.3.0 候选指南已读回个人归属、完整正文、两张图和九页导出排版,正式发布后原地更新实际产物与状态。首次 no-clone smoke 在隔离 HOME 清理失败,fresh-clone/update 分别通过;最终 full-public、完整 pytest 和精确 commit 资格仍须完成,真实 PostgreSQL 未被称为通过。

Compare: v1.2.4...v1.3.0

Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant