Skip to content

feat(benchmark): default LoopX entry to planned and document ablations - #5769

Open
loopx-agent wants to merge 1 commit into
mainfrom
codex/benchmark-planned-default
Open

loopx-agent wants to merge 1 commit into
mainfrom
codex/benchmark-planned-default

Conversation

@loopx-agent

Copy link
Copy Markdown
Collaborator

New LoopX benchmark runs now use loopx-planned when task entry is omitted, so initial task decomposition goes through the existing planning checkpoint. Explicit seeded-todo remains available for compatibility and controlled ablation. Official, single and native-goal behavior stays unchanged.

The shared Execution owner resolves the default for Harbor, EdgeBench, worker environment and LHTB rendering/preflight. Runtime receipts record the resolved selection. The LHTB launcher and SWE-Marathon example follow it. No historical study configuration is rewritten.

benchmark/runtime/SETTINGS.md distinguishes defaults from recommended study settings and supplies matched entry, Explore, turn-envelope, effective-turn cadence, and no-LoopX controls. Effective-turn cadence remains explicit; this default change makes no causal score claim.

Validation: 133 SForge/shared-runtime checks passed with real Harbor and SForge packages; 12 focused bootstrap, resolved-receipt, real CLI planning and adapter-budget checks passed. Four real LHTB renderer readbacks and shell syntax passed. Initial failures from seed-dependent tests were corrected by naming explicit seeded entry; receipt checks now exercise the real worker. Ruff and diff checks passed. No model experiment is part of this PR validation.

Placement: existing benchmark provider adapters, with one shared settings owner; no new control-plane decision rule or frontend capability. Bounded refactor: remove duplicate task-entry validation/defaults in callers. Existing entry and settings interfaces remain reversible through an explicit seeded option. Runtime merge remains maintainer-owned.

…lations

Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer: model_agent; gpt-6.1-sol; OpenAI; runtime_reported; reasoning_effort=xhigh

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

Exact head: 8107a6d

动机

通过 Harbor、EdgeBench 或 LHTB 启动新 LoopX benchmark 任务的研究操作者。
旧版本省略 task entry 时直接创建通用 Todo;新版本先通过现有规划入口分解任务,显式 seeded-todo 仍保留旧行为。
同一输入在八个 LoopX 默认入口从 seeded-todo 变为 loopx-planned;其它模式、显式选项、拒绝信息和完整生成配置保持一致。
这只资格化新运行的默认入口与配置读回,不证明模型效果或分数改善,不改写旧实验或正在运行的 worker。
S11/E3 的模型采用、持续工作和 matched-run 效果仍由既有研究 owner 用新 attempt 验证。

改动思路

省略选项只是调用意图,不能让各 provider 各自解释默认值。已有 Execution 根据 mode 一次性解析 omission:heartbeat、turn、loopx-goal 使用 planned;plain/native-goal 保留 inert seeded 值,然后继续既有合法输入检查。Harbor、SForge、worker、LHTB renderer/preflight 都复用这个配置 owner,EdgeBench 收据记录 worker 的有效选项。

选择 planned 仍进入既有 product planning checkpoint,检查真实 Todo 身份、任务摘要、owner 和可运行状态;budget、等待、pending Turn 与重试规则没有新 owner。显式 seeded 可以回退。与每个 caller 复制默认值相比,该改动更容易修改和验证;Python 位于专门 benchmark adapter,不新增泛用 Python 控制面决策源。

具体改动

完整 15-path diff 已审查:Execution、Harbor、SForge、worker 四个 runtime 文件;EdgeBench 入口和有效收据;LHTB shell、renderer、preflight 与 README;SWE 示例配置;RUNTIME、SETTINGS 和三个既有测试模块。未添加执行模块、评分规则或新的 CLI 选项。SETTINGS.md 占最大新增量,区分操作默认值与可选研究建议,要求显式 pins、matched factors、新 attempt 和历史不可变。

规格:benchmark/runtime/RUNTIME.md,spec_revision 8251ec80e0c327d28d13a1fd64e3f343a2aa3c0e,在读 diff 前读取该不可变版本:

  • Task entry and planning ablation:implemented。入口与 mode 分开;新的 omission default 是披露的有意变化,显式 seeded/planned 保留原语义;规划时间仍消耗既有总预算。
  • Native SForge task entry:implemented。非 heartbeat 拒绝 planned;heartbeat worker 和 runtime 收据都使用已解析的 entry,规划失败或 stale/blocked readback 不进入执行。
  • Each phase keeps an immutable task document:implemented。既有 phase、pending Turn 和等待路径未改写,相关真实 CLI/adapter 测试通过。
  • Migration and qualification:deferred。模型采用、matched study 和 benchmark score 不由该操作默认变更证明,继续由既有 S11/E3 研究 owner 验证。

关键代码讲解

  1. benchmark/runtime/codex.py:28 Execution.__post_init__:先验证 mode/context,只把 task_entry is None 按 existing uses_loopx 解析;empty、unknown 和非 LoopX planned 仍拒绝。
  2. benchmark/runtime/sforge.py:79 SForgeWorker.__init__:去掉平行的默认与 TASK_ENTRIES 检查,复用 Execution;profile、timeout、envelope、cadence 保护保留,official/single/native-goal 的 continuation authority 不变。
  3. benchmark/runtime/worker.py:242 run_once:从环境读取 omission,解析后写入 wake receipt;错误 plan stage、非法 envelope 等既有拒绝路径保留,规划 checkpoint 不作为完成或扣额。
  4. benchmark/LHTB/scripts/render_config.py:15 main:CLI omission 交给共享 owner,并把有效 entry 写进完整 YAML;run.sh 不再强行传 seeded,preflight 同样读取解析后的配置。

对主干的风险

最大风险是跨版本实验省略 entry 后 treatment 已变化,却被误当作只改一个模型/分数。文档明确旧/新默认、显式 rollback、receipt 与 source pins;Explore、envelope、effective-turn cadence 仍独立 opt-in,推荐矩阵不自动启用它们。已有历史配置和 active worker 没有被改写,非 LoopX 模式未获得规划权限。

独立 head 163 项测试通过,source-asserted 不可变 base 148 项通过,安装实际 Harbor/SForge 包且无 dependency skip。另以同一输入运行 58 组 mode/profile/entry/真实 LHTB renderer 对照:八组 omission 按披露变为 planned,除此之外完整有效配置和拒绝语义一致;仅规范化源目录与 traceback 行号,没有丢弃诊断文字。覆盖 explicit seeded、plain/native/heartbeat、非法/empty entry;task-entry suite 还验证真实 CLI 首次规划/复用、fabricated/blocked plan、phase 等待、失败 handoff、预算耗尽和恢复。

独立 renderer 探针第一版遗漏 source PYTHONPATH,base/head 均报同一导入失败;按公开 runner 的 source-checkout 约定补上后,八组 renderer 均完成正/负断言。最初 semantic smoke 命令路径误写,保留该 setup 失败,随后用仓库现有全树脚本通过。Ruff、shell syntax、diff 和 advisory 后的 semantic smoke 通过。模型、Docker/native remote solve 与 scorer 不在本次资格化内,不查询或等待 CI。

语义与 CI 对齐

复用现有 MODES/TASK_ENTRIES,不引入新的共享词汇;None 表达 parser omission,并在一个 owner 解析为原合法集合。默认行为是明确改变,不能声称整个 LoopX omission 路径 feature-off parity;仍保留既有 optional mechanism 的关闭分支和显式 seeded compatibility。研究建议与 machine-enforced mode/entry/budget 条件分别说明。

我的整体评价

APPROVE,交付判断为 justified_increment:新的操作默认在所有实际 caller 中一致且可读回,剩余模型效果交由现有 matched-run 验收。long_horizon preserved:预算、等待、历史输入和规划恢复保持;user_experience improved:调用者省略选项时得到一致有效值,回退只需一个显式 entry 设置,没有新增必须重复输入的信息或确认步骤。

未来变更的有界重构已应用:共享 Execution 取代重复 caller defaults 与 SForge 校验,未增加猜测性 abstraction。默认变化对新 attempt 生效,rollback 使用 explicit seeded 或旧 pinned revision,不改旧收据。未发现当前阻塞;approval 是这个 exact head 的证据结论,CI、平台 aggregate、closeout 和维护者 merge 授权分别判断。同作者账户以 COMMENTED 保存本结论,不能称为 GitHub formal APPROVED。

English verdict: APPROVE - 8107a6d. One existing Execution owner resolves omitted LoopX entry to planned; explicit compatibility, non-LoopX modes and native planning/budget guards remain. Independent head163/base148 tests and58 paired settings/renderer cases pass. Model benefit remains unqualified; CI was not consulted.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant