e2e/slash-command-menu.spec.ts:166 › an open menu keeps its container and skills group across projection refreshes fails intermittently on unrelated branches, and it is currently the single most frequent cause of a red test check.
Evidence
Across the 60 most recent failed CI runs, Desktop e2e was the failing step in 9. Five of those were this one test, on five different branches:
Every one reports 1 failed, 61 passed, and every one fails at slash-command-menu.spec.ts:237 — the expect.poll(...).toBe(0) on the mutation counter.
Note the branches have nothing in common; the main occurrence is the same failure on a commit whose own PR head was green.
Why the assertion is wider than its intent
The test arms a MutationObserver before three same-content projection refreshes and asserts that no [role="listbox"] or [role="group"] node was removed. That is the right property to guard (#2667). Two details make it report more than that property:
- The counter only increases. Once
removals is non-zero it can never return to 0, so expect.poll(..., { timeout: 3_000 }) cannot retry into success — it just waits 3 seconds and then fails. The poll adds latency, not tolerance.
- The observation window is not scoped to the refreshes. The observer runs on
document.body with subtree: true for the whole poll window, and it counts any removed element that contains a listbox or group anywhere beneath it. A toast, tooltip, or popup unmounting during those 3 seconds is counted the same as the regression the test is guarding.
So a failure here means "something containing a group was removed somewhere in the document during a 3-second window", which is a superset of "the skills group was torn down by the projection refresh".
What not to do
apps/desktop/e2e/playwright.config.ts sets retries: 0 with an explicit rationale — flakes should fail loudly. That is a deliberate choice and this issue is not a request to change it. Adding retries would hide this rather than fix it.
Suggested direction
Scope the observation to what the test actually means:
- Start the observer immediately before the refresh rounds and disconnect it as soon as they complete, instead of leaving it running for the poll window.
- Observe the menu container rather than
document.body, so unrelated overlays cannot contribute.
- Assert the count once after the refreshes have settled, rather than polling a monotonically increasing counter.
Anyone picking this up should confirm the failure is genuinely the observation window and not a real intermittent teardown — the trace and video are retained on failure (trace: 'retain-on-failure'), and the artifacts on the runs above are the place to start.
简体中文
e2e/slash-command-menu.spec.ts:166 › an open menu keeps its container and skills group across projection refreshes 会在互不相干的分支上间歇性失败,目前是导致 test 检查变红的最主要单一原因。
证据
在最近 60 次失败的 CI 运行中,Desktop e2e 是失败步骤的有 9 次。其中 5 次是同一个用例,分布在五个不同分支上:
每一次都是 1 failed, 61 passed,且每一次都失败在 slash-command-menu.spec.ts:237——也就是对 mutation 计数器的 expect.poll(...).toBe(0)。
请注意这些分支之间毫无共同点;main 上那次是同一个失败,而它对应的 PR head 自身是绿的。
为什么这个断言比它的意图更宽
该测试在三次同内容的 projection 刷新之前装上一个 MutationObserver,断言没有任何 [role="listbox"] 或 [role="group"] 节点被移除。要守的性质是对的(#2667)。但有两处细节让它报出的东西超出了这个性质:
- 计数器只增不减。
removals 一旦非零就再也回不到 0,因此 expect.poll(..., { timeout: 3_000 }) 不可能重试成真——它只是等满 3 秒然后失败。这个 poll 带来的是延迟,不是容忍度。
- 观察窗口没有限定在刷新期间。 观察器以
subtree: true 挂在 document.body 上、贯穿整个 poll 窗口,并且只要被移除的元素内部任何位置含有 listbox 或 group 就计数。在那 3 秒里卸载的一个 toast、tooltip 或弹层,与测试真正要守的那个回归被同等对待。
所以这里的失败意味着"在 3 秒窗口内,文档中某处移除了一个含 group 的元素"——这是"skills group 被 projection 刷新拆掉"的一个超集。
不要做什么
apps/desktop/e2e/playwright.config.ts 设的是 retries: 0,并写明了理由——flake 应当响亮地失败。那是有意的选择,本 issue 不是要求改它。加重试是把问题盖住,不是修好。
建议方向
把观察范围收敛到测试真正想表达的意思:
- 在刷新轮次开始前才装上观察器,刷新一结束就 disconnect,而不是让它在整个 poll 窗口里一直跑。
- 观察菜单容器本身而不是
document.body,这样无关的浮层无法贡献计数。
- 在刷新稳定之后一次性断言计数,而不是去轮询一个单调递增的计数器。
接手的人应当先确认这确实是观察窗口的问题,而不是一次真实的间歇性拆除——失败时 trace 与 video 都会保留(trace: 'retain-on-failure'),上面几次运行的产物就是起点。
e2e/slash-command-menu.spec.ts:166 › an open menu keeps its container and skills group across projection refreshesfails intermittently on unrelated branches, and it is currently the single most frequent cause of a redtestcheck.Evidence
Across the 60 most recent failed CI runs,
Desktop e2ewas the failing step in 9. Five of those were this one test, on five different branches:main(push)fix/align-usage-request-countsfeat/runtime-host-update-reconciliationConnect-Custom-relay-fetch-modelsfeat/github-copilot-device-flow-loginEvery one reports
1 failed, 61 passed, and every one fails atslash-command-menu.spec.ts:237— theexpect.poll(...).toBe(0)on the mutation counter.Note the branches have nothing in common; the
mainoccurrence is the same failure on a commit whose own PR head was green.Why the assertion is wider than its intent
The test arms a
MutationObserverbefore three same-content projection refreshes and asserts that no[role="listbox"]or[role="group"]node was removed. That is the right property to guard (#2667). Two details make it report more than that property:removalsis non-zero it can never return to 0, soexpect.poll(..., { timeout: 3_000 })cannot retry into success — it just waits 3 seconds and then fails. The poll adds latency, not tolerance.document.bodywithsubtree: truefor the whole poll window, and it counts any removed element that contains a listbox or group anywhere beneath it. A toast, tooltip, or popup unmounting during those 3 seconds is counted the same as the regression the test is guarding.So a failure here means "something containing a group was removed somewhere in the document during a 3-second window", which is a superset of "the skills group was torn down by the projection refresh".
What not to do
apps/desktop/e2e/playwright.config.tssetsretries: 0with an explicit rationale — flakes should fail loudly. That is a deliberate choice and this issue is not a request to change it. Adding retries would hide this rather than fix it.Suggested direction
Scope the observation to what the test actually means:
document.body, so unrelated overlays cannot contribute.Anyone picking this up should confirm the failure is genuinely the observation window and not a real intermittent teardown — the trace and video are retained on failure (
trace: 'retain-on-failure'), and the artifacts on the runs above are the place to start.简体中文
e2e/slash-command-menu.spec.ts:166 › an open menu keeps its container and skills group across projection refreshes会在互不相干的分支上间歇性失败,目前是导致test检查变红的最主要单一原因。证据
在最近 60 次失败的 CI 运行中,
Desktop e2e是失败步骤的有 9 次。其中 5 次是同一个用例,分布在五个不同分支上:main(push)fix/align-usage-request-countsfeat/runtime-host-update-reconciliationConnect-Custom-relay-fetch-modelsfeat/github-copilot-device-flow-login每一次都是
1 failed, 61 passed,且每一次都失败在slash-command-menu.spec.ts:237——也就是对 mutation 计数器的expect.poll(...).toBe(0)。请注意这些分支之间毫无共同点;
main上那次是同一个失败,而它对应的 PR head 自身是绿的。为什么这个断言比它的意图更宽
该测试在三次同内容的 projection 刷新之前装上一个
MutationObserver,断言没有任何[role="listbox"]或[role="group"]节点被移除。要守的性质是对的(#2667)。但有两处细节让它报出的东西超出了这个性质:removals一旦非零就再也回不到 0,因此expect.poll(..., { timeout: 3_000 })不可能重试成真——它只是等满 3 秒然后失败。这个 poll 带来的是延迟,不是容忍度。subtree: true挂在document.body上、贯穿整个 poll 窗口,并且只要被移除的元素内部任何位置含有 listbox 或 group 就计数。在那 3 秒里卸载的一个 toast、tooltip 或弹层,与测试真正要守的那个回归被同等对待。所以这里的失败意味着"在 3 秒窗口内,文档中某处移除了一个含 group 的元素"——这是"skills group 被 projection 刷新拆掉"的一个超集。
不要做什么
apps/desktop/e2e/playwright.config.ts设的是retries: 0,并写明了理由——flake 应当响亮地失败。那是有意的选择,本 issue 不是要求改它。加重试是把问题盖住,不是修好。建议方向
把观察范围收敛到测试真正想表达的意思:
document.body,这样无关的浮层无法贡献计数。接手的人应当先确认这确实是观察窗口的问题,而不是一次真实的间歇性拆除——失败时 trace 与 video 都会保留(
trace: 'retain-on-failure'),上面几次运行的产物就是起点。