From 899612faf897ac1248c207a17cabeca171f1609e Mon Sep 17 00:00:00 2001 From: Cto Date: Thu, 24 Sep 2026 16:25:54 +0000 Subject: [PATCH] test(heartbeat): lift waitForHeartbeatIdle past the runs it waits on (PEN-3508) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit `waitForHeartbeatIdle` in issue-monitor-scheduler.test.ts defaulted to a 3s budget while waiting on real agent runs measured at 2s / 5s / 6s on a loaded host — two of three already past the budget that guards them. It therefore failed on load rather than on any defect in the code under test, which is why it passes in CI and in isolation and goes red on a contended shared box. This is the remedy #1792 already applied under BLO-22985, which widened a same-named `waitForHeartbeatIdle` from 5s to 15s in the heartbeat-workspace suites. This copy was missed and sits tighter still. server/vitest.config.ts documents the same policy for BLO-17026: "the tight per-test overrides in the embedded-postgres suites are raised alongside this lift". Raising the default rather than passing at the single call site (:646) matches that config comment's own rationale — lifting the baseline avoids retrofitting every suite — and protects future call sites. 30s is ~5x the observed 6s worst case, above #1792's ~2x rule because the costs are asymmetric: too long only makes a true hang surface slower, and the 60s testTimeout still bounds it, while too short produces false reds. Confirmed pre-existing on master by a matched two-arm run: master alone (85/86) and master + #1910 (93/94) failed on the same test with the same error, so the failure is unrelated to that PR. Deliberately not folded into #1910 — it is not in that PR's files and pushing would restart its CI cycle and queue position for a defect it did not cause. Refs PEN-3508 Signed-off-by: Cto --- .../__tests__/issue-monitor-scheduler.test.ts | 17 ++++++++++++++++- 1 file changed, 16 insertions(+), 1 deletion(-) diff --git a/server/src/__tests__/issue-monitor-scheduler.test.ts b/server/src/__tests__/issue-monitor-scheduler.test.ts index edccffac6780..bf9e897c0029 100644 --- a/server/src/__tests__/issue-monitor-scheduler.test.ts +++ b/server/src/__tests__/issue-monitor-scheduler.test.ts @@ -59,7 +59,22 @@ describeEmbeddedPostgres("issue monitor scheduler", () => { db = createDb(tempDb.connectionString); }); - async function waitForHeartbeatIdle(timeoutMs = 3_000) { + // CI-load margin (PEN-3508), the same remedy #1792 applied under BLO-22985 — + // which widened a same-named `waitForHeartbeatIdle` from 5s to 15s in the + // heartbeat-workspace suites. This copy was missed and sits tighter still. + // + // The wait guards REAL agent runs against embedded Postgres. Timed from this + // suite's own agent.run.started/finished lines on a loaded host: 2s / 5s / 6s + // for the three runs it spawns — two of three already past a 3s budget, so it + // failed on load rather than on any defect in the code under test. Confirmed + // pre-existing by a matched two-arm run (master alone, and master + #1910: + // same test, same error, both arms). + // + // 30s is ~5x the observed 6s worst case, above #1792's ~2x rule because the + // costs are asymmetric: too long only makes a true hang surface slower (and + // the 60s testTimeout still bounds it), while too short produces false reds — + // the failure vitest.config.ts blames for ~21h of skipped deploys. + async function waitForHeartbeatIdle(timeoutMs = 30_000) { const deadline = Date.now() + timeoutMs; while (Date.now() < deadline) { const active = await db