docs(code-index): serving-current latency after seat keep - #1567
ScriptedAlchemy wants to merge 4 commits into
Conversation
A retryable graph activation used to erase the prepared serving candidate, and an unfinished clone-fingerprint successor withheld the same seat after exact and lexical owners were ready. Keep the candidate in both cases so search can move off the predecessor while graph retries. Co-authored-by: Zack Jackson <ScriptedAlchemy@users.noreply.github.com>
Co-authored-by: Zack Jackson <ScriptedAlchemy@users.noreply.github.com>
|
Co-authored-by: Zack Jackson <ScriptedAlchemy@users.noreply.github.com>
Cross-linkAgrees with keeping Also see:
Flag: if status-census wakes requeue reconcile for ~30–60s after proof expiry, land that cut even if #1568 removes successor from the receipt — otherwise search-can-hit / status-not-fresh can still strand the helper. No merge from this comment. |
Performance Comparison
|
|
Closing as stale. Evidence from the triage: git diff --stat origin/master...origin/cursor/graph-rebuild-current-latency-fb0d = docs/code-index-serving-current-latency.md +83 PLUS the 3 files of b80dd58; doc-only once #1562's fix is attributed to #1562; merge-base 918 behind Reopen with a rebase if the remaining value is wanted. |
Summary
Investigation only. Draft. Do not merge. Does not change the 90s receipt timeout.
Stacks on #1562 (
b80dd58). The new commits are the note indocs/code-index-serving-current-latency.md.Local receipt: after the seat-candidate keep, the serving seat matches and search can hit inside ~90s, but
wait_for_current_generationstill hitsRECEIPT_TIMEOUT. Timeout status hasseated_generation_age_seconds≈ 89. A mid-run sample puts the test thread in that helper atgraph_rebuild_status_test.rs:174(thetimeout(...).awaitaround the loop at148-172), inside status/tool MCP calls. The prior fail did ~16k status polls. Search is not that park: it runs only after status already looks current (153-160).Age is seal time (
status.rs:293-296), so the replacement sealed about a second into the wait. Estimates are structural, not a profiled re-run.Wait helper vs production current
The helper (
graph_rebuild_status_test.rs:140-179) matches productioncurrent(status.rs:451-489): coveragecomplete, stalenessfresh, graphReady, seat match (registry.rs:1205-1231). Search only needs the seated exact/lexical owners. Do not treat seat match plus hits ascurrent, and do not raiseRECEIPT_TIMEOUT(graph_rebuild_status_test.rs:26). Do not add a sleep to thin the polls.Status poll hits capture and project
tracedecay_statusalways snapshots the generation census (crates/tracedecay/src/mcp/tools/handlers/dispatch_groups.rs:533-541, read atstatus.rs:257) even though this test already disables branch diagnostics, storage health, session ingest, and staleness (graph_rebuild_status_test.rs:103-109). The census reader isProjectCodeGraphServingAuthorityV1::project(project_reads.rs:86-139).projecttries the ready-decoded seat first (serving_reads.rs:965-988). That callsready_without_stat(serving_reads.rs:928) →GitMetadataFingerprintV1::capture(reconcile.rs:540-561,identity.rs:167-176) and, on a miss,note_wake_if_idle(serving_reads.rs:924-930). The next arm,retained_text_owner_freshness_for_scope, captures again (serving_reads.rs:1291-1308). Only then does it use the O(1) seat (serving_reads.rs:1342-1364), which neither captures nor wakes.captureis a few git-metadata stats (identity.rs:167-176). ~16k of them are about 1–5s of filesystem, not the 89s. The activation coupling is the wake.note_wake_if_idlecoalesces (registry.rs:2133-2151), but once a pass drains the arrival the next poll posts another. The proof expires at 30s (code_index_scheduler.rs:36, checked atreconcile.rs:562). After that, an unchanged seat still failsready_without_stat, and the census keeps requeueing the worker intoReconcilePassGuard(mount.rs:539). Freshness stays offfreshwhile search can already hit. The freshness ladder the helper reads does not callcapture(serving_reads.rs:324-333).Post-seal serialization
Published pass, one worker:
mount.rs:1109-1141, stop atregistry.rs:2330-2332). Clone pages are not in that await.serving.rs:3204-3217and2650-2658) after dropping the text reservation (serving.rs:3188). Graph then runs under the 128 MiB clone reservation. Full overlap with the text reservation is the RSS failure atmount.rs:1052-1057.mount.rs:1463-1490).activateis awaited before the swap (mount.rs:1601-1706, swap at1818). Retryable failure keeps the candidate (1676-1681). Resident-memoryBudgetExhaustedis a refusal (reconcile.rs:360-367), sodashboard_generation_is_readystays false.mount.rs:1747-1782). The swap already bindspass_proves_latest(1862-1867).mount.rs:1735-1740) contradicts the post-swap idle rule (1978-1991) and re-enters the pass guard before source reconcile (539-541, drop at1172-1174).load_active_sharedunder the scheduler mutex and the serving write lock (mount.rs:1854-1870,ignored_dependencies.rs:158-165).This test compiles the 50ms activation floor (
registry.rs:87-90), not the production 30s.Proposals
#1562-eligible
reconcile_in_progressacross source reconcile for a successor-only continuation (mount.rs:539-541vs1172-1174). About 2–15s of post-seat non-current. Not the whole 89s.mount.rs:1735-1740), but yield once after the swap with the pass guard down. Fixing only this wake is not enough: the status census posts the same wake (serving_reads.rs:924-930and1296-1308).Wait helper
project(project_reads.rs:86-96) should take graph statistics from the seated generation (serving_reads.rs:1342-1364) and must not callready_without_stat,GitMetadataFingerprintV1::capture, ornote_wake_if_idle. Query admission can keep the wake. Up to ~30–60s ofverifyingafter the 30s proof expiry (code_index_scheduler.rs:36), which is the activation stall the sample explains. The 1–5s of fingerprint syscalls are not the 89s by themselves.dispatch_groups.rs:533-541). This test already turns the other sections off (graph_rebuild_status_test.rs:103-109) and still paysproject. The helper does not readgraph_statistics. Combined with (3), the test thread stops parking incaptureandprojectatgraph_rebuild_status_test.rs:174.Follow-up
begin_clone_successorin the owner-installing advance (serving.rs:3212-3217and2654-2658). If the dump shows a resident-memory refusal, this is never-currentversus one publish. Otherwise a few seconds.mount.rs:1754-1782). About 1–8s.builder.rs:32-70) with text advances. Keepclaim_build(code_graph.rs:1244-1316) afterserving.rs:3188. About 5–20s. Do not fully overlap without an RSS measurement.load_active_sharedonly on a cache miss. About 0, or 5–20s on a miss.Do not change
captureandproject, andprojectstill wakes the worker.current.Checklist