fix(daemon): serve cold-daemon reads once their owner is ready - #2612
Conversation
A git owner published before project open installed its authority answered every read as a policy denial; until the authority lands the owner is still mounting, so reads get the retryable mounting problem the CLI re-sends. A whole-generation code read's own demand starts the deferred decode of a published text owner; it now waits for that seat within its resolution deadline instead of answering the first read unavailable. Test fixtures: the surface transport test waits on graph_ready before reading the graph, the runtime fixtures' hermetic PATH regains the toolchain's $CARGO_HOME/bin so the daemon can route rust-analyzer, and the host-event acceptance test asserts the session attribution hook rows carry since #2480 while still refusing it in any other field.
…ssues-master-reds # Conflicts: # crates/tracedecay-code-index-runtime/src/code_index_scheduler/registry/serving_reads.rs
|
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
…ssues-master-reds
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 78e2f28946
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| worktree | ||
| .complete_generation_requested | ||
| .load(Ordering::Acquire) | ||
| && !worktree.memory_retry.waiting() | ||
| && !worktree.shutting_down.load(Ordering::Acquire) |
There was a problem hiding this comment.
Stop waiting after terminal convergence parks
When a retained text owner has no decoded seat and its decode ends in one of the worker's terminal convergence-park paths (for example, reproducible input failure or unrecoverable publication corruption), complete_generation_requested remains permanently true while memory_retry is false and shutdown has not begun. This predicate therefore continues to report work as pending after the worker's final notification; the loop consumes the entire generation-resolution deadline—up to 30 seconds—before returning the already-known unavailable result. Include the convergence park or actual worker/decode activity in this check so terminal background repair cannot block an ordinary code-facet read.
AGENTS.md reference: AGENTS.md:L222-L224
Useful? React with 👍 / 👎.
…ssues-master-reds
Fixes #2482
Fixes #2487
Cold-daemon git reads (#2482)
The git owner registry published an owner before project open installed its authority, so a read in that window got a terminal policy denial (
not_found_or_not_authorized). An owner whose authority is not installed yet is now reported as mounting, which returns the retryableapplication.runtime.mountingproblem that the CLI re-sends. The surface transport test now waits on the product'sgraph_readystatus signal before it reads the graph.Loop proof under
CPUQuota=100%, 8 runs each:core_cli_suite::tool_surface_transport_testpassed 14/14 three times.runtime_acceptance_suite reds (#2487)
wall_microsparity was already fixed on master by test(runtime): align transport fixtures with typed daemon protocol #2489.production_lsp_negotiates_and_projects_canonical_context: this was a fixture regression. The hermetic PATH from test: build fixture child PATH from an explicit allowlist #2447 hid rust-analyzer; the fixture now puts$CARGO_HOME/binback on PATH.every_cursor_carrying_code_operation_mints_and_spends_a_continuation(code_facets): this was a production regression from fix(code-index): serve the first index without the whole decode #2332's deferred decode. The first whole-generation read triggered the decode and then answeredUnavailable. It now waits, within its resolution deadline, for the seat its own demand requested, and re-reads if the seat lands between the miss and the pending check.native_host_event_fixtures_execute_provider_admission_paths: this test pinned the pre-fix(hooks): drain spools promptly and attribute every hook row #2480 behavior. fix(hooks): drain spools promptly and attribute every hook row #2480 intentionally made hook rows carry the session id. The test now asserts the session id is present in that field and still refuses it in every other field. This is the host-event half of test: two lib/runtime_acceptance tests red on master (RMCP disconnect write, Codex host-event privacy) #2569; the RMCP-disconnect half of test: two lib/runtime_acceptance tests red on master (RMCP disconnect write, Codex host-event privacy) #2569 is not addressed here.Built-CLI journey: a cold daemon,
init, thegraph_readywait, then the firstcode_facetsread.unavailable20/20completed20/20, with 6 path facetsSuites:
runtime_acceptance_suite, pre-merge: 136/136 twice, then 136/136 again with the seat race fix at a 15-minute load average of about 320. Thecode_facetstest looped 8/8.runtime_acceptance_suite, merged tree: 135/137. The loom restart test passes onceCARGOis set; outside cargo that variable is missing.workflow_json_preserves_a_typed_application_problem_envelopefails in the workflow CLI path, which this PR does not touch. It passed pre-merge and arrived red with the master merge.tracedecay-code-index-runtimelib: 540/540 pre-merge. Merged tree: 544/545; the failure is the RSS-watermarkmemory_teststest, which passes alone 3/3.core_cli_suite: 159/159 pre-merge.tracedecay-applicationlib: 481/481.application_suite: 66/66.mcp_suite: failures only in the harness's 20 s wait for code-index publication under host load, with the same failure shape on master: 44 failures on master against 41 on the fix for the same 85-test batch. That wait is tracked by test(mcp): graph_query fixtures miss the 20s publication wait under load #2502.-D warnings(tracedecay-code-index-runtime, tracedecay, tracedecay-cli, all targets) andcargo fmt --checkare clean.#2430, not addressed
This PR does not address the SIGTERM exit bound. The measured timeline under a CPU cap shows the cooperative drain expiring while project open is inside the session-store mount's schema install. That store-runtime open is detached and not cancellable, and shutdown closes only idle Ready runtimes. Cutting that wait needs interruptible schema install with typed rollback, which belongs in a separate change.