fix(code-index): end fresh and ready waits after the graph tail - #2365
Merged
Merged
Conversation
A first publication serves from its text owner before the worker's graph tail seats the decoded generation. That seat step takes a counted pass, so a status wait that ended in the gap reported fresh and ready and the next search read verifying for the same, unchanged generation. Fresh and ready waits now also require the worker to be back at a wait, so the reading they return is not followed by a tail of its own pass. The status behavior pins read that readiness through `wait_for` instead of polling plain status until it first looks current.
The one-file publication often seals before the follow-up `branch list` runs, so the pin on a transient `indexing` label raced it. The listing must now read as pending with its pending exact index, or as synced on provenance that is already sealed, and after the seal as synced.
A freed ephemeral port returns to the kernel pool, so rebinding it lost to any process that took the port first. On Linux the test now proves the service's own listening socket is gone from the kernel socket table.
tracing-core caches callsite interest process-wide, and while only one dispatcher is registered a callsite first hit on another thread takes that thread's empty default. The failing recovery-loop test hits the same warning callsite without a subscriber, so under parallel runs the capture lost every event. A registered global dispatcher keeps later registrations consulting the capture.
|
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
This was referenced Sep 28, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #2356
Fixes #2360
Fixes #2342
Four master-red flakes: one production fix (#2356) and three tests pinned to state they could not own. #2347 is already green on master (closed separately with evidence).
#2356:
fresh/readywaits ended before the worker's graph tail (production)Cause. A first publication serves from its text owner before the worker's graph tail seats the decoded generation. The pass counter drops between graph-tail steps by design, so a freshness reading in that gap is
freshwith graphready, andwait_for readyreturned. The seat swap then takes a counted step (lock_scheduler_for_graph_step), so the next read reportedverifyingfor the same unchanged generation. It also missed thedecoded_generationresident owner that the status pins expect.Instrumented failing run on master (timestamps µs):
wait Ready shortcut reached Some(Fresh)at t=98098976, thenswap outcome installs=trueat t=98105575, thensupport.rs:373failed withstaleness_state: "verifying",served_generation == latest_generation. Readiness was declared after the source proof (sealed_currency=true), but before the tail that binds that proof to the seat.Change.
CodeIndexSchedulerRegistryV1::wait_for_readiness(the one authority behindtracedecay_status wait_for, built ondashboard_freshness_readand #2361's reached reading) reachesfresh/readyonly when the owner signals it already subscribes to reportpass_finished(): no pass running and the worker back at a wait.graph_readykeeps "whatever the source freshness". The status behavior tests now usewait_for readyinstead of polling plain status until it first looks current.Deterministic test.
code_index_scheduler::tests::reconcile::ready_wait_ends_only_after_the_graph_tail_seats_the_generationholds the worker right before the serving swap (a newcfg(test)gate besidepause_next_published_text_projection). It asserts the text owner already readsFresh/Reachedwith no seat, then that areadywait is not reached. After release, the wait is reached, the advertised generation is seated, and a later read is stillfresh.ready must not be reached while the graph tail is held: Reached { reading: … staleness_state: Some(Fresh) … code_graph_serving: Some(Ready) … }(FAILED)1 passedLoops (
context_related_order_test::context_ranks_related_symbols…+status_behavior_test::tracedecay_status_reports_the_sealed_branch…+status_behavior_test::project_status_lists_only_its_own_memory_owners…; cold start = binary evicted from the page cache before each run, which is what reproduced the red runs):#2360:
branch addpin raced its own publication (test)Cause. The one-file publication often seals before the follow-up
branch list, so theindexinglabel pin raced it. Observed truthful listings:[…indexing…] … exact index pending,feature/new [current], 4.0 KB (from main), synced just now(sealed, daemon not yet switched to it), and[current, serving] … synced.Change. Assert that set with its evidence. A synced listing must already rest on sealed
graph_sourcein branch meta (sealing never reverts). After the seal, the listing must read synced with noindexing.Loops (20 runs, 4 parallel): original pin 9/20 failed (
admitted branch must be durably visible as indexing); new test 0/24, rebased 0/20.#2342: loopback release proof raced port reuse (test)
Cause. Rebinding a freed ephemeral port loses to any process that takes it first.
Change. On Linux the test records the service listener's socket inode from
/proc/net/tcp(asserting exactly one) and proves that inode is gone after shutdown. Other platforms keep the rebind.Loops. A thief in a private network namespace (ephemeral range 40000-40001; it spin-binds the listened port the instant it is freed) stole the port in every run: master 20/20 failed with the issue's
AddrInUsepanic; this branch 0/20 (rebased 0/20), plain 0/20.daemon-service
identical_failure_logs_are_byte_and_event_bounded(test, no issue)Cause. tracing-core 0.1.36 caches callsite interest process-wide. While exactly one dispatcher is registered (
has_just_one), a callsite's first registration takes its interest from the calling thread's default.repeated_failures_back_off_and_rate_limit_warningshits the samedurable recovery attempt failedcallsite from a subscriber-less thread, so it could cacheneverwhile the capturing test's dispatcher was the only one, and the capture lost the events (left: 0/left: 1,right: 11).Change. One
captured_tracinghelper for both capture tests registers a globalNoSubscriberonce, so later registrations consult the registered set, which includes the capture.Loops (the three
recovery_scheduletests together, 16 parallel): master 2/200 and 7/1000 failed; this branch 0/1000. Full daemon-service lib 20× (5 parallel): 0/20, 322 passed each.Runtime journey (debug CLI from this branch, isolated
HOME/TRACEDECAY_DATA_DIR, one daemon undersystemd-run --user --scope -p MemoryMax=6G -p MemorySwapMax=1G, corpus = a git copy ofcrates/tracedecay-daemon-service, 118 files)Suites
tracedecay-code-index-runtime --lib: 530 passed, 2 ignoredtracedecay --test mcp_suite(full,TRACEDECAY_TEST_BINfrom this branch): 587/587 on the second run. The first run had 1 failure independency_hint_test::search_explicit_lazy_admission…(ignored-dependency-scheduler-unavailable), a path that never calls the readiness wait. It was 0/40 isolated and 0/40 cold on both master and this branch. A full run of the master binary failed a different test (relation_page_cost, at the test(mcp): status and context pins read a transient verifying state #2356support.rs:373assertion). Filed as test(mcp): lazy dependency admission refuses in the pre-seat window #2364.tracedecay --test daemon_suite code_index_journey indexing_lifecycle: 6 passedtracedecay-daemon-service --lib: 322 passedcargo clippy -p tracedecay-code-index-runtime -p tracedecay-daemon-service -p tracedecay -p tracedecay-cli --all-targets --features tracedecay/test-transport,tracedecay/test-helpers -- -D warnings: cleancargo fmt --all -- --check: cleanWindows: the #2342 inode proof is
cfg(target_os = "linux"); non-Linux keeps the prior rebind proof unchanged.