Skip to content

fix(code-index): end fresh and ready waits after the graph tail - #2365

Merged
ScriptedAlchemy merged 4 commits into
masterfrom
fleet/master-red-2356
Sep 28, 2026
Merged

ScriptedAlchemy merged 4 commits into
masterfrom
fleet/master-red-2356

Conversation

@ScriptedAlchemy

Copy link
Copy Markdown
Owner

Fixes #2356
Fixes #2360
Fixes #2342

Four master-red flakes: one production fix (#2356) and three tests pinned to state they could not own. #2347 is already green on master (closed separately with evidence).

#2356: fresh/ready waits ended before the worker's graph tail (production)

Cause. A first publication serves from its text owner before the worker's graph tail seats the decoded generation. The pass counter drops between graph-tail steps by design, so a freshness reading in that gap is fresh with graph ready, and wait_for ready returned. The seat swap then takes a counted step (lock_scheduler_for_graph_step), so the next read reported verifying for the same unchanged generation. It also missed the decoded_generation resident owner that the status pins expect.

Instrumented failing run on master (timestamps µs): wait Ready shortcut reached Some(Fresh) at t=98098976, then swap outcome installs=true at t=98105575, then support.rs:373 failed with staleness_state: "verifying", served_generation == latest_generation. Readiness was declared after the source proof (sealed_currency=true), but before the tail that binds that proof to the seat.

Change. CodeIndexSchedulerRegistryV1::wait_for_readiness (the one authority behind tracedecay_status wait_for, built on dashboard_freshness_read and #2361's reached reading) reaches fresh/ready only when the owner signals it already subscribes to report pass_finished(): no pass running and the worker back at a wait. graph_ready keeps "whatever the source freshness". The status behavior tests now use wait_for ready instead of polling plain status until it first looks current.

Deterministic test. code_index_scheduler::tests::reconcile::ready_wait_ends_only_after_the_graph_tail_seats_the_generation holds the worker right before the serving swap (a new cfg(test) gate beside pause_next_published_text_projection). It asserts the text owner already reads Fresh/Reached with no seat, then that a ready wait is not reached. After release, the wait is reached, the advertised generation is seated, and a later read is still fresh.

  • Fix reverted: ready must not be reached while the graph tail is held: Reached { reading: … staleness_state: Some(Fresh) … code_graph_serving: Some(Ready) … } (FAILED)
  • With fix: 1 passed

Loops (context_related_order_test::context_ranks_related_symbols… + status_behavior_test::tracedecay_status_reports_the_sealed_branch… + status_behavior_test::project_status_lists_only_its_own_memory_owners…; cold start = binary evicted from the page cache before each run, which is what reproduced the red runs):

binary cold, 4 parallel warm, 8 parallel
master 4ada4eb 7/40 failed (lines 204, 344) 3/20 earlier, 0/40 later (load-dependent)
this branch 0/40, rebased 0/24 0/40

#2360: branch add pin raced its own publication (test)

Cause. The one-file publication often seals before the follow-up branch list, so the indexing label pin raced it. Observed truthful listings: […indexing…] … exact index pending, feature/new [current], 4.0 KB (from main), synced just now (sealed, daemon not yet switched to it), and [current, serving] … synced.

Change. Assert that set with its evidence. A synced listing must already rest on sealed graph_source in branch meta (sealing never reverts). After the seal, the listing must read synced with no indexing.

Loops (20 runs, 4 parallel): original pin 9/20 failed (admitted branch must be durably visible as indexing); new test 0/24, rebased 0/20.

#2342: loopback release proof raced port reuse (test)

Cause. Rebinding a freed ephemeral port loses to any process that takes it first.

Change. On Linux the test records the service listener's socket inode from /proc/net/tcp (asserting exactly one) and proves that inode is gone after shutdown. Other platforms keep the rebind.

Loops. A thief in a private network namespace (ephemeral range 40000-40001; it spin-binds the listened port the instant it is freed) stole the port in every run: master 20/20 failed with the issue's AddrInUse panic; this branch 0/20 (rebased 0/20), plain 0/20.

daemon-service identical_failure_logs_are_byte_and_event_bounded (test, no issue)

Cause. tracing-core 0.1.36 caches callsite interest process-wide. While exactly one dispatcher is registered (has_just_one), a callsite's first registration takes its interest from the calling thread's default. repeated_failures_back_off_and_rate_limit_warnings hits the same durable recovery attempt failed callsite from a subscriber-less thread, so it could cache never while the capturing test's dispatcher was the only one, and the capture lost the events (left: 0 / left: 1, right: 11).

Change. One captured_tracing helper for both capture tests registers a global NoSubscriber once, so later registrations consult the registered set, which includes the capture.

Loops (the three recovery_schedule tests together, 16 parallel): master 2/200 and 7/1000 failed; this branch 0/1000. Full daemon-service lib 20× (5 parallel): 0/20, 322 passed each.

Runtime journey (debug CLI from this branch, isolated HOME/TRACEDECAY_DATA_DIR, one daemon under systemd-run --user --scope -p MemoryMax=6G -p MemorySwapMax=1G, corpus = a git copy of crates/tracedecay-daemon-service, 118 files)

tracedecay init
tracedecay tool status --args '{"format":"json","wait_for":{"state":"ready","timeout_ms":60000}}' --json
tracedecay tool search --args '{"query":"run_recovery_loop","limit":1,"format":"json"}' --json
tracedecay tool status --json
status wait: {'outcome': 'reached'} | freshness.status: current | staleness_state: fresh | graph: {'state': 'ready'}
memory owners: ['graph_catalog', 'decoded_generation', 'graph_engine']
search freshness: {'state': 'fresh'} | code_generation == latest: True | results: 1
plain status: ['**code_index_freshness.status:** current']

Suites

  • tracedecay-code-index-runtime --lib: 530 passed, 2 ignored
  • tracedecay --test mcp_suite (full, TRACEDECAY_TEST_BIN from this branch): 587/587 on the second run. The first run had 1 failure in dependency_hint_test::search_explicit_lazy_admission… (ignored-dependency-scheduler-unavailable), a path that never calls the readiness wait. It was 0/40 isolated and 0/40 cold on both master and this branch. A full run of the master binary failed a different test (relation_page_cost, at the test(mcp): status and context pins read a transient verifying state #2356 support.rs:373 assertion). Filed as test(mcp): lazy dependency admission refuses in the pre-seat window #2364.
  • tracedecay --test daemon_suite code_index_journey indexing_lifecycle: 6 passed
  • tracedecay-daemon-service --lib: 322 passed
  • cargo clippy -p tracedecay-code-index-runtime -p tracedecay-daemon-service -p tracedecay -p tracedecay-cli --all-targets --features tracedecay/test-transport,tracedecay/test-helpers -- -D warnings: clean
  • cargo fmt --all -- --check: clean

Windows: the #2342 inode proof is cfg(target_os = "linux"); non-Linux keeps the prior rebind proof unchanged.

A first publication serves from its text owner before the worker's graph
tail seats the decoded generation. That seat step takes a counted pass,
so a status wait that ended in the gap reported fresh and ready and the
next search read verifying for the same, unchanged generation. Fresh and
ready waits now also require the worker to be back at a wait, so the
reading they return is not followed by a tail of its own pass.

The status behavior pins read that readiness through `wait_for` instead
of polling plain status until it first looks current.
The one-file publication often seals before the follow-up `branch list`
runs, so the pin on a transient `indexing` label raced it. The listing
must now read as pending with its pending exact index, or as synced on
provenance that is already sealed, and after the seal as synced.
A freed ephemeral port returns to the kernel pool, so rebinding it lost
to any process that took the port first. On Linux the test now proves
the service's own listening socket is gone from the kernel socket table.
tracing-core caches callsite interest process-wide, and while only one
dispatcher is registered a callsite first hit on another thread takes
that thread's empty default. The failing recovery-loop test hits the
same warning callsite without a subscriber, so under parallel runs the
capture lost every event. A registered global dispatcher keeps later
registrations consulting the capture.
@changeset-bot

changeset-bot Bot commented Sep 28, 2026

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: 0dfe68a

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-28T00:08:17.257247Z 0dfe68a PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@ScriptedAlchemy
ScriptedAlchemy merged commit fd10010 into master Sep 28, 2026
1 check passed
@ScriptedAlchemy
ScriptedAlchemy deleted the fleet/master-red-2356 branch September 28, 2026 00:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant