Skip to content

fix(code-index): serve the first index without the whole decode - #2332

Merged
ScriptedAlchemy merged 1 commit into
masterfrom
fleet/six-gb-serving
Sep 27, 2026
Merged

ScriptedAlchemy merged 1 commit into
masterfrom
fleet/six-gb-serving

Conversation

@ScriptedAlchemy

Copy link
Copy Markdown
Owner

Refs #2123. A first index of this repository under MemoryMax=6G MemorySwapMax=1G now reaches ready (fresh, graph serving) without an OOM kill and without the serving decode. On the final build, 3 of 4 runs had no park. The fourth still parked for 435 s at the sealed graph build's admission before it recovered, so this PR does not close the issue. The remaining gap is at the end.

What was wrong

  • Admission counted clean file pages. The decode admission, the sealed graph build admission, the text build admission and the pressure latch all measured VmRSS. That includes the mapped sealed container and graph store pages, which the kernel drops on demand before the cgroup kill line.
  • Ready required the whole decode. A publication always decoded the whole generation (2.8 GB retained, 1.3–1.9 GB more at its peak) to seat it, and freshness stayed non-terminal until a decoded seat named the advertised generation. Search already serves exact, lexical and graph from the text owner when no seat is held.
  • Project-open owners demanded the decode for no reader. The query authority mount, the advisory mount, the TypeScript diagnostics producer and the registry's feedback-identity fallback all requested the complete generation. They read only the sealed manifest.
  • Refusals parked forever. A refused decode parked convergence even after the graph served. A graph build refused while the same pass's text build held the memory also parked, until a backoff timer fired.

What changed

  • Honest admission (runtime-core):
    • ResidentMemoryPressureV1 samples through an injectable ProcessResidentSamplerV1. The production sampler parses /proc/self/status into VmRSS and RssAnon + RssShmem. Admission, the text-artifact budget and the pressure latch use the unreclaimable bytes.
    • daemon.process.resident_bytes stays VmRSS. A new daemon.process.unreclaimable_resident_bytes gauge carries the admission figure.
    • A refused decode, graph build or text build returns the pool workers' freed allocator pages before it re-measures, whether or not an owner was shed. Before, it re-measured only when an owner was shed.
  • Serving model (code-index-runtime):
    • After publish_sealed_graph, the published pass seats the head on the text owner with recover_verified_head, the same path a restart uses. The whole decode runs only when a reader demanded it (code_index_serving_decode_deferred), and a stale predecessor seat is released.
    • An empty seat is terminal, because search then serves the advertised text owner. released_seat is deleted.
    • An owner that already serves its graph is neither rebuilt nor re-recovered when a decode is demanded.
  • Readers moved off the decode.
    • Graph reads and the generation census already had a text-owner path; with no decode at init they use it.
    • The project-open LSP census reads metadata().snapshot().files.
    • latest_feedback_generation_for_scope no longer awaits a decode.
    • The four eager request_complete_generation / latest_complete_ready* demands are removed.
    • The text owner signals serving_generation_changed when its projection finishes, so owners that waited on the seat still wake.
  • Refusals no longer park.
    • A graph build refused while its own text build runs retries at the text join.
    • A refused decode after the graph serves schedules a memory retry and does not park convergence.
  • Decode charge.
    • A decode samples unreclaimable bytes at every pass boundary (DecodePeakProbeV1) and records decode_peak_growth_bytes. The next decode of that generation is charged max(retained, measured peak) (daemon.code_index.generation.decode.peak_growth_bytes).
    • The sealed graph build keeps the retained charge.
    • Cross-file edges now resolve before the per-file edge copies, into one exact-capacity vector. This shrinks retained edges (fixture: −4,016 B per generation), but it does not measurably bound the transient; see below.

Admission view at the post-text-build retry (6 GB)

observation available below 5.80 GB graph build charge result
#2295 master (base6a) VmRSS 3.32–3.47 GB (anon 3.00–3.06, file 0.17–0.24) 2.33–2.48 GB 2.82 GB refused 25× over 13 min, parked
this PR (fin6c) RssAnon+RssShmem 2.93 GB (VmRSS 3.18) 2.86 GB 2.80 GB admitted

Why #2295 parks where master completed (measured)

Both binaries ran on the same clone at 6 GB.

Per-phase peaks, 6 GB (cgroup RSS+swap / anon, GB)

run extract+seal text artifact graph rows sealed store ready whole-run peak Hotpath VmRSS max
#2295 master 6.83 / 6.42 3.45 / 3.22 refused, parked – never 6.83 4.98
pre-#2295 6.84 / 6.39 3.41 / 2.89 5.74 / 4.91 5.28 / 3.91 OOM in the decode (7.52) 7.52 6.39
acc6a 6.83 / 6.40 3.56 / 3.10 6.11 / 5.34 5.45 / 3.82 322 s 6.83 4.76
acc6b 6.69 / 6.40 3.54 / 3.15 6.09 / 5.39 5.45 / 4.20 344 s 6.69 5.16
acc6c 6.56 / 6.39 3.68 / 3.25 6.27 / 5.21 5.66 / 4.04 724 s (435 s park) 6.56 5.40
dec6 6.71 / 6.40 3.53 / 2.69 6.08 / 4.92 5.44 / 3.68 330 s 6.71 5.24

No run had oom_kill. The earlier builds of this branch reached ready in all 5 runs (322–455 s). acc6a, acc6b, acc6c and dec6 ran on the rebased build before the last guard (no rebuild or re-recovery of an already-serving graph). That guard changes only retained passes under decode demand, not init.

12 GB time to ready (same clone, same host)

runs time to ready peak RSS+swap
master base12a / base12b / base12c 310 / 289 / 243 s 10.95 / 10.96 / 11.38 GB
this PR fin12a / fin12b / dec12 280 / 288 / 281 s 7.13 / 7.17 / 7.14 GB

Time to ready is within noise: master averaged 281 s with a 67 s spread, this PR 283 s. Peak memory dropped by about 3.8 GB because init no longer decodes.

Decode transient

  • Before, from per-pass probes on 12 GB:
    • the decode rose from 4.39 to 9.13 GB, peaking at edge evidence, and settled at 7.79 GB;
    • that is 4.74 GB of peak growth against 3.40 GB retained.
  • After: the decode no longer runs during init. When doctor demands it, decode.peak_growth_bytes reads 5.19 GB. The process-wide probe also counts concurrent work.
  • Counting-allocator fixture (resident_accounting): the transient is 2,407,636 B before and after.
  • Not achieved: the edge restructure does not measurably bound the transient. What this PR does is make the next decode charge the measured peak instead of the retained figure.

Retrieval identity

The #2185 query set (search, context, callers, callees per symbol) was compared against master at 697755b. Per-run identity is normalized: profile cursor keys, handles, timestamps and the time-sealed generation id.

  • j768: 42/42 responses identical.
  • rsbuild: 40/40 responses identical.

Tests

  • memory_tests::a_decode_is_admitted_against_unreclaimable_bytes_not_clean_file_pages
    • The injected view has resident bytes one byte under the watermark and 2 GiB unreclaimable; the decode runs and decodes src/lib.rs::retained_generation.
    • When the unreclaimable bytes fill the headroom, the same decode is refused and nothing is decoded.
    • Fails with the admission reading resident_bytes: "clean file pages do not refuse a decode that fits in anonymous headroom".
  • daemon::code_index_runtime_graph_activation_tests::first_index_serves_graph_reads_without_decoding_the_generation
    • Persistent graph runtime. wait_for_readiness(Ready) reaches, with sealed_decode_count == 0.
    • A graph read returns ["alpha"] with Current freshness. The census reports 28 source bytes.
    • Fails on master: sealed_decode_count is 1.
  • resident_memory::tests::process_status_splits_clean_file_pages_from_unreclaimable_bytes: the /proc/self/status split, with literals.
  • Updated tests:
    • The corrupt-graph restart is renamed and asserts a rebuild from the sealed segments with no decode.
    • The residency literals follow exact-capacity edges.
    • The empty-seat readiness assertion now expects terminal.
    • ProductionProjectCompositionHarnessV1 hands a composition over once its graph catalog is warm.

Proof

  • perf(code-index): make the seal the build's one durable point #2186 kill tests (daemon_suite::sealed_generation_crash_test, against the built CLI): 2 passed.
  • code-index: lib 259, code_index_suite 172, resident_accounting 2.
  • code-index-runtime: lib 525 (2 ignored) after the final rebase onto 41fb9be.
  • runtime-core: lib 441, runtime_core_suite 14, git_repository_authority 16, and the smaller binaries.
  • query: lib 267 (1 ignored), canonical_execution_equivalence 5, retrieval_contract_spine 2, search_quality_suite 70 of 71. reader_refuses_historical_v10_writer_artifact_as_incompatible also fails on master (test(query): artifact reader tests fail on master after bounded restore #2313).
  • graph-db: lib 144 (1 ignored), graph_db_suite 153 (2 ignored).
  • application: lib 477, application_suite 64, pr_tracking 7.
  • tracedecay lib: after the final rebase, the graph-activation and project-open-owner tests pass (25). In the full run on f1d4e56, 765 of 766 passed; the one failure, concurrent_same_identity_worktrees_keep_exact_server_and_scheduler_bindings, fails on master as well.
  • mcp_suite: 585 of 585 after the final rebase, which includes test(mcp): pin owner refusals as typed isError problems #2319. Earlier, move_symbol_test raced the background catalog warm. With the guard and the harness change it passed 15/15 in 6 of 6 runs.
  • Lint: cargo clippy -p tracedecay-runtime-core -p tracedecay-code-index -p tracedecay-code-index-runtime -p tracedecay -p tracedecay-graph-db --all-targets -- -D warnings is clean with and without tracedecay/test-transport. cargo fmt --all -- --check is clean.

Open

  • The sealed graph build is still charged the decode's retained bytes (2.80 GB). Its measured growth at 6 GB is 2.1–2.4 GB. With the base anon near 3.0 GB after the text build, the headroom is 2.67–2.86 GB. In acc6c the build was refused 18 times, short by 30–120 MB, and parked 435 s before a retry fit.
    • Closing this needs a pre-run graph-build charge that is honest. The structural resolution-inputs measure I tried under-counted the fixture's measured peak (3.5 MB against 12.7 MB), so it was dropped.
    • The other route is to run the graph build before the text build, while the base is at about 2.4 GB.
  • Lazy per-partition decode is not implemented. symbol_search, qualified_name, signature_search, source_metadata, facets, timeline, workspace diagnostics, doctor and branch publication still take the whole decode on demand.
  • madvise of mapped pages is not added. With admission counting only anon and shmem, clean mapped pages no longer affect admission. MADV_DONTNEED on a shared file mapping would not uncharge the page cache from the cgroup anyway.

A first index at a 6 GB cap parked forever: after the text build, the
sealed graph build (charged the decode's 2.82 GB retained bytes) and
then the serving decode were measured against VmRSS, which counts
clean mapped container pages, and ready waited on a decoded seat.

- Admission reads RssAnon + RssShmem through the pressure cell's
  sampler (injectable), and a refused decode or text build returns the
  pool workers' freed allocator pages before re-measuring, whether or
  not an owner was shed.
- A publication seats its graph head from the text owner's mapped store
  (verified-head recovery), so ready no longer needs the decode. With
  no decoded seat, search serves the text owner, so an empty seat is
  terminal; the released-seat bookkeeping is gone.
- Project-open owners (query authority, advisory, TypeScript producer),
  the LSP census and feedback identity read the sealed manifest and no
  longer demand the whole decode.
- A graph build refused while its own text build holds the memory
  retries at the text join instead of parking; a refused decode after
  the graph serves does not park convergence.
- A decode records its measured peak growth and its next decode is
  charged that, not the retained figure. Cross-file edges resolve
  before per-file copies into one exact-capacity vector.
@ScriptedAlchemy
ScriptedAlchemy merged commit 60eec1f into master Sep 27, 2026
1 check passed
@changeset-bot

changeset-bot Bot commented Sep 27, 2026

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: 20efb3f

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-27T13:03:22.048063Z 20efb3f PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 20efb3f130

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +1063 to +1064
if (generation_ready || text_graph_catalog_warm.is_some())
&& text_graph_catalog_warm != Some(false)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Keep the readiness probe decode-free

When ProductionProjectCompositionHarnessV1 waits here, the generation_ready probe immediately above calls latest_complete_ready_for_scope, whose AwaitDecode path sets complete_generation_requested. That flag makes registry/mount.rs decline defer_serving_decode, so every default harness open still rehydrates the whole generation before this new text-graph readiness branch can hand over the composition, reproducing the memory pressure this change is intended to avoid. Use a decode-free readiness probe in this path.

AGENTS.md reference: AGENTS.md:L9-L12

Useful? React with 👍 / 👎.

.is_some_and(|text| text.interactive_graph_store().is_ok());
if (generation_ready || recovered_text_ready)
.and_then(|text| text.interactive_graph_store().ok())
.map(|store| store.interactive_catalog_is_warm().unwrap_or(false));

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Propagate catalog warm failures

If the background catalog warm enters InteractiveCatalogState::Failed, the catalog lock is poisoned, or the warm-state check otherwise returns a typed error, interactive_catalog_is_warm() returns Err; converting every error to false makes this loop treat a permanent failure as ordinary warming until the outer 20-second timeout and then report the misleading “did not publish” error. Preserve the typed failure here, retrying only explicitly transient contention if appropriate.

AGENTS.md reference: AGENTS.md:L189-L190

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant