Skip to content

perf(runtime-core): close idle stores concurrently at shutdown - #2669

Merged
ScriptedAlchemy merged 4 commits into
masterfrom
fleet/fix-unowned-issues-shutdown-close
Sep 29, 2026
Merged

ScriptedAlchemy merged 4 commits into
masterfrom
fleet/fix-unowned-issues-shutdown-close

Conversation

@ScriptedAlchemy

@ScriptedAlchemy ScriptedAlchemy commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Refs #2430

Root cause

The terminal memory_graph_reconciliation store close ends in StoreRuntimeRegistry::close_idle_for_shutdown, which closed every idle SQLite runtime one after another inside a single spawn_blocking. Each close drains its readers, joins its writer, and runs the writer's shutdown TRUNCATE checkpoint. The stores are independent: separate files, writer threads, and reader pools. So the shutdown waited for the sum of the checkpoints and thread hand-offs.

Measured on the #2430 harness (details below, temporary scaffolding since removed). In the test, every close copies the WAL that the fixture's SIGKILLed init daemon left behind. That WAL is the whole store: 634 frames (2.5 MB) for global.db, user-sessions.db, and the project sessions.db. Per store, the thread CPU was:

  • TRUNCATE checkpoint: 5–6 ms, with 2 fsyncs
  • writer connection close: 1–2.7 ms
  • four reader connection closes: about 0.5 ms each, freeing the 741-object schema

At 3–5% quota every blocking hand-off waits for the next 100 ms CFS period, so the serial chain stacked periods. At 5%, one run took 500 ms for store 0 and then 400 ms for store 1. The drain quiesce poll never looped (polls=0 in every run), so the 5 ms poll interval is not a factor. Neither checkpoint fsync is redundant, because a TRUNCATE can only drop the WAL after the database file is durable.

Change

close_idle_for_shutdown gives each reserved runtime its own blocking close and joins all of them. The first failure is kept in reservation order, a join error is typed per close, and dropped handles detach rather than abort, as before. No deadline, drain window, or bound changed.

Evidence

Behavior test, fail-before/pass-after: shard_runtime::registry::tests::attachment::shutdown_close_drains_every_idle_runtime_concurrently. It holds both idle runtimes' drains until both have entered. With origin/master's close.rs swapped in, it fails: every idle runtime drains while the others are still draining left: 1 right: 2. With this branch, it passes. The production-path test terminal_shutdown_truncates_every_released_session_store_wal still passes: two real session stores now close concurrently, and each keeps a 0-byte WAL, so the checkpoint still returns every committed frame.

Throttled SIGTERM loop (the #2430 / #2644 harness, real debug CLI). The test is daemon_sigterm_exits_while_authenticated_project_client_is_connected, driven through a TRACEDECAY_TEST_BIN wrapper:

  • every daemon runs under systemd-run --user --scope -p MemoryMax=6G -p MemorySwapMax=1G
  • the test's daemon gets CPUQuota once it logs event=daemon_ready
  • stderr is timestamped, and the scope's cpu.stat is sampled every 5 ms from outside the scope
  • the host load average was about 80–205 throughout

Before and after use the same core_cli_suite test binary. "master" is 8d4fd5e and "fix" is this change on the same base.

quota master pass fix pass
7% 12/12 12/12
5% 12/12 12/12, plus 7/8 on the final merged binary
4% 19/20 19/20
3% 7/22 (4/12, then 3/10 interleaved) 19/22 (10/12, then 9/10 interleaved), plus 10/12 on the final merged binary

The 5% miss on the final binary spent 1000 ms in background_drain (the git_watcher and session_sync owners) and never reached store close.

Store-close phase at 3%, over all 22 runs of each:

wall median wall max cgroup CPU median
master 1220 ms 1900 ms 35.7 ms
fix 999 ms 1690 ms 26.0 ms

In passing runs, SIGTERM-to-exit CPU was 82 ms median on master and 62 ms on the fix.

Hotpath (--features hotpath, 5%, 4 runs each). On master, runtime_core.registry.close_idle_for_shutdown equals the sum of its rusqlite.attachment.drain calls: 1.00/1.40/1.50/2.02 s against 1.00/1.40/1.50/1.92 s, i.e. serial. On the fix, the drain sum exceeds the enclosing wall time: 1.80/2.00/2.10/3.09 s of drains inside 1.00/1.00/1.20/1.59 s, i.e. concurrent. The daemon.shutdown.store_close median dropped from 1550 to 1150 ms.

At 3% the Hotpath build does not separate the two: its own collector threads (hp-threads, hp-functions, hp-cpu-baseline, …) used about 22 ms of the store-close window's CPU. For that reason the plain binary's pass rates and cgroup CPU are the 3% evidence.

A final one-run journey on the merged binary (final-q3-1, 3%): client drain 302 ms, store close 901 ms (28.1 ms CPU), exited within the 3 s bound.

$ loop.sh final-q3 target/debug/tracedecay core_cli_suite 3% 12
final-q3 quota=3% pass=10 fail=2
# final-q3-1 daemon stderr / test output
[tracedecay] event=daemon_shutdown outcome=client_drain_complete elapsed_ms=300 deadline_remaining_ms=1699
[tracedecay] event=daemon_shutdown outcome=store_close_complete elapsed_ms=901 deadline_remaining_ms=11098
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 168 filtered out; finished in 7.36s

Focused suites (merged with origin/master 154a0eb):

  • tracedecay-global-db --lib: 387 passed
  • tracedecay-runtime-core --lib: 455 passed
  • tracedecay-store-runtime --lib: 124 passed (1 ignored, pre-existing)
  • tracedecay --lib -- daemon::store_runtime_tests daemon::tests::bootstrap daemon::engine project_open: 138 passed
  • core_cli_suite -- tool_daemon_test: 41 passed
  • cargo clippy -p tracedecay-runtime-core -p tracedecay-store-runtime --all-targets -- -D warnings: clean
  • cargo fmt --all -- --check: clean

After merging origin/master 410e8f8, these re-ran and passed:

  • daemon_suite -- store_shutdown_checkpoint_test: 3 passed. This is the production stop journey: graceful stop, and stop during a project open, both truncate the profile and project store WALs.
  • tracedecay-runtime-core --lib shard_runtime::registry: 65 passed
  • core_cli_suite -- tool_daemon_test: 41 passed

ripwire:

  • --edit-check=close_idle_for_shutdown: unchanged, 3 callers compatible.
  • --quality-delta=origin/master..HEAD: gating=0. One minor verbosity row on close_idle_for_shutdown (64→68 LOC; it was already over the bar). One dead-code row on the new #[tokio::test], which ripwire does not see as registered.

Why this is Refs, not Fixes

The issue's measurement used 5% and 3% quotas. 5% holds (31/32 on the fix across both binaries; the one miss was in background_drain). 3% is 85% (29/34), not reliable. Two things remain.

  1. Store close is now at its CPU floor. At 3% the daemon gets 3 ms of CPU per 100 ms period. The close needs about 26–40 ms of CPU, so it takes at least 0.9 s. Most of that CPU is the copy-back of the WAL inherited from the fixture's SIGKILLed init daemon.
    • Pre-checkpointing every inherited WAL before the daemon starts gave 12/12 at 3% with master.
    • Pre-checkpointing only the two profile stores gave 4/12, the same as baseline. So the cost is the project store the in-flight open mounts.
    • I also tried returning an inherited WAL at the writer's first idle point. It only moved that copy into the client-drain window, which is where the in-flight project open mounts the store. 3% improved to 8/12, but 7% and 5% each lost a run. I did not keep it.
  2. The other misses come from phases outside store close. Examples: client_drain at 1.2–2.0 s, where the in-flight project open's join costs 37–54 ms CPU, and background_drain at 1.0 s. The exit tail after store_close_complete also takes 0.4–1.0 s.

The whole SIGTERM-to-exit path needs about 60–80 ms of CPU against the 90 ms that 3 s at 3% allows. Any of those phases can push a run past 3 s.

@changeset-bot

changeset-bot Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: 61e2f21

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@ScriptedAlchemy
ScriptedAlchemy merged commit 180200b into master Sep 29, 2026
4 checks passed
@ScriptedAlchemy
ScriptedAlchemy deleted the fleet/fix-unowned-issues-shutdown-close branch September 29, 2026 19:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant