Skip to content

perf(daemon): stop capping glibc malloc arenas in the service unit - #1594

Merged
ScriptedAlchemy merged 1 commit into
masterfrom
perf/daemon-malloc-arenas
Sep 18, 2026
Merged

ScriptedAlchemy merged 1 commit into
masterfrom
perf/daemon-malloc-arenas

Conversation

@ScriptedAlchemy

Copy link
Copy Markdown
Owner

Finding

After #1570 removed the token clones and the PikeVM fallback, the daemon (local build of master, 12 projects warming, 20 index workers) still spent 87.6% of index-worker cycles in lock waits, 60% of all daemon CPU in __pv_queued_spin_lock_slowpath. Every wait chain ended in __lll_lock_wait_private — glibc's malloc arena lock — under a dozen unrelated allocation sites (RawVec::finish_grow, serde_json deserialize_string, prepare_page String::clone/free, encode_ngram_bitmap, str::to_lowercase).

That pattern only happens when arenas are shared. The daemon's environment: MALLOC_ARENA_MAX=2. The only process on the host carrying it was the daemon, and the only place in the tree setting it is the systemd user unit tracedecay-daemon-control renders — a knob chosen to bound retained memory. The beta.40 daemon under the same unit reached 15 GB RSS anyway, so it bought nothing it was meant to buy.

Change

  • Remove Environment="MALLOC_ARENA_MAX=2" from the rendered unit.
  • Fix the stale allocator comment in tracedecay-cli/src/main.rs (it also claimed a malloc_trim that is not called anywhere).
  • systemd_unit_does_not_cap_malloc_arenas keeps the unit from re-acquiring the cap.

Evidence (same host, same 12 projects, same binary, 10 s perf record -F 199 -g on the daemon)

index-worker cycles cap = 2 cap removed
lock (kernel spin/futex) 87.6% 26.7% — and those are now the admit_units throttle and rayon sleep, not malloc (2 samples)
extraction (tracedecay*) 2.2% 23.1%
regex (sanitizer, now on the lazy DFA) 0% 18.1%
sha256 / SQLite / memmove ~1% ~13%
__pv_queued_spin_lock_slowpath self, whole daemon 60.0% 4.7%

Retained memory: the daemon under the cap sat at 12–15 GB RSS; the sampling window here shows 7.9 GB two minutes into warming. If retained memory becomes a real problem, the answer is an allocator (alloc-mimalloc/alloc-jemalloc features already exist), not a two-arena serializer.

The installed systemd unit set MALLOC_ARENA_MAX=2 to bound retained memory.
On a 20-worker index the two arenas became the daemon's hot lock: perf on
beta.41 put 60% of daemon CPU in __pv_queued_spin_lock_slowpath, every wait
under __lll_lock_wait_private from ordinary Vec/String growth, serde_json
string decode, and page preparation, while RSS still reached 15 GB. With the
cap removed on the same host and corpus the spinlock fell to 4.7% and
extraction's share of worker cycles rose from 2% to 23%. The unit leaves the
allocator alone; a test keeps it that way.
@changeset-bot

changeset-bot Bot commented Sep 18, 2026

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: 4d72053

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@ScriptedAlchemy
ScriptedAlchemy merged commit 717c223 into master Sep 18, 2026
4 of 6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant