perf(daemon): stop capping glibc malloc arenas in the service unit - #1594
Merged
Merged
Conversation
The installed systemd unit set MALLOC_ARENA_MAX=2 to bound retained memory. On a 20-worker index the two arenas became the daemon's hot lock: perf on beta.41 put 60% of daemon CPU in __pv_queued_spin_lock_slowpath, every wait under __lll_lock_wait_private from ordinary Vec/String growth, serde_json string decode, and page preparation, while RSS still reached 15 GB. With the cap removed on the same host and corpus the spinlock fell to 4.7% and extraction's share of worker cycles rose from 2% to 23%. The unit leaves the allocator alone; a test keeps it that way.
|
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Finding
After #1570 removed the token clones and the PikeVM fallback, the daemon (local build of master, 12 projects warming, 20 index workers) still spent 87.6% of index-worker cycles in lock waits, 60% of all daemon CPU in
__pv_queued_spin_lock_slowpath. Every wait chain ended in__lll_lock_wait_private— glibc's malloc arena lock — under a dozen unrelated allocation sites (RawVec::finish_grow, serde_jsondeserialize_string,prepare_pageString::clone/free,encode_ngram_bitmap,str::to_lowercase).That pattern only happens when arenas are shared. The daemon's environment:
MALLOC_ARENA_MAX=2. The only process on the host carrying it was the daemon, and the only place in the tree setting it is the systemd user unittracedecay-daemon-controlrenders — a knob chosen to bound retained memory. The beta.40 daemon under the same unit reached 15 GB RSS anyway, so it bought nothing it was meant to buy.Change
Environment="MALLOC_ARENA_MAX=2"from the rendered unit.tracedecay-cli/src/main.rs(it also claimed amalloc_trimthat is not called anywhere).systemd_unit_does_not_cap_malloc_arenaskeeps the unit from re-acquiring the cap.Evidence (same host, same 12 projects, same binary, 10 s
perf record -F 199 -gon the daemon)admit_unitsthrottle and rayon sleep, not malloc (2 samples)tracedecay*)__pv_queued_spin_lock_slowpathself, whole daemonRetained memory: the daemon under the cap sat at 12–15 GB RSS; the sampling window here shows 7.9 GB two minutes into warming. If retained memory becomes a real problem, the answer is an allocator (
alloc-mimalloc/alloc-jemallocfeatures already exist), not a two-arena serializer.