memtable: reduce range-deletion fragment rebuild allocations - #6207
Open
tdeebswihart wants to merge 3 commits into
Open
tdeebswihart wants to merge 3 commits into
tdeebswihart wants to merge 3 commits into
Conversation
Member
tdeebswihart
force-pushed
the
tim/keyspanfrash-cheaper-rebuild-clean
branch
from
October 2, 2026 18:55
029fbf7 to
89efad0
Compare
tdeebswihart
marked this pull request as ready for review
October 4, 2026 17:44
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Each batch containing a range deletion invalidates the memtable's cached fragments, causing the next reader to rebuild them. While the cache is being built all readers are blocked, which is causing nasty spikes in the 99th percentile tail latency of our critical workloads.
Rebuilding previously allocated keys for each tombstone and fragment, even when the tombstones were disjoint. This PR keeps the lazy cache design and reduces its rebuild allocations:
Metrics.MemTableRangeDelCacheto understand the dynamics of the cache.pebble bench rangedelto approximate our workload and measure the impact of this optimization.Fragmenter. Range-key rebuilding stays unchanged.Early testing on our production workload shows 99th percentile read latencies at around one third of their previous level.
Earlier rebuild microbenchmarks on an AMD EPYC 9R45, with 1,000 and 10,000 tombstones, showed disjoint rebuilds taking 57-73% less time and mostly-disjoint rebuilds taking 44-69% less time. Disjoint rebuild allocations fell from roughly 2,000-20,000 to 5-6 per operation. Mostly-disjoint allocations fell by about 90%, and allocations fell by about 33% when range deletes overlap.
Future Work
Given how much we rely on range deletion, as a follow-up I'm prototyping an incrementally-maintained rangedel cache. That's a much more invasive change and I've a lot more work to do on it before I waste y'all time on it, but is that something you'd be interested in in the first place?