You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
perf(query): narrow glob pruning, cache footprint and trace-candidate reads #281
PR #272 review follow-ups from the e818a8d review (#272 (review)). None of these affect correctness. Each item below links its original thread.
Captured-ID glob filter does not prune files (medium, measured): feat(telemetry)!: adopt DuckDB 2 and upgrade dependencies #272 (comment) internal/query/file_snapshot.go filters the broad-selection glob with substr(filename, …) IN (ids). DuckDB's multi-file reader does not push that down, so every file matching the glob is opened and scanned. With 100 batches and 80 selected, profiles show 101 files read (explicit list: 80; filename IN ('<full path>', …): 81). This path also loses newest-first ordering for top-N reads. Options: filter on full paths, or raise the explicit-list threshold, and measure statement size against files read.
Read-cache catalog is about as large as the Parquet data (medium-low, measured): feat(telemetry)!: adopt DuckDB 2 and upgrade dependencies #272 (comment)
After a 300 s run at 120k offered rows/s (telemetrygen, two spans per trace), catalog.duckdb was 762 MiB against 529 MiB of Parquet: about 474 MiB of used blocks and 288 MiB free. read_trace_parts and read_trace_candidates each held about 3.6M rows. Document that cache size scales with trace count and that free blocks are not returned without a rebuild, and consider compaction of the cache file.
Trace candidates re-aggregate every overlapping part when any cached file straddles the window (low, plausible): feat(telemetry)!: adopt DuckDB 2 and upgrade dependencies #272 (comment)
Under continuous ingest the newest cached file usually straddles the window end, so most trace reads group read_trace_parts for all overlapping batches instead of reading read_trace_candidates. Restrict the parts aggregation to trace_id IN (partial) and take fully contained traces from the index, as the other branch does. Add a parity test and measure trace read cost as retention grows.
The candidate index is evaluated twice in the contained-files branch (low): feat(telemetry)!: adopt DuckDB 2 and upgrade dependencies #272 (comment)
The branch added in e7dbc79 reads index once for the fully contained rows and again inside partial. On scoped reads with mixed-scope traces this repeats the read_trace_parts union. Results are unaffected. Bind it once, or share it with the rewrite in item 3.
Related: #277 (ingest CPU per row), #278 (sustained read capacity and peak RSS).
PR #272 review follow-ups from the e818a8d review (#272 (review)). None of these affect correctness. Each item below links its original thread.
Captured-ID glob filter does not prune files (medium, measured): feat(telemetry)!: adopt DuckDB 2 and upgrade dependencies #272 (comment)
internal/query/file_snapshot.gofilters the broad-selection glob withsubstr(filename, …) IN (ids). DuckDB's multi-file reader does not push that down, so every file matching the glob is opened and scanned. With 100 batches and 80 selected, profiles show 101 files read (explicit list: 80;filename IN ('<full path>', …): 81). This path also loses newest-first ordering for top-N reads. Options: filter on full paths, or raise the explicit-list threshold, and measure statement size against files read.Read-cache catalog is about as large as the Parquet data (medium-low, measured): feat(telemetry)!: adopt DuckDB 2 and upgrade dependencies #272 (comment)
After a 300 s run at 120k offered rows/s (telemetrygen, two spans per trace),
catalog.duckdbwas 762 MiB against 529 MiB of Parquet: about 474 MiB of used blocks and 288 MiB free.read_trace_partsandread_trace_candidateseach held about 3.6M rows. Document that cache size scales with trace count and that free blocks are not returned without a rebuild, and consider compaction of the cache file.Trace candidates re-aggregate every overlapping part when any cached file straddles the window (low, plausible): feat(telemetry)!: adopt DuckDB 2 and upgrade dependencies #272 (comment)
Under continuous ingest the newest cached file usually straddles the window end, so most trace reads group
read_trace_partsfor all overlapping batches instead of readingread_trace_candidates. Restrict the parts aggregation totrace_id IN (partial)and take fully contained traces from the index, as the other branch does. Add a parity test and measure trace read cost as retention grows.The candidate index is evaluated twice in the contained-files branch (low): feat(telemetry)!: adopt DuckDB 2 and upgrade dependencies #272 (comment)
The branch added in e7dbc79 reads
indexonce for the fully contained rows and again insidepartial. On scoped reads with mixed-scope traces this repeats theread_trace_partsunion. Results are unaffected. Bind it once, or share it with the rewrite in item 3.Related: #277 (ingest CPU per row), #278 (sustained read capacity and peak RSS).