Skip to content

perf(query): sustain dashboard reads and bound peak RSS during mixed ingest #278

Description

@vishr

PR #272 review follow-up: #272 (comment) and #272 (comment)

The completed-batch cache now keeps up with sustained ingestion, but the full dashboard-read capacity gate and peak process memory remain unresolved. Do not equate cache backlog stability or a smaller persisted catalog with sustained read capacity or lower peak RSS.

The fixed Linux ARM64 trial uses 4 CPUs / 6 GiB, same-host telemetrygen 0.157.0, nominal 120k offered rows/s, 20 open-loop reads/s, a 300.154-second mixed phase, two-second requested sampling and thirty-second cooldown. Fanout accepted 66,086 rows/s; the target offered rate was not actually sustained. Source and binary identities, complete harness patch, raw measured JSON, and failed iterations are retained in the review report.

  • Pending batches: minute medians 7, 11, 12, 16, 20.5; sampled maximum 63; last sustained sample 16; manual cooldown observation 0. The preceding iteration ended at 1,690 pending batches.
  • Successful reads: 4,931 (16.4/s); 19 failed and 1,099 shed. Successful-client p50/p95/p99: 88.12ms / 5.955s / 10.630s.
  • Server-handler p50 deteriorates from 32.9ms in minute one to 4.45s in minute five despite the stable cache backlog. Those values come from request logs grouped by request start, not client latencies.
  • Peak process RSS: 3.32 GiB; cooldown RSS: 0.98 GiB. The preceding failed iteration peaked at 3.03 GiB. The smaller cache does not establish a lower process peak.
  • Identical persisted-cache fixture: catalog 180,891,648 -> 70,004,736 bytes (-61.3%), with identical 52,287,828-byte Parquet data. The final catalog remains 1.34x the compressible fixture.
  • Generator exits, dropped-row counters, restarts and full authoritative storage verification pass; the sustained-query-load gate fails.

Profile the remaining ranking/file-binding/scanning work, analytical service/edge work, publication scheduling and retained transaction memory before assigning a cause. Use the fixed five-minute preset with the gauge and actual timestamps, matched accepted data where possible, and a separate-generator run to distinguish contention from backend limits. Preserve exact scoped/clipped results, stable snapshots, typed format 3 and bounded storage publication. Success requires the sustained 20-read/s gate to pass with zero errors/shedding and cache backlog bounded; separately report peak/cooldown RSS, accepted rate, CPU per row and query latency. Do not improve a benchmark by changing its admission/workload limits without disclosing that change.

This issue is separate from #277's ingestion encoding CPU regression. Both remain open after the cache follow-up; neither should be closed by merging #272.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions