Skip to content

feat: export engine, KV-cache, and accelerator telemetry - #28

Merged
github-actions[bot] merged 1 commit into
mainfrom
feat/15-engine-telemetry
Oct 3, 2026
Merged

github-actions[bot] merged 1 commit into
mainfrom
feat/15-engine-telemetry

Conversation

@Yash-Chindam

Copy link
Copy Markdown
Owner

First of several PRs closing real gaps between the design spec and the implementation.

Gap this closes

§15 names KV-cache occupancy, current/average batch size, GPU utilization and memory, and model-load duration; §7.4 makes exporting engine metrics an engine responsibility. None of it reached the gateway:

  • the engine's own /metrics was never scraped — only the gateway's view of a request existed
  • router_model_load_seconds was declared and never observed (a dead collector)
  • quality was only ever predicted; §15's "predicted versus observed quality" had no observed side

What changed

  • engine_stats.py — parses Prometheus exposition from the serving path (vLLM's vllm:* series plus DCGM_FI_DEV_* from a GPU exporter beside it), converting DCGM's percent/MiB into §15's ratio/bytes. Tolerant by design: comments, NaN, +Inf, and malformed lines are skipped, replicas reporting separately are summed, and a missing sample stays None rather than becoming zero.
  • 7 new metrics — router_engine_running_requests (the live batch), router_engine_batch_size (histogram, so average is _sum/_count), router_engine_waiting_requests, router_engine_kv_cache_occupancy_ratio, router_engine_preemptions_total, router_gpu_utilization_ratio, router_gpu_memory_{used,total}_bytes.
  • Pull-through on scrape — engine state is sampled when /metrics is collected, so the gateway stays the single scrape target and no background poller runs when nobody is collecting. An unreachable engine costs its series, never an error.
  • Cold start is now measured, at the readiness transition — the one place that observes the engine going from loading to serving. The window reopens on later recoveries, so a reload after an OOM eviction or lost node is measured too, not just first start (§13 asks for this to be documented rather than assumed).
  • Observed quality, honestly scoped — live traffic is ungraded, so the only production-observable signal is whether a request declaring the new routing.structured flag returned parseable JSON. Recorded as observed quality 1/0 and compared against the model's predicted quality. The metric's help text and the README both state this limitation; full task-level quality stays with the offline §16 harness. (routing.structured also starts closing §7.2's "structured-output requirement" routing feature.)

Test plan

  • 22 new tests (13 unit for parsing/units/tolerance/cold-start, 9 integration for scrape republishing, graceful degradation, and validity recording on both buffered and streamed paths)
  • ruff format --check ., ruff check ., mypy clean
  • pytest tests/unit tests/integration — 163 passed (was 135), coverage 97.2% vs 90% required

🤖 Generated with Claude Code

Section 15 asks for KV-cache occupancy, batch size, GPU utilization and
memory, and model-load duration, and section 7.4 makes exporting engine
metrics an engine responsibility. None of it was reaching the gateway:
the model-load collector was declared and never observed, and quality was
only ever predicted, never compared with anything.

The gateway now reads the serving path's own exposition (vLLM's vllm:*
series plus DCGM_FI_DEV_* from a GPU exporter beside it) on each scrape
and republishes the subset section 15 names, so one scrape answers for
the whole path and no poller runs when nobody collects. A missing sample
stays missing rather than being reported as zero, and an unreachable
engine costs its series, not an error.

Cold start is measured at the readiness transition instead of configured,
and reopens on later recoveries so a reload after an eviction or a lost
node is measured too. Live traffic is ungraded, so observed quality uses
the one signal production exposes: whether a request declaring
routing.structured returned parseable JSON, compared against the quality
routing predicted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@github-actions github-actions Bot added documentation Improvements or additions to documentation area/api area/tests labels Oct 3, 2026
@github-actions
github-actions Bot merged commit 3449e21 into main Oct 3, 2026
6 checks passed
@github-actions
github-actions Bot deleted the feat/15-engine-telemetry branch October 3, 2026 13:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/api area/tests documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant