feat: export engine, KV-cache, and accelerator telemetry - #28
Merged
Merged
Conversation
Section 15 asks for KV-cache occupancy, batch size, GPU utilization and memory, and model-load duration, and section 7.4 makes exporting engine metrics an engine responsibility. None of it was reaching the gateway: the model-load collector was declared and never observed, and quality was only ever predicted, never compared with anything. The gateway now reads the serving path's own exposition (vLLM's vllm:* series plus DCGM_FI_DEV_* from a GPU exporter beside it) on each scrape and republishes the subset section 15 names, so one scrape answers for the whole path and no poller runs when nobody collects. A missing sample stays missing rather than being reported as zero, and an unreachable engine costs its series, not an error. Cold start is measured at the readiness transition instead of configured, and reopens on later recoveries so a reload after an eviction or a lost node is measured too. Live traffic is ungraded, so observed quality uses the one signal production exposes: whether a request declaring routing.structured returned parseable JSON, compared against the quality routing predicted. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
First of several PRs closing real gaps between the design spec and the implementation.
Gap this closes
§15 names KV-cache occupancy, current/average batch size, GPU utilization and memory, and model-load duration; §7.4 makes exporting engine metrics an engine responsibility. None of it reached the gateway:
/metricswas never scraped — only the gateway's view of a request existedrouter_model_load_secondswas declared and never observed (a dead collector)What changed
engine_stats.py— parses Prometheus exposition from the serving path (vLLM'svllm:*series plusDCGM_FI_DEV_*from a GPU exporter beside it), converting DCGM's percent/MiB into §15's ratio/bytes. Tolerant by design: comments, NaN,+Inf, and malformed lines are skipped, replicas reporting separately are summed, and a missing sample staysNonerather than becoming zero.router_engine_running_requests(the live batch),router_engine_batch_size(histogram, so average is_sum/_count),router_engine_waiting_requests,router_engine_kv_cache_occupancy_ratio,router_engine_preemptions_total,router_gpu_utilization_ratio,router_gpu_memory_{used,total}_bytes./metricsis collected, so the gateway stays the single scrape target and no background poller runs when nobody is collecting. An unreachable engine costs its series, never an error.routing.structuredflag returned parseable JSON. Recorded as observed quality 1/0 and compared against the model's predicted quality. The metric's help text and the README both state this limitation; full task-level quality stays with the offline §16 harness. (routing.structuredalso starts closing §7.2's "structured-output requirement" routing feature.)Test plan
ruff format --check .,ruff check .,mypycleanpytest tests/unit tests/integration— 163 passed (was 135), coverage 97.2% vs 90% required🤖 Generated with Claude Code