design/telemetry.md says a deployment's replicas are merged before their metrics leave the cluster, and the guide tells an operator that pod labels are gone "because replicas are summed". They are grouped. They are not summed, and the series that reaches a backend carries one replica's value as though it were the fleet's.
Measured on GKE
Two replicas of one ModelDeployment, on two GPU nodes, both scraped (the GPU series arrive from both nodes, so discovery reaches both). Unequal traffic to each, then read each engine's own counter:
10.1.4.11 vllm:prefix_cache_queries_total 300
10.1.0.25 vllm:prefix_cache_queries_total 4640
true total 4940
What Prometheus holds for the merged series:
modelplane_prefix_cache_lookups_total{cluster="gke-us-central",deployment="qwen-demo",
engine="qwen",role="Standalone"} = 300
Stable across repeated samples, so it isn't a race that averages out. 4640 of 4940 is discarded, silently — no duplicate-sample errors, no gap in the graph, just a number that is 6% of the truth and looks entirely plausible.
Why
The pipeline groups with groupbyattrs, which moves the identity attributes onto the resource and merges resources that then match. It does not touch datapoints. Two replicas whose attributes are now identical produce two datapoints per interval under one series, and whichever the exporter writes last is the one that survives.
Nothing in the pipeline aggregates. The design names aggregate_on_attributes for this and the composed config has no equivalent.
What a fix has to decide
The aggregation isn't one function. From the design's own surface:
- counters —
modelplane_prefix_cache_lookups_total, modelplane_requests_preempted_total, modelplane_energy_joules_total — sum
- queue gauges —
modelplane_requests_running, modelplane_requests_waiting — sum
- ratios —
modelplane_kv_cache_utilization_ratio — mean, and the design also wants a _max alongside, because a mean hides the hot replica
- histograms —
modelplane_request_ttft_seconds and the rest — merge the buckets, which is only sound where the boundaries agree (the reason SGLang's latency histograms aren't renamed)
So the aggregation is a property of the metric, and MetricMapping has no field for it. Either the built-ins carry one and an operator's mapping gets a default, or the kind grows a field.
Scope today
Correct with one replica per deployment, which is what every test before this one ran. Wrong from the second replica onward, and the error grows with the fleet.
Found in #470 while reading a Grafana dashboard against a two-replica deployment.
design/telemetry.mdsays a deployment's replicas are merged before their metrics leave the cluster, and the guide tells an operator that pod labels are gone "because replicas are summed". They are grouped. They are not summed, and the series that reaches a backend carries one replica's value as though it were the fleet's.Measured on GKE
Two replicas of one
ModelDeployment, on two GPU nodes, both scraped (the GPU series arrive from both nodes, so discovery reaches both). Unequal traffic to each, then read each engine's own counter:What Prometheus holds for the merged series:
Stable across repeated samples, so it isn't a race that averages out. 4640 of 4940 is discarded, silently — no duplicate-sample errors, no gap in the graph, just a number that is 6% of the truth and looks entirely plausible.
Why
The pipeline groups with
groupbyattrs, which moves the identity attributes onto the resource and merges resources that then match. It does not touch datapoints. Two replicas whose attributes are now identical produce two datapoints per interval under one series, and whichever the exporter writes last is the one that survives.Nothing in the pipeline aggregates. The design names
aggregate_on_attributesfor this and the composed config has no equivalent.What a fix has to decide
The aggregation isn't one function. From the design's own surface:
modelplane_prefix_cache_lookups_total,modelplane_requests_preempted_total,modelplane_energy_joules_total— summodelplane_requests_running,modelplane_requests_waiting— summodelplane_kv_cache_utilization_ratio— mean, and the design also wants a_maxalongside, because a mean hides the hot replicamodelplane_request_ttft_secondsand the rest — merge the buckets, which is only sound where the boundaries agree (the reason SGLang's latency histograms aren't renamed)So the aggregation is a property of the metric, and
MetricMappinghas no field for it. Either the built-ins carry one and an operator's mapping gets a default, or the kind grows a field.Scope today
Correct with one replica per deployment, which is what every test before this one ran. Wrong from the second replica onward, and the error grows with the fleet.
Found in #470 while reading a Grafana dashboard against a two-replica deployment.