Skip to content

Replicas are grouped but not aggregated, so a deployment reports one replica's metrics #483

Description

@dennis-upbound

design/telemetry.md says a deployment's replicas are merged before their metrics leave the cluster, and the guide tells an operator that pod labels are gone "because replicas are summed". They are grouped. They are not summed, and the series that reaches a backend carries one replica's value as though it were the fleet's.

Measured on GKE

Two replicas of one ModelDeployment, on two GPU nodes, both scraped (the GPU series arrive from both nodes, so discovery reaches both). Unequal traffic to each, then read each engine's own counter:

10.1.4.11  vllm:prefix_cache_queries_total  300
10.1.0.25  vllm:prefix_cache_queries_total  4640
                                     true total  4940

What Prometheus holds for the merged series:

modelplane_prefix_cache_lookups_total{cluster="gke-us-central",deployment="qwen-demo",
  engine="qwen",role="Standalone"}  =  300

Stable across repeated samples, so it isn't a race that averages out. 4640 of 4940 is discarded, silently — no duplicate-sample errors, no gap in the graph, just a number that is 6% of the truth and looks entirely plausible.

Why

The pipeline groups with groupbyattrs, which moves the identity attributes onto the resource and merges resources that then match. It does not touch datapoints. Two replicas whose attributes are now identical produce two datapoints per interval under one series, and whichever the exporter writes last is the one that survives.

Nothing in the pipeline aggregates. The design names aggregate_on_attributes for this and the composed config has no equivalent.

What a fix has to decide

The aggregation isn't one function. From the design's own surface:

  • counters — modelplane_prefix_cache_lookups_total, modelplane_requests_preempted_total, modelplane_energy_joules_total — sum
  • queue gauges — modelplane_requests_running, modelplane_requests_waiting — sum
  • ratios — modelplane_kv_cache_utilization_ratio — mean, and the design also wants a _max alongside, because a mean hides the hot replica
  • histograms — modelplane_request_ttft_seconds and the rest — merge the buckets, which is only sound where the boundaries agree (the reason SGLang's latency histograms aren't renamed)

So the aggregation is a property of the metric, and MetricMapping has no field for it. Either the built-ins carry one and an operator's mapping gets a default, or the kind grows a field.

Scope today

Correct with one replica per deployment, which is what every test before this one ran. Wrong from the second replica onward, and the error grows with the fleet.

Found in #470 while reading a Grafana dashboard against a two-replica deployment.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions