Cross-cluster routing does not work on GKE. Every request through an InferenceGateway to a model on a GKE InferenceCluster returns:
HTTP 503 no healthy upstream
Root cause
compose-inference-gateway publishes a cluster's gateway address as a selectorless headless Service plus an EndpointSlice (fn.py:839, "cluster DNS answers with the EndpointSlice's addresses").
That holds on CoreDNS. GKE runs kube-dns, which builds DNS records from the legacy Endpoints API and was never taught to read EndpointSlice. So the name does not resolve:
$ nslookup gateway-gke-us-central-90145.modelplane-system.svc.cluster.local
Server: 10.2.0.10
** server can't find ...: NXDOMAIN
$ kubectl -n modelplane-system get endpoints gateway-gke-us-central-90145
Error from server (NotFound)
$ kubectl -n modelplane-system get endpointslice -l kubernetes.io/service-name=gateway-gke-us-central-90145
gateway-gke-us-central-90145 IPv4 443 34.67.90.252
Envoy therefore resolves no endpoints for the backend, which is visible in its own statistics:
envoy_cluster_membership_healthy{envoy_cluster_name="httproute/.../qwen/rule/0"} 0
envoy_cluster_ssl_connection_error{...} 0
Zero healthy members and zero TLS errors together mean it never had an address to try, rather than failing to connect to one.
Why the e2e doesn't catch it
kind ships CoreDNS, which reads EndpointSlice, so the same path works locally and in CI. The difference is the DNS implementation, not the cloud.
Scope
Any source: GKE cluster, for every model. The cluster gateway itself is healthy — it refuses an uncertified caller mid-handshake exactly as designed, and the engines behind it serve correctly when addressed directly.
Fix
Compose a legacy Endpoints object alongside the EndpointSlice for the gateway Service, or address the gateway by IP rather than by name. The first keeps the current shape and costs one more object; Endpoints is deprecated in 1.33+ but still what kube-dns reads, and GKE still defaults to kube-dns.
Found while verifying telemetry on a real GKE cluster for #470. It blocks modelplane_frontend_*, which can only be produced once a request reaches the AI gateway's ext-proc.
Cross-cluster routing does not work on GKE. Every request through an
InferenceGatewayto a model on a GKEInferenceClusterreturns:Root cause
compose-inference-gatewaypublishes a cluster's gateway address as a selectorless headless Service plus an EndpointSlice (fn.py:839, "cluster DNS answers with the EndpointSlice's addresses").That holds on CoreDNS. GKE runs kube-dns, which builds DNS records from the legacy
EndpointsAPI and was never taught to readEndpointSlice. So the name does not resolve:Envoy therefore resolves no endpoints for the backend, which is visible in its own statistics:
Zero healthy members and zero TLS errors together mean it never had an address to try, rather than failing to connect to one.
Why the e2e doesn't catch it
kindships CoreDNS, which reads EndpointSlice, so the same path works locally and in CI. The difference is the DNS implementation, not the cloud.Scope
Any
source: GKEcluster, for every model. The cluster gateway itself is healthy — it refuses an uncertified caller mid-handshake exactly as designed, and the engines behind it serve correctly when addressed directly.Fix
Compose a legacy
Endpointsobject alongside the EndpointSlice for the gateway Service, or address the gateway by IP rather than by name. The first keeps the current shape and costs one more object;Endpointsis deprecated in 1.33+ but still what kube-dns reads, and GKE still defaults to kube-dns.Found while verifying telemetry on a real GKE cluster for #470. It blocks
modelplane_frontend_*, which can only be produced once a request reaches the AI gateway's ext-proc.