Skip to content

About

Terraform module for deploying SIE on Google Cloud GKE

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

SIE GKE Terraform Module

One command to get a GPU-ready GKE cluster for SIE (Search Inference Engine). The module creates the underlying GCP resources (VPC, GKE, GPU node pools, Artifact Registry, IAM, a model-cache + payload-store GCS bucket created by default); the SIE application itself - gateway, sie-config, workers, KEDA, Prometheus, Grafana, Loki, NATS - is deployed on top via the sie-cluster Helm chart.

  • GPU node pools sized for scale-to-zero via KEDA (configured in the Helm chart)
  • Artifact Registry with cleanup policies
  • Workload Identity for GCS access

What you get

  • GKE cluster with VPC-native networking, private nodes, and Cloud NAT
  • GPU node pools - L4, T4, A100, or A100-80GB, with automatic driver installation
  • Scale-to-zero - GPU nodes scale down to zero when idle, so you only pay when running inference
  • Node Auto-Provisioning (NAP) - GKE automatically creates node pools to fit pending workloads
  • Artifact Registry - private Docker registry with automatic cleanup policies for dev images
  • Workload Identity - pods authenticate to GCP without service account keys
  • Observability-ready - outputs wired for the Helm chart's Prometheus, Grafana, Loki, and KEDA integration
  • Paired with the sie-cluster Helm chart - Kubernetes workloads (gateway, sie-config, workers, NATS, ingress, auth) are installed on top of this cluster via Helm

Module structure

Layer Path What it creates
Infrastructure infra/ GCP resources only: VPC, GKE cluster, node pools, IAM, Artifact Registry, a model-cache + payload-store GCS bucket (created by default). Can be applied without a running cluster.
Application sie-cluster Helm chart Kubernetes resources: sie-config, gateway, workers, NATS, KEDA, Prometheus, Grafana, Loki, optional ingress + oauth2-proxy. Applied after the cluster is up.

Examples in examples/ use the infra/ submodule directly and deploy K8s resources via the Helm chart in a follow-up step.

Quick start

cd examples/dev-l4-spot
export TF_VAR_project_id="your-project-id"
# CIDRs allowed to reach the Kubernetes API; include this machine's egress
# address (for example the /32 of `curl -s https://checkip.amazonaws.com`).
export TF_VAR_api_server_authorized_ip_ranges='["203.0.113.10/32"]'
terraform init
terraform plan
terraform apply

203.0.113.10/32 is a documentation placeholder. The module rejects documentation ranges, so replace it with your own address. See Kubernetes API access for the private-endpoint mode.

After apply, configure kubectl and deploy SIE with chart 0.9.0. The chart selects the published v0.9.0 service images and v0.9.0-cuda12-default worker image for the GKE overlay. The Terraform module version is independent of the SIE application version:

# Point kubectl at the new cluster
$(terraform output -raw kubectl_command)

# Fetch the matching GKE overlay and deploy the published chart
curl -fsSL -o values-gke.yaml \
  https://raw.githubusercontent.com/superlinked/sie/v0.9.0/deploy/helm/sie-cluster/values-gke.yaml
helm upgrade --install sie-cluster oci://ghcr.io/superlinked/charts/sie-cluster \
  --version 0.9.0 -f values-gke.yaml \
  --create-namespace -n sie \
  --set-string "serviceAccount.annotations.iam\\.gke\\.io/gcp-service-account=$(terraform output -raw sie_workload_service_account)" \
  $(terraform output -raw model_cache_helm_args)

This creates no Ingress: the gateway Service is ClusterIP. To expose the gateway outside the cluster, enable the Ingress together with gateway authentication and TLS, as described in the chart's Ingress section.

To reach the gateway from your machine without an Ingress, forward a local port to its ClusterIP Service:

kubectl -n sie port-forward svc/sie-gateway 8080:8080

Upgrading to SIE 0.9.0

Chart 0.9.0 has breaking changes. Read the SIE 0.9.0 release notes before upgrading an existing release. For a release installed with the command above:

  • NATS authentication is on by default. The upgrade restarts NATS and rolls sie-config, the gateway, and the workers. With the default memory-backed work queues, queued and in-flight work is lost, as on any NATS restart. A gateway, sie-config, or worker pod that has not been replaced yet has no credentials, and NATS refuses it: requests can fail with 503, and workers that have not restarted take no work until they do. To avoid that gap, run the command above twice: first with --set nats.auth.allowAnonymous=true added, then, once every pod has restarted, with --set nats.auth.allowAnonymous=false. Between the two steps NATS also accepts anonymous clients, with unrestricted permissions, and the chart ships no NetworkPolicy for the NATS pods, so allow only trusted workloads to reach NATS. The first step still restarts NATS, so the two steps do not prevent the loss of queued and in-flight work. See the chart's NATS authentication section.

  • Pass values explicitly. helm upgrade --reuse-values now fails to render. Re-run the full command above, which passes the values file with -f, or use --reset-then-reuse-values (Helm 3.14 or later).

  • The GKE values file no longer enables the gateway Ingress. The upgrade removes the host-less, plain-HTTP Ingress that earlier releases created. To keep external access, enable the Ingress with gateway authentication and TLS. To keep the previous unauthenticated catch-all Ingress, set ingress.enabled=true together with ingress.allowUnauthenticated=true and ingress.allowPlaintext=true. A LoadBalancer or NodePort gateway Service needs gateway authentication or gateway.service.allowUnauthenticated=true, and also gateway.service.allowPlaintext=true, because the gateway serves plain HTTP.

  • sie-config tokens are split. The upgrade generates a sie-config admin token and a separate read token for the gateway and the worker sidecars. The GKE values file sets fullnameOverride: sie, so the Secrets are sie-config-admin-token and sie-config-read-token, not the sie-cluster-config-* names in the chart's examples. sie-config then requires a token on every /v1/configs request, so give the admin token to tooling that writes model configs. Read it with:

    kubectl get secret -n sie sie-config-admin-token \
      -o jsonpath='{.data.SIE_ADMIN_TOKEN}' | base64 -d

    The gateway no longer receives that admin token: with gateway authentication enabled, its admin routes (POST, PUT, and DELETE under /v1/pools, /v1/admin, and /v1/configs) answer 403 until gateway.auth.adminTokenSecretName names a separate Secret. Run the sie-config, gateway, and worker sidecar images of the same release. See the chart's sie-config tokens section.

For existing installations with custom model profiles, also review the SIE 0.8.0 breaking changes for adapter options and launch arguments before upgrading from 0.7.x.

Examples

Example GPU Description
dev-l4-spot L4 (g2-standard-8) Spot instances, scale 0-5 nodes, minimal cost for development

Prerequisites

  1. GCP project with billing enabled
  2. GPU quota in your target region - check with: gcloud compute regions describe REGION --format="table(quotas.filter(metric:NVIDIA))". Request increases at IAM & Admin > Quotas.
  3. APIs enabled: container.googleapis.com, compute.googleapis.com, artifactregistry.googleapis.com
  4. Terraform >= 1.14

Bootstrap (CI/CD)

For CI/CD pipelines, create a deployer service account with the required IAM roles:

cd bootstrap
export TF_VAR_project_id="your-project-id"
terraform init
terraform apply

This creates a service account with the minimum roles needed to deploy SIE infrastructure. See bootstrap/main.tf for details.

Variables

Required

Variable Description
project_id GCP project ID
region GCP region (e.g., us-central1, europe-west4)

You must also choose how the Kubernetes API is reached; see Kubernetes API access.

Cluster

Variable Default Description
cluster_name sie-cluster GKE cluster name
deletion_protection true Prevent accidental deletion (set false for dev)
kubernetes_version null (latest) Pin Kubernetes version, or let GKE manage it
release_channel REGULAR RAPID, REGULAR, STABLE, or UNSPECIFIED
deployer_service_account "" Email of the SA running Terraform (auto-detected in CI/CD)

GPU configuration

Variable Default Description
gpu_node_pools 1x L4 spot pool List of GPU node pool configurations (see below)
cpu_node_pool e2-standard-4 CPU pool for system workloads (kube-system, monitoring)
kubelet_container_log_max_size 20Mi Per-container kubelet log file size before rotation
kubelet_container_log_max_files 30 Rotated files retained per container; kubelet retention is size/count based, not hourly

Each entry in gpu_node_pools supports:

Field Required Default Description
name yes n/a Pool name (e.g., l4-spot)
machine_type yes n/a GCE machine type
gpu_type yes n/a Accelerator type
gpu_count yes n/a GPUs per node
min_node_count yes n/a Minimum nodes (0 = scale-to-zero)
max_node_count yes n/a Maximum nodes
spot no false Use spot VMs (~60-91% savings)
disk_size_gb no 100 Boot disk size
disk_type no pd-ssd Boot disk type
local_ssd_count no 0 NVMe local SSDs for model cache
zones no all Restrict to specific zones
taints no [] Kubernetes taints for GPU isolation
labels no {} Additional node labels

For a multi-GPU worker pod, set the node pool gpu_count to at least the matching Helm workers.pools.<name>.gpu.count. Kubernetes can only schedule a pod requesting N GPUs onto a node that advertises N allocatable GPUs.

GPU machine cheat sheet:

GPU Machine Type VRAM Approx. spot/hr Best for
L4 g2-standard-8 24 GB ~$0.50 Development, small/medium models
T4 n1-standard-8 16 GB ~$0.35 Budget inference
A100 40GB a2-highgpu-1g 40 GB ~$3.60 Large models, production
A100 80GB a2-ultragpu-1g 80 GB ~$5.10 Maximum VRAM

Network

Variable Default Description
create_network true Create VPC and subnet (set false to use existing)
network sie-network VPC name
subnetwork sie-subnet Subnetwork name
subnet_cidr 10.0.0.0/20 CIDR range for the subnetwork
pods_cidr 10.1.0.0/16 Secondary CIDR range for pods
services_cidr 10.2.0.0/20 Secondary CIDR range for services
enable_private_nodes true No public IPs on nodes (Cloud NAT for egress)
master_ipv4_cidr_block 172.16.0.0/28 CIDR block for the master network

Kubernetes API access

The GKE control plane is never open to the whole Internet unless you ask for it. The plan fails until you choose one mode:

Variable Default Description
authorized_networks [] CIDRs (with display names) allowed to reach the Kubernetes API through master authorized networks. Include every machine that runs kubectl or helm against the cluster. The list restricts the public endpoint, and also the private endpoint in private-endpoint mode.
enable_private_endpoint false Disable the public endpoint and serve the API only on the private endpoint. Requires enable_private_nodes. The private endpoint is reachable from the cluster's VPC network in the cluster's region; this module does not enable access from other regions. A non-empty authorized_networks (for example a VPN range) is then also enforced on the private endpoint; with an empty list any address that reaches the private endpoint is admitted. Enforcement needs a control plane at GKE 1.28.10-gke.1058000 or later with Envoy enabled; otherwise GKE rejects the update and access stays unchanged.
allow_public_api_server false Explicit opt-in to accept any Internet address. With an empty authorized_networks the module leaves master authorized networks unmanaged.

Rules for authorized_networks:

  • Entries must be IPv4 CIDR blocks, because the module creates an IPv4 cluster.
  • At most 100 entries, the GKE limit on authorized networks.
  • With the public endpoint enabled, the entries together may cover at most 16,777,216 addresses, the size of one /8. 0.0.0.0/0, split halves such as two /1 blocks, and several broad ranges are rejected unless allow_public_api_server = true. In private-endpoint mode the list holds internal ranges (for example all three RFC 1918 ranges) and has no total limit.
  • Entries inside a documentation range (192.0.2.0/24, 198.51.100.0/24, 203.0.113.0/24) are rejected, so an unedited placeholder fails at plan time. Broader entries that contain one need allow_public_api_server = true.

In the public-endpoint mode the list does not restrict the private endpoint, which stays reachable from the cluster's VPC network in its region. When master authorized networks are managed, access from Google Cloud public IP addresses is disabled. Network restrictions are in addition to Kubernetes API authentication and authorization. Apart from the unauthenticated health and version endpoints (such as /healthz, /readyz, and /version), every request must still pass them.

This module installs nothing in the cluster, so if the list stops including your address, correct authorized_networks and apply again to restore access.

Upgrading from 0.x.

  • Earlier versions left the public endpoint open to any address when authorized_networks was empty. That configuration now fails the plan with a message asking you to choose.
  • Configurations that already set authorized_networks keep working if the entries are IPv4, not documentation ranges, at most 100, and no broader than one /8 in total (otherwise set allow_public_api_server = true). The plan may show gcp_public_cidrs_access_enabled = false if it was enabled outside Terraform. GKE documents that this change can take several hours to be enforced, so verify the effective access before treating the endpoint as restricted.
  • To restrict an open cluster, set authorized_networks. The plan shows an in-place update that adds master_authorized_networks_config.
  • To keep the previous behaviour explicitly, set allow_public_api_server = true. The plan shows no change to the endpoint.
  • enable_private_endpoint is also an in-place update.

Node Auto-Provisioning (NAP)

Variable Default Description
enable_node_auto_provisioning true Let GKE auto-create node pools for pending pods
nap_max_cpu 1000 Maximum CPU cores NAP can provision
nap_max_memory_gb 4000 Maximum memory NAP can provision

Application layer

The infra/ module only creates GCP resources (VPC, GKE, node pools, IAM, Artifact Registry). The SIE application - gateway, sie-config, workers, observability stack, NATS, optional ingress + auth - is deployed separately via the sie-cluster Helm chart. All install_*, sie_*, and nats_* knobs live on the Helm values file (see the 0.9.0 chart values), not on this Terraform module.

Outputs

After terraform apply, use these outputs to connect and deploy:

Output Description
kubectl_config_command Run this to configure kubectl
cluster_name GKE cluster name
cluster_endpoint GKE cluster API endpoint (sensitive)
artifact_registry_url Where to push Docker images
artifact_registry_server_repository_url Where to push sie-server images
artifact_registry_gateway_repository_url Where to push sie-gateway images
artifact_registry_config_repository_url Where to push sie-config images
sie_workload_service_account GCP service account email for the Helm Workload Identity annotation
workload_identity_annotation Precomposed iam.gke.io/gcp-service-account=email pair
gpu_node_pools GPU pool configs (for Helm worker pool mapping)
gpu_node_pool_disk_sizes_gb Boot disk size per configured GPU node pool

Architecture

                      +----------------------------------------------------------+
                      |                    GCP Project                           |
                      |                                                          |
+----------+          |  +----------------------------------------------------+  |
|          |  HTTPS   |  |              VPC (private nodes + Cloud NAT)       |  |
|  Client  |--------> |  |                                                    |  |
|          |          |  |  +----------------------------------------------+  |  |
+----------+          |  |  |     GKE Cluster                              |  |  |
                      |  |  |                                              |  |  |
                      |  |  |  +------------+    +----------------------+  |  |  |
                      |  |  |  |   Gateway  |--->|    GPU Workers       |  |  |  |
                      |  |  |  |  (consumer)|    |  (L4 / A100 / T4)    |  |  |  |
                      |  |  |  +------+-----+    +----------------------+  |  |  |
                      |  |  |         |                    |               |  |  |
                      |  |  |  +------+-----+              |               |  |  |
                      |  |  |  | sie-config |  (writes + NATS deltas)      |  |  |
                      |  |  |  +------------+              |               |  |  |
                      |  |  |                              |               |  |  |
                      |  |  |  +--------------------------------------------+  |  |
                      |  |  |  |  KEDA . Prometheus . Grafana . Loki . NATS  |  |  |
                      |  |  |  +--------------------------------------------+  |  |
                      |  |  |                                              |  |  |
                      |  |  |  +--------------+  +----------------------+  |  |  |
                      |  |  |  |  CPU Pool    |  |  GPU Pool(s)         |  |  |  |
                      |  |  |  | (e2-std-4)   |  |  (g2/a2/n1 + spot)   |  |  |  |
                      |  |  |  +--------------+  +----------------------+  |  |  |
                      |  |  +----------------------------------------------+  |  |
                      |  |                                                    |  |
                      |  |  +----------------+  +------------+  +---------+   |  |
                      |  |  |  Artifact Reg. |  |  Cloud NAT |  |   IAM   |   |  |
                      |  |  |  (images)      |  |  (egress)  |  |  (WI)   |   |  |
                      |  |  +----------------+  +------------+  +---------+   |  |
                      |  +----------------------------------------------------+  |
                      +----------------------------------------------------------+

Pushing images to Artifact Registry

This is optional, because the official images are available under ghcr.io/superlinked/.

After terraform apply, mirror the published images used by the GKE overlay into your Artifact Registry, preserving their versioned tags. When upgrading, mirror the v0.9.0 images before running helm upgrade: sie-config, the gateway, and the worker sidecars must run the same release:

# Authenticate Docker to Artifact Registry
gcloud auth configure-docker $(terraform output -raw artifact_registry_url | cut -d/ -f1)

# Mirror the GKE overlay's default CUDA 12 worker image
docker pull --platform linux/amd64 ghcr.io/superlinked/sie-server:v0.9.0-cuda12-default
docker tag ghcr.io/superlinked/sie-server:v0.9.0-cuda12-default "$(terraform output -raw artifact_registry_server_repository_url):v0.9.0-cuda12-default"
docker push "$(terraform output -raw artifact_registry_server_repository_url):v0.9.0-cuda12-default"

# Mirror the gateway image
docker pull --platform linux/amd64 ghcr.io/superlinked/sie-gateway:v0.9.0
docker tag ghcr.io/superlinked/sie-gateway:v0.9.0 "$(terraform output -raw artifact_registry_gateway_repository_url):v0.9.0"
docker push "$(terraform output -raw artifact_registry_gateway_repository_url):v0.9.0"

# Mirror the configuration-service image
docker pull --platform linux/amd64 ghcr.io/superlinked/sie-config:v0.9.0
docker tag ghcr.io/superlinked/sie-config:v0.9.0 "$(terraform output -raw artifact_registry_config_repository_url):v0.9.0"
docker push "$(terraform output -raw artifact_registry_config_repository_url):v0.9.0"

Set workers.common.image.repository, gateway.image.repository, and config.image.repository to the corresponding Artifact Registry outputs when installing the chart. Leave their tags unset so chart 0.9.0 selects the tags above; it appends -cuda12-default to the worker tag. The sidecar continues to use ghcr.io/superlinked/sie-server-sidecar:v0.9.0. If you enable additional worker bundles or platforms, mirror each matching versioned image before overriding the common worker repository.

Model cache and payload store

SIE clusters benefit from two object-store backed features that share a single GCS bucket:

  • Model cache: pre-staged model weights at gs://<bucket>/models/, so workers cold-start from object storage rather than re-downloading from Hugging Face on every pod spin-up.
  • Payload store: large work-item payloads (images, long documents that exceed the 1 MiB NATS in-band budget) at gs://<bucket>/payloads/, written by the gateway and read once by the worker. Garbage-collected by a runtime TTL plus a bucket lifecycle rule.

Because the payload store is required for >1 MiB work items, the shared bucket is created by default (create_model_cache = true). With it enabled, the module:

  1. Provisions a managed GCS bucket with uniform bucket-level access, public-access prevention enforced, and a lifecycle rule that deletes objects under the payloads/ prefix after one day (configurable via model_cache_payload_expiration_days).
  2. Defines two custom IAM roles (sie_model_cache_reader, sie_payload_store_writer) with the minimum permission set each side needs.
  3. Binds both roles to the SIE workload service account with IAM Conditions that scope each role to its own top-level prefix (models/ for read, payloads/ for read/write/delete). Workers can read weights but cannot delete or overwrite them; the gateway can write and delete payload refs but cannot touch weights.

After apply, pass the bucket into Helm with one terraform output:

curl -fsSL -o values-gke.yaml \
  https://raw.githubusercontent.com/superlinked/sie/v0.9.0/deploy/helm/sie-cluster/values-gke.yaml
helm upgrade --install sie-cluster oci://ghcr.io/superlinked/charts/sie-cluster \
  --version 0.9.0 -f values-gke.yaml \
  --create-namespace -n sie \
  --set-string "serviceAccount.annotations.iam\\.gke\\.io/gcp-service-account=$(terraform output -raw sie_workload_service_account)" \
  $(terraform output -raw model_cache_helm_args)

The chart auto-derives payloadStore.url from workers.common.clusterCache.url, so a single --set for the cache covers both the optional weights cache (models/) and the payload store (payloads/); the payload_store_url output is exposed for visibility and can be wired explicitly via --set payloadStore.url=.... On the chart side payloadStore.enabled defaults to true, decoupled from the optional workers.common.clusterCache. create_model_cache (managed bucket) and gcs_bucket_name (BYO) are mutually exclusive: to bring your own bucket, set create_model_cache = false and pass gcs_bucket_name instead; that path keeps the broader roles/storage.objectViewer binding for backward compatibility, but you forgo the prefix-scoped roles and the lifecycle rule, and you must wire payloadStore.url yourself or work items larger than 1 MiB (e.g. images) fail.

See infra/gcs_model_cache.tf and infra/iam.tf for the resource definitions and condition expressions.

Security features

This module follows GCP security best practices out of the box:

  • Restricted control plane - the Kubernetes API accepts only authorized_networks, or only the private endpoint with enable_private_endpoint; any-address access needs allow_public_api_server
  • Private nodes - worker nodes have no public IPs; egress via Cloud NAT
  • Shielded nodes - Secure Boot and Integrity Monitoring on all node pools
  • Workload Identity - pods use GCP service accounts, no JSON key files
  • Least-privilege IAM - node SA has only logging, monitoring, and Artifact Registry reader
  • VPC-native networking - pod and service CIDRs use secondary IP ranges (alias IPs)
  • GPU taints - GPU nodes are tainted so only GPU workloads schedule on them
  • Image streaming - GCFS enabled for fast container startup
  • Registry cleanup - automatic deletion of dev/test images after 14 days, untagged after 30 days
  • Legacy endpoints disabled - metadata concealment on all nodes

Bring-your-own components

Some pieces of a production deployment are intentionally not turnkey - either because they're cluster-wide / cross-stack concerns (registry, OIDC) or because they require domains and DNS records that only you can own (TLS, DNS). This module lets you opt out where it makes sense and points at the right knobs.

  • Container registry - optional. The module manages a regional Artifact Registry by default (create_artifact_registry = true, see infra/variables.tf). Set create_artifact_registry = false to reuse a registry managed by another stack; in that mode artifact_registry_url and the per-image repository outputs stay null, so point the Helm chart at the external registry via gateway.image.repository, workers.common.image.repository, and config.image.repository. The worker-sidecar uses the chart's ghcr.io/superlinked/sie-server-sidecar default.

  • TLS certificate - BYO by default. Set ingress.tlsConfig.mode to one of:

    • byo - supply your own kubernetes.io/tls Secret.
    • cert-manager - install cert-manager once in the cluster; the chart annotates the Ingress for automated Let's Encrypt issuance via HTTP-01.
    • self-signed - for air-gapped clusters; set certManagerBundle.certManager.install: true to bundle cert-manager (single-tenant clusters only).

    See the chart README's TLS / HTTPS section. DNS-01 / wildcard / Google-managed certificate paths are out of scope for the chart.

  • DNS / domain - always BYO. This module does not provision Cloud DNS zones or records. After terraform apply, take the ingress controller's LoadBalancer IP (kubectl -n ingress-nginx get svc ingress-nginx-controller) and create an A/AAAA record pointing at it under a domain you control.

  • OIDC provider - BYO. When auth.enabled: true in the chart, set auth.oauth2Proxy.oidcIssuerUrl and the corresponding client ID / secret to your existing identity provider (Okta, Auth0, Google Workspace, Azure AD, ...). The module does not create an IdP.

Cleanup

terraform destroy

Important: GPU nodes can be expensive. Always destroy dev/test clusters when not in use. Spot VMs (spot = true) save 60-91% but may be preempted.

If deletion_protection = true (default for production), you must first disable it:

terraform apply -var="deletion_protection=false"
terraform destroy

About

Terraform module for deploying SIE on Google Cloud GKE

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages