Skip to content

Repository files navigation

SIE EKS Terraform Module

One command to get a GPU-ready EKS cluster for SIE (Search Inference Engine). The module creates everything you need - VPC, EKS, GPU nodes, container registry, autoscaling - so you can focus on running inference, not managing infrastructure.

What you get

  • EKS cluster (Kubernetes 1.35) with private networking and KMS-encrypted secrets
  • GPU node group - pick your GPU: g6 (L4), g5 (A10G), p4d (A100), or p5 (H100)
  • Scale-to-zero - GPU nodes scale down to zero when idle, so you only pay when running inference
  • Cluster Autoscaler - automatically scales node groups based on pending pod demand
  • NVIDIA device plugin - pre-installed so GPU pods schedule immediately
  • ECR repositories (opt-in) - private container registries for customer-built images (<project_name>/sie-server, <project_name>/sie-gateway, <project_name>/sie-config). Off by default; set create_ecr_repositories = true to opt in. The worker-sidecar image stays on the chart's public GHCR default.
  • IRSA (IAM Roles for Service Accounts) - pods authenticate to AWS without stored credentials
  • VPC endpoints - private connectivity to ECR, S3, STS, and other AWS services
  • EBS CSI driver - persistent volumes work out of the box

Quick start

cd examples/dev-g6-spot
export AWS_REGION="eu-central-1"   # or your preferred region
# CIDRs allowed to reach the Kubernetes API; include this machine's egress
# address (for example the /32 of `curl -s https://checkip.amazonaws.com`).
export TF_VAR_api_server_authorized_ip_ranges='["203.0.113.10/32"]'
terraform init
terraform plan
terraform apply

203.0.113.10/32 is a documentation placeholder. The module rejects documentation ranges, so replace it with your own address. See Kubernetes API access for the private-endpoint mode.

After apply, configure kubectl and deploy SIE with chart 0.9.0. Its default image tags select the matching v0.9.0 gateway, config, worker, and sidecar images. The AWS values file is pinned to the same release:

# Point kubectl at the new cluster
$(terraform output -raw kubectl_config_command)

# Deploy SIE (gateway, workers, KEDA, Prometheus, Grafana)
helm upgrade --install sie-cluster oci://ghcr.io/superlinked/charts/sie-cluster --version 0.9.0 \
  -f https://raw.githubusercontent.com/superlinked/sie/v0.9.0/deploy/helm/sie-cluster/values-aws.yaml \
  --create-namespace -n sie \
  --set serviceAccount.annotations."eks\.amazonaws\.com/role-arn"="$(terraform output -raw sie_irsa_role_arn)" \
  $(terraform output -raw model_cache_helm_args)

This creates no Ingress: the gateway Service is ClusterIP. To expose the gateway outside the cluster, enable the Ingress together with gateway authentication and TLS, as described in the chart's Ingress section.

To reach the gateway from your machine without an Ingress, forward a local port to its ClusterIP Service:

kubectl -n sie port-forward svc/sie-gateway 8080:8080

Upgrading to SIE 0.9.0

Chart 0.9.0 has breaking changes. Read the SIE 0.9.0 release notes before upgrading an existing release. For a release installed with the command above:

  • NATS authentication is on by default. The upgrade restarts NATS and rolls sie-config, the gateway, and the workers. With the default memory-backed work queues, queued and in-flight work is lost, as on any NATS restart. A gateway, sie-config, or worker pod that has not been replaced yet has no credentials, and NATS refuses it: requests can fail with 503, and workers that have not restarted take no work until they do. To avoid that gap, run the command above twice: first with --set nats.auth.allowAnonymous=true added, then, once every pod has restarted, with --set nats.auth.allowAnonymous=false. Between the two steps NATS also accepts anonymous clients, with unrestricted permissions, and the chart ships no NetworkPolicy for the NATS pods, so allow only trusted workloads to reach NATS. The first step still restarts NATS, so the two steps do not prevent the loss of queued and in-flight work. See the chart's NATS authentication section.

  • Pass values explicitly. helm upgrade --reuse-values now fails to render. Re-run the full command above, which passes the values file with -f, or use --reset-then-reuse-values (Helm 3.14 or later).

  • The AWS values file no longer enables the gateway Ingress. The upgrade removes the host-less, plain-HTTP Ingress that earlier releases created. To keep external access, enable the Ingress with gateway authentication and TLS. To keep the previous unauthenticated catch-all Ingress, set ingress.enabled=true together with ingress.allowUnauthenticated=true and ingress.allowPlaintext=true. A LoadBalancer or NodePort gateway Service needs gateway authentication or gateway.service.allowUnauthenticated=true, and also gateway.service.allowPlaintext=true, because the gateway serves plain HTTP.

  • sie-config tokens are split. The upgrade generates a sie-config admin token and a separate read token for the gateway and the worker sidecars. The AWS values file sets fullnameOverride: sie, so the Secrets are sie-config-admin-token and sie-config-read-token, not the sie-cluster-config-* names in the chart's examples. sie-config then requires a token on every /v1/configs request, so give the admin token to tooling that writes model configs. Read it with:

    kubectl get secret -n sie sie-config-admin-token \
      -o jsonpath='{.data.SIE_ADMIN_TOKEN}' | base64 -d

    The gateway no longer receives that admin token: with gateway authentication enabled, its admin routes (POST, PUT, and DELETE under /v1/pools, /v1/admin, and /v1/configs) answer 403 until gateway.auth.adminTokenSecretName names a separate Secret. Run the sie-config, gateway, and worker sidecar images of the same release. See the chart's sie-config tokens section.

For existing installations with custom model profiles, also review the SIE 0.8.0 breaking changes for adapter options and launch arguments before upgrading from 0.7.x.

Examples

Example GPU Cost Description
dev-g6-spot L4 (g6.xlarge) ~$0.30/hr Spot instances, scale 0-5 nodes, minimal cost for development

Prerequisites

  1. AWS credentials configured (aws configure, environment variables, or IAM role)
  2. GPU quota in your target region - check EC2 limits for your chosen instance type
  3. Terraform >= 1.14

Variables

Required

Choose how the Kubernetes API is reached (see Kubernetes API access); the plan fails until you do. The other variables have defaults. Override these for your environment:

Variable Default Description
aws_region eu-central-1 AWS region to deploy in
project_name sie Name prefix for all resources (EKS cluster, IAM roles, etc.)

Kubernetes API access

The EKS API endpoint is never open to the whole Internet unless you ask for it. Choose one mode:

Variable Default Description
api_server_authorized_ip_ranges [] CIDRs allowed to reach the public API endpoint. Include every machine that runs terraform, kubectl, or helm against the cluster: the module installs Helm releases (cluster autoscaler, NVIDIA device plugin) during apply.
enable_private_endpoint false Disable the public endpoint and serve the API only inside the VPC. Run Terraform, kubectl, and Helm from a network that reaches the VPC (VPN, peering, or a runner in the VPC).
allow_public_api_server false Explicit opt-in to accept any Internet address. With an empty allowlist the endpoint allows 0.0.0.0/0.

Rules for api_server_authorized_ip_ranges:

  • Entries must be IPv4 CIDR blocks, because the module creates an IPv4 cluster and EKS accepts IPv6 public access CIDRs only for IPv6 clusters.
  • At most 40 entries, the EKS limit per cluster, which is not adjustable.
  • Together the entries may cover at most 16,777,216 addresses, the size of one /8. 0.0.0.0/0, split halves such as two /1 blocks, and several broad ranges are rejected unless allow_public_api_server = true.
  • Entries inside a documentation range (192.0.2.0/24, 198.51.100.0/24, 203.0.113.0/24) are rejected, so an unedited placeholder fails at plan time. Broader entries that contain one need allow_public_api_server = true.

The private endpoint stays enabled in every mode, so nodes always reach the API inside the VPC. Network restrictions are in addition to Kubernetes API authentication and authorization. Apart from the unauthenticated health and version endpoints (such as /healthz, /readyz, and /version), every request must still pass them. The api_server_access output shows the effective endpoint settings.

Recovering from an allowlist that excludes Terraform. The module refreshes its Helm and Kubernetes resources on every plan, so a plan fails once the machine running Terraform is no longer in the allowlist. Widen the list with the AWS CLI from any machine with EKS permissions, then set api_server_authorized_ip_ranges to the same list and apply:

aws eks update-cluster-config --region <region> --name <cluster> \
  --resources-vpc-config endpointPublicAccess=true,publicAccessCidrs="<cidr-1>,<cidr-2>"

Alternatively, correct the variable and run terraform apply -refresh=false -target=module.eks, which updates the endpoint without reading the in-cluster resources, then run a normal plan.

Upgrading from 0.x.

  • Earlier versions opened the public endpoint to 0.0.0.0/0 without an input. The plan now fails with a message asking you to choose.
  • To restrict access, set api_server_authorized_ip_ranges. The plan shows an in-place update of the cluster's public_access_cidrs. Include the address Terraform runs from, or the refresh of the module's Helm releases cannot reach the cluster.
  • To keep the previous behaviour explicitly, set allow_public_api_server = true. The plan shows no change to the endpoint.

GPU configuration

Variable Default Description
gpu_instance_type g6.xlarge EC2 instance type for GPU nodes
gpu_capacity_type ON_DEMAND ON_DEMAND or SPOT (spot saves ~60-70%)
gpu_min_size 1 Minimum GPU nodes - set to 0 for scale-to-zero
gpu_max_size 10 Maximum GPU nodes
gpu_disk_size_gb 500 Root EBS volume size for the legacy single GPU node group
gpu_disk_type gp3 Root EBS volume type for the legacy single GPU node group

For multi-pool clusters, set gpu_node_groups[*].disk_size_gb and gpu_node_groups[*].disk_type per pool. This mirrors the GCP module's gpu_node_pools[*].disk_size_gb / disk_type shape and is the knob that backs Kubernetes emptyDir model caches on EKS nodes.

Keep the disk above the Helm chart's workers.common.cacheStorageSize (default 300Gi). That cache is an emptyDir on the node root volume, so its sizeLimit caps the cache without reserving the space: on a smaller disk a worker lazily loading several large models fills the root volume first and kubelet evicts it on disk pressure mid-inference, the replacement re-downloads the weights, and the cycle repeats.

Size for kubelet's eviction threshold, not just raw capacity. At the default nodefs.available<10% only ~90% of the disk is usable before eviction starts, so the 300Gi cache plus ~100GiB of OS, container images, and logs needs ceil(400 / 0.9) = 445GiB. The 500 defaults round that up to leave room for filesystem overhead.

gpu_node_groups configures node groups, not Helm worker pod GPU count. For a multi-GPU worker pod, choose an instance type with enough GPUs, then set the matching Helm workers.pools.<name>.gpu.count to the number of GPUs the pod should consume on one node.

Networking

Variable Default Description
vpc_cidr 10.0.0.0/16 CIDR block for the EKS VPC
private_subnet_prefix_length 20 Private worker subnet size; default creates /20 private subnets for EKS pod IP headroom
public_subnet_prefix_length 24 Public subnet size; default creates /24 public subnets for load balancers/NAT

Changing VPC/subnet sizing is intentionally breaking for existing clusters because AWS subnet CIDRs are replacement-sensitive. Recreate ephemeral clusters or plan a migration window for persistent clusters.

Node log rotation

Variable Default Description
kubelet_container_log_max_size 20Mi Per-container kubelet log file size before rotation
kubelet_container_log_max_files 30 Rotated files retained per container; kubelet retention is size/count based, not hourly

GPU instance cheat sheet:

Instance GPU VRAM Approx. on-demand/hr Best for
g6.xlarge 1x L4 24 GB $0.80 Development, small models
g5.xlarge 1x A10G 24 GB $1.00 Development, medium models
p4d.24xlarge 8x A100 320 GB $32.77 Large models, production
p5.48xlarge 8x H100 640 GB $98.32 Maximum throughput

Container registry

Variable Default Description
server_ecr_repository_name sie-server ECR repo name for the inference server
gateway_ecr_repository_name sie-gateway ECR repo name for the request gateway
config_ecr_repository_name sie-config ECR repo name for the sie-config control plane image
create_ecr_repositories false Whether this module manages the ECR repos. Default false matches the chart's GHCR-by-default behaviour and avoids RepositoryAlreadyExistsException on accounts where the repos already exist. Set true to opt in. The ecr_*_repository_url outputs are emitted regardless.
ecr_repository_prefix null -> <project_name> Namespace prefix for ECR repo names; final names become <prefix>/<repo_name>. Set to "" to disable prefixing (bare names) for accounts where ECR is externally managed.

Workload identity

Variable Default Description
sie_namespace sie Kubernetes namespace for SIE workloads
sie_service_account_name sie-server K8s ServiceAccount that assumes the IRSA role

Outputs

After terraform apply, use these outputs to connect and deploy:

Output Description
kubectl_config_command Run this to configure kubectl
api_server_access Effective API endpoint access: public endpoint on or off and the CIDRs it accepts
cluster_name EKS cluster name
cluster_endpoint Kubernetes API endpoint (sensitive)
ecr_server_repository_url Where to push sie-server images
ecr_gateway_repository_url Where to push sie-gateway images
ecr_config_repository_url Where to push sie-config images
sie_irsa_role_arn Pass to Helm for workload identity
cluster_autoscaler_irsa_role_arn Cluster autoscaler IAM role
gpu_instance_type Confirm which GPU type is deployed
gpu_capacity_type Confirm ON_DEMAND vs SPOT
gpu_node_group_disk_sizes_gb Root EBS volume size per effective GPU node group

Architecture

                         ┌─────────────────────────────────────────────────────┐
                         │                    AWS Region                       │
                         │                                                     │
┌──────────┐             │  ┌───────────────────────────────────────────────┐  │
│          │   HTTPS     │  │                 VPC (10.0.0.0/16)             │  │
│  Client  │────────────▶│  │                                               │  │
│          │             │  │  ┌──────────────────────────────────────────┐ │  │
└──────────┘             │  │  │     EKS Cluster (private + public)       │ │  │
                         │  │  │                                          │ │  │
                         │  │  │  ┌────────────┐    ┌─────────────────┐   │ │  │
                         │  │  │  │   Gateway   │───▶│  GPU Workers    │   │ │  │
                         │  │  │  │            │    │  (L4/A10G/A100) │   │ │  │
                         │  │  │  └─────┬──────┘    └─────────────────┘   │ │  │
                         │  │  │        │                    │            │ │  │
                         │  │  │  ┌─────┴──────┐              │            │ │  │
                         │  │  │  │ sie-config │ (control plane, NATS)    │ │  │
                         │  │  │  └────────────┘              │            │ │  │
                         │  │  │                              │            │ │  │
                         │  │  │  ┌────────────────────────────────────┐   │ │  │
                         │  │  │  │  KEDA · Prometheus · Grafana       │   │ │  │
                         │  │  │  └────────────────────────────────────┘   │ │  │
                         │  │  │                                          │ │  │
                         │  │  │  ┌──────────────┐  ┌─────────────────┐   │ │  │
                         │  │  │  │  CPU Nodes   │  │  GPU Nodes      │   │ │  │
                         │  │  │  │  (t3.xlarge) │  │  (g6/g5/p4d/p5) │   │ │  │
                         │  │  │  └──────────────┘  └─────────────────┘   │ │  │
                         │  │  └──────────────────────────────────────────┘ │  │
                         │  │                                               │  │
                         │  │  ┌───────────┐  ┌───────────┐  ┌──────────┐   │  │
                         │  │  │    ECR    │  │   KMS     │  │  NAT GW  │   │  │
                         │  │  │ (images)  │  │ (secrets) │  │ (egress) │   │  │
                         │  │  └───────────┘  └───────────┘  └──────────┘   │  │
                         │  └───────────────────────────────────────────────┘  │
                         └─────────────────────────────────────────────────────┘

Pushing images to ECR

This is optional, because the official images are available under ghcr.io/superlinked/.

Requires create_ecr_repositories = true (or repos managed by another stack - see ecr_repository_prefix).

After terraform apply, mirror the released images into ECR. The AWS values file enables the default and sglang worker images (its sglang-vision-extract bundle reuses the sglang image); mirror both when overriding workers.common.image.repository. When upgrading, mirror the v0.9.0 images before running helm upgrade: sie-config, the gateway, and the worker sidecars must run the same release:

# Authenticate Docker to ECR
aws ecr get-login-password --region $(terraform output -raw aws_region 2>/dev/null || echo $AWS_REGION) \
  | docker login --username AWS --password-stdin $(terraform output -raw ecr_server_repository_url | cut -d/ -f1)

# Mirror both worker images selected by values-aws.yaml
for bundle in default sglang; do
  tag="v0.9.0-cuda12-${bundle}"
  docker pull --platform linux/amd64 "ghcr.io/superlinked/sie-server:${tag}"
  docker tag "ghcr.io/superlinked/sie-server:${tag}" "$(terraform output -raw ecr_server_repository_url):${tag}"
  docker push "$(terraform output -raw ecr_server_repository_url):${tag}"
done

# Mirror gateway and config images without changing their release tags
docker pull --platform linux/amd64 ghcr.io/superlinked/sie-gateway:v0.9.0
docker tag ghcr.io/superlinked/sie-gateway:v0.9.0 "$(terraform output -raw ecr_gateway_repository_url):v0.9.0"
docker push "$(terraform output -raw ecr_gateway_repository_url):v0.9.0"

docker pull --platform linux/amd64 ghcr.io/superlinked/sie-config:v0.9.0
docker tag ghcr.io/superlinked/sie-config:v0.9.0 "$(terraform output -raw ecr_config_repository_url):v0.9.0"
docker push "$(terraform output -raw ecr_config_repository_url):v0.9.0"

Model cache and payload store

SIE clusters benefit from two object-store backed features that share a single S3 bucket:

  • Model cache: pre-staged model weights at s3://<bucket>/models/, so workers cold-start from object storage rather than re-downloading from Hugging Face on every pod spin-up.
  • Payload store: large work-item payloads (images, long documents that exceed the 1 MiB NATS in-band budget) at s3://<bucket>/payloads/, written by the gateway and read once by the worker. Garbage-collected by a runtime TTL plus a bucket lifecycle rule.

Because the payload store is required for >1 MiB work items, the shared bucket is created by default (create_model_cache = true). With it enabled, the module:

  1. Provisions a managed S3 bucket with versioning, abort-incomplete-multipart, and a lifecycle rule that deletes objects under the payloads/ prefix after one day.
  2. Attaches two scoped inline policies to the SIE workload IRSA role: read-only on the cache, and s3:Get/Put/Delete/AbortMultipartUpload constrained to the payloads/* prefix, with a ListBucket prefix condition.
  3. KMS-encrypted buckets get matching kms:Decrypt/Encrypt/GenerateDataKey grants.

After apply, pass the bucket into Helm with one terraform output:

helm upgrade --install sie-cluster oci://ghcr.io/superlinked/charts/sie-cluster --version 0.9.0 \
  -f https://raw.githubusercontent.com/superlinked/sie/v0.9.0/deploy/helm/sie-cluster/values-aws.yaml \
  --namespace sie --create-namespace \
  --set serviceAccount.annotations."eks\.amazonaws\.com/role-arn"="$(terraform output -raw sie_irsa_role_arn)" \
  $(terraform output -raw model_cache_helm_args)

The chart auto-derives payloadStore.url from workers.common.clusterCache.url, so a single --set for the cache covers both the optional weights cache (models/) and the payload store (payloads/); the payload_store_url output is exposed for visibility and can be wired explicitly via --set payloadStore.url=... for the rare override case. On the chart side payloadStore.enabled defaults to true, decoupled from the optional workers.common.clusterCache. Operators who bring their own object-storage bucket can opt out (create_model_cache = false) and wire payloadStore.url themselves; skipping the payload store entirely means work items larger than 1 MiB (e.g. images) fail.

See infra/s3_model_cache.tf and infra/irsa.tf for the resource definitions.

Security features

This module follows AWS security best practices out of the box:

  • KMS encryption - EKS secrets encrypted at rest with a dedicated, auto-rotating KMS key
  • Restricted API endpoint - the public Kubernetes API endpoint accepts only api_server_authorized_ip_ranges, or is disabled with enable_private_endpoint; any-address access needs allow_public_api_server
  • Private subnets - worker nodes run in private subnets with no public IPs
  • NAT gateway - outbound internet via NAT (one per AZ for high availability)
  • VPC endpoints - private access to ECR, S3, STS, EC2, CloudWatch, and other services
  • IRSA - pods use IAM roles instead of long-lived credentials
  • GPU taints - GPU nodes are tainted so only GPU workloads schedule on them
  • Image scanning - ECR scans images on push for known vulnerabilities
  • Audit logging - all EKS control plane log types enabled

Bring-your-own components

Some pieces of a production deployment are intentionally not turnkey - either because they're cluster-wide / cross-stack concerns (registry, OIDC) or because they require domains and DNS records that only you can own (TLS, DNS). This module lets you opt out where it makes sense and points at the right knobs.

  • Container registry - optional. The module does not create ECR repos by default (create_ecr_repositories = false, see infra/variables.tf) - this matches the chart's GHCR-by-default behaviour and avoids RepositoryAlreadyExistsException on accounts where repos already exist. Set create_ecr_repositories = true to opt in to terraform-managed ECR; the module will create project-scoped repos (<project_name>/sie-server, <project_name>/sie-gateway, <project_name>/sie-config). Override the namespace via ecr_repository_prefix - set to "" to disable prefixing for accounts where ECR is externally managed under bare names. The module always emits ecr_*_repository_url outputs (composed from caller identity + repo names) so IRSA / Helm wiring is unchanged whether you opt in or not. The worker-sidecar uses the chart's ghcr.io/superlinked/sie-server-sidecar default; to use an external registry for the other runtime images, point the Helm chart at it via gateway.image.repository, workers.common.image.repository, and config.image.repository.

  • TLS certificate - BYO by default. Set ingress.tlsConfig.mode to one of:

    • byo - supply your own kubernetes.io/tls Secret.
    • cert-manager - install cert-manager once in the cluster; the chart annotates the Ingress for automated Let's Encrypt issuance via HTTP-01.
    • self-signed - for air-gapped clusters; set certManagerBundle.certManager.install: true to bundle cert-manager (single-tenant clusters only).

    See the chart README's TLS / HTTPS section. DNS-01 / wildcard / ACM paths are out of scope for the chart.

  • DNS / domain - always BYO. This module does not provision Route53 zones or records. After terraform apply, take the ingress controller's LoadBalancer hostname (kubectl -n ingress-nginx get svc ingress-nginx-controller) and create an A/AAAA record pointing at it under a domain you control.

  • OIDC provider - BYO. When auth.enabled: true in the chart, set auth.oauth2Proxy.oidcIssuerUrl and the corresponding client ID / secret to your existing identity provider (Okta, Auth0, Google Workspace, Azure AD, ...). The module does not create an IdP.

Cleanup

terraform destroy

Important: GPU instances can be expensive. Always destroy dev/test clusters when not in use. Spot instances (gpu_capacity_type = "SPOT") save 60-70% but may be interrupted.

About

Terraform module for deploying SIE on Amazon EKS

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages