One command to get a GPU-ready EKS cluster for SIE (Search Inference Engine). The module creates everything you need - VPC, EKS, GPU nodes, container registry, autoscaling - so you can focus on running inference, not managing infrastructure.
- EKS cluster (Kubernetes 1.35) with private networking and KMS-encrypted secrets
- GPU node group - pick your GPU: g6 (L4), g5 (A10G), p4d (A100), or p5 (H100)
- Scale-to-zero - GPU nodes scale down to zero when idle, so you only pay when running inference
- Cluster Autoscaler - automatically scales node groups based on pending pod demand
- NVIDIA device plugin - pre-installed so GPU pods schedule immediately
- ECR repositories (opt-in) - private container registries for customer-built images (
<project_name>/sie-server,<project_name>/sie-gateway,<project_name>/sie-config). Off by default; setcreate_ecr_repositories = trueto opt in. The worker-sidecar image stays on the chart's public GHCR default. - IRSA (IAM Roles for Service Accounts) - pods authenticate to AWS without stored credentials
- VPC endpoints - private connectivity to ECR, S3, STS, and other AWS services
- EBS CSI driver - persistent volumes work out of the box
cd examples/dev-g6-spot
export AWS_REGION="eu-central-1" # or your preferred region
# CIDRs allowed to reach the Kubernetes API; include this machine's egress
# address (for example the /32 of `curl -s https://checkip.amazonaws.com`).
export TF_VAR_api_server_authorized_ip_ranges='["203.0.113.10/32"]'
terraform init
terraform plan
terraform apply203.0.113.10/32 is a documentation placeholder. The module rejects
documentation ranges, so replace it with your own address. See
Kubernetes API access for the private-endpoint mode.
After apply, configure kubectl and deploy SIE with chart 0.9.0. Its default
image tags select the matching v0.9.0 gateway, config, worker, and sidecar
images. The AWS values file is pinned to the same release:
# Point kubectl at the new cluster
$(terraform output -raw kubectl_config_command)
# Deploy SIE (gateway, workers, KEDA, Prometheus, Grafana)
helm upgrade --install sie-cluster oci://ghcr.io/superlinked/charts/sie-cluster --version 0.9.0 \
-f https://raw.githubusercontent.com/superlinked/sie/v0.9.0/deploy/helm/sie-cluster/values-aws.yaml \
--create-namespace -n sie \
--set serviceAccount.annotations."eks\.amazonaws\.com/role-arn"="$(terraform output -raw sie_irsa_role_arn)" \
$(terraform output -raw model_cache_helm_args)This creates no Ingress: the gateway Service is ClusterIP. To expose the
gateway outside the cluster, enable the Ingress together with gateway
authentication and TLS, as described in the chart's
Ingress section.
To reach the gateway from your machine without an Ingress, forward a local
port to its ClusterIP Service:
kubectl -n sie port-forward svc/sie-gateway 8080:8080Chart 0.9.0 has breaking changes. Read the
SIE 0.9.0 release notes
before upgrading an existing release. For a release installed with the command
above:
-
NATS authentication is on by default. The upgrade restarts NATS and rolls sie-config, the gateway, and the workers. With the default memory-backed work queues, queued and in-flight work is lost, as on any NATS restart. A gateway, sie-config, or worker pod that has not been replaced yet has no credentials, and NATS refuses it: requests can fail with
503, and workers that have not restarted take no work until they do. To avoid that gap, run the command above twice: first with--set nats.auth.allowAnonymous=trueadded, then, once every pod has restarted, with--set nats.auth.allowAnonymous=false. Between the two steps NATS also accepts anonymous clients, with unrestricted permissions, and the chart ships no NetworkPolicy for the NATS pods, so allow only trusted workloads to reach NATS. The first step still restarts NATS, so the two steps do not prevent the loss of queued and in-flight work. See the chart's NATS authentication section. -
Pass values explicitly.
helm upgrade --reuse-valuesnow fails to render. Re-run the full command above, which passes the values file with-f, or use--reset-then-reuse-values(Helm 3.14 or later). -
The AWS values file no longer enables the gateway Ingress. The upgrade removes the host-less, plain-HTTP Ingress that earlier releases created. To keep external access, enable the Ingress with gateway authentication and TLS. To keep the previous unauthenticated catch-all Ingress, set
ingress.enabled=truetogether withingress.allowUnauthenticated=trueandingress.allowPlaintext=true. ALoadBalancerorNodePortgateway Service needs gateway authentication orgateway.service.allowUnauthenticated=true, and alsogateway.service.allowPlaintext=true, because the gateway serves plain HTTP. -
sie-config tokens are split. The upgrade generates a sie-config admin token and a separate read token for the gateway and the worker sidecars. The AWS values file sets
fullnameOverride: sie, so the Secrets aresie-config-admin-tokenandsie-config-read-token, not thesie-cluster-config-*names in the chart's examples. sie-config then requires a token on every/v1/configsrequest, so give the admin token to tooling that writes model configs. Read it with:kubectl get secret -n sie sie-config-admin-token \ -o jsonpath='{.data.SIE_ADMIN_TOKEN}' | base64 -d
The gateway no longer receives that admin token: with gateway authentication enabled, its admin routes (
POST,PUT, andDELETEunder/v1/pools,/v1/admin, and/v1/configs) answer403untilgateway.auth.adminTokenSecretNamenames a separate Secret. Run the sie-config, gateway, and worker sidecar images of the same release. See the chart's sie-config tokens section.
For existing installations with custom model profiles, also review the SIE 0.8.0 breaking changes for adapter options and launch arguments before upgrading from 0.7.x.
| Example | GPU | Cost | Description |
|---|---|---|---|
dev-g6-spot |
L4 (g6.xlarge) | ~$0.30/hr | Spot instances, scale 0-5 nodes, minimal cost for development |
- AWS credentials configured (
aws configure, environment variables, or IAM role) - GPU quota in your target region - check EC2 limits for your chosen instance type
- Terraform >= 1.14
Choose how the Kubernetes API is reached (see Kubernetes API access); the plan fails until you do. The other variables have defaults. Override these for your environment:
| Variable | Default | Description |
|---|---|---|
aws_region |
eu-central-1 |
AWS region to deploy in |
project_name |
sie |
Name prefix for all resources (EKS cluster, IAM roles, etc.) |
The EKS API endpoint is never open to the whole Internet unless you ask for it. Choose one mode:
| Variable | Default | Description |
|---|---|---|
api_server_authorized_ip_ranges |
[] |
CIDRs allowed to reach the public API endpoint. Include every machine that runs terraform, kubectl, or helm against the cluster: the module installs Helm releases (cluster autoscaler, NVIDIA device plugin) during apply. |
enable_private_endpoint |
false |
Disable the public endpoint and serve the API only inside the VPC. Run Terraform, kubectl, and Helm from a network that reaches the VPC (VPN, peering, or a runner in the VPC). |
allow_public_api_server |
false |
Explicit opt-in to accept any Internet address. With an empty allowlist the endpoint allows 0.0.0.0/0. |
Rules for api_server_authorized_ip_ranges:
- Entries must be IPv4 CIDR blocks, because the module creates an IPv4 cluster and EKS accepts IPv6 public access CIDRs only for IPv6 clusters.
- At most 40 entries, the EKS limit per cluster, which is not adjustable.
- Together the entries may cover at most 16,777,216 addresses, the size of one
/8.0.0.0.0/0, split halves such as two/1blocks, and several broad ranges are rejected unlessallow_public_api_server = true. - Entries inside a documentation range (
192.0.2.0/24,198.51.100.0/24,203.0.113.0/24) are rejected, so an unedited placeholder fails at plan time. Broader entries that contain one needallow_public_api_server = true.
The private endpoint stays enabled in every mode, so nodes always reach the API
inside the VPC. Network restrictions are in addition to Kubernetes API
authentication and authorization. Apart from the unauthenticated health and
version endpoints (such as /healthz, /readyz, and /version), every
request must still pass them. The
api_server_access output shows the effective endpoint settings.
Recovering from an allowlist that excludes Terraform. The module refreshes
its Helm and Kubernetes resources on every plan, so a plan fails once the
machine running Terraform is no longer in the allowlist. Widen the list with
the AWS CLI from any machine with EKS permissions, then set
api_server_authorized_ip_ranges to the same list and apply:
aws eks update-cluster-config --region <region> --name <cluster> \
--resources-vpc-config endpointPublicAccess=true,publicAccessCidrs="<cidr-1>,<cidr-2>"Alternatively, correct the variable and run
terraform apply -refresh=false -target=module.eks, which updates the
endpoint without reading the in-cluster resources, then run a normal plan.
Upgrading from 0.x.
- Earlier versions opened the public endpoint to
0.0.0.0/0without an input. The plan now fails with a message asking you to choose. - To restrict access, set
api_server_authorized_ip_ranges. The plan shows an in-place update of the cluster'spublic_access_cidrs. Include the address Terraform runs from, or the refresh of the module's Helm releases cannot reach the cluster. - To keep the previous behaviour explicitly, set
allow_public_api_server = true. The plan shows no change to the endpoint.
| Variable | Default | Description |
|---|---|---|
gpu_instance_type |
g6.xlarge |
EC2 instance type for GPU nodes |
gpu_capacity_type |
ON_DEMAND |
ON_DEMAND or SPOT (spot saves ~60-70%) |
gpu_min_size |
1 |
Minimum GPU nodes - set to 0 for scale-to-zero |
gpu_max_size |
10 |
Maximum GPU nodes |
gpu_disk_size_gb |
500 |
Root EBS volume size for the legacy single GPU node group |
gpu_disk_type |
gp3 |
Root EBS volume type for the legacy single GPU node group |
For multi-pool clusters, set gpu_node_groups[*].disk_size_gb and
gpu_node_groups[*].disk_type per pool. This mirrors the GCP module's
gpu_node_pools[*].disk_size_gb / disk_type shape and is the knob that
backs Kubernetes emptyDir model caches on EKS nodes.
Keep the disk above the Helm chart's workers.common.cacheStorageSize
(default 300Gi). That cache is an emptyDir on the node root volume, so
its sizeLimit caps the cache without reserving the space: on a smaller disk a
worker lazily loading several large models fills the root volume first and
kubelet evicts it on disk pressure mid-inference, the replacement re-downloads
the weights, and the cycle repeats.
Size for kubelet's eviction threshold, not just raw capacity. At the default
nodefs.available<10% only ~90% of the disk is usable before eviction starts,
so the 300Gi cache plus ~100GiB of OS, container images, and logs needs
ceil(400 / 0.9) = 445GiB. The 500 defaults round that up to leave room for
filesystem overhead.
gpu_node_groups configures node groups, not Helm worker pod GPU count. For a
multi-GPU worker pod, choose an instance type with enough GPUs, then set the
matching Helm workers.pools.<name>.gpu.count to the number of GPUs the pod
should consume on one node.
| Variable | Default | Description |
|---|---|---|
vpc_cidr |
10.0.0.0/16 |
CIDR block for the EKS VPC |
private_subnet_prefix_length |
20 |
Private worker subnet size; default creates /20 private subnets for EKS pod IP headroom |
public_subnet_prefix_length |
24 |
Public subnet size; default creates /24 public subnets for load balancers/NAT |
Changing VPC/subnet sizing is intentionally breaking for existing clusters because AWS subnet CIDRs are replacement-sensitive. Recreate ephemeral clusters or plan a migration window for persistent clusters.
| Variable | Default | Description |
|---|---|---|
kubelet_container_log_max_size |
20Mi |
Per-container kubelet log file size before rotation |
kubelet_container_log_max_files |
30 |
Rotated files retained per container; kubelet retention is size/count based, not hourly |
GPU instance cheat sheet:
| Instance | GPU | VRAM | Approx. on-demand/hr | Best for |
|---|---|---|---|---|
g6.xlarge |
1x L4 | 24 GB | $0.80 | Development, small models |
g5.xlarge |
1x A10G | 24 GB | $1.00 | Development, medium models |
p4d.24xlarge |
8x A100 | 320 GB | $32.77 | Large models, production |
p5.48xlarge |
8x H100 | 640 GB | $98.32 | Maximum throughput |
| Variable | Default | Description |
|---|---|---|
server_ecr_repository_name |
sie-server |
ECR repo name for the inference server |
gateway_ecr_repository_name |
sie-gateway |
ECR repo name for the request gateway |
config_ecr_repository_name |
sie-config |
ECR repo name for the sie-config control plane image |
create_ecr_repositories |
false |
Whether this module manages the ECR repos. Default false matches the chart's GHCR-by-default behaviour and avoids RepositoryAlreadyExistsException on accounts where the repos already exist. Set true to opt in. The ecr_*_repository_url outputs are emitted regardless. |
ecr_repository_prefix |
null -> <project_name> |
Namespace prefix for ECR repo names; final names become <prefix>/<repo_name>. Set to "" to disable prefixing (bare names) for accounts where ECR is externally managed. |
| Variable | Default | Description |
|---|---|---|
sie_namespace |
sie |
Kubernetes namespace for SIE workloads |
sie_service_account_name |
sie-server |
K8s ServiceAccount that assumes the IRSA role |
After terraform apply, use these outputs to connect and deploy:
| Output | Description |
|---|---|
kubectl_config_command |
Run this to configure kubectl |
api_server_access |
Effective API endpoint access: public endpoint on or off and the CIDRs it accepts |
cluster_name |
EKS cluster name |
cluster_endpoint |
Kubernetes API endpoint (sensitive) |
ecr_server_repository_url |
Where to push sie-server images |
ecr_gateway_repository_url |
Where to push sie-gateway images |
ecr_config_repository_url |
Where to push sie-config images |
sie_irsa_role_arn |
Pass to Helm for workload identity |
cluster_autoscaler_irsa_role_arn |
Cluster autoscaler IAM role |
gpu_instance_type |
Confirm which GPU type is deployed |
gpu_capacity_type |
Confirm ON_DEMAND vs SPOT |
gpu_node_group_disk_sizes_gb |
Root EBS volume size per effective GPU node group |
┌─────────────────────────────────────────────────────┐
│ AWS Region │
│ │
┌──────────┐ │ ┌───────────────────────────────────────────────┐ │
│ │ HTTPS │ │ VPC (10.0.0.0/16) │ │
│ Client │────────────▶│ │ │ │
│ │ │ │ ┌──────────────────────────────────────────┐ │ │
└──────────┘ │ │ │ EKS Cluster (private + public) │ │ │
│ │ │ │ │ │
│ │ │ ┌────────────┐ ┌─────────────────┐ │ │ │
│ │ │ │ Gateway │───▶│ GPU Workers │ │ │ │
│ │ │ │ │ │ (L4/A10G/A100) │ │ │ │
│ │ │ └─────┬──────┘ └─────────────────┘ │ │ │
│ │ │ │ │ │ │ │
│ │ │ ┌─────┴──────┐ │ │ │ │
│ │ │ │ sie-config │ (control plane, NATS) │ │ │
│ │ │ └────────────┘ │ │ │ │
│ │ │ │ │ │ │
│ │ │ ┌────────────────────────────────────┐ │ │ │
│ │ │ │ KEDA · Prometheus · Grafana │ │ │ │
│ │ │ └────────────────────────────────────┘ │ │ │
│ │ │ │ │ │
│ │ │ ┌──────────────┐ ┌─────────────────┐ │ │ │
│ │ │ │ CPU Nodes │ │ GPU Nodes │ │ │ │
│ │ │ │ (t3.xlarge) │ │ (g6/g5/p4d/p5) │ │ │ │
│ │ │ └──────────────┘ └─────────────────┘ │ │ │
│ │ └──────────────────────────────────────────┘ │ │
│ │ │ │
│ │ ┌───────────┐ ┌───────────┐ ┌──────────┐ │ │
│ │ │ ECR │ │ KMS │ │ NAT GW │ │ │
│ │ │ (images) │ │ (secrets) │ │ (egress) │ │ │
│ │ └───────────┘ └───────────┘ └──────────┘ │ │
│ └───────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────┘
This is optional, because the official images are available under
ghcr.io/superlinked/.
Requires create_ecr_repositories = true (or repos managed by another stack - see ecr_repository_prefix).
After terraform apply, mirror the released images into ECR. The AWS values
file enables the default and sglang worker images (its
sglang-vision-extract bundle reuses the sglang image); mirror both when
overriding workers.common.image.repository. When upgrading, mirror the
v0.9.0 images before running helm upgrade: sie-config, the gateway, and the
worker sidecars must run the same release:
# Authenticate Docker to ECR
aws ecr get-login-password --region $(terraform output -raw aws_region 2>/dev/null || echo $AWS_REGION) \
| docker login --username AWS --password-stdin $(terraform output -raw ecr_server_repository_url | cut -d/ -f1)
# Mirror both worker images selected by values-aws.yaml
for bundle in default sglang; do
tag="v0.9.0-cuda12-${bundle}"
docker pull --platform linux/amd64 "ghcr.io/superlinked/sie-server:${tag}"
docker tag "ghcr.io/superlinked/sie-server:${tag}" "$(terraform output -raw ecr_server_repository_url):${tag}"
docker push "$(terraform output -raw ecr_server_repository_url):${tag}"
done
# Mirror gateway and config images without changing their release tags
docker pull --platform linux/amd64 ghcr.io/superlinked/sie-gateway:v0.9.0
docker tag ghcr.io/superlinked/sie-gateway:v0.9.0 "$(terraform output -raw ecr_gateway_repository_url):v0.9.0"
docker push "$(terraform output -raw ecr_gateway_repository_url):v0.9.0"
docker pull --platform linux/amd64 ghcr.io/superlinked/sie-config:v0.9.0
docker tag ghcr.io/superlinked/sie-config:v0.9.0 "$(terraform output -raw ecr_config_repository_url):v0.9.0"
docker push "$(terraform output -raw ecr_config_repository_url):v0.9.0"SIE clusters benefit from two object-store backed features that share a single S3 bucket:
- Model cache: pre-staged model weights at
s3://<bucket>/models/, so workers cold-start from object storage rather than re-downloading from Hugging Face on every pod spin-up. - Payload store: large work-item payloads (images, long documents that exceed the 1 MiB NATS in-band budget) at
s3://<bucket>/payloads/, written by the gateway and read once by the worker. Garbage-collected by a runtime TTL plus a bucket lifecycle rule.
Because the payload store is required for >1 MiB work items, the shared bucket is created by default (create_model_cache = true). With it enabled, the module:
- Provisions a managed S3 bucket with versioning, abort-incomplete-multipart, and a lifecycle rule that deletes objects under the
payloads/prefix after one day. - Attaches two scoped inline policies to the SIE workload IRSA role: read-only on the cache, and
s3:Get/Put/Delete/AbortMultipartUploadconstrained to thepayloads/*prefix, with aListBucketprefix condition. - KMS-encrypted buckets get matching
kms:Decrypt/Encrypt/GenerateDataKeygrants.
After apply, pass the bucket into Helm with one terraform output:
helm upgrade --install sie-cluster oci://ghcr.io/superlinked/charts/sie-cluster --version 0.9.0 \
-f https://raw.githubusercontent.com/superlinked/sie/v0.9.0/deploy/helm/sie-cluster/values-aws.yaml \
--namespace sie --create-namespace \
--set serviceAccount.annotations."eks\.amazonaws\.com/role-arn"="$(terraform output -raw sie_irsa_role_arn)" \
$(terraform output -raw model_cache_helm_args)The chart auto-derives payloadStore.url from workers.common.clusterCache.url, so a single --set for the cache covers both the optional weights cache (models/) and the payload store (payloads/); the payload_store_url output is exposed for visibility and can be wired explicitly via --set payloadStore.url=... for the rare override case. On the chart side payloadStore.enabled defaults to true, decoupled from the optional workers.common.clusterCache. Operators who bring their own object-storage bucket can opt out (create_model_cache = false) and wire payloadStore.url themselves; skipping the payload store entirely means work items larger than 1 MiB (e.g. images) fail.
See infra/s3_model_cache.tf and infra/irsa.tf for the resource definitions.
This module follows AWS security best practices out of the box:
- KMS encryption - EKS secrets encrypted at rest with a dedicated, auto-rotating KMS key
- Restricted API endpoint - the public Kubernetes API endpoint accepts only
api_server_authorized_ip_ranges, or is disabled withenable_private_endpoint; any-address access needsallow_public_api_server - Private subnets - worker nodes run in private subnets with no public IPs
- NAT gateway - outbound internet via NAT (one per AZ for high availability)
- VPC endpoints - private access to ECR, S3, STS, EC2, CloudWatch, and other services
- IRSA - pods use IAM roles instead of long-lived credentials
- GPU taints - GPU nodes are tainted so only GPU workloads schedule on them
- Image scanning - ECR scans images on push for known vulnerabilities
- Audit logging - all EKS control plane log types enabled
Some pieces of a production deployment are intentionally not turnkey - either because they're cluster-wide / cross-stack concerns (registry, OIDC) or because they require domains and DNS records that only you can own (TLS, DNS). This module lets you opt out where it makes sense and points at the right knobs.
-
Container registry - optional. The module does not create ECR repos by default (
create_ecr_repositories = false, seeinfra/variables.tf) - this matches the chart's GHCR-by-default behaviour and avoidsRepositoryAlreadyExistsExceptionon accounts where repos already exist. Setcreate_ecr_repositories = trueto opt in to terraform-managed ECR; the module will create project-scoped repos (<project_name>/sie-server,<project_name>/sie-gateway,<project_name>/sie-config). Override the namespace viaecr_repository_prefix- set to""to disable prefixing for accounts where ECR is externally managed under bare names. The module always emitsecr_*_repository_urloutputs (composed from caller identity + repo names) so IRSA / Helm wiring is unchanged whether you opt in or not. The worker-sidecar uses the chart'sghcr.io/superlinked/sie-server-sidecardefault; to use an external registry for the other runtime images, point the Helm chart at it viagateway.image.repository,workers.common.image.repository, andconfig.image.repository. -
TLS certificate - BYO by default. Set
ingress.tlsConfig.modeto one of:byo- supply your ownkubernetes.io/tlsSecret.cert-manager- install cert-manager once in the cluster; the chart annotates the Ingress for automated Let's Encrypt issuance via HTTP-01.self-signed- for air-gapped clusters; setcertManagerBundle.certManager.install: trueto bundle cert-manager (single-tenant clusters only).
See the chart README's TLS / HTTPS section. DNS-01 / wildcard / ACM paths are out of scope for the chart.
-
DNS / domain - always BYO. This module does not provision Route53 zones or records. After
terraform apply, take the ingress controller's LoadBalancer hostname (kubectl -n ingress-nginx get svc ingress-nginx-controller) and create an A/AAAA record pointing at it under a domain you control. -
OIDC provider - BYO. When
auth.enabled: truein the chart, setauth.oauth2Proxy.oidcIssuerUrland the corresponding client ID / secret to your existing identity provider (Okta, Auth0, Google Workspace, Azure AD, ...). The module does not create an IdP.
terraform destroyImportant: GPU instances can be expensive. Always destroy dev/test clusters when not in use. Spot instances (gpu_capacity_type = "SPOT") save 60-70% but may be interrupted.