From 57828deb219cf7dca0c3ab81f5acdb0b765bce2a Mon Sep 17 00:00:00 2001 From: sam Date: Sat, 3 Oct 2026 10:04:20 +0800 Subject: [PATCH 1/3] Deploy the control plane to Kubernetes from one source commit Core and the Web console get a Kubernetes deployment beside the installer, for an operator who already runs PostgreSQL and a cluster. The installer keeps owning the single-host installation and its config.json; the manifests set only the documented Core and Web environment, so no setting gains a second home. One run builds both images from one commit and rolls them out together, so the console never talks to a Core of another release. Core runs as a single replica with the Recreate strategy because it takes a PostgreSQL lease that admits one execution service per database, and a surging Pod cannot take over from a running one. Web runs the same way because it holds console sign-in sessions in process memory, where a second replica would reject a cookie the other Pod issued. An init container prepares each secret as an owner-only file for the service account, because Kubernetes owns Secret volume files as root and Web refuses a Core key file that grants group or other access. A second one applies the embedded migrations before Core opens the database. The adapter state volume is opt-in. Only the E2B adapter writes there, for the receipts that let Core clean up, observe and verify ownership of sandboxes in E2B's cloud, so an installation with no E2B deployment needs no claim at all. The workflow validates every setting and secret before it touches the cluster, verifies the rolled-out image and ready endpoints, and, when it applies the Ingress, that /v1 reaches Core rather than the console. --- .github/actions/setup-kubectl/action.yaml | 76 ++++ .github/workflows/deploy.yml | 437 +++++++++++++++++++++ CONTRIBUTING.md | 1 + deploy/kubernetes/README.md | 89 +++++ deploy/kubernetes/prod/core-env.yaml.tpl | 30 ++ deploy/kubernetes/prod/core-state.yaml.tpl | 21 + deploy/kubernetes/prod/core.yaml.tpl | 199 ++++++++++ deploy/kubernetes/prod/ingress.yaml.tpl | 54 +++ deploy/kubernetes/prod/web.yaml.tpl | 164 ++++++++ docs/getting-started/install-options.md | 2 + docs/zh/getting-started/install-options.md | 4 +- scripts/ci_plan.py | 3 + 12 files changed, 1079 insertions(+), 1 deletion(-) create mode 100644 .github/actions/setup-kubectl/action.yaml create mode 100644 .github/workflows/deploy.yml create mode 100644 deploy/kubernetes/README.md create mode 100644 deploy/kubernetes/prod/core-env.yaml.tpl create mode 100644 deploy/kubernetes/prod/core-state.yaml.tpl create mode 100644 deploy/kubernetes/prod/core.yaml.tpl create mode 100644 deploy/kubernetes/prod/ingress.yaml.tpl create mode 100644 deploy/kubernetes/prod/web.yaml.tpl diff --git a/.github/actions/setup-kubectl/action.yaml b/.github/actions/setup-kubectl/action.yaml new file mode 100644 index 000000000..184eb055b --- /dev/null +++ b/.github/actions/setup-kubectl/action.yaml @@ -0,0 +1,76 @@ +name: Setup kubectl +description: Install kubectl and point it at the deployment cluster + +inputs: + kubeconfig: + description: Base64-encoded kubeconfig + required: true + server: + description: Override the cluster's API server URL; empty keeps the kubeconfig's own + required: false + default: '' + insecure-skip-tls-verify: + description: Skip API server certificate verification; only for a server address the certificate does not name + required: false + default: 'false' + version: + description: kubectl version to install when the runner has none + required: false + default: v1.31.4 + +runs: + using: composite + steps: + - shell: bash + env: + KUBECONFIG_CONTENT: ${{ inputs.kubeconfig }} + SERVER: ${{ inputs.server }} + INSECURE: ${{ inputs.insecure-skip-tls-verify }} + VERSION: ${{ inputs.version }} + run: | + set -euo pipefail + if [ -z "$KUBECONFIG_CONTENT" ]; then + echo "::error::The kubeconfig secret is not configured for this GitHub Environment." + exit 1 + fi + if ! command -v kubectl >/dev/null 2>&1; then + curl --fail --show-error --silent --location \ + -o "$RUNNER_TEMP/kubectl" \ + "https://dl.k8s.io/release/${VERSION}/bin/linux/amd64/kubectl" + install -m 0755 "$RUNNER_TEMP/kubectl" /usr/local/bin/kubectl + fi + + # The kubeconfig is a credential: write it privately and keep it out of the log. + install -d -m 0700 "$HOME/.kube" + umask 077 + printf '%s' "$KUBECONFIG_CONTENT" | base64 -d > "$HOME/.kube/config" + chmod 600 "$HOME/.kube/config" + + # Every later kubectl call uses the current context, so patch that + # context's cluster rather than whichever one comes first. + context="$(kubectl config current-context)" + cluster="$(kubectl config view -o jsonpath="{.contexts[?(@.name=='$context')].context.cluster}")" + if [ -z "$cluster" ]; then + echo "::error::The kubeconfig's current context names no cluster." + exit 1 + fi + echo "Using context $context, cluster $cluster" + arguments=() + if [ -n "$SERVER" ]; then + arguments+=("--server=$SERVER") + fi + if [ "$INSECURE" = "true" ]; then + arguments+=(--insecure-skip-tls-verify=true) + fi + if [ "${#arguments[@]}" -gt 0 ]; then + kubectl config set-cluster "$cluster" "${arguments[@]}" + fi + if [ "$INSECURE" = "true" ]; then + # kubectl rejects a configuration that carries both a certificate authority + # and the insecure flag. A property path splits on dots, so a cluster name + # containing one has to escape them. + escaped="${cluster//./\\.}" + kubectl config unset "clusters.${escaped}.certificate-authority-data" >/dev/null + kubectl config unset "clusters.${escaped}.certificate-authority" >/dev/null + fi + kubectl version --output=yaml diff --git a/.github/workflows/deploy.yml b/.github/workflows/deploy.yml new file mode 100644 index 000000000..8f7d6f554 --- /dev/null +++ b/.github/workflows/deploy.yml @@ -0,0 +1,437 @@ +name: core-deploy + +# Production deployment of the control plane: Core and the Web console, built from +# one source commit into two images and rolled out to Kubernetes. Both always come +# from the same commit, so the console never talks to a Core of another release. +# PostgreSQL is an existing external database reached through OAC_DATABASE_URL. +# Nodes, self-hosted machines and Runtime images are installed from a release, not +# here. deploy/kubernetes/README.md holds the prerequisites and every setting. + +on: + workflow_dispatch: + inputs: + ref: + description: Release tag or full commit SHA to deploy + required: true + type: string + apply_ingress: + description: Also apply the Ingress and verify the public route split + default: false + type: boolean + +permissions: + contents: read + +concurrency: + group: core-deploy-production + cancel-in-progress: false + +jobs: + prepare: + runs-on: ${{ vars.OAC_DEPLOY_RUNNER || 'ubuntu-22.04' }} + outputs: + revision: ${{ steps.source.outputs.revision }} + steps: + - uses: actions/checkout@v7 + with: + ref: ${{ inputs.ref }} + fetch-depth: 0 + persist-credentials: false + - name: Pin the source commit + id: source + run: | + set -euo pipefail + revision="$(git rev-parse HEAD)" + echo "revision=$revision" >> "$GITHUB_OUTPUT" + echo "Deploying \`$revision\`" >> "$GITHUB_STEP_SUMMARY" + + build: + needs: prepare + runs-on: ${{ vars.OAC_DEPLOY_RUNNER || 'ubuntu-22.04' }} + environment: production + timeout-minutes: 90 + strategy: + fail-fast: true + matrix: + component: [core, web] + env: + IMAGE: ${{ vars.OAC_REGISTRY }}/oac-${{ matrix.component }}:sha-${{ needs.prepare.outputs.revision }} + REVISION: ${{ needs.prepare.outputs.revision }} + steps: + - name: Prepare the build home + # The build scripts keep their caches and outputs under ~/.oac and refuse a + # directory elsewhere, so give the job a HOME it owns. + run: | + set -euo pipefail + mkdir -p "$RUNNER_TEMP/home/.oac" + echo "HOME=$RUNNER_TEMP/home" >> "$GITHUB_ENV" + - uses: actions/checkout@v7 + with: + ref: ${{ needs.prepare.outputs.revision }} + persist-credentials: false + - name: Check the registry setting + env: + REGISTRY: ${{ vars.OAC_REGISTRY }} + run: | + set -euo pipefail + if [ -z "$REGISTRY" ]; then + echo "::error::OAC_REGISTRY is not configured for the production environment." + exit 1 + fi + - name: Select shared Go caches + run: | + set -euo pipefail + echo "GOCACHE=$HOME/.oac/cache/go-build" >> "$GITHUB_ENV" + echo "GOMODCACHE=$HOME/.oac/cache/go-mod" >> "$GITHUB_ENV" + - uses: actions/setup-go@v7 + with: + go-version-file: go.mod + cache: false + - uses: actions/cache@v6 + with: + path: | + ${{ runner.temp }}/home/.oac/cache/go-build + ${{ runner.temp }}/home/.oac/cache/go-mod + key: core-deploy-go-${{ runner.os }}-${{ runner.arch }}-${{ hashFiles('**/go.mod', '**/go.sum') }}-${{ matrix.component }} + restore-keys: | + core-deploy-go-${{ runner.os }}-${{ runner.arch }}-${{ hashFiles('**/go.mod', '**/go.sum') }}- + - uses: actions/setup-node@v6 + if: matrix.component == 'web' + with: + node-version: '22' + - name: Install the frontend package manager + if: matrix.component == 'web' + run: npm install --global pnpm@10.30.3 + - name: Log in to the image registry + env: + REGISTRY_HOST: ${{ vars.OAC_REGISTRY_HOST || vars.OAC_REGISTRY }} + REGISTRY_USERNAME: ${{ secrets.OAC_REGISTRY_USERNAME }} + REGISTRY_PASSWORD: ${{ secrets.OAC_REGISTRY_PASSWORD }} + run: | + set -euo pipefail + for name in REGISTRY_USERNAME REGISTRY_PASSWORD; do + if [ -z "${!name:-}" ]; then + echo "::error::OAC_$name is not configured for the production environment." + exit 1 + fi + done + printf '%s' "$REGISTRY_PASSWORD" \ + | docker login "${REGISTRY_HOST%%/*}" --username "$REGISTRY_USERNAME" --password-stdin + - name: Build the Core image + if: matrix.component == 'core' + run: | + set -euo pipefail + context="$HOME/.oac/build/core-image-context" + rm -rf "$context" + mkdir -p "$context" + # Builds the Core commands for linux/amd64 and the E2B helper, then copies + # deploy/distribution/Dockerfile into the context. + ./scripts/build-core-image-context.sh "$context" + docker build --platform linux/amd64 \ + --label "org.opencontainers.image.revision=$REVISION" \ + --tag "$IMAGE" "$context" + - name: Build the Web image + if: matrix.component == 'web' + run: | + set -euo pipefail + context="$HOME/.oac/build/web-image-context" + rm -rf "$context" + mkdir -p "$context" + # build-web.sh compiles for the host unless told otherwise, and the image + # the cluster pulls is linux/amd64. + GOOS=linux GOARCH=amd64 OAC_DEV_WEB_BUILD_DIR="$context" ./scripts/build-web.sh + make web-deps + OAC_WEB_OPENAI_HOSTED_SESSIONS=1 OAC_WEB_ENVIRONMENT_FILES=1 pnpm build:web + cp -R apps/web/dist "$context/dist" + cp services/web/Dockerfile "$context/Dockerfile" + docker build --platform linux/amd64 \ + --label "org.opencontainers.image.revision=$REVISION" \ + --tag "$IMAGE" "$context" + - name: Push the image + run: docker push "$IMAGE" + - name: Discard the registry credential + if: always() + env: + REGISTRY_HOST: ${{ vars.OAC_REGISTRY_HOST || vars.OAC_REGISTRY }} + # A self-hosted runner keeps its home directory for the next job. + run: docker logout "${REGISTRY_HOST%%/*}" || true + + deploy: + needs: [prepare, build] + runs-on: ${{ vars.OAC_DEPLOY_RUNNER || 'ubuntu-22.04' }} + environment: production + timeout-minutes: 60 + env: + NS: ${{ vars.OAC_NAMESPACE || 'openagentcore' }} + OAC_CORE_IMAGE: ${{ vars.OAC_REGISTRY }}/oac-core:sha-${{ needs.prepare.outputs.revision }} + OAC_WEB_IMAGE: ${{ vars.OAC_REGISTRY }}/oac-web:sha-${{ needs.prepare.outputs.revision }} + steps: + - uses: actions/checkout@v7 + with: + ref: ${{ needs.prepare.outputs.revision }} + persist-credentials: false + - uses: ./.github/actions/setup-kubectl + with: + kubeconfig: ${{ secrets.OAC_KUBECONFIG }} + server: ${{ vars.OAC_K8S_SERVER }} + insecure-skip-tls-verify: ${{ vars.OAC_K8S_INSECURE || 'false' }} + + - name: Check the deployment settings + env: + OAC_PUBLIC_URL: ${{ vars.OAC_PUBLIC_URL }} + OAC_DATABASE_URL: ${{ vars.OAC_DATABASE_URL }} + OAC_INSTALLATION_ID: ${{ vars.OAC_INSTALLATION_ID }} + OAC_DATABASE_PASSWORD: ${{ secrets.OAC_DATABASE_PASSWORD }} + OAC_CREDENTIAL_KEY: ${{ secrets.OAC_CREDENTIAL_KEY }} + OAC_CORE_KEY: ${{ secrets.OAC_CORE_KEY }} + run: | + set -euo pipefail + if ! command -v envsubst >/dev/null 2>&1; then + echo "::error::envsubst is missing on this runner; install the gettext-base package." + exit 1 + fi + missing=0 + for name in OAC_PUBLIC_URL OAC_DATABASE_URL OAC_INSTALLATION_ID \ + OAC_DATABASE_PASSWORD OAC_CREDENTIAL_KEY OAC_CORE_KEY; do + if [ -z "${!name:-}" ]; then + echo "::error::$name is not configured for the production environment." + missing=1 + fi + done + [ "$missing" -eq 0 ] + # Core derives the daemon WebSocket address, the sandbox core_url and the + # self-hosted remote_url from the public URL, and Web serves only its own + # origin. Both reject a path, a query or a fragment. + case "$OAC_PUBLIC_URL" in + https://*/ | https://*/* | *\?* | *\#*) + echo "::error::OAC_PUBLIC_URL must be an HTTPS origin without a path, such as https://core.example." + exit 1 ;; + https://*) ;; + *) echo "::error::OAC_PUBLIC_URL must use HTTPS."; exit 1 ;; + esac + # The password belongs only in OAC_DATABASE_PASSWORD; Core refuses a URL + # that also carries one, and the URL must name the database user. + python3 - "$OAC_DATABASE_URL" <<'PY' + import sys, urllib.parse + parsed = urllib.parse.urlsplit(sys.argv[1]) + if parsed.scheme not in ("postgres", "postgresql") or not parsed.username: + sys.exit("OAC_DATABASE_URL must be postgres://USER@HOST:PORT/DATABASE") + # Go reports an empty password as present, and Core then refuses the URL. + if parsed.password is not None: + sys.exit("Set the database password only in the OAC_DATABASE_PASSWORD secret") + PY + bytes="$(printf '%s' "$OAC_CREDENTIAL_KEY" | base64 -d 2>/dev/null | wc -c | tr -d '[:space:]')" || bytes=0 + if [ "$bytes" != "32" ]; then + echo "::error::OAC_CREDENTIAL_KEY must be the base64 encoding of exactly 32 random bytes." + exit 1 + fi + if [ "${#OAC_CORE_KEY}" -lt 32 ]; then + echo "::error::OAC_CORE_KEY must have at least 32 characters." + exit 1 + fi + if printf '%s' "$OAC_CORE_KEY" | grep -q '[[:space:]]'; then + echo "::error::OAC_CORE_KEY must contain no whitespace." + exit 1 + fi + if ! printf '%s' "$OAC_INSTALLATION_ID" \ + | grep -Eqx '[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}'; then + echo "::error::OAC_INSTALLATION_ID must be a canonical lowercase UUID." + exit 1 + fi + + - name: Ensure the namespace + run: | + set -euo pipefail + kubectl get namespace "$NS" >/dev/null 2>&1 || kubectl create namespace "$NS" + + - name: Apply the registry pull secret + env: + REGISTRY_HOST: ${{ vars.OAC_REGISTRY_HOST || vars.OAC_REGISTRY }} + REGISTRY_USERNAME: ${{ secrets.OAC_REGISTRY_USERNAME }} + REGISTRY_PASSWORD: ${{ secrets.OAC_REGISTRY_PASSWORD }} + run: | + set -euo pipefail + kubectl create secret docker-registry oac-registry -n "$NS" \ + --docker-server="${REGISTRY_HOST%%/*}" \ + --docker-username="$REGISTRY_USERNAME" \ + --docker-password="$REGISTRY_PASSWORD" \ + --dry-run=client -o yaml | kubectl apply -f - + + - name: Apply the secrets and the environment + id: configuration + env: + OAC_PUBLIC_URL: ${{ vars.OAC_PUBLIC_URL }} + OAC_DATABASE_URL: ${{ vars.OAC_DATABASE_URL }} + OAC_INSTALLATION_ID: ${{ vars.OAC_INSTALLATION_ID }} + # Process settings carry no default here; docs/configuration.md owns + # them and Core applies its own when a key is absent. + OAC_HARNESSES: ${{ vars.OAC_HARNESSES }} + OAC_DEFAULT_HARNESS: ${{ vars.OAC_DEFAULT_HARNESS }} + OAC_EXECUTION_CONCURRENCY: ${{ vars.OAC_EXECUTION_CONCURRENCY }} + OAC_WRITE_AUDIT_RETENTION: ${{ vars.OAC_WRITE_AUDIT_RETENTION }} + OAC_LOG_LEVEL: ${{ vars.OAC_LOG_LEVEL }} + OAC_DATABASE_PASSWORD: ${{ secrets.OAC_DATABASE_PASSWORD }} + OAC_CREDENTIAL_KEY: ${{ secrets.OAC_CREDENTIAL_KEY }} + OAC_CORE_KEY: ${{ secrets.OAC_CORE_KEY }} + run: | + set -euo pipefail + umask 077 + work="$(mktemp -d)" + trap 'rm -rf "$work"' EXIT + + # Core authenticates /core/v1 against the digest, never against the key. + digest="$(printf '%s' "$OAC_CORE_KEY" | sha256sum | cut -d' ' -f1)" + printf '%s' "$OAC_DATABASE_PASSWORD" > "$work/database.password" + printf '%s' "$OAC_CREDENTIAL_KEY" > "$work/credential.key" + printf '["%s"]\n' "$digest" > "$work/core-key-digests.json" + printf '%s' "$OAC_CORE_KEY" > "$work/core.key" + + kubectl create secret generic oac-core-secrets -n "$NS" \ + --from-file="database.password=$work/database.password" \ + --from-file="credential.key=$work/credential.key" \ + --from-file="core-key-digests.json=$work/core-key-digests.json" \ + --dry-run=client -o yaml | kubectl apply -f - + kubectl create secret generic oac-web-core-key -n "$NS" \ + --from-file="core.key=$work/core.key" \ + --dry-run=client -o yaml | kubectl apply -f - + + # shellcheck disable=SC2016 # envsubst takes the variable list literally + envsubst '$OAC_PUBLIC_URL $OAC_INSTALLATION_ID $OAC_DATABASE_URL' \ + < deploy/kubernetes/prod/core-env.yaml.tpl > "$work/core-env.yaml" + for name in OAC_HARNESSES OAC_DEFAULT_HARNESS OAC_EXECUTION_CONCURRENCY \ + OAC_WRITE_AUDIT_RETENTION OAC_LOG_LEVEL; do + value="${!name:-}" + if [ -n "$value" ]; then printf ' %s: "%s"\n' "$name" "$value" >> "$work/core-env.yaml"; fi + done + kubectl apply -n "$NS" -f "$work/core-env.yaml" + + # Core and Web read their environment and secret files once at startup. + # This digest enters both Pod templates, so a settings or secret change + # replaces the Pods even when the images are unchanged. + digest_input="$(cat "$work/core-env.yaml"; printf '%s' "$digest$OAC_DATABASE_PASSWORD$OAC_CREDENTIAL_KEY")" + printf 'digest=%s\n' "$(printf '%s' "$digest_input" | sha256sum | cut -d' ' -f1)" >> "$GITHUB_OUTPUT" + + - name: Roll out Core + env: + OAC_CONFIGURATION_DIGEST: ${{ steps.configuration.outputs.digest }} + OAC_STORAGE_CLASS: ${{ vars.OAC_STORAGE_CLASS }} + OAC_STATE_SIZE: ${{ vars.OAC_STATE_SIZE || '10Gi' }} + OAC_CORE_CPU_REQUEST: ${{ vars.OAC_CORE_CPU_REQUEST || '500m' }} + OAC_CORE_CPU_LIMIT: ${{ vars.OAC_CORE_CPU_LIMIT || '2' }} + OAC_CORE_MEMORY_REQUEST: ${{ vars.OAC_CORE_MEMORY_REQUEST || '1Gi' }} + OAC_CORE_MEMORY_LIMIT: ${{ vars.OAC_CORE_MEMORY_LIMIT || '4Gi' }} + run: | + set -euo pipefail + # Only an E2B deployment writes to the state root. Without a storage + # class the volume lives as long as the Pod, which is correct until + # then; with one, Core keeps its E2B receipts across rollouts. + if [ -n "$OAC_STORAGE_CLASS" ]; then + # shellcheck disable=SC2016 # envsubst takes the variable list literally + envsubst '$OAC_STORAGE_CLASS $OAC_STATE_SIZE' \ + < deploy/kubernetes/prod/core-state.yaml.tpl | kubectl apply -n "$NS" -f - + OAC_STATE_VOLUME='persistentVolumeClaim: {claimName: oac-core-state}' + else + OAC_STATE_VOLUME='emptyDir: {}' + echo "::notice::No OAC_STORAGE_CLASS: the Core state root is a Pod-lifetime directory. Set one before an E2B deployment exists, or every rollout loses its receipts." + fi + export OAC_STATE_VOLUME + # shellcheck disable=SC2016 # envsubst takes the variable list literally + envsubst '$OAC_CORE_IMAGE $OAC_CONFIGURATION_DIGEST $OAC_STATE_VOLUME $OAC_CORE_CPU_REQUEST $OAC_CORE_CPU_LIMIT $OAC_CORE_MEMORY_REQUEST $OAC_CORE_MEMORY_LIMIT' \ + < deploy/kubernetes/prod/core.yaml.tpl | kubectl apply -n "$NS" -f - + # Recreate replaces the single Pod, so the rollout spans the old Pod's + # shutdown, the schema migration and Core's own startup. + kubectl rollout status deployment/oac-core -n "$NS" --timeout=15m + live="$(kubectl get deployment oac-core -n "$NS" \ + -o jsonpath='{.spec.template.spec.containers[?(@.name=="core")].image}')" + test "$live" = "$OAC_CORE_IMAGE" + + - name: Roll out Web + env: + OAC_CONFIGURATION_DIGEST: ${{ steps.configuration.outputs.digest }} + OAC_PUBLIC_URL: ${{ vars.OAC_PUBLIC_URL }} + OAC_LOG_LEVEL: ${{ vars.OAC_LOG_LEVEL || 'info' }} + run: | + set -euo pipefail + # shellcheck disable=SC2016 # envsubst takes the variable list literally + envsubst '$OAC_WEB_IMAGE $OAC_CORE_IMAGE $OAC_CONFIGURATION_DIGEST $OAC_PUBLIC_URL $OAC_LOG_LEVEL' \ + < deploy/kubernetes/prod/web.yaml.tpl | kubectl apply -n "$NS" -f - + kubectl rollout status deployment/oac-web -n "$NS" --timeout=10m + live="$(kubectl get deployment oac-web -n "$NS" \ + -o jsonpath='{.spec.template.spec.containers[?(@.name=="web")].image}')" + test "$live" = "$OAC_WEB_IMAGE" + + - name: Verify the Services have ready endpoints + run: | + set -euo pipefail + for service in oac-core oac-web; do + addresses="$(kubectl get endpoints "$service" -n "$NS" \ + -o jsonpath='{.subsets[*].addresses[*].ip}')" + if [ -z "$addresses" ]; then + echo "::error::Service $service has no ready endpoint." + exit 1 + fi + echo "$service: $addresses" + done + + - name: Apply the Ingress + if: inputs.apply_ingress + env: + OAC_PUBLIC_URL: ${{ vars.OAC_PUBLIC_URL }} + OAC_INGRESS_CLASS: ${{ vars.OAC_INGRESS_CLASS || 'nginx' }} + OAC_TLS_SECRET: ${{ vars.OAC_TLS_SECRET || 'oac-tls' }} + run: | + set -euo pipefail + # A missing class or certificate otherwise surfaces only as the route + # check timing out, which blames the path split instead. + if ! kubectl get ingressclass "$OAC_INGRESS_CLASS" >/dev/null 2>&1; then + echo "::error::No IngressClass named $OAC_INGRESS_CLASS; set OAC_INGRESS_CLASS to one of: $(kubectl get ingressclass -o jsonpath='{.items[*].metadata.name}')" + exit 1 + fi + if ! kubectl get secret "$OAC_TLS_SECRET" -n "$NS" >/dev/null 2>&1; then + echo "::error::No TLS Secret named $OAC_TLS_SECRET in namespace $NS; create it from the public host's certificate, or set OAC_TLS_SECRET." + exit 1 + fi + OAC_PUBLIC_HOST="${OAC_PUBLIC_URL#https://}" + export OAC_PUBLIC_HOST + # shellcheck disable=SC2016 # envsubst takes the variable list literally + envsubst '$OAC_PUBLIC_HOST $OAC_INGRESS_CLASS $OAC_TLS_SECRET' \ + < deploy/kubernetes/prod/ingress.yaml.tpl | kubectl apply -n "$NS" -f - + + - name: Verify the public route split + if: inputs.apply_ingress + env: + OAC_PUBLIC_URL: ${{ vars.OAC_PUBLIC_URL }} + run: | + set -euo pipefail + # 401 proves /v1 reached Core, which asks for a Project API key. 404 means + # it reached Web instead: every application call and node connection would + # then fail. A new Ingress needs a moment to program its load balancer. + for attempt in $(seq 1 30); do + code="$(curl --show-error --silent -o /dev/null -w '%{http_code}' \ + -H 'OpenAI-Beta: agents=v1' "$OAC_PUBLIC_URL/v1/agents" || echo 000)" + if [ "$code" = "401" ]; then + echo "GET /v1/agents -> 401: /v1 reaches Core" + exit 0 + fi + echo "attempt $attempt: GET /v1/agents -> $code" + sleep 10 + done + echo "::error::/v1 does not reach Core; check the Ingress path split." + exit 1 + + - name: Report the live state on failure + if: failure() + run: | + set +e + for app in oac-core oac-web; do + echo "=== $app" + kubectl get pods -n "$NS" -l "app=$app" -o wide + kubectl describe pods -n "$NS" -l "app=$app" | tail -60 + kubectl logs -n "$NS" -l "app=$app" --all-containers --tail=120 + done + kubectl get events -n "$NS" --sort-by=.lastTimestamp | tail -40 + + - name: Discard the cluster credential + if: always() + # A self-hosted runner keeps its home directory for the next job. + run: rm -f "$HOME/.kube/config" diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 929787bb8..7e4718bf0 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -34,6 +34,7 @@ This guide owns how to work in the repository: documentation ownership, the repo | Website: landing page, bilingual documentation maintenance, documentation site build and GitHub Pages publication | [Website guide](website/README.md) | | Self-hosted Runtime installation, recovery and local operation | [Self-hosted execution](docs/getting-started/self-hosted.md) | | Installer lifecycle, locking, generated state, managed HTTPS and downloads | [Installer design rules](deploy/install/README.md) | +| Kubernetes control-plane deployment, its manifests and its workflow | [Kubernetes deployment](deploy/kubernetes/README.md) | | Operator installation and alternatives | [Installation](docs/getting-started/install.md), [installation options](docs/getting-started/install-options.md) | | Settings, defaults, files and installation layout | [Configuration](docs/configuration.md) | | Operator commands, keys, backup and version policy | [Operations](docs/getting-started/operations.md) | diff --git a/deploy/kubernetes/README.md b/deploy/kubernetes/README.md new file mode 100644 index 000000000..f9ccdb25d --- /dev/null +++ b/deploy/kubernetes/README.md @@ -0,0 +1,89 @@ +# Kubernetes deployment + +This directory deploys the control plane — Core and the Web console — to a Kubernetes cluster from a source commit, and the `core-deploy` workflow drives it. It is an alternative to the [installer](../../docs/getting-started/install.md), for an operator who already runs PostgreSQL and a cluster. The installer owns the single-host installation and its `config.json`; nothing here reads or writes those files. Core and Web take their settings from the environment, as [Configuration](../../docs/configuration.md#appendix-core-environment-without-the-installer) defines it, and this deployment is the one place that sets those variables for the cluster. + +## What it deploys + +| Object | From | Notes | +| --- | --- | --- | +| Deployment and Service `oac-core` | `prod/core.yaml.tpl` | Core on port 8091. One replica: Core takes a PostgreSQL lease that gives one execution service per database, so a second replica exits at startup. An init container prepares the adapter state volume, then `oac-core-migrate` applies the schema | +| ConfigMap `oac-core-env` | `prod/core-env.yaml.tpl` | Core's process environment; no secret | +| PersistentVolumeClaim `oac-core-state` | `prod/core-state.yaml.tpl` | `OAC_PROVIDER_STATE_ROOT`, applied only when `OAC_STORAGE_CLASS` is set; see [the state volume](#the-state-volume). Back it up with the database and the credential key | +| Deployment and Service `oac-web` | `prod/web.yaml.tpl` | The console on port 8080. One replica: Web holds sign-in sessions in process memory, so a cookie is valid only on the Pod that issued it | +| Ingress `oac` | `prod/ingress.yaml.tpl` | Applied only when the run selects **Also apply the Ingress**. One origin, split by path | +| Secrets `oac-core-secrets`, `oac-web-core-key`, `oac-registry` | GitHub Environment secrets | Created by the workflow from stdin; never written to a file in the repository | + +PostgreSQL is yours. So is anything installed on a machine: nodes, self-hosted machines and the Runtime images come from a [release](../../docs/getting-started/nodes.md), not from this workflow. + +Both images always come from one commit, so the console never talks to a Core of another release. + +## Cluster prerequisites + +- A PostgreSQL database and an account of its own, reachable from the cluster. Core needs no extension and migrates the schema itself. +- For an E2B deployment only: a StorageClass that provisions a ReadWriteOnce volume. See [the state volume](#the-state-volume). +- An image registry the cluster can pull from, with a user the workflow can push as. +- For the Ingress: a controller and a TLS Secret for the public host. The manifest carries [ingress-nginx](https://kubernetes.github.io/ingress-nginx/) annotations; another controller needs its own equivalents for the four requirements in [HTTPS and the reverse proxy](../../docs/getting-started/install-options.md#https-and-the-reverse-proxy). Leave the Ingress out of the run and keep your own routing if you prefer. +- The runner needs `docker`, `envsubst` and outbound access to the registry and the cluster's API server. Set `OAC_DEPLOY_RUNNER` to a self-hosted label when the API server is private. + +## Deploy + +Run **Actions** → **core-deploy** → **Run workflow**, give it a release tag or a full commit SHA, and select **Also apply the Ingress** the first time or after the routing changes. The run builds both images, pushes them, applies the secrets and the environment, rolls Core out and then Web, and checks that both Services have a ready endpoint. With the Ingress it also checks that `/v1` answers `401` — proof that the path split reaches Core and not the console. + +**A deployment is an outage.** Core's single replica stops before its replacement starts, and the replacement migrates the schema before it listens, so the Agents API, the machine routes and every running Session are unavailable for the rollout. Deploy in a window you can afford to lose. The Pod that takes over also needs the previous one's leased database connection to be gone; until it is, `AcquireLease` fails and Core exits, and the rollout depends on a restart landing inside the 15-minute deadline. + +To roll back, run the workflow again with the earlier commit. A rollback across a schema migration is not automatic: the earlier Core runs against the newer schema. + +## Settings + +Set these on the repository's `production` GitHub Environment. Variables are not secret; secrets are. The workflow stops and names anything missing or malformed before it touches the cluster. + +### Secrets + +| Secret | Value | +| --- | --- | +| `OAC_KUBECONFIG` | Base64 of a kubeconfig for the cluster: `base64 -w0 < kubeconfig` | +| `OAC_REGISTRY_USERNAME`, `OAC_REGISTRY_PASSWORD` | The registry account the workflow pushes as and the cluster pulls with | +| `OAC_DATABASE_PASSWORD` | The database account's password. Keep it out of `OAC_DATABASE_URL`; Core refuses a URL that also carries one | +| `OAC_CREDENTIAL_KEY` | Base64 of exactly 32 random bytes, which encrypts stored credentials: `openssl rand -base64 32`. Losing it loses every stored credential | +| `OAC_CORE_KEY` | The Core key: sign-in to the console and the credential for `/core/v1`. At least 32 characters, no whitespace: `openssl rand -hex 32`. Core stores only its SHA-256 | + +### Variables + +| Variable | Value | +| --- | --- | +| `OAC_REGISTRY` | Image repository prefix, such as `registry.example.com/openagentcore`. The workflow pushes `oac-core` and `oac-web` beneath it | +| `OAC_REGISTRY_HOST` | Registry host for `docker login` and the pull Secret. Optional; the host of `OAC_REGISTRY` by default | +| `OAC_PUBLIC_URL` | The canonical HTTPS origin, such as `https://core.example`, with no path, query or fragment. Core derives the daemon WebSocket address, the sandbox `core_url` and the self-hosted `remote_url` from it, and Web serves only this origin. Without it Core executes no Session | +| `OAC_DATABASE_URL` | `postgres://USER@HOST:PORT/DATABASE`, with `sslmode` and any `pool_*` parameter in the query and no password | +| `OAC_INSTALLATION_ID` | A canonical lowercase UUID, generated once with `uuidgen \| tr 'A-Z' 'a-z'`. Core refuses an ID other than the one its database recorded, so it belongs to the database: keep the two together and restore them together | +| `OAC_NAMESPACE` | Namespace; `openagentcore` by default. The workflow creates it if it is missing | +| `OAC_STORAGE_CLASS`, `OAC_STATE_SIZE` | StorageClass and size of the Core state volume. Unset means no volume; see [the state volume](#the-state-volume). `10Gi` by default | +| `OAC_K8S_SERVER` | API server URL that overrides the kubeconfig's. Optional | +| `OAC_K8S_INSECURE` | `true` skips API server certificate verification, for an address the certificate does not name. Leave it unset otherwise | +| `OAC_DEPLOY_RUNNER` | Runner label; `ubuntu-22.04` by default. Set this one at repository level: the job that pins the commit runs before the environment is entered, so an environment-scoped value does not reach it | +| `OAC_INGRESS_CLASS`, `OAC_TLS_SECRET` | `ingressClassName` and the TLS Secret of the public host; `nginx` and `oac-tls` by default | +| `OAC_CORE_CPU_REQUEST`, `OAC_CORE_CPU_LIMIT`, `OAC_CORE_MEMORY_REQUEST`, `OAC_CORE_MEMORY_LIMIT` | Core's resources; `500m`, `2`, `1Gi` and `4Gi` by default | +| `OAC_HARNESSES`, `OAC_DEFAULT_HARNESS`, `OAC_EXECUTION_CONCURRENCY`, `OAC_WRITE_AUDIT_RETENTION`, `OAC_LOG_LEVEL` | The matching [process settings](../../docs/configuration.md#settings), which define their values and defaults | + +A change to any of these takes effect on the next run: a digest of the environment and the secrets enters both Pod templates, so the Pods are replaced even when the images are unchanged. + +## The state volume + +`OAC_PROVIDER_STATE_ROOT` is the private root each Sandbox Provider adapter keeps its own state under, and today exactly one adapter uses it: E2B stores its receipts in `/state/e2b`. The helper writes a receipt before each remote `Create`, because a helper's exit never proves that the remote call settled, and Core needs the receipt afterwards to clean up, observe and verify ownership of a sandbox living in E2B's cloud. Each receipt is at most 64 KiB, they are never pruned, and losing them orphans sandboxes that E2B keeps billing and Core can no longer destroy. [The E2B helper](../../services/core/tools/e2b-provider/README.md#receipts-and-state-directory) owns these rules. + +So the volume is needed exactly when an E2B deployment exists, and only then: + +- **No `OAC_STORAGE_CLASS`:** no claim is created and `/state` lives as long as the Pod. Correct for an installation with no E2B deployment, which writes nothing there. +- **`OAC_STORAGE_CLASS` set:** the claim is created and the receipts survive rollouts. Set it before selecting E2B, not after — every rollout in between discards the receipts written since the last one. + +Size it for the claim's lifetime rather than for today, because a StorageClass without `allowVolumeExpansion` cannot grow and a `Delete` reclaim policy destroys the receipts with the claim. The default `10Gi` also clears the minimum that a block-storage provisioner may impose. + +A claim's storage class and size are immutable. Changing either variable once the claim exists makes the run fail on the claim; migrate the receipts to a new claim yourself and delete the old one. + +## Differences from an installed installation + +- **E2B is the only backend that executes anything.** Web serves the node installer from the release bundle's node payload, which these images do not carry, so `/node-install/*` is unavailable and Add node is hidden. The Core image's native installer catalog is empty too — only the release build fills it — so `POST /api/v1/agent-daemon/installation` answers `503 installation_unavailable` and a self-hosted Session's install command cannot resolve. Until the node payload and the native catalog are built into these images, select [E2B](../../docs/getting-started/install-options.md), which needs neither. +- **No startup settings panel.** `OAC_SETTINGS_FILE` is the installer's snapshot, so `GET /core/v1/installation` reports none and the console shows no process settings. Runtime settings in Web are unaffected; they live in Core's database. +- **No domain setup in Web.** There is no installer socket, so the public origin comes from `OAC_PUBLIC_URL` and the Ingress, not from the console. +- **No `oac` command.** The maintenance commands ship in the Core image: `kubectl exec deploy/oac-core -- oac-core-device …`. [Operations](../../docs/getting-started/operations.md) covers keys and backup. +- **A rollout signs operators out** of the console, because Web holds its sessions in memory. diff --git a/deploy/kubernetes/prod/core-env.yaml.tpl b/deploy/kubernetes/prod/core-env.yaml.tpl new file mode 100644 index 000000000..67ac8d208 --- /dev/null +++ b/deploy/kubernetes/prod/core-env.yaml.tpl @@ -0,0 +1,30 @@ +# Core's process environment. Rendered by the deploy workflow with envsubst. +# Every variable here is documented in docs/configuration.md; this file holds no +# secret. The database password, the credential key and the Core key digest live +# in the oac-core-secrets Secret the workflow creates from GitHub Environment +# secrets, and Core reads each of them from a file. +apiVersion: v1 +kind: ConfigMap +metadata: + name: oac-core-env + labels: + app.kubernetes.io/name: oac-core + app.kubernetes.io/part-of: openagentcore +data: + OAC_ADDR: ":8091" + OAC_PUBLIC_URL: "${OAC_PUBLIC_URL}" + OAC_INSTALLATION_ID: "${OAC_INSTALLATION_ID}" + + OAC_DATABASE_URL: "${OAC_DATABASE_URL}" + OAC_DATABASE_PASSWORD_FILE: "/run/oac/database.password" + OAC_CREDENTIAL_KEY_FILE: "/run/oac/credential.key" + OAC_CORE_KEY_DIGESTS_FILE: "/run/oac/core-key-digests.json" + + # Adapter artifacts ship in the image; adapter state lives on the Core volume. + OAC_PROVIDER_ROOT: "/opt/oac" + OAC_PROVIDER_STATE_ROOT: "/state" + + OAC_LOG_FORMAT: "json" + # The deploy workflow appends the process settings the operator set here. A + # setting nobody set stays absent, so Core applies its own documented default + # instead of a copy of it. diff --git a/deploy/kubernetes/prod/core-state.yaml.tpl b/deploy/kubernetes/prod/core-state.yaml.tpl new file mode 100644 index 000000000..f60fa1a94 --- /dev/null +++ b/deploy/kubernetes/prod/core-state.yaml.tpl @@ -0,0 +1,21 @@ +# Core's adapter state volume, applied only when OAC_STORAGE_CLASS is set. +# +# Only the E2B adapter uses this volume today: it holds the receipts Core writes +# before each remote Create and needs afterwards to clean up, observe and verify +# ownership of sandboxes in E2B's cloud. They are small, they are never pruned, +# and losing them orphans billed sandboxes Core can no longer destroy. An +# installation with no E2B deployment writes nothing here and needs no claim. +apiVersion: v1 +kind: PersistentVolumeClaim +metadata: + name: oac-core-state + labels: + app.kubernetes.io/name: oac-core + app.kubernetes.io/part-of: openagentcore +spec: + accessModes: + - ReadWriteOnce + storageClassName: "${OAC_STORAGE_CLASS}" + resources: + requests: + storage: "${OAC_STATE_SIZE}" diff --git a/deploy/kubernetes/prod/core.yaml.tpl b/deploy/kubernetes/prod/core.yaml.tpl new file mode 100644 index 000000000..319afc6c2 --- /dev/null +++ b/deploy/kubernetes/prod/core.yaml.tpl @@ -0,0 +1,199 @@ +# Core: the Agents API, the Core administration API and the machine routes, plus +# the execution worker. Rendered by the deploy workflow with envsubst. +apiVersion: apps/v1 +kind: Deployment +metadata: + name: oac-core + labels: + app.kubernetes.io/name: oac-core + app.kubernetes.io/part-of: openagentcore +spec: + # Core takes a PostgreSQL lease that gives one execution service per database, + # so a second replica exits at startup and a surging Pod cannot take over from + # a running one. One replica, replaced only after the previous one stops. + replicas: 1 + strategy: + type: Recreate + progressDeadlineSeconds: 900 + selector: + matchLabels: + app: oac-core + template: + metadata: + labels: + app: oac-core + app.kubernetes.io/name: oac-core + app.kubernetes.io/part-of: openagentcore + annotations: + # Core reads its environment and its secret files once, so a settings or + # secret change must replace the Pod even when the image is unchanged. + io.oac/configuration-digest: "${OAC_CONFIGURATION_DIGEST}" + spec: + imagePullSecrets: + - name: oac-registry + terminationGracePeriodSeconds: 90 + securityContext: + seccompProfile: + type: RuntimeDefault + volumes: + # Kubernetes owns the files of a Secret volume as root, and Core's user + # cannot read an owner-only file it does not own. The preparation step + # copies each one to Core's user with owner-only access. + - name: secret-source + secret: + secretName: oac-core-secrets + defaultMode: 0400 + - name: secrets + emptyDir: + medium: Memory + # The adapter state root. Only the E2B adapter uses it today, for the + # receipts that let Core clean up its remote sandboxes, and those must + # outlive the Pod. Without a claim it is a Pod-lifetime directory, which + # is correct until an E2B deployment exists. core-state.yaml.tpl owns the + # claim; deploy/kubernetes/README.md explains when to turn it on. + - {name: state, ${OAC_STATE_VOLUME}} + - name: tmp + emptyDir: {} + initContainers: + # The volume arrives owned by root, and the E2B adapter requires its state + # directory to exist with no group or other access. + - name: prepare + image: "${OAC_CORE_IMAGE}" + imagePullPolicy: IfNotPresent + command: + - /bin/sh + - -eu + - -c + - | + mkdir -p /state/e2b + chown -R 65532:65532 /state + chmod 700 /state/e2b + chown 65532:65532 /run/oac + chmod 700 /run/oac + for name in database.password credential.key core-key-digests.json; do + install -o 65532 -g 65532 -m 0400 "/run/oac-source/$name" "/run/oac/$name" + done + securityContext: + runAsUser: 0 + runAsGroup: 0 + allowPrivilegeEscalation: false + readOnlyRootFilesystem: true + capabilities: + drop: ["ALL"] + add: ["CHOWN", "FOWNER", "DAC_OVERRIDE"] + volumeMounts: + - name: secret-source + mountPath: /run/oac-source + readOnly: true + - name: secrets + mountPath: /run/oac + - name: state + mountPath: /state + resources: + requests: + cpu: 10m + memory: 32Mi + limits: + cpu: 200m + memory: 128Mi + # Core never migrates its own schema. The schema must match the image + # before Core opens the database. + - name: migrate + image: "${OAC_CORE_IMAGE}" + imagePullPolicy: IfNotPresent + command: ["/usr/local/bin/oac-core-migrate"] + envFrom: + - configMapRef: + name: oac-core-env + securityContext: + runAsUser: 65532 + runAsGroup: 65532 + runAsNonRoot: true + allowPrivilegeEscalation: false + readOnlyRootFilesystem: true + capabilities: + drop: ["ALL"] + volumeMounts: + - name: secrets + mountPath: /run/oac + readOnly: true + - name: tmp + mountPath: /tmp + resources: + requests: + cpu: 100m + memory: 128Mi + limits: + cpu: "1" + memory: 512Mi + containers: + - name: core + image: "${OAC_CORE_IMAGE}" + imagePullPolicy: IfNotPresent + ports: + - name: http + containerPort: 8091 + protocol: TCP + envFrom: + - configMapRef: + name: oac-core-env + securityContext: + runAsUser: 65532 + runAsGroup: 65532 + runAsNonRoot: true + allowPrivilegeEscalation: false + readOnlyRootFilesystem: true + capabilities: + drop: ["ALL"] + volumeMounts: + - name: secrets + mountPath: /run/oac + readOnly: true + - name: state + mountPath: /state + - name: tmp + mountPath: /tmp + resources: + requests: + cpu: "${OAC_CORE_CPU_REQUEST}" + memory: "${OAC_CORE_MEMORY_REQUEST}" + limits: + cpu: "${OAC_CORE_CPU_LIMIT}" + memory: "${OAC_CORE_MEMORY_LIMIT}" + startupProbe: + httpGet: + path: /healthz + port: http + periodSeconds: 3 + timeoutSeconds: 3 + failureThreshold: 60 + livenessProbe: + httpGet: + path: /healthz + port: http + periodSeconds: 30 + timeoutSeconds: 5 + failureThreshold: 3 + readinessProbe: + httpGet: + path: /healthz + port: http + periodSeconds: 10 + timeoutSeconds: 5 + failureThreshold: 3 +--- +apiVersion: v1 +kind: Service +metadata: + name: oac-core + labels: + app.kubernetes.io/name: oac-core + app.kubernetes.io/part-of: openagentcore +spec: + type: ClusterIP + selector: + app: oac-core + ports: + - name: http + port: 8091 + targetPort: http diff --git a/deploy/kubernetes/prod/ingress.yaml.tpl b/deploy/kubernetes/prod/ingress.yaml.tpl new file mode 100644 index 000000000..87789ef58 --- /dev/null +++ b/deploy/kubernetes/prod/ingress.yaml.tpl @@ -0,0 +1,54 @@ +# One public origin for Core and Web, split by path. The annotations below are +# ingress-nginx's; another controller needs its own equivalents for the four +# requirements in docs/getting-started/install-options.md#https-and-the-reverse-proxy: +# preserve Host, pass WebSocket upgrades on /api/v1, never buffer or time out a +# stream, and accept large uploads. +apiVersion: networking.k8s.io/v1 +kind: Ingress +metadata: + name: oac + labels: + app.kubernetes.io/part-of: openagentcore + annotations: + # /v1 streams Session events and /api/v1 carries long-lived WebSockets. + nginx.ingress.kubernetes.io/proxy-buffering: "off" + nginx.ingress.kubernetes.io/proxy-request-buffering: "off" + nginx.ingress.kubernetes.io/proxy-read-timeout: "3600" + nginx.ingress.kubernetes.io/proxy-send-timeout: "3600" + # Core enforces its own upload limits; source files may reach 512 MiB. + nginx.ingress.kubernetes.io/proxy-body-size: "0" + nginx.ingress.kubernetes.io/ssl-redirect: "true" +spec: + ingressClassName: "${OAC_INGRESS_CLASS}" + tls: + - hosts: + - "${OAC_PUBLIC_HOST}" + secretName: "${OAC_TLS_SECRET}" + rules: + - host: "${OAC_PUBLIC_HOST}" + http: + paths: + # Applications, with a Project API key. + - path: /v1 + pathType: Prefix + backend: + service: + name: oac-core + port: + number: 8091 + # Nodes, sandboxes and self-hosted machines. Uses WebSockets. + - path: /api/v1 + pathType: Prefix + backend: + service: + name: oac-core + port: + number: 8091 + # Browsers. + - path: / + pathType: Prefix + backend: + service: + name: oac-web + port: + number: 8080 diff --git a/deploy/kubernetes/prod/web.yaml.tpl b/deploy/kubernetes/prod/web.yaml.tpl new file mode 100644 index 000000000..286110eb6 --- /dev/null +++ b/deploy/kubernetes/prod/web.yaml.tpl @@ -0,0 +1,164 @@ +# Web: the console. It serves the built frontend from the image and forwards +# signed-in /core/v1 requests to Core with the Core key. Rendered with envsubst. +apiVersion: apps/v1 +kind: Deployment +metadata: + name: oac-web + labels: + app.kubernetes.io/name: oac-web + app.kubernetes.io/part-of: openagentcore +spec: + # Web holds console sign-in sessions in process memory, so a cookie is valid + # only on the Pod that issued it. One replica, and no surge that would send a + # signed-in operator to a Pod that does not know the cookie. A rollout signs + # operators out; they sign in again with the Core key. + replicas: 1 + strategy: + type: Recreate + progressDeadlineSeconds: 600 + selector: + matchLabels: + app: oac-web + template: + metadata: + labels: + app: oac-web + app.kubernetes.io/name: oac-web + app.kubernetes.io/part-of: openagentcore + annotations: + io.oac/configuration-digest: "${OAC_CONFIGURATION_DIGEST}" + spec: + imagePullSecrets: + - name: oac-registry + terminationGracePeriodSeconds: 30 + securityContext: + seccompProfile: + type: RuntimeDefault + volumes: + # Web rejects a Core key file that grants group or other access, and + # Kubernetes owns the files of a Secret volume as root. The preparation + # step copies the key to Web's user with owner-only access. + - name: secret-source + secret: + secretName: oac-web-core-key + defaultMode: 0400 + - name: secrets + emptyDir: + medium: Memory + - name: tmp + emptyDir: {} + initContainers: + # The Web image is distroless and has no shell, so the preparation step + # runs Core's image from the same release. + - name: prepare + image: "${OAC_CORE_IMAGE}" + imagePullPolicy: IfNotPresent + command: + - /bin/sh + - -eu + - -c + - | + chown 65532:65532 /run/oac + chmod 700 /run/oac + install -o 65532 -g 65532 -m 0400 /run/oac-source/core.key /run/oac/core.key + securityContext: + runAsUser: 0 + runAsGroup: 0 + allowPrivilegeEscalation: false + readOnlyRootFilesystem: true + capabilities: + drop: ["ALL"] + add: ["CHOWN", "FOWNER", "DAC_OVERRIDE"] + volumeMounts: + - name: secret-source + mountPath: /run/oac-source + readOnly: true + - name: secrets + mountPath: /run/oac + resources: + requests: + cpu: 10m + memory: 32Mi + limits: + cpu: 200m + memory: 128Mi + containers: + - name: web + image: "${OAC_WEB_IMAGE}" + imagePullPolicy: IfNotPresent + ports: + - name: http + containerPort: 8080 + protocol: TCP + env: + - name: OAC_WEB_ADDR + value: ":8080" + # The exact browser-facing origin. Web serves no other host. + - name: OAC_WEB_ORIGIN + value: "${OAC_PUBLIC_URL}" + - name: OAC_WEB_UPSTREAM + value: "http://oac-core:8091" + - name: OAC_WEB_CORE_KEY_FILE + value: "/run/oac/core.key" + - name: OAC_LOG_LEVEL + value: "${OAC_LOG_LEVEL}" + - name: OAC_LOG_FORMAT + value: "json" + securityContext: + runAsUser: 65532 + runAsGroup: 65532 + runAsNonRoot: true + allowPrivilegeEscalation: false + readOnlyRootFilesystem: true + capabilities: + drop: ["ALL"] + volumeMounts: + - name: secrets + mountPath: /run/oac + readOnly: true + - name: tmp + mountPath: /tmp + resources: + requests: + cpu: 50m + memory: 64Mi + limits: + cpu: 500m + memory: 256Mi + startupProbe: + httpGet: + path: /healthz + port: http + periodSeconds: 2 + timeoutSeconds: 2 + failureThreshold: 30 + livenessProbe: + httpGet: + path: /healthz + port: http + periodSeconds: 30 + timeoutSeconds: 5 + failureThreshold: 3 + readinessProbe: + httpGet: + path: /healthz + port: http + periodSeconds: 5 + timeoutSeconds: 3 + failureThreshold: 3 +--- +apiVersion: v1 +kind: Service +metadata: + name: oac-web + labels: + app.kubernetes.io/name: oac-web + app.kubernetes.io/part-of: openagentcore +spec: + type: ClusterIP + selector: + app: oac-web + ports: + - name: http + port: 8080 + targetPort: http diff --git a/docs/getting-started/install-options.md b/docs/getting-started/install-options.md index 2e5063800..138c3f946 100644 --- a/docs/getting-started/install-options.md +++ b/docs/getting-started/install-options.md @@ -4,6 +4,8 @@ title: "Installation options and advanced deployments" The [default installation](./install.md) needs no options. Use this page to run behind an existing reverse proxy or install without internet access. +To run Core and the console on Kubernetes with your own PostgreSQL instead of the installer, see the [Kubernetes deployment](../../deploy/kubernetes/README.md). + Pass options to the downloaded script: ```sh diff --git a/docs/zh/getting-started/install-options.md b/docs/zh/getting-started/install-options.md index 03b553f3e..ca4fdc3ba 100644 --- a/docs/zh/getting-started/install-options.md +++ b/docs/zh/getting-started/install-options.md @@ -1,11 +1,13 @@ --- title: "安装选项与高级部署" source: docs/getting-started/install-options.md -source_hash: 7178baf93a8d8b6e7f086d73033afe4ea14afcf41c88dbcc6031a3a6dd20ca17 +source_hash: 4e726974eaa85c8ec64c242ed4947faee0c08bb5fdcfa527e43c6284af96c851 --- [默认安装](install.md)无需任何选项。使用本页可以在现有反向代理后运行,或者在无法访问互联网时进行安装。 +若要用自己的 PostgreSQL 在 Kubernetes 上运行 Core 和控制台,而不使用安装器,参见 [Kubernetes 部署](../../../deploy/kubernetes/README.md)。 + 向下载的脚本传递选项: ```sh diff --git a/scripts/ci_plan.py b/scripts/ci_plan.py index 2b3ef9d67..f512d172f 100644 --- a/scripts/ci_plan.py +++ b/scripts/ci_plan.py @@ -19,6 +19,9 @@ ".github/workflows/actionlint.yml": ("lint",), ".github/actionlint.yaml": ("lint",), ".github/workflows/ci-review.yml": ("lint",), + # Deployment inputs: no check builds or tests them, so only workflow syntax applies. + ".github/workflows/deploy.yml": ("lint",), + ".github/actions/setup-kubectl/action.yaml": ("lint",), ".github/workflows/website.yml": ("website", "lint"), ".github/actions/node/action.yml": (*NODE_JOBS, "lint"), "scripts/ci_plan.py": JOBS, From 0465f4c115e2775249948f930a8b17c16c9da9b4 Mon Sep 17 00:00:00 2001 From: sam Date: Sat, 3 Oct 2026 11:47:06 +0800 Subject: [PATCH 2/3] Keep fork changes limited to production deployment --- .github/actions/setup-kubectl/action.yaml | 11 +- .github/workflows/deploy.yml | 141 ++++++++++----------- CONTRIBUTING.md | 1 - deploy/kubernetes/README.md | 26 ++-- deploy/kubernetes/prod/core-state.yaml.tpl | 2 +- deploy/kubernetes/prod/core.yaml.tpl | 7 +- deploy/kubernetes/prod/ingress.yaml.tpl | 54 -------- docs/getting-started/install-options.md | 2 - docs/zh/getting-started/install-options.md | 4 +- scripts/ci_plan.py | 3 - 10 files changed, 93 insertions(+), 158 deletions(-) delete mode 100644 deploy/kubernetes/prod/ingress.yaml.tpl diff --git a/.github/actions/setup-kubectl/action.yaml b/.github/actions/setup-kubectl/action.yaml index 184eb055b..e45c638e3 100644 --- a/.github/actions/setup-kubectl/action.yaml +++ b/.github/actions/setup-kubectl/action.yaml @@ -5,6 +5,9 @@ inputs: kubeconfig: description: Base64-encoded kubeconfig required: true + expected-cluster-uid: + description: UID of the target cluster kube-system namespace + required: true server: description: Override the cluster's API server URL; empty keeps the kubeconfig's own required: false @@ -16,7 +19,7 @@ inputs: version: description: kubectl version to install when the runner has none required: false - default: v1.31.4 + default: v1.34.1 runs: using: composite @@ -24,6 +27,7 @@ runs: - shell: bash env: KUBECONFIG_CONTENT: ${{ inputs.kubeconfig }} + EXPECTED_CLUSTER_UID: ${{ inputs.expected-cluster-uid }} SERVER: ${{ inputs.server }} INSECURE: ${{ inputs.insecure-skip-tls-verify }} VERSION: ${{ inputs.version }} @@ -73,4 +77,9 @@ runs: kubectl config unset "clusters.${escaped}.certificate-authority-data" >/dev/null kubectl config unset "clusters.${escaped}.certificate-authority" >/dev/null fi + actual_uid="$(kubectl get namespace kube-system -o jsonpath='{.metadata.uid}')" + if [ -z "$EXPECTED_CLUSTER_UID" ] || [ "$actual_uid" != "$EXPECTED_CLUSTER_UID" ]; then + echo "::error::The kubeconfig does not identify the expected production cluster." + exit 1 + fi kubectl version --output=yaml diff --git a/.github/workflows/deploy.yml b/.github/workflows/deploy.yml index 8f7d6f554..b8ccb65d6 100644 --- a/.github/workflows/deploy.yml +++ b/.github/workflows/deploy.yml @@ -11,13 +11,9 @@ on: workflow_dispatch: inputs: ref: - description: Release tag or full commit SHA to deploy + description: Release tag or full commit SHA already merged into main required: true type: string - apply_ingress: - description: Also apply the Ingress and verify the public route split - default: false - type: boolean permissions: contents: read @@ -39,9 +35,19 @@ jobs: persist-credentials: false - name: Pin the source commit id: source + env: + WORKFLOW_REF: ${{ github.ref }} run: | set -euo pipefail + if [ "$WORKFLOW_REF" != "refs/heads/main" ]; then + echo "::error::Run this production workflow from main." + exit 1 + fi revision="$(git rev-parse HEAD)" + if ! git merge-base --is-ancestor "$revision" origin/main; then + echo "::error::The deployment revision must already be merged into main." + exit 1 + fi echo "revision=$revision" >> "$GITHUB_OUTPUT" echo "Deploying \`$revision\`" >> "$GITHUB_STEP_SUMMARY" @@ -173,11 +179,13 @@ jobs: - uses: ./.github/actions/setup-kubectl with: kubeconfig: ${{ secrets.OAC_KUBECONFIG }} + expected-cluster-uid: ${{ vars.OAC_CLUSTER_UID }} server: ${{ vars.OAC_K8S_SERVER }} insecure-skip-tls-verify: ${{ vars.OAC_K8S_INSECURE || 'false' }} - name: Check the deployment settings env: + OAC_STORAGE_CLASS: ${{ vars.OAC_STORAGE_CLASS }} OAC_PUBLIC_URL: ${{ vars.OAC_PUBLIC_URL }} OAC_DATABASE_URL: ${{ vars.OAC_DATABASE_URL }} OAC_INSTALLATION_ID: ${{ vars.OAC_INSTALLATION_ID }} @@ -191,7 +199,7 @@ jobs: exit 1 fi missing=0 - for name in OAC_PUBLIC_URL OAC_DATABASE_URL OAC_INSTALLATION_ID \ + for name in OAC_STORAGE_CLASS OAC_PUBLIC_URL OAC_DATABASE_URL OAC_INSTALLATION_ID \ OAC_DATABASE_PASSWORD OAC_CREDENTIAL_KEY OAC_CORE_KEY; do if [ -z "${!name:-}" ]; then echo "::error::$name is not configured for the production environment." @@ -199,16 +207,32 @@ jobs: fi done [ "$missing" -eq 0 ] - # Core derives the daemon WebSocket address, the sandbox core_url and the - # self-hosted remote_url from the public URL, and Web serves only its own - # origin. Both reject a path, a query or a fragment. - case "$OAC_PUBLIC_URL" in - https://*/ | https://*/* | *\?* | *\#*) - echo "::error::OAC_PUBLIC_URL must be an HTTPS origin without a path, such as https://core.example." - exit 1 ;; - https://*) ;; - *) echo "::error::OAC_PUBLIC_URL must use HTTPS."; exit 1 ;; - esac + python3 - "$OAC_PUBLIC_URL" <<'PYTHON' + import ipaddress, re, sys, urllib.parse + value = sys.argv[1] + try: + parsed = urllib.parse.urlsplit(value) + host, port = parsed.hostname, parsed.port + valid = (parsed.scheme == "https" and bool(host) + and parsed.username is None and parsed.password is None + and not parsed.path and not parsed.query and not parsed.fragment + and parsed.netloc == parsed.netloc.lower() + and not any(c in value for c in "\\% \t\r\n?#") + and not parsed.netloc.endswith(":")) + if host: + try: + ipaddress.ip_address(host) + except ValueError: + valid = valid and len(host) <= 253 and all( + re.fullmatch(r"[a-z0-9](?:[a-z0-9-]{0,61}[a-z0-9])?", label) + for label in host.split(".")) + if port is not None: + valid = valid and 1 <= port <= 65535 and parsed.netloc.endswith(":" + str(port)) + except ValueError: + valid = False + if not valid: + sys.exit("OAC_PUBLIC_URL must be a canonical HTTPS origin without credentials, path, query or fragment") + PYTHON # The password belongs only in OAC_DATABASE_PASSWORD; Core refuses a URL # that also carries one, and the URL must name the database user. python3 - "$OAC_DATABASE_URL" <<'PY' @@ -239,6 +263,28 @@ jobs: exit 1 fi + - name: Check the existing state volume + env: + OAC_STORAGE_CLASS: ${{ vars.OAC_STORAGE_CLASS }} + OAC_STATE_SIZE: ${{ vars.OAC_STATE_SIZE || '10Gi' }} + run: | + set -euo pipefail + claim="$(mktemp)" + trap 'rm -f "$claim"' EXIT + kubectl get pvc oac-core-state -n "$NS" --ignore-not-found -o json > "$claim" + if [ -s "$claim" ]; then + python3 - "$claim" "$OAC_STORAGE_CLASS" "$OAC_STATE_SIZE" <<'PY' + import json, sys + with open(sys.argv[1]) as source: + claim = json.load(source) + if (claim['spec'].get('storageClassName') != sys.argv[2] + or claim['spec']['resources']['requests']['storage'] != sys.argv[3] + or claim['spec'].get('accessModes') != ['ReadWriteOnce'] + or claim.get('status', {}).get('phase') != 'Bound'): + sys.exit('The existing oac-core-state claim must be Bound with the configured class, size and ReadWriteOnce access') + PY + fi + - name: Ensure the namespace run: | set -euo pipefail @@ -322,18 +368,11 @@ jobs: OAC_CORE_MEMORY_LIMIT: ${{ vars.OAC_CORE_MEMORY_LIMIT || '4Gi' }} run: | set -euo pipefail - # Only an E2B deployment writes to the state root. Without a storage - # class the volume lives as long as the Pod, which is correct until - # then; with one, Core keeps its E2B receipts across rollouts. - if [ -n "$OAC_STORAGE_CLASS" ]; then - # shellcheck disable=SC2016 # envsubst takes the variable list literally - envsubst '$OAC_STORAGE_CLASS $OAC_STATE_SIZE' \ - < deploy/kubernetes/prod/core-state.yaml.tpl | kubectl apply -n "$NS" -f - - OAC_STATE_VOLUME='persistentVolumeClaim: {claimName: oac-core-state}' - else - OAC_STATE_VOLUME='emptyDir: {}' - echo "::notice::No OAC_STORAGE_CLASS: the Core state root is a Pod-lifetime directory. Set one before an E2B deployment exists, or every rollout loses its receipts." - fi + # E2B receipts must survive every rollout. + # shellcheck disable=SC2016 # envsubst takes the variable list literally + envsubst '$OAC_STORAGE_CLASS $OAC_STATE_SIZE' \ + < deploy/kubernetes/prod/core-state.yaml.tpl | kubectl apply -n "$NS" -f - + OAC_STATE_VOLUME='persistentVolumeClaim: {claimName: oac-core-state}' export OAC_STATE_VOLUME # shellcheck disable=SC2016 # envsubst takes the variable list literally envsubst '$OAC_CORE_IMAGE $OAC_CONFIGURATION_DIGEST $OAC_STATE_VOLUME $OAC_CORE_CPU_REQUEST $OAC_CORE_CPU_LIMIT $OAC_CORE_MEMORY_REQUEST $OAC_CORE_MEMORY_LIMIT' \ @@ -373,52 +412,6 @@ jobs: echo "$service: $addresses" done - - name: Apply the Ingress - if: inputs.apply_ingress - env: - OAC_PUBLIC_URL: ${{ vars.OAC_PUBLIC_URL }} - OAC_INGRESS_CLASS: ${{ vars.OAC_INGRESS_CLASS || 'nginx' }} - OAC_TLS_SECRET: ${{ vars.OAC_TLS_SECRET || 'oac-tls' }} - run: | - set -euo pipefail - # A missing class or certificate otherwise surfaces only as the route - # check timing out, which blames the path split instead. - if ! kubectl get ingressclass "$OAC_INGRESS_CLASS" >/dev/null 2>&1; then - echo "::error::No IngressClass named $OAC_INGRESS_CLASS; set OAC_INGRESS_CLASS to one of: $(kubectl get ingressclass -o jsonpath='{.items[*].metadata.name}')" - exit 1 - fi - if ! kubectl get secret "$OAC_TLS_SECRET" -n "$NS" >/dev/null 2>&1; then - echo "::error::No TLS Secret named $OAC_TLS_SECRET in namespace $NS; create it from the public host's certificate, or set OAC_TLS_SECRET." - exit 1 - fi - OAC_PUBLIC_HOST="${OAC_PUBLIC_URL#https://}" - export OAC_PUBLIC_HOST - # shellcheck disable=SC2016 # envsubst takes the variable list literally - envsubst '$OAC_PUBLIC_HOST $OAC_INGRESS_CLASS $OAC_TLS_SECRET' \ - < deploy/kubernetes/prod/ingress.yaml.tpl | kubectl apply -n "$NS" -f - - - - name: Verify the public route split - if: inputs.apply_ingress - env: - OAC_PUBLIC_URL: ${{ vars.OAC_PUBLIC_URL }} - run: | - set -euo pipefail - # 401 proves /v1 reached Core, which asks for a Project API key. 404 means - # it reached Web instead: every application call and node connection would - # then fail. A new Ingress needs a moment to program its load balancer. - for attempt in $(seq 1 30); do - code="$(curl --show-error --silent -o /dev/null -w '%{http_code}' \ - -H 'OpenAI-Beta: agents=v1' "$OAC_PUBLIC_URL/v1/agents" || echo 000)" - if [ "$code" = "401" ]; then - echo "GET /v1/agents -> 401: /v1 reaches Core" - exit 0 - fi - echo "attempt $attempt: GET /v1/agents -> $code" - sleep 10 - done - echo "::error::/v1 does not reach Core; check the Ingress path split." - exit 1 - - name: Report the live state on failure if: failure() run: | diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 7e4718bf0..929787bb8 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -34,7 +34,6 @@ This guide owns how to work in the repository: documentation ownership, the repo | Website: landing page, bilingual documentation maintenance, documentation site build and GitHub Pages publication | [Website guide](website/README.md) | | Self-hosted Runtime installation, recovery and local operation | [Self-hosted execution](docs/getting-started/self-hosted.md) | | Installer lifecycle, locking, generated state, managed HTTPS and downloads | [Installer design rules](deploy/install/README.md) | -| Kubernetes control-plane deployment, its manifests and its workflow | [Kubernetes deployment](deploy/kubernetes/README.md) | | Operator installation and alternatives | [Installation](docs/getting-started/install.md), [installation options](docs/getting-started/install-options.md) | | Settings, defaults, files and installation layout | [Configuration](docs/configuration.md) | | Operator commands, keys, backup and version policy | [Operations](docs/getting-started/operations.md) | diff --git a/deploy/kubernetes/README.md b/deploy/kubernetes/README.md index f9ccdb25d..92ca0de76 100644 --- a/deploy/kubernetes/README.md +++ b/deploy/kubernetes/README.md @@ -1,5 +1,7 @@ # Kubernetes deployment +This fork keeps its Kubernetes deployment in `.github/workflows/deploy.yml`, `.github/actions/setup-kubectl/` and this directory. Upstream application code, documentation and CI selection remain unchanged. + This directory deploys the control plane — Core and the Web console — to a Kubernetes cluster from a source commit, and the `core-deploy` workflow drives it. It is an alternative to the [installer](../../docs/getting-started/install.md), for an operator who already runs PostgreSQL and a cluster. The installer owns the single-host installation and its `config.json`; nothing here reads or writes those files. Core and Web take their settings from the environment, as [Configuration](../../docs/configuration.md#appendix-core-environment-without-the-installer) defines it, and this deployment is the one place that sets those variables for the cluster. ## What it deploys @@ -8,9 +10,8 @@ This directory deploys the control plane — Core and the Web console — to a K | --- | --- | --- | | Deployment and Service `oac-core` | `prod/core.yaml.tpl` | Core on port 8091. One replica: Core takes a PostgreSQL lease that gives one execution service per database, so a second replica exits at startup. An init container prepares the adapter state volume, then `oac-core-migrate` applies the schema | | ConfigMap `oac-core-env` | `prod/core-env.yaml.tpl` | Core's process environment; no secret | -| PersistentVolumeClaim `oac-core-state` | `prod/core-state.yaml.tpl` | `OAC_PROVIDER_STATE_ROOT`, applied only when `OAC_STORAGE_CLASS` is set; see [the state volume](#the-state-volume). Back it up with the database and the credential key | +| PersistentVolumeClaim `oac-core-state` | `prod/core-state.yaml.tpl` | `OAC_PROVIDER_STATE_ROOT`, required for E2B; see [the state volume](#the-state-volume). Back it up with the database and the credential key | | Deployment and Service `oac-web` | `prod/web.yaml.tpl` | The console on port 8080. One replica: Web holds sign-in sessions in process memory, so a cookie is valid only on the Pod that issued it | -| Ingress `oac` | `prod/ingress.yaml.tpl` | Applied only when the run selects **Also apply the Ingress**. One origin, split by path | | Secrets `oac-core-secrets`, `oac-web-core-key`, `oac-registry` | GitHub Environment secrets | Created by the workflow from stdin; never written to a file in the repository | PostgreSQL is yours. So is anything installed on a machine: nodes, self-hosted machines and the Runtime images come from a [release](../../docs/getting-started/nodes.md), not from this workflow. @@ -20,14 +21,14 @@ Both images always come from one commit, so the console never talks to a Core of ## Cluster prerequisites - A PostgreSQL database and an account of its own, reachable from the cluster. Core needs no extension and migrates the schema itself. -- For an E2B deployment only: a StorageClass that provisions a ReadWriteOnce volume. See [the state volume](#the-state-volume). +- A StorageClass that provisions a ReadWriteOnce volume. See [the state volume](#the-state-volume). - An image registry the cluster can pull from, with a user the workflow can push as. -- For the Ingress: a controller and a TLS Secret for the public host. The manifest carries [ingress-nginx](https://kubernetes.github.io/ingress-nginx/) annotations; another controller needs its own equivalents for the four requirements in [HTTPS and the reverse proxy](../../docs/getting-started/install-options.md#https-and-the-reverse-proxy). Leave the Ingress out of the run and keep your own routing if you prefer. +- Configure DNS, certificates and public ingress separately. Core and Web can start before the domain is configured; E2B execution needs the public address to reach Core. - The runner needs `docker`, `envsubst` and outbound access to the registry and the cluster's API server. Set `OAC_DEPLOY_RUNNER` to a self-hosted label when the API server is private. ## Deploy -Run **Actions** → **core-deploy** → **Run workflow**, give it a release tag or a full commit SHA, and select **Also apply the Ingress** the first time or after the routing changes. The run builds both images, pushes them, applies the secrets and the environment, rolls Core out and then Web, and checks that both Services have a ready endpoint. With the Ingress it also checks that `/v1` answers `401` — proof that the path split reaches Core and not the console. +Merge the deployment workflow into `main`. Run **Actions** → **core-deploy** → **Run workflow** from `main` and give it a release tag or a full commit SHA already merged into `main`. The run builds both images, pushes them, applies the secrets and the environment, rolls Core out and then Web, and checks that both Services have a ready endpoint. It does not create an Ingress or configure DNS or certificates. Route `/v1` and `/api/v1` to `oac-core:8091`, and all other paths to `oac-web:8080`, in the `openagentcore` namespace. After configuring HTTPS, check `/healthz`, verify that unauthenticated `/v1/agents` returns `401`, and qualify a fresh E2B Session and Turn. **A deployment is an outage.** Core's single replica stops before its replacement starts, and the replacement migrates the schema before it listens, so the Agents API, the machine routes and every running Session are unavailable for the rollout. Deploy in a window you can afford to lose. The Pod that takes over also needs the previous one's leased database connection to be gone; until it is, `AcquireLease` fails and Core exits, and the rollout depends on a restart landing inside the 15-minute deadline. @@ -35,7 +36,7 @@ To roll back, run the workflow again with the earlier commit. A rollback across ## Settings -Set these on the repository's `production` GitHub Environment. Variables are not secret; secrets are. The workflow stops and names anything missing or malformed before it touches the cluster. +Restrict the repository’s `production` GitHub Environment to deployments from the `main` branch. Set these on that environment. Variables are not secret; secrets are. The workflow stops and names anything missing or malformed before it touches the cluster. ### Secrets @@ -57,11 +58,11 @@ Set these on the repository's `production` GitHub Environment. Variables are not | `OAC_DATABASE_URL` | `postgres://USER@HOST:PORT/DATABASE`, with `sslmode` and any `pool_*` parameter in the query and no password | | `OAC_INSTALLATION_ID` | A canonical lowercase UUID, generated once with `uuidgen \| tr 'A-Z' 'a-z'`. Core refuses an ID other than the one its database recorded, so it belongs to the database: keep the two together and restore them together | | `OAC_NAMESPACE` | Namespace; `openagentcore` by default. The workflow creates it if it is missing | -| `OAC_STORAGE_CLASS`, `OAC_STATE_SIZE` | StorageClass and size of the Core state volume. Unset means no volume; see [the state volume](#the-state-volume). `10Gi` by default | +| `OAC_STORAGE_CLASS`, `OAC_STATE_SIZE` | StorageClass and size of the required Core state volume. See [the state volume](#the-state-volume). `10Gi` by default | +| `OAC_CLUSTER_UID` | UID of the target cluster’s `kube-system` namespace. The workflow verifies it before applying resources | | `OAC_K8S_SERVER` | API server URL that overrides the kubeconfig's. Optional | | `OAC_K8S_INSECURE` | `true` skips API server certificate verification, for an address the certificate does not name. Leave it unset otherwise | | `OAC_DEPLOY_RUNNER` | Runner label; `ubuntu-22.04` by default. Set this one at repository level: the job that pins the commit runs before the environment is entered, so an environment-scoped value does not reach it | -| `OAC_INGRESS_CLASS`, `OAC_TLS_SECRET` | `ingressClassName` and the TLS Secret of the public host; `nginx` and `oac-tls` by default | | `OAC_CORE_CPU_REQUEST`, `OAC_CORE_CPU_LIMIT`, `OAC_CORE_MEMORY_REQUEST`, `OAC_CORE_MEMORY_LIMIT` | Core's resources; `500m`, `2`, `1Gi` and `4Gi` by default | | `OAC_HARNESSES`, `OAC_DEFAULT_HARNESS`, `OAC_EXECUTION_CONCURRENCY`, `OAC_WRITE_AUDIT_RETENTION`, `OAC_LOG_LEVEL` | The matching [process settings](../../docs/configuration.md#settings), which define their values and defaults | @@ -71,19 +72,16 @@ A change to any of these takes effect on the next run: a digest of the environme `OAC_PROVIDER_STATE_ROOT` is the private root each Sandbox Provider adapter keeps its own state under, and today exactly one adapter uses it: E2B stores its receipts in `/state/e2b`. The helper writes a receipt before each remote `Create`, because a helper's exit never proves that the remote call settled, and Core needs the receipt afterwards to clean up, observe and verify ownership of a sandbox living in E2B's cloud. Each receipt is at most 64 KiB, they are never pruned, and losing them orphans sandboxes that E2B keeps billing and Core can no longer destroy. [The E2B helper](../../services/core/tools/e2b-provider/README.md#receipts-and-state-directory) owns these rules. -So the volume is needed exactly when an E2B deployment exists, and only then: - -- **No `OAC_STORAGE_CLASS`:** no claim is created and `/state` lives as long as the Pod. Correct for an installation with no E2B deployment, which writes nothing there. -- **`OAC_STORAGE_CLASS` set:** the claim is created and the receipts survive rollouts. Set it before selecting E2B, not after — every rollout in between discards the receipts written since the last one. +This deployment requires the persistent claim. Set `OAC_STORAGE_CLASS` before deploying; an unset value stops the workflow before it applies resources. Every Core rollout mounts `oac-core-state`, preserving the E2B receipts. Size it for the claim's lifetime rather than for today, because a StorageClass without `allowVolumeExpansion` cannot grow and a `Delete` reclaim policy destroys the receipts with the claim. The default `10Gi` also clears the minimum that a block-storage provisioner may impose. -A claim's storage class and size are immutable. Changing either variable once the claim exists makes the run fail on the claim; migrate the receipts to a new claim yourself and delete the old one. +Keep the existing claim's class and size. This cluster's `cbs` StorageClass does not allow expansion; changing the class or enlarging this claim fails. To change storage, migrate the receipts to a new claim. ## Differences from an installed installation - **E2B is the only backend that executes anything.** Web serves the node installer from the release bundle's node payload, which these images do not carry, so `/node-install/*` is unavailable and Add node is hidden. The Core image's native installer catalog is empty too — only the release build fills it — so `POST /api/v1/agent-daemon/installation` answers `503 installation_unavailable` and a self-hosted Session's install command cannot resolve. Until the node payload and the native catalog are built into these images, select [E2B](../../docs/getting-started/install-options.md), which needs neither. - **No startup settings panel.** `OAC_SETTINGS_FILE` is the installer's snapshot, so `GET /core/v1/installation` reports none and the console shows no process settings. Runtime settings in Web are unaffected; they live in Core's database. -- **No domain setup in Web.** There is no installer socket, so the public origin comes from `OAC_PUBLIC_URL` and the Ingress, not from the console. +- **No domain setup in Web.** There is no installer socket, so the public origin comes from `OAC_PUBLIC_URL` and your external ingress, not from the console. - **No `oac` command.** The maintenance commands ship in the Core image: `kubectl exec deploy/oac-core -- oac-core-device …`. [Operations](../../docs/getting-started/operations.md) covers keys and backup. - **A rollout signs operators out** of the console, because Web holds its sessions in memory. diff --git a/deploy/kubernetes/prod/core-state.yaml.tpl b/deploy/kubernetes/prod/core-state.yaml.tpl index f60fa1a94..b801bdb15 100644 --- a/deploy/kubernetes/prod/core-state.yaml.tpl +++ b/deploy/kubernetes/prod/core-state.yaml.tpl @@ -1,4 +1,4 @@ -# Core's adapter state volume, applied only when OAC_STORAGE_CLASS is set. +# Core's persistent E2B adapter state volume. # # Only the E2B adapter uses this volume today: it holds the receipts Core writes # before each remote Create and needs afterwards to clean up, observe and verify diff --git a/deploy/kubernetes/prod/core.yaml.tpl b/deploy/kubernetes/prod/core.yaml.tpl index 319afc6c2..5b766fce8 100644 --- a/deploy/kubernetes/prod/core.yaml.tpl +++ b/deploy/kubernetes/prod/core.yaml.tpl @@ -46,11 +46,8 @@ spec: - name: secrets emptyDir: medium: Memory - # The adapter state root. Only the E2B adapter uses it today, for the - # receipts that let Core clean up its remote sandboxes, and those must - # outlive the Pod. Without a claim it is a Pod-lifetime directory, which - # is correct until an E2B deployment exists. core-state.yaml.tpl owns the - # claim; deploy/kubernetes/README.md explains when to turn it on. + # E2B receipts must outlive this Pod. The deploy workflow always mounts + # the claim owned by core-state.yaml.tpl. - {name: state, ${OAC_STATE_VOLUME}} - name: tmp emptyDir: {} diff --git a/deploy/kubernetes/prod/ingress.yaml.tpl b/deploy/kubernetes/prod/ingress.yaml.tpl deleted file mode 100644 index 87789ef58..000000000 --- a/deploy/kubernetes/prod/ingress.yaml.tpl +++ /dev/null @@ -1,54 +0,0 @@ -# One public origin for Core and Web, split by path. The annotations below are -# ingress-nginx's; another controller needs its own equivalents for the four -# requirements in docs/getting-started/install-options.md#https-and-the-reverse-proxy: -# preserve Host, pass WebSocket upgrades on /api/v1, never buffer or time out a -# stream, and accept large uploads. -apiVersion: networking.k8s.io/v1 -kind: Ingress -metadata: - name: oac - labels: - app.kubernetes.io/part-of: openagentcore - annotations: - # /v1 streams Session events and /api/v1 carries long-lived WebSockets. - nginx.ingress.kubernetes.io/proxy-buffering: "off" - nginx.ingress.kubernetes.io/proxy-request-buffering: "off" - nginx.ingress.kubernetes.io/proxy-read-timeout: "3600" - nginx.ingress.kubernetes.io/proxy-send-timeout: "3600" - # Core enforces its own upload limits; source files may reach 512 MiB. - nginx.ingress.kubernetes.io/proxy-body-size: "0" - nginx.ingress.kubernetes.io/ssl-redirect: "true" -spec: - ingressClassName: "${OAC_INGRESS_CLASS}" - tls: - - hosts: - - "${OAC_PUBLIC_HOST}" - secretName: "${OAC_TLS_SECRET}" - rules: - - host: "${OAC_PUBLIC_HOST}" - http: - paths: - # Applications, with a Project API key. - - path: /v1 - pathType: Prefix - backend: - service: - name: oac-core - port: - number: 8091 - # Nodes, sandboxes and self-hosted machines. Uses WebSockets. - - path: /api/v1 - pathType: Prefix - backend: - service: - name: oac-core - port: - number: 8091 - # Browsers. - - path: / - pathType: Prefix - backend: - service: - name: oac-web - port: - number: 8080 diff --git a/docs/getting-started/install-options.md b/docs/getting-started/install-options.md index 138c3f946..2e5063800 100644 --- a/docs/getting-started/install-options.md +++ b/docs/getting-started/install-options.md @@ -4,8 +4,6 @@ title: "Installation options and advanced deployments" The [default installation](./install.md) needs no options. Use this page to run behind an existing reverse proxy or install without internet access. -To run Core and the console on Kubernetes with your own PostgreSQL instead of the installer, see the [Kubernetes deployment](../../deploy/kubernetes/README.md). - Pass options to the downloaded script: ```sh diff --git a/docs/zh/getting-started/install-options.md b/docs/zh/getting-started/install-options.md index ca4fdc3ba..03b553f3e 100644 --- a/docs/zh/getting-started/install-options.md +++ b/docs/zh/getting-started/install-options.md @@ -1,13 +1,11 @@ --- title: "安装选项与高级部署" source: docs/getting-started/install-options.md -source_hash: 4e726974eaa85c8ec64c242ed4947faee0c08bb5fdcfa527e43c6284af96c851 +source_hash: 7178baf93a8d8b6e7f086d73033afe4ea14afcf41c88dbcc6031a3a6dd20ca17 --- [默认安装](install.md)无需任何选项。使用本页可以在现有反向代理后运行,或者在无法访问互联网时进行安装。 -若要用自己的 PostgreSQL 在 Kubernetes 上运行 Core 和控制台,而不使用安装器,参见 [Kubernetes 部署](../../../deploy/kubernetes/README.md)。 - 向下载的脚本传递选项: ```sh diff --git a/scripts/ci_plan.py b/scripts/ci_plan.py index f512d172f..2b3ef9d67 100644 --- a/scripts/ci_plan.py +++ b/scripts/ci_plan.py @@ -19,9 +19,6 @@ ".github/workflows/actionlint.yml": ("lint",), ".github/actionlint.yaml": ("lint",), ".github/workflows/ci-review.yml": ("lint",), - # Deployment inputs: no check builds or tests them, so only workflow syntax applies. - ".github/workflows/deploy.yml": ("lint",), - ".github/actions/setup-kubectl/action.yaml": ("lint",), ".github/workflows/website.yml": ("website", "lint"), ".github/actions/node/action.yml": (*NODE_JOBS, "lint"), "scripts/ci_plan.py": JOBS, From a85823c548fc80312eee13063838f2933d4c857b Mon Sep 17 00:00:00 2001 From: sam Date: Sat, 3 Oct 2026 11:48:13 +0800 Subject: [PATCH 3/3] Require the existing E2B state claim before deployment --- .github/workflows/deploy.yml | 11 ++--------- deploy/kubernetes/README.md | 6 +++--- 2 files changed, 5 insertions(+), 12 deletions(-) diff --git a/.github/workflows/deploy.yml b/.github/workflows/deploy.yml index b8ccb65d6..1dc969f7a 100644 --- a/.github/workflows/deploy.yml +++ b/.github/workflows/deploy.yml @@ -271,9 +271,8 @@ jobs: set -euo pipefail claim="$(mktemp)" trap 'rm -f "$claim"' EXIT - kubectl get pvc oac-core-state -n "$NS" --ignore-not-found -o json > "$claim" - if [ -s "$claim" ]; then - python3 - "$claim" "$OAC_STORAGE_CLASS" "$OAC_STATE_SIZE" <<'PY' + kubectl get pvc oac-core-state -n "$NS" -o json > "$claim" + python3 - "$claim" "$OAC_STORAGE_CLASS" "$OAC_STATE_SIZE" <<'PY' import json, sys with open(sys.argv[1]) as source: claim = json.load(source) @@ -283,12 +282,6 @@ jobs: or claim.get('status', {}).get('phase') != 'Bound'): sys.exit('The existing oac-core-state claim must be Bound with the configured class, size and ReadWriteOnce access') PY - fi - - - name: Ensure the namespace - run: | - set -euo pipefail - kubectl get namespace "$NS" >/dev/null 2>&1 || kubectl create namespace "$NS" - name: Apply the registry pull secret env: diff --git a/deploy/kubernetes/README.md b/deploy/kubernetes/README.md index 92ca0de76..e75a929e7 100644 --- a/deploy/kubernetes/README.md +++ b/deploy/kubernetes/README.md @@ -21,7 +21,7 @@ Both images always come from one commit, so the console never talks to a Core of ## Cluster prerequisites - A PostgreSQL database and an account of its own, reachable from the cluster. Core needs no extension and migrates the schema itself. -- A StorageClass that provisions a ReadWriteOnce volume. See [the state volume](#the-state-volume). +- An existing namespace and a Bound ReadWriteOnce `oac-core-state` claim. Prepare it once using `prod/core-state.yaml.tpl` with `OAC_STORAGE_CLASS` and `OAC_STATE_SIZE`; see [the state volume](#the-state-volume). - An image registry the cluster can pull from, with a user the workflow can push as. - Configure DNS, certificates and public ingress separately. Core and Web can start before the domain is configured; E2B execution needs the public address to reach Core. - The runner needs `docker`, `envsubst` and outbound access to the registry and the cluster's API server. Set `OAC_DEPLOY_RUNNER` to a self-hosted label when the API server is private. @@ -57,7 +57,7 @@ Restrict the repository’s `production` GitHub Environment to deployments from | `OAC_PUBLIC_URL` | The canonical HTTPS origin, such as `https://core.example`, with no path, query or fragment. Core derives the daemon WebSocket address, the sandbox `core_url` and the self-hosted `remote_url` from it, and Web serves only this origin. Without it Core executes no Session | | `OAC_DATABASE_URL` | `postgres://USER@HOST:PORT/DATABASE`, with `sslmode` and any `pool_*` parameter in the query and no password | | `OAC_INSTALLATION_ID` | A canonical lowercase UUID, generated once with `uuidgen \| tr 'A-Z' 'a-z'`. Core refuses an ID other than the one its database recorded, so it belongs to the database: keep the two together and restore them together | -| `OAC_NAMESPACE` | Namespace; `openagentcore` by default. The workflow creates it if it is missing | +| `OAC_NAMESPACE` | Existing namespace; `openagentcore` by default | | `OAC_STORAGE_CLASS`, `OAC_STATE_SIZE` | StorageClass and size of the required Core state volume. See [the state volume](#the-state-volume). `10Gi` by default | | `OAC_CLUSTER_UID` | UID of the target cluster’s `kube-system` namespace. The workflow verifies it before applying resources | | `OAC_K8S_SERVER` | API server URL that overrides the kubeconfig's. Optional | @@ -72,7 +72,7 @@ A change to any of these takes effect on the next run: a digest of the environme `OAC_PROVIDER_STATE_ROOT` is the private root each Sandbox Provider adapter keeps its own state under, and today exactly one adapter uses it: E2B stores its receipts in `/state/e2b`. The helper writes a receipt before each remote `Create`, because a helper's exit never proves that the remote call settled, and Core needs the receipt afterwards to clean up, observe and verify ownership of a sandbox living in E2B's cloud. Each receipt is at most 64 KiB, they are never pruned, and losing them orphans sandboxes that E2B keeps billing and Core can no longer destroy. [The E2B helper](../../services/core/tools/e2b-provider/README.md#receipts-and-state-directory) owns these rules. -This deployment requires the persistent claim. Set `OAC_STORAGE_CLASS` before deploying; an unset value stops the workflow before it applies resources. Every Core rollout mounts `oac-core-state`, preserving the E2B receipts. +This deployment requires the persistent claim. Set `OAC_STORAGE_CLASS` before deploying; an unset value or missing claim stops the workflow before it applies resources. Every Core rollout mounts `oac-core-state`, preserving the E2B receipts. Size it for the claim's lifetime rather than for today, because a StorageClass without `allowVolumeExpansion` cannot grow and a `Delete` reclaim policy destroys the receipts with the claim. The default `10Gi` also clears the minimum that a block-storage provisioner may impose.