Deploy the control plane to Kubernetes from one source commit - #1
Conversation
Core and the Web console get a Kubernetes deployment beside the installer, for an operator who already runs PostgreSQL and a cluster. The installer keeps owning the single-host installation and its config.json; the manifests set only the documented Core and Web environment, so no setting gains a second home. One run builds both images from one commit and rolls them out together, so the console never talks to a Core of another release. Core runs as a single replica with the Recreate strategy because it takes a PostgreSQL lease that admits one execution service per database, and a surging Pod cannot take over from a running one. Web runs the same way because it holds console sign-in sessions in process memory, where a second replica would reject a cookie the other Pod issued. An init container prepares each secret as an owner-only file for the service account, because Kubernetes owns Secret volume files as root and Web refuses a Core key file that grants group or other access. A second one applies the embedded migrations before Core opens the database. The adapter state volume is opt-in. Only the E2B adapter writes there, for the receipts that let Core clean up, observe and verify ownership of sandboxes in E2B's cloud, so an installation with no E2B deployment needs no claim at all. The workflow validates every setting and secret before it touches the cluster, verifies the rolled-out image and ready endpoints, and, when it applies the Ingress, that /v1 reaches Core rather than the console.
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 57828deb21
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| - uses: ./.github/actions/setup-kubectl | ||
| with: | ||
| kubeconfig: ${{ secrets.OAC_KUBECONFIG }} |
There was a problem hiding this comment.
Keep production credentials out of the selected revision
When ref identifies an unreviewed commit or tag, the preceding checkout replaces this local action with code from that revision and then executes it with the production kubeconfig. A collaborator who can dispatch the workflow can therefore bypass branch protection by pushing a commit that changes setup-kubectl and selecting its SHA, gaining the cluster credential; the build job similarly runs selected-revision scripts after registry login. Validate that the revision belongs to a trusted protected ref before entering the production environment, or execute credential-bearing deployment logic from a separately trusted checkout.
Useful? React with 👍 / 👎.
|
|
||
| ## Deploy | ||
|
|
||
| Run **Actions** → **core-deploy** → **Run workflow**, give it a release tag or a full commit SHA, and select **Also apply the Ingress** the first time or after the routing changes. The run builds both images, pushes them, applies the secrets and the environment, rolls Core out and then Web, and checks that both Services have a ready endpoint. With the Ingress it also checks that `/v1` answers `401` — proof that the path split reaches Core and not the console. |
There was a problem hiding this comment.
Move operator steps into the owned bilingual guide
This is the actual operator procedure, and the remainder of this README also owns all Kubernetes settings, despite CONTRIBUTING.md still assigning operator installation and alternatives to docs/getting-started/. Because deploy/** is excluded from the documentation site, the new English and Chinese installation pages only send readers to an unpublished, English-only component README. Move the operator-facing procedure and settings into the owned English guide and its Chinese mirror, leaving implementation-specific manifest and workflow rules here.
AGENTS.md reference: AGENTS.md:L56-L61
Useful? React with 👍 / 👎.
| case "$OAC_PUBLIC_URL" in | ||
| https://*/ | https://*/* | *\?* | *\#*) | ||
| echo "::error::OAC_PUBLIC_URL must be an HTTPS origin without a path, such as https://core.example." | ||
| exit 1 ;; | ||
| https://*) ;; | ||
| *) echo "::error::OAC_PUBLIC_URL must use HTTPS."; exit 1 ;; |
There was a problem hiding this comment.
Validate the public URL as an origin before applying it
On an existing deployment, a value such as https://user@example.com passes this shell case because it has no path, query, or fragment, but Core's deployment.ValidateCoreURL rejects embedded credentials. The workflow consequently updates configuration and applies the Recreate Deployment, terminating the healthy Core pod before the replacement fails at startup. Use the same canonical-origin checks as Core—including a nonempty hostname, no user info, canonical host/port, and valid port range—during this preflight.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: a85823c548
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| parsed = urllib.parse.urlsplit(sys.argv[1]) | ||
| if parsed.scheme not in ("postgres", "postgresql") or not parsed.username: | ||
| sys.exit("OAC_DATABASE_URL must be postgres://USER@HOST:PORT/DATABASE") | ||
| # Go reports an empty password as present, and Core then refuses the URL. | ||
| if parsed.password is not None: | ||
| sys.exit("Set the database password only in the OAC_DATABASE_PASSWORD secret") |
There was a problem hiding this comment.
Validate every field of the database URL
On an existing deployment, values such as postgres://user@host:notaport/db or postgres://user@/db pass this check because it only examines the scheme, username, and password. The workflow then applies the Recreate Deployment, stops the healthy Core pod, and leaves the migration init container unable to connect; a URL without a database path may instead select PostgreSQL's default database. Validate the hostname, port, database path, and supported query parameters before modifying the cluster.
Useful? React with 👍 / 👎.
| for name in OAC_HARNESSES OAC_DEFAULT_HARNESS OAC_EXECUTION_CONCURRENCY \ | ||
| OAC_WRITE_AUDIT_RETENTION OAC_LOG_LEVEL; do | ||
| value="${!name:-}" | ||
| if [ -n "$value" ]; then printf ' %s: "%s"\n' "$name" "$value" >> "$work/core-env.yaml"; fi | ||
| done |
There was a problem hiding this comment.
Reject invalid process settings before rollout
When an optional setting is malformed—for example, OAC_EXECUTION_CONCURRENCY=abc, OAC_WRITE_AUDIT_RETENTION=5m, or an unknown OAC_DEFAULT_HARNESS—this loop blindly adds it to the ConfigMap. Kubernetes accepts the manifest, the Recreate strategy terminates the healthy Core pod, and the replacement exits during Core's startup validation, leaving production unavailable until the setting is corrected and the workflow reruns. Validate these values before applying any cluster changes.
Useful? React with 👍 / 👎.
| if ! command -v kubectl >/dev/null 2>&1; then | ||
| curl --fail --show-error --silent --location \ | ||
| -o "$RUNNER_TEMP/kubectl" \ | ||
| "https://dl.k8s.io/release/${VERSION}/bin/linux/amd64/kubectl" | ||
| install -m 0755 "$RUNNER_TEMP/kubectl" /usr/local/bin/kubectl |
There was a problem hiding this comment.
Install kubectl in a writable runner directory
When a self-hosted runner does not already have kubectl, this fallback runs as the runner account but writes directly to root-owned /usr/local/bin. A normal non-root runner therefore fails with permission denied before it can connect to the private cluster, despite this action claiming to install the missing binary. Install it under a runner-owned directory such as $RUNNER_TEMP and add that directory to PATH, or explicitly use the runner's privilege mechanism.
Useful? React with 👍 / 👎.
Deploys Core and Web to Kubernetes from one source revision, using external PostgreSQL and a persistent E2B receipt volume. The fork adds only the deployment workflow, kubectl action and Kubernetes files; upstream application code, documentation, contributor rules and CI selection are unchanged.
The workflow runs from main and accepts only revisions already merged into main. The production environment is restricted to main. It verifies the target cluster's kube-system UID, validates settings and checks an existing state claim's class, size, RWO access and Bound state before changing services. Core always mounts oac-core-state; missing storage configuration fails instead of selecting temporary storage.
DNS, HTTPS and ingress are configured separately. Route /v1 and /api/v1 to oac-core:8091, and other paths to oac-web:8080. Both services use one replica and Recreate, so later deployments interrupt running sessions.
Validation: actionlint 1.7.12 with shellcheck, composite-action shellcheck, YAML rendering, public-origin validation cases, git diff --check, and server dry-run against sandbase-prod. Independent static review is required before merge. The existing openagentcore/oac-core-state claim is Bound at 10Gi on cbs.