Repository navigation
Fix what a newcomer install actually hits: source builds, gateway routing, and URLs that name the wrong host - #10
Merged
Conversation
…ort mapping A public deployment on any cluster but the reference trial reported `http://127.0.0.1:1809x` in `sites status`, in the CR's `status.url`, in the database, and behind the console's Open action. Measured on kind: nodePort 30080 was reported as 18090, which nothing was listening on. `HOST_PORT_BASE` defaulted to 18090, the host port the repository's own Lima relay forwards 30080..30088 to. That is a fact about the machine in front of the cluster, not about the cluster, and the Chart hardcoded `SITES_PUBLIC_URL_HOST` with no value at all. The wrong answer is invisible: the deployment is healthy and server-side verification passes because it probes the in-cluster address. The mapping is now declared or absent, for the same reason `clusterNetwork.podCIDR` has no default. Undeclared, `public_url` returns None instead of guessing. New Chart values `nodePort.hostPortBase` (no default) and `nodePort.publicUrlHost`; `cluster.sh` gains matching `SITES_HOST_PORT_BASE` and `SITES_PUBLIC_URL_HOST` seams. `host_port_base()` reads the environment per call for the same reason `backend()` does. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
`make quickstart-doctor` reported "local ports: available" while 18090 and 18091 were held by a root-owned process. Two-sided control on one address, 127.0.0.1:18092: held by the invoking user the doctor refuses correctly; held by root it passes. Only the owner differs. `lsof` lists sockets owned by the user running it, so a port held by anyone else returns non-zero and reads as free. Lima binds these forwards on 127.0.0.1, so this is a guaranteed failure several minutes later -- exactly what the doctor's own comment says it exists to prevent. Replaced with a bind probe: it answers whether a server can take the address, for any owner, and sets SO_REUSEADDR so a lingering TIME_WAIT socket does not read as a conflict. Both preflights use it, `lsof` drops out of the prerequisites, and the fixed-port list is derived from `host_port_base` instead of repeating 18090 a second time in the same file. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…evidence README points at `deploy-specs/` for bundle examples and neither file could be submitted. `example-single.json` was a flat single-deployment shape the bundle endpoint refuses outright (`unsupported bundle fields`); `example-stack.json` carried a `command` field that is not a deployment field. Nothing had ever submitted them, so both drifted out of the contract they illustrate. Submitting them then showed a second gap: bundle components reported `verification: null` while the operator had written the evidence onto every CR. docs/AGENT_CONTRACT.md makes `status.verification` the only success criterion and bundles are one of its deployment entry points, so the projection was hiding the one thing a caller is told to check. Both examples are now real bundles that reach Running with 200s on a live cluster, their component names no longer collide so either can be submitted first, and the bundle projection carries `verification`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…README limitation A gate per fix, each verified by mutation (restore the old behaviour, watch it go red): - no declared host mapping means no public URL; the default render must not set SITES_HOST_PORT_BASE; the trial's declared base must equal its Lima forward - the port preflight must bind rather than list sockets it owns - every file in deploy-specs/ passes the endpoint's own validator, and no two examples claim the same component name - bundle components carry verification Also drops "Historical UIDs remain in the Git history": the repository is public, its history is 22 commits, and scanning all 324 blobs finds no UID or Aliyun identifier -- the only `aliyuncs.com` occurrences are the OSS endpoint check and its tests. The console's `The only one in the merchant` placeholder, a machine translation, becomes `Unique within the merchant`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…w the install namespace `sites build submit` always failed, and the only signal after five minutes was "Job was active longer than specified deadline". The build Pod never started: `MountVolume.SetUp failed for volume "registry-auth": references non-existent secret key: config.json`. builds.py selects `config.json` out of `sites-registry-auth` -- DOCKER_CONFIG points at that mount and Docker looks for that filename -- and nothing in the repository ever created the key. Selecting a missing key is not a soft failure: kubelet refuses to set the volume up, so the entire source-build entry point was dead on every installation that followed this repository's own instructions, and the error named neither the Secret nor the key. Same area, second defect: SITES_REGISTRY_PUSH_HOST and SITES_REGISTRY_API defaulted to a hardcoded `sites-local` namespace and the Chart never set them, so any other namespace pushed builds at, and resolved digests against, a Service that is not there. bootstrap-standalone-secrets.sh now generates the docker config from the same registry password (host defaults to `sites-registry.<namespace>.svc:5000`, overridable), requires the key when reusing an existing Secret so an older installation is told which key is missing, and the Chart renders both registry addresses from `namespaces.control` for the API and the operator. Verified on a fresh install: build, push, deploy, verified 200 in about 30 seconds. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
`--health-path` defaults to `/` for `deploy` and `deploy-static` and `/healthz` for `build submit`, and none of the three said so. Taking the sibling default on faith produced a build that pushed a correct image and then spent the whole readiness window failing a probe on a path the image does not serve, reported as "Deployment does not have minimum availability" -- which points at the workload rather than at the flag. `--port` and `--health-path` now print their real defaults, and the source-build one says why it differs. The defaults themselves are unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…ch group has `sites admin merchants create --help` read "only the summary is stored in the library". Two machine translations in one sentence: a digest is not a summary (the security meaning differs) and "library" is a literal rendering of "database". The tenant equivalent said it differently again. Both now say the control plane stores only a digest, and each names the rotate command its own group actually has. A gate checks that: the first attempt at this message said `rotate-token`, which tenants do not have, and a substring check would have accepted it because `rotate` is its prefix. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…write tool requires Following AGENT_CONTRACT's own client configuration, all six MCP write tools answer `deployment_authorization_required`, and the message says only "do not retry" -- nothing an integrator can implement. The reserved `_agent_deployment_authorization` argument is deliberately absent from every tools/list schema, so the model never learns it exists. That leaves the document as the only place a host can read its shape, and the shape was not there. Adds the field table (version must be 1, runId 1-128, nonce at least 24 chars, expiresAt in the future, allowInternal for internal exposure) and the three rules that make it worth having. The gate checks each documented field against the validator rather than a copy of it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…on it names Observed on a live cluster: a dynamic site's version 2 never became ready, and the same `sites status` response carried `phase: Failed` next to `verification.ok: true, httpStatus: 200` -- version 1's evidence. The retention is deliberate; the operator comment says so, and the evidence names the revision it was collected for. What was missing is the rule: AGENT_CONTRACT made `verification.ok=true` the success criterion without saying to compare `verification.revision` with the deployment's `revision`, so a caller reading `ok` alone reports a rollout that never came up as a live site. Both halves are now gated: the response must carry both revisions (or the comparison is impossible) and the contract must state it (or nobody makes it). Also fixes the `scaleToZero` refusal text, where two translated fragments were concatenated without a space and reached the caller as "there will be noThe link can receive the request". A new sweep checks all 408 refusal messages for that shape; its judge is self-tested in both directions, because the first version filtered out the very word it was looking for and reported zero. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…ng a timeout The fifth site in a tenant is accepted with a 200 and fails 120 seconds later with "Deployment was not ready within 120s (Deployment does not have minimum availability.)". The cause was only in a ReplicaSet event: `exceeded quota: sites-tenant-quota, requested: limits.cpu=1, used: limits.cpu=4`. `_rollout_stall_reason` exists to separate OOMKilled from CrashLoopBackOff from ImagePullBackOff, and reads the Available condition and pod container states. A namespace ResourceQuota rejects creation at the ReplicaSet, so there is no pod to inspect -- the one stall it cannot see is the one whose Available message is content-free. Kubernetes copies the reason onto the Deployment as a ReplicaFailure condition, which the operator already holds. Two more of the same shape: bundle components reported `Failed` with no message at all, and the two tenant limits disagree -- admission counts `maxDeployments` (10) while the namespace quota allows four concurrent sites. The mismatch is now in Known limitations, with a gate that asks for the limitation to be deleted rather than updated if the numbers ever agree. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…ed from A site whose root returns 403 reported `phase: Running`, `ready: true`, `verification.ok: true`, `httpStatus: 200`, and a public URL that answers 403. The probe requests `spec.healthPath`, not the site root, and writes that address into `verification.url`. `deploy` and `deploy-static` default to `/` so the two have always coincided; `build submit` defaults to `/healthz`, and there they are different resources. AGENT_CONTRACT nonetheless tells callers to compare the public URL's digest with `bodySha256`, which can only hold at `/`, and the quickstart's "public URL body: matches verification digest" assertion passes only because it takes the default. The bound is now stated where the criterion is. Gated on all three sides: the URL is built from healthPath, the document says so, and the trial passes no `--health-path` (its own assertion would fail if it did). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…ster.sh already does `scripts/cluster_benchmark.py` could not run on any cluster whose Pod network is not the reference topology's. It installs through `standalone.sh`, which hardcoded `values-dev.yaml` and therefore podCIDR 10.201.0.0/16; the install succeeds and the operator then refuses to start, by design, because its own address falls outside the declaration. `cluster.sh` has had SITES_HELM_VALUES and SITES_CLUSTER_POD_CIDR all along. Adding the same two seams to its sibling makes the benchmark -- the script whose report invites reproduction -- runnable anywhere. Default behaviour is unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…ale count README and CONFIGURATION both said "45 SITES_* variables ... appear in no document or chart". Measured on that sentence's own definition: 116 are read, 86 appear in documents, 40 in the Chart or scripts, and 8 appear in neither. Even counting documents alone gives 30. Same shape as the "historical UIDs" bullet removed earlier: a limitation that overstates a problem which has largely been closed, drifting in the safe direction so nothing surfaces it. With eight left they are simply documented, the count goes to zero, and the README bullet now states what is genuinely still true -- these variables are not schema-validated. A gate requires every variable the code reads to be named in a document or the Chart, and self-tests its own scan so an empty read set cannot pass as success. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…elease exclusive Installing a second release into another namespace fails with Helm's ownership error about `meta.helm.sh/release-namespace` -- a message about Helm annotations that mentions neither Site nor the limit it is enforcing. Both CRDs are cluster-scoped and rendered as ordinary templates, so the first release owns them, while `--namespace` on both install scripts implies the opposite. Documented, with the real error text, plus a gate that also pins the premise: if the CRDs ever move to the Chart's `crds/` directory the exclusivity disappears and the note should be deleted rather than edited. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…xins from the code The per-module table listed 43 of the package's 45 modules. Missing were `api_mcp.py` -- the `POST /mcp` endpoint that README and AGENT_CONTRACT both treat as a headline surface -- and `http_kit.py`. The composition-root sentence counted eight endpoint mixins while `Handler` combines nine. A hand-maintained table drifts silently because adding a module gives nobody a reason to reopen it. Gated three ways: every module has a row, every row names a file that exists, and the mixin count and its spelled-out word are derived from which `api_*.py` files actually define a Mixin. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…template literals The three quantities that compose a gateway URL -- domain suffix, scheme, host port -- were literals in the template, and the comment said the last two are the local Lima forward. Same shape as the NodePort backend's HOST_PORT_BASE on the other backend: routing works, only the address handed to the caller is wrong, and server-side verification passes either way because it uses the in-cluster address. The suffix was worse than hardcoded: it appeared twice, in the ConfigMap and in the listener wildcard, with the file instructing operators to change both by hand and a test to catch them if they did not. One value renders both now, so they cannot be changed apart. Defaults are byte-identical, so the reference topology is unaffected. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
… created On the gateway backend a site reported Running with `verification.ok: true` and a public URL that answered 404, and its HTTPRoute had no parent status at all. The Gateway object is named after `namespaces.gateway`, while SITES_GATEWAY_NAME and SITES_GATEWAY_NAMESPACE defaulted to the literal `sites-gateway` and the Chart never set them. They agree only when the release happens to be installed into a namespace of that name -- and both install scripts set `namespaces.gateway` to the release namespace, which defaults to `sites-local`. Gateway routing has therefore been broken on every installation made through this repository's own scripts. A second consequence from the same missing pair: the tenant NetworkPolicy admits `owning-gateway-name: <GATEWAY_NAME>`, so it allowed a label the real data-plane Pods do not carry -- precisely the symptom exposure.py's comment predicts, where the policy is valid, applies cleanly, and the site still receives rejected traffic. None of the three layers reports an error. The Chart now renders both variables from `namespaces.gateway`, so one source feeds the parentRef and the policy label. Measured before and after on kind with Envoy Gateway v1.2.4 and KEDA 2.20.2, installed into `sites-gw`: parentRef `sites-gateway/sites-gateway` with no parent status and a 404, then parentRef `sites-gw/sites-gw`, Accepted=True, ResolvedRefs=True, policy label `sites-gw`, and a 200 whose body digest matches the recorded verification digest. The existing identity test renders the defaults, where both sides read `sites-gateway` and agree regardless of what the Chart does; the new one uses a different namespace. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Sixteen fixes and one test-only commit, found by installing this repository from a
clean clone as a newcomer would and then following its own documentation: the
development path, the "install it into an existing cluster" path, all five
deployment forms, the console, the database, both exposure backends, and the
60-trial cluster benchmark.
Every fix is verified against a live cluster where that is possible, and every one
carries a gate that was checked by mutation — reintroduce the old behaviour, watch
the new test go red, restore, watch it go green.
The three that break documented paths outright
Source builds were dead on every installation. The BuildKit Job selects
config.jsonout ofsites-registry-authand nothing in the repository evercreated that key. A missing selected key is not a soft failure: kubelet refuses to
set the volume up, the Pod never starts, and after five minutes the caller is told
"Job was active longer than specified deadline" — naming neither the Secret nor the
key.
sites build submit,source_deployand the source-build row of thedeployment-forms table were unreachable for anyone following the documentation.
Gateway routing was broken on every installation made through the repository's
own scripts. The Gateway object is named after
namespaces.gateway;SITES_GATEWAY_NAME/SITES_GATEWAY_NAMESPACEdefaulted to the literalsites-gatewayand the Chart never set them. Both install scripts setnamespaces.gatewayto the release namespace, which defaults tosites-local, sothe two agreed only by coincidence. Every HTTPRoute referenced a Gateway that does
not exist (no parent status at all), and the tenant NetworkPolicy admitted a
data-plane Pod label nothing carries. All three layers stay quiet: the site is
Running,verification.okis true because the probe uses the in-cluster address,and the public URL answers 404.
Public URLs were assembled from the reference trial's host port. On any cluster
but the kubeadm trial,
status.url, the CR, the database and the console's Openaction all named
127.0.0.1:1809x, which is a fact about the machine in front ofthat one cluster. The mapping is now declared or absent, for the same reason
clusterNetwork.podCIDRhas no default. The gateway backend had the same threeliterals; they are Chart values now, with the domain suffix feeding both the
ConfigMap and the listener wildcard so they cannot be changed apart.
Evidence that says less than it appears to
healthPath, not the site root. With the/default the twocoincide;
sites build submitdefaults to/healthz, and a site whose rootreturns 403 was
Running,ready,verification.ok: true,httpStatus: 200,with a public URL answering 403. AGENT_CONTRACT told callers to compare that
URL's digest with
bodySha256, which can only hold at/.roll out, so
phase: Failedandverification.ok: trueappear together. Thefields to compare were both in the response; the rule to compare them was not.
verificationnormessage, so a caller sawRunningorFailedand nothing else.ResourceQuotarejects the pod at the ReplicaSet, so thereis no pod to inspect and the stall reported only "Deployment does not have
minimum availability". The reason was already on the Deployment as a
ReplicaFailurecondition.Things a newcomer runs and finds broken or absent
deploy-specs/, which README points at, were rejected by theendpoint they illustrate. Nothing had ever submitted them.
_agent_deployment_authorization, whoseshape is deliberately absent from every tool schema and was absent from the
documentation too — so the whole write surface could not be implemented from the
contract.
make quickstart-doctorpassed while quickstart ports were held by another user:lsofonly lists sockets it owns, and Lima binds the same addresses. Now a bindprobe, which answers the question for any owner.
standalone.shhardcoded the values overlay, socluster_benchmark.py— thescript whose report invites reproduction — could only run on a cluster whose Pod
network matched the reference topology.
the Chart's cluster-scoped CRDs make a release exclusive.
--health-pathdefaults differ between sibling commands and none printed itsdefault; two admin help strings were machine translations ("summary" for digest);
one refusal message reached callers as "there will be noThe link can receive the
request".
Claims that had drifted
SITES_*variables appear in no document or chart" — measured on thatsentence's own definition, eight. Documented; the count is zero and gated.
api_mcp.pyandhttp_kit.py, and counted eight endpoint mixins whereHandlercombines nine.324 blobs are clean.
Verification
make test1031 → 1060, all green, with 29 new gates. Console, chart lint andrender, contract benchmark, homepage check and
docker build --checkall pass froma clean clone.
The 60-trial cluster benchmark passes on a single-node kind cluster
(
valid: true, passed: true, 60/60 trials, 300/300 stages, revision-match andcleanup rates 1.0), which also required the
standalone.shseam above:The gateway scale-to-zero and activator cold-start chain — which the published
benchmark never executes, since every trial runs
exposure: internal— wasexercised by hand on Envoy Gateway v1.2.4 and KEDA 2.20.2: from zero replicas a
single request returned 200 with the verified digest after 4.38 s, against 49 ms
warm.
🤖 Generated with Claude Code
https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is