Skip to content

Fix what a newcomer install actually hits: source builds, gateway routing, and URLs that name the wrong host - #10

Merged
convee merged 17 commits into
mainfrom
bugfix/20260905-newcomer-onboarding/main
Sep 6, 2026
Merged

convee merged 17 commits into
mainfrom
bugfix/20260905-newcomer-onboarding/main

Conversation

@convee

@convee convee commented Sep 6, 2026

Copy link
Copy Markdown
Contributor

Sixteen fixes and one test-only commit, found by installing this repository from a
clean clone as a newcomer would and then following its own documentation: the
development path, the "install it into an existing cluster" path, all five
deployment forms, the console, the database, both exposure backends, and the
60-trial cluster benchmark.

Every fix is verified against a live cluster where that is possible, and every one
carries a gate that was checked by mutation — reintroduce the old behaviour, watch
the new test go red, restore, watch it go green.

The three that break documented paths outright

Source builds were dead on every installation. The BuildKit Job selects
config.json out of sites-registry-auth and nothing in the repository ever
created that key. A missing selected key is not a soft failure: kubelet refuses to
set the volume up, the Pod never starts, and after five minutes the caller is told
"Job was active longer than specified deadline" — naming neither the Secret nor the
key. sites build submit, source_deploy and the source-build row of the
deployment-forms table were unreachable for anyone following the documentation.

Gateway routing was broken on every installation made through the repository's
own scripts.
The Gateway object is named after namespaces.gateway;
SITES_GATEWAY_NAME / SITES_GATEWAY_NAMESPACE defaulted to the literal
sites-gateway and the Chart never set them. Both install scripts set
namespaces.gateway to the release namespace, which defaults to sites-local, so
the two agreed only by coincidence. Every HTTPRoute referenced a Gateway that does
not exist (no parent status at all), and the tenant NetworkPolicy admitted a
data-plane Pod label nothing carries. All three layers stay quiet: the site is
Running, verification.ok is true because the probe uses the in-cluster address,
and the public URL answers 404.

Public URLs were assembled from the reference trial's host port. On any cluster
but the kubeadm trial, status.url, the CR, the database and the console's Open
action all named 127.0.0.1:1809x, which is a fact about the machine in front of
that one cluster. The mapping is now declared or absent, for the same reason
clusterNetwork.podCIDR has no default. The gateway backend had the same three
literals; they are Chart values now, with the domain suffix feeding both the
ConfigMap and the listener wildcard so they cannot be changed apart.

Evidence that says less than it appears to

  • Verification probes healthPath, not the site root. With the / default the two
    coincide; sites build submit defaults to /healthz, and a site whose root
    returns 403 was Running, ready, verification.ok: true, httpStatus: 200,
    with a public URL answering 403. AGENT_CONTRACT told callers to compare that
    URL's digest with bodySha256, which can only hold at /.
  • Evidence for an earlier revision is deliberately kept when a later one fails to
    roll out, so phase: Failed and verification.ok: true appear together. The
    fields to compare were both in the response; the rule to compare them was not.
  • Bundle components carried neither verification nor message, so a caller saw
    Running or Failed and nothing else.
  • An exhausted tenant ResourceQuota rejects the pod at the ReplicaSet, so there
    is no pod to inspect and the stall reported only "Deployment does not have
    minimum availability". The reason was already on the Deployment as a
    ReplicaFailure condition.

Things a newcomer runs and finds broken or absent

  • Both files in deploy-specs/, which README points at, were rejected by the
    endpoint they illustrate. Nothing had ever submitted them.
  • The MCP write tools all refuse without _agent_deployment_authorization, whose
    shape is deliberately absent from every tool schema and was absent from the
    documentation too — so the whole write surface could not be implemented from the
    contract.
  • make quickstart-doctor passed while quickstart ports were held by another user:
    lsof only lists sockets it owns, and Lima binds the same addresses. Now a bind
    probe, which answers the question for any owner.
  • standalone.sh hardcoded the values overlay, so cluster_benchmark.py — the
    script whose report invites reproduction — could only run on a cluster whose Pod
    network matched the reference topology.
  • Installing a second release fails with Helm's ownership error and nothing said
    the Chart's cluster-scoped CRDs make a release exclusive.
  • --health-path defaults differ between sibling commands and none printed its
    default; two admin help strings were machine translations ("summary" for digest);
    one refusal message reached callers as "there will be noThe link can receive the
    request".

Claims that had drifted

  • "45 SITES_* variables appear in no document or chart" — measured on that
    sentence's own definition, eight. Documented; the count is zero and gated.
  • The per-module table listed 43 of 45 modules, missing api_mcp.py and
    http_kit.py, and counted eight endpoint mixins where Handler combines nine.
  • "Historical UIDs remain in the Git history" — the history is 22 commits and all
    324 blobs are clean.

Verification

make test 1031 → 1060, all green, with 29 new gates. Console, chart lint and
render, contract benchmark, homepage check and docker build --check all pass from
a clean clone.

The 60-trial cluster benchmark passes on a single-node kind cluster
(valid: true, passed: true, 60/60 trials, 300/300 stages, revision-match and
cleanup rates 1.0), which also required the standalone.sh seam above:

stage published report (3-VM kubeadm) this branch (single-node kind)
static publish p50 6.668 s / p95 8.090 s p50 6.524 s / p95 6.886 s
dynamic publish p50 6.908 s / p95 8.085 s p50 6.832 s / p95 7.732 s
rollback and recovery p50 15.862 s / p95 77.826 s p50 16.156 s / p95 47.875 s

The gateway scale-to-zero and activator cold-start chain — which the published
benchmark never executes, since every trial runs exposure: internal — was
exercised by hand on Envoy Gateway v1.2.4 and KEDA 2.20.2: from zero replicas a
single request returned 200 with the verified digest after 4.38 s, against 49 ms
warm.

🤖 Generated with Claude Code

https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is

convee and others added 17 commits September 5, 2026 18:20
…ort mapping

A public deployment on any cluster but the reference trial reported
`http://127.0.0.1:1809x` in `sites status`, in the CR's `status.url`, in the
database, and behind the console's Open action. Measured on kind: nodePort 30080
was reported as 18090, which nothing was listening on.

`HOST_PORT_BASE` defaulted to 18090, the host port the repository's own Lima
relay forwards 30080..30088 to. That is a fact about the machine in front of the
cluster, not about the cluster, and the Chart hardcoded `SITES_PUBLIC_URL_HOST`
with no value at all. The wrong answer is invisible: the deployment is healthy
and server-side verification passes because it probes the in-cluster address.

The mapping is now declared or absent, for the same reason
`clusterNetwork.podCIDR` has no default. Undeclared, `public_url` returns None
instead of guessing. New Chart values `nodePort.hostPortBase` (no default) and
`nodePort.publicUrlHost`; `cluster.sh` gains matching `SITES_HOST_PORT_BASE` and
`SITES_PUBLIC_URL_HOST` seams. `host_port_base()` reads the environment per call
for the same reason `backend()` does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
`make quickstart-doctor` reported "local ports: available" while 18090 and 18091
were held by a root-owned process. Two-sided control on one address,
127.0.0.1:18092: held by the invoking user the doctor refuses correctly; held by
root it passes. Only the owner differs.

`lsof` lists sockets owned by the user running it, so a port held by anyone else
returns non-zero and reads as free. Lima binds these forwards on 127.0.0.1, so
this is a guaranteed failure several minutes later -- exactly what the doctor's
own comment says it exists to prevent.

Replaced with a bind probe: it answers whether a server can take the address,
for any owner, and sets SO_REUSEADDR so a lingering TIME_WAIT socket does not
read as a conflict. Both preflights use it, `lsof` drops out of the prerequisites,
and the fixed-port list is derived from `host_port_base` instead of repeating
18090 a second time in the same file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…evidence

README points at `deploy-specs/` for bundle examples and neither file could be
submitted. `example-single.json` was a flat single-deployment shape the bundle
endpoint refuses outright (`unsupported bundle fields`); `example-stack.json`
carried a `command` field that is not a deployment field. Nothing had ever
submitted them, so both drifted out of the contract they illustrate.

Submitting them then showed a second gap: bundle components reported
`verification: null` while the operator had written the evidence onto every CR.
docs/AGENT_CONTRACT.md makes `status.verification` the only success criterion and
bundles are one of its deployment entry points, so the projection was hiding the
one thing a caller is told to check.

Both examples are now real bundles that reach Running with 200s on a live
cluster, their component names no longer collide so either can be submitted
first, and the bundle projection carries `verification`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…README limitation

A gate per fix, each verified by mutation (restore the old behaviour, watch it
go red):
- no declared host mapping means no public URL; the default render must not set
  SITES_HOST_PORT_BASE; the trial's declared base must equal its Lima forward
- the port preflight must bind rather than list sockets it owns
- every file in deploy-specs/ passes the endpoint's own validator, and no two
  examples claim the same component name
- bundle components carry verification

Also drops "Historical UIDs remain in the Git history": the repository is public,
its history is 22 commits, and scanning all 324 blobs finds no UID or Aliyun
identifier -- the only `aliyuncs.com` occurrences are the OSS endpoint check and
its tests. The console's `The only one in the merchant` placeholder, a machine
translation, becomes `Unique within the merchant`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…w the install namespace

`sites build submit` always failed, and the only signal after five minutes was
"Job was active longer than specified deadline". The build Pod never started:
`MountVolume.SetUp failed for volume "registry-auth": references non-existent
secret key: config.json`.

builds.py selects `config.json` out of `sites-registry-auth` -- DOCKER_CONFIG
points at that mount and Docker looks for that filename -- and nothing in the
repository ever created the key. Selecting a missing key is not a soft failure:
kubelet refuses to set the volume up, so the entire source-build entry point was
dead on every installation that followed this repository's own instructions, and
the error named neither the Secret nor the key.

Same area, second defect: SITES_REGISTRY_PUSH_HOST and SITES_REGISTRY_API
defaulted to a hardcoded `sites-local` namespace and the Chart never set them,
so any other namespace pushed builds at, and resolved digests against, a Service
that is not there.

bootstrap-standalone-secrets.sh now generates the docker config from the same
registry password (host defaults to `sites-registry.<namespace>.svc:5000`,
overridable), requires the key when reusing an existing Secret so an older
installation is told which key is missing, and the Chart renders both registry
addresses from `namespaces.control` for the API and the operator. Verified on a
fresh install: build, push, deploy, verified 200 in about 30 seconds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
`--health-path` defaults to `/` for `deploy` and `deploy-static` and `/healthz`
for `build submit`, and none of the three said so. Taking the sibling default on
faith produced a build that pushed a correct image and then spent the whole
readiness window failing a probe on a path the image does not serve, reported as
"Deployment does not have minimum availability" -- which points at the workload
rather than at the flag.

`--port` and `--health-path` now print their real defaults, and the source-build
one says why it differs. The defaults themselves are unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…ch group has

`sites admin merchants create --help` read "only the summary is stored in the
library". Two machine translations in one sentence: a digest is not a summary
(the security meaning differs) and "library" is a literal rendering of "database".
The tenant equivalent said it differently again.

Both now say the control plane stores only a digest, and each names the rotate
command its own group actually has. A gate checks that: the first attempt at this
message said `rotate-token`, which tenants do not have, and a substring check
would have accepted it because `rotate` is its prefix.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…write tool requires

Following AGENT_CONTRACT's own client configuration, all six MCP write tools
answer `deployment_authorization_required`, and the message says only "do not
retry" -- nothing an integrator can implement.

The reserved `_agent_deployment_authorization` argument is deliberately absent
from every tools/list schema, so the model never learns it exists. That leaves
the document as the only place a host can read its shape, and the shape was not
there.

Adds the field table (version must be 1, runId 1-128, nonce at least 24 chars,
expiresAt in the future, allowInternal for internal exposure) and the three rules
that make it worth having. The gate checks each documented field against the
validator rather than a copy of it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…on it names

Observed on a live cluster: a dynamic site's version 2 never became ready, and
the same `sites status` response carried `phase: Failed` next to
`verification.ok: true, httpStatus: 200` -- version 1's evidence.

The retention is deliberate; the operator comment says so, and the evidence names
the revision it was collected for. What was missing is the rule: AGENT_CONTRACT
made `verification.ok=true` the success criterion without saying to compare
`verification.revision` with the deployment's `revision`, so a caller reading
`ok` alone reports a rollout that never came up as a live site.

Both halves are now gated: the response must carry both revisions (or the
comparison is impossible) and the contract must state it (or nobody makes it).

Also fixes the `scaleToZero` refusal text, where two translated fragments were
concatenated without a space and reached the caller as "there will be noThe link
can receive the request". A new sweep checks all 408 refusal messages for that
shape; its judge is self-tested in both directions, because the first version
filtered out the very word it was looking for and reported zero.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…ng a timeout

The fifth site in a tenant is accepted with a 200 and fails 120 seconds later
with "Deployment was not ready within 120s (Deployment does not have minimum
availability.)". The cause was only in a ReplicaSet event: `exceeded quota:
sites-tenant-quota, requested: limits.cpu=1, used: limits.cpu=4`.

`_rollout_stall_reason` exists to separate OOMKilled from CrashLoopBackOff from
ImagePullBackOff, and reads the Available condition and pod container states. A
namespace ResourceQuota rejects creation at the ReplicaSet, so there is no pod to
inspect -- the one stall it cannot see is the one whose Available message is
content-free. Kubernetes copies the reason onto the Deployment as a
ReplicaFailure condition, which the operator already holds.

Two more of the same shape: bundle components reported `Failed` with no message
at all, and the two tenant limits disagree -- admission counts `maxDeployments`
(10) while the namespace quota allows four concurrent sites. The mismatch is now
in Known limitations, with a gate that asks for the limitation to be deleted
rather than updated if the numbers ever agree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…ed from

A site whose root returns 403 reported `phase: Running`, `ready: true`,
`verification.ok: true`, `httpStatus: 200`, and a public URL that answers 403.

The probe requests `spec.healthPath`, not the site root, and writes that address
into `verification.url`. `deploy` and `deploy-static` default to `/` so the two
have always coincided; `build submit` defaults to `/healthz`, and there they are
different resources. AGENT_CONTRACT nonetheless tells callers to compare the
public URL's digest with `bodySha256`, which can only hold at `/`, and the
quickstart's "public URL body: matches verification digest" assertion passes only
because it takes the default.

The bound is now stated where the criterion is. Gated on all three sides: the URL
is built from healthPath, the document says so, and the trial passes no
`--health-path` (its own assertion would fail if it did).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…ster.sh already does

`scripts/cluster_benchmark.py` could not run on any cluster whose Pod network is
not the reference topology's. It installs through `standalone.sh`, which
hardcoded `values-dev.yaml` and therefore podCIDR 10.201.0.0/16; the install
succeeds and the operator then refuses to start, by design, because its own
address falls outside the declaration.

`cluster.sh` has had SITES_HELM_VALUES and SITES_CLUSTER_POD_CIDR all along.
Adding the same two seams to its sibling makes the benchmark -- the script whose
report invites reproduction -- runnable anywhere. Default behaviour is unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…ale count

README and CONFIGURATION both said "45 SITES_* variables ... appear in no
document or chart". Measured on that sentence's own definition: 116 are read, 86
appear in documents, 40 in the Chart or scripts, and 8 appear in neither. Even
counting documents alone gives 30.

Same shape as the "historical UIDs" bullet removed earlier: a limitation that
overstates a problem which has largely been closed, drifting in the safe
direction so nothing surfaces it.

With eight left they are simply documented, the count goes to zero, and the
README bullet now states what is genuinely still true -- these variables are not
schema-validated. A gate requires every variable the code reads to be named in a
document or the Chart, and self-tests its own scan so an empty read set cannot
pass as success.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…elease exclusive

Installing a second release into another namespace fails with Helm's ownership
error about `meta.helm.sh/release-namespace` -- a message about Helm annotations
that mentions neither Site nor the limit it is enforcing.

Both CRDs are cluster-scoped and rendered as ordinary templates, so the first
release owns them, while `--namespace` on both install scripts implies the
opposite. Documented, with the real error text, plus a gate that also pins the
premise: if the CRDs ever move to the Chart's `crds/` directory the exclusivity
disappears and the note should be deleted rather than edited.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…xins from the code

The per-module table listed 43 of the package's 45 modules. Missing were
`api_mcp.py` -- the `POST /mcp` endpoint that README and AGENT_CONTRACT both
treat as a headline surface -- and `http_kit.py`. The composition-root sentence
counted eight endpoint mixins while `Handler` combines nine.

A hand-maintained table drifts silently because adding a module gives nobody a
reason to reopen it. Gated three ways: every module has a row, every row names a
file that exists, and the mixin count and its spelled-out word are derived from
which `api_*.py` files actually define a Mixin.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
…template literals

The three quantities that compose a gateway URL -- domain suffix, scheme, host
port -- were literals in the template, and the comment said the last two are the
local Lima forward. Same shape as the NodePort backend's HOST_PORT_BASE on the
other backend: routing works, only the address handed to the caller is wrong, and
server-side verification passes either way because it uses the in-cluster address.

The suffix was worse than hardcoded: it appeared twice, in the ConfigMap and in
the listener wildcard, with the file instructing operators to change both by hand
and a test to catch them if they did not. One value renders both now, so they
cannot be changed apart. Defaults are byte-identical, so the reference topology
is unaffected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
… created

On the gateway backend a site reported Running with `verification.ok: true` and a
public URL that answered 404, and its HTTPRoute had no parent status at all.

The Gateway object is named after `namespaces.gateway`, while
SITES_GATEWAY_NAME and SITES_GATEWAY_NAMESPACE defaulted to the literal
`sites-gateway` and the Chart never set them. They agree only when the release
happens to be installed into a namespace of that name -- and both install scripts
set `namespaces.gateway` to the release namespace, which defaults to
`sites-local`. Gateway routing has therefore been broken on every installation
made through this repository's own scripts.

A second consequence from the same missing pair: the tenant NetworkPolicy admits
`owning-gateway-name: <GATEWAY_NAME>`, so it allowed a label the real data-plane
Pods do not carry -- precisely the symptom exposure.py's comment predicts, where
the policy is valid, applies cleanly, and the site still receives rejected
traffic.

None of the three layers reports an error. The Chart now renders both variables
from `namespaces.gateway`, so one source feeds the parentRef and the policy label.
Measured before and after on kind with Envoy Gateway v1.2.4 and KEDA 2.20.2,
installed into `sites-gw`: parentRef `sites-gateway/sites-gateway` with no parent
status and a 404, then parentRef `sites-gw/sites-gw`, Accepted=True,
ResolvedRefs=True, policy label `sites-gw`, and a 200 whose body digest matches
the recorded verification digest.

The existing identity test renders the defaults, where both sides read
`sites-gateway` and agree regardless of what the Chart does; the new one uses a
different namespace.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01365xzFrkTZV7Fr9WmNe6is
@convee
convee merged commit 4361bcc into main Sep 6, 2026
16 checks passed
@convee
convee deleted the bugfix/20260905-newcomer-onboarding/main branch September 6, 2026 02:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant