Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 6 additions & 5 deletions .claude/skills/verify/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,14 +5,14 @@
go build -o /tmp/pulse-operator ./cmd/main.go
```

## envtest (unit tests with fake cluster)
## envtest (local API server and etcd, without workload controllers)
```bash
# Install once
go install sigs.k8s.io/controller-runtime/tools/setup-envtest@latest
$(go env GOPATH)/bin/setup-envtest use 1.31 --bin-dir /tmp/kubebuilder-bin
export KUBEBUILDER_ASSETS="$("$(go env GOPATH)/bin/setup-envtest" use 1.31 -p path)"

# Run
KUBEBUILDER_ASSETS=/tmp/kubebuilder-bin/k8s/1.31.0-darwin-arm64 go test ./...
go test -count=1 ./...
```

## Local run against live OCP cluster
Expand All @@ -38,8 +38,9 @@ oc apply -f examples/pulse.yaml
agent Deployment (removed — see agent_reconciler_test.go "Deployment is created even
while memory PVC is Pending"); Kubernetes just holds the pod Pending until it binds.
- `--metrics-bind-address` defaults to `:8082`. Specify a different port if 8082 is in use.
- controller-runtime v0.24.1: startup sequence logs "Stopping and waiting..." during init,
not during shutdown — ignore these; wait for "Reconciling OpenShiftPulse" log line.
- Inspect errors and process exit status when startup logs say "Stopping and waiting";
do not treat a shutdown message as proof that initialization is healthy. Confirm
readiness and successful reconciliation separately.
- Finalizer: delete CR only after operator is running so finalizer cleanup fires correctly.

## Sample CR
Expand Down
7 changes: 6 additions & 1 deletion .github/dependabot.yml
Original file line number Diff line number Diff line change
Expand Up @@ -5,11 +5,16 @@ updates:
schedule:
interval: "weekly"
open-pull-requests-limit: 10
groups:
kubernetes:
patterns:
- "k8s.io/*"
- "sigs.k8s.io/controller-runtime"
- package-ecosystem: "github-actions"
directory: "/"
schedule:
interval: "weekly"
- package-ecosystem: "docker"
directory: "/"
schedule:
interval: "weekly"
interval: "daily"
4 changes: 2 additions & 2 deletions Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@
# themselves regardless of the Go version baked into this base image. This image's
# default WORKDIR is /opt/app-root/src, owned by its non-root default user (uid 1001),
# so we build there instead of /workspace.
FROM registry.access.redhat.com/ubi9/go-toolset:1.26@sha256:1a9bbbfa854931a97dbff276bd69dc0e32b36cb2fbce3b9813b2cf9892aa8d43 AS builder
FROM registry.access.redhat.com/ubi9/go-toolset:1.26@sha256:8cf89835994846ca0dffb9078e3a5638c57ec6175750f0af02fbe9c9942696d3 AS builder
WORKDIR /opt/app-root/src
# Unlike Docker Hub's golang images, Red Hat's go-toolset builds Go with
# GOTOOLCHAIN defaulting to "local" instead of upstream's "auto". Without this,
Expand All @@ -26,7 +26,7 @@ COPY internal/ internal/
RUN CGO_ENABLED=0 GOOS=linux GOARCH=amd64 go build -a -o manager cmd/main.go

# Runtime stage
FROM registry.access.redhat.com/ubi9/ubi-minimal:latest@sha256:8eb2830d0936237fc13a1f2f7e45aecf90d69043380ad167fad0343632937f41
FROM registry.access.redhat.com/ubi9/ubi-minimal:latest@sha256:7fbeae18dc9476399f565e68255f602a3374ea8614ba3d14843565131a13ff93
WORKDIR /
# The base digest is pinned for reproducibility, which also pins its CVEs.
# Pull in published package fixes at build time so a rebuild picks up errata
Expand Down
4 changes: 2 additions & 2 deletions Dockerfile.catalog
Original file line number Diff line number Diff line change
Expand Up @@ -29,14 +29,14 @@
# fails once deployed: "exec container process `/bin/opm`: Exec format error"
# on any real (amd64) cluster node. Get the correct per-arch digest with:
# podman manifest inspect quay.io/operator-framework/opm:latest
FROM quay.io/operator-framework/opm@sha256:c6dc739b9f630ae8efb9cb0e6f4306a55ca738dffc4fa6ef376b313d90928dfa AS builder
FROM quay.io/operator-framework/opm@sha256:b32d3891616662620da08d7f0ec42c2e69fa2de43427dc975d35b12f7a969a0f AS builder

# Copy the FBC root (built by README.md's render step into ./catalog) into the
# image at /configs and pre-populate the serve cache.
ADD catalog /configs
RUN ["/bin/opm", "serve", "/configs", "--cache-dir=/tmp/cache", "--cache-only"]

FROM quay.io/operator-framework/opm@sha256:c6dc739b9f630ae8efb9cb0e6f4306a55ca738dffc4fa6ef376b313d90928dfa
FROM quay.io/operator-framework/opm@sha256:b32d3891616662620da08d7f0ec42c2e69fa2de43427dc975d35b12f7a969a0f
ENTRYPOINT ["/bin/opm"]
CMD ["serve", "/configs", "--cache-dir=/tmp/cache"]

Expand Down
60 changes: 32 additions & 28 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,10 +43,10 @@ One `OpenShiftPulse` CR drives the full lifecycle:
|---|---|
| **Agent** | ClusterRole (read-only cluster access), WS token Secret, memory PVC, Deployment, Service |
| **PostgreSQL** | StatefulSet (pg-data PVC retained on delete), pg-auth Secret (also retained — see below), ClusterIP + headless Services |
| **UI** | nginx ConfigMap, oauth-proxy Deployment (TLS on 8443), Service, Route, OAuthClient |
| **UI** | nginx configuration Secret, Deployment with nginx and oauth-proxy (TLS on 8443), Service, Route, OAuthClient |
| **Monitoring** | ServiceMonitor (agent `/metrics`), PrometheusRule (`PulseAgentDown`, `PulsePostgreSQLDown`) |
| **MCP** | MCP server ServiceAccount + ClusterRole (read-only) + ClusterRoleBinding, Deployment, Service (optional, `spec.agent.mcp.enabled: true`) |
| **Network** | Per-component ingress-only NetworkPolicies: UI (OCP ingress + Prometheus), PostgreSQL (agent pod only), agent (UI pod + Prometheus), MCP server (agent pod only) |
| **Network** | Per-component ingress-only NetworkPolicies: UI (OCP ingress + Prometheus), PostgreSQL (agent and enabled Temporal pods), agent (UI pod + Prometheus), MCP server (agent pod only) |
| **Cluster detect** | Reads ingress domain, oauth-proxy image digest, ACM availability on first reconcile |

A `pulse.ai/cleanup` finalizer ensures ClusterRoles and OAuthClient are removed when the CR is deleted — no orphans on uninstall.
Expand Down Expand Up @@ -89,9 +89,9 @@ no skew rather than guessing.

### Why the operator's version differs

The operator versions independently of the Pulse application it deploys. As of
this release the operator is **v0.7.0** while the agent and UI ship **v2.27.0**
— that gap is deliberate, not drift:
The operator versions independently of the Pulse application it deploys.
Read `OperatorVersion` in `internal/controller/compat.go` for this checkout's
operator version, and the CR's image fields for the deployed application versions:

- The operator's version tracks *its own* API and reconcile behaviour. The CRD
is still `v1alpha1`, and a 0.x version says so honestly.
Expand Down Expand Up @@ -334,12 +334,13 @@ spec:
# ── Agent ───────────────────────────────────────────────────────────────────
agent:
image: quay.io/amobrem/pulse-agent:latest
trustLevel: 2 # 0=observe · 1=suggest · 2=confirm · 3=batch · 4=autonomous
allowWriteOperations: false # adds delete(pods), patch(deployments) to agent ClusterRole
trustLevel: 2 # monitor: 0/1=no remediation, 2=propose, 3/4=execute
adminUsers: "kube:admin" # restrict administrator endpoints; empty allows any authenticated identity
allowWriteOperations: false # adds delete(pods), patch/update(deployments,statefulsets) to agent ClusterRole
allowSecretAccess: false # adds get/list/watch(secrets) to agent ClusterRole
resources: {} # corev1.ResourceRequirements
mcp:
enabled: false # deploys MCP server sidecar for tool extension
enabled: false # deploys a separate MCP server Deployment for tool extension
minOperatorVersion: "" # optional semver floor for this operator build; unset (default) = inert — see "Agent-version compatibility gate" below

# ── UI ──────────────────────────────────────────────────────────────────────
Expand All @@ -364,11 +365,17 @@ spec:

| Level | Behaviour |
|---|---|
| `0` — observe | Read-only. Agent answers questions but takes no action. |
| `1` — suggest | Proposes actions in the UI, user approves each one. |
| `2` — confirm | Default. Agent executes after a single user confirmation. |
| `3` — batch | Executes batches of low-risk actions with one confirmation. |
| `4` — autonomous | Executes without confirmation. Use with caution. |
| `0` / `1` | Background monitor does not enter remediation; investigations can still run. |
| `2` | Background monitor creates proposed actions for approval. |
| `3` / `4` | Background monitor executes eligible remediation without the level-2 approval gate. |

This table describes the background monitor, not a universal authorization
boundary for chat, MCP tools, or plans. Kubernetes RBAC still applies. The
current monitor uses the configured level as a floor and caps browser requests
at that same level, so browser trust selection cannot lower its authority.
Its enabled categories start with all server handlers; browser category
checkboxes currently cannot restrict that set. Configure the CR and agent
permissions deliberately rather than relying on browser preferences.

---

Expand Down Expand Up @@ -523,12 +530,7 @@ that outage window actually took, and holds its previous value between upgrades
first one ever completes).

The agent Deployment uses the `Recreate` strategy, not `RollingUpdate` — a deliberate choice, not
an unexamined default. Its memory-cache PVC is `ReadWriteOnce`, and `pulse-agent`'s own Helm chart
runs `Recreate` for the identical reason (`chart/values.yaml`: *"Required because the memory PVC
is ReadWriteOnce (RWO) and cannot be mounted by two pods simultaneously"*) — a `RollingUpdate`
overlap here would leave the surging pod's volume attach stuck `Pending`
(`FailedAttachVolume`), arguably a worse failure mode than today's brief, bounded stop-then-start
outage. The agent also runs forward-only, no-rollback DB migrations automatically on startup, so
an unexamined default. Its memory-cache PVC is `ReadWriteOnce`; overlapping pods on different nodes can fail volume attachment. The agent also runs forward-only, no-rollback DB migrations automatically on startup, so
two concurrent agent versions sharing one Postgres instance is a second, independent reason to
avoid the overlap. `lastUpgradeDurationSeconds` exists to make that outage window's size visible
and measured, not to eliminate it — see the self-heal coverage above for what happens if the new
Expand Down Expand Up @@ -597,14 +599,16 @@ scripts/olm-uninstall.py authorino --yes # actually uninstall
### Prerequisites

```bash
# Install the Go version required by go.mod first.
go install sigs.k8s.io/controller-runtime/tools/setup-envtest@latest
setup-envtest use 1.31 --bin-dir /tmp/kubebuilder-bin
export KUBEBUILDER_ASSETS="$(setup-envtest use 1.31 -p path)"
```

### Run tests

```bash
KUBEBUILDER_ASSETS=/tmp/kubebuilder-bin/k8s/1.31.0-darwin-arm64 make test
make test
make vet
```

### Run locally against a live cluster
Expand Down Expand Up @@ -774,7 +778,7 @@ pulse-operator-system/
│ ├── {ns}/{name}-openshiftpulse (ServiceAccount)
│ ├── {ns}-{name}-openshiftpulse-reader (ClusterRole + ClusterRoleBinding [+ -auth-delegator] — cluster-scoped)
│ ├── {ns}/{name}-oauth-secrets (Secret — client-secret + cookie-secret)
│ ├── {ns}/{name}-nginx (ConfigMap — nginx.conf, root /opt/app-root/src)
│ ├── {ns}/{name}-nginx (Secret — nginx.conf, root /opt/app-root/src)
│ ├── {ns}/{name}-openshiftpulse (Deployment: nginx + oauth-proxy sidecars)
│ ├── {ns}/{name}-openshiftpulse (Service :8443)
│ ├── {ns}/{name}-openshiftpulse (Route — reencrypt, OCP assigns hostname)
Expand All @@ -791,7 +795,7 @@ pulse-operator-system/
│
└── NetworkPolicyReconciler
├── {ns}/{name}-openshiftpulse (UI: ingress from OCP router + Prometheus only)
├── {ns}/{name}-pg-access (PG: ingress from the agent pod only)
├── {ns}/{name}-pg-access (PG: ingress from agent and enabled Temporal pods)
└── {ns}/{name}-agent-access (Agent: ingress from the UI pod + Prometheus only)
```

Expand Down Expand Up @@ -826,10 +830,10 @@ oc delete pods -n openshiftpulse -l app=pulse-openshiftpulse

**Symptom:** Default nginx welcome page at the route URL.

**Cause:** Stale ConfigMap or nginx not pointing at `/opt/app-root/src`. Force reconcile:
**Possible cause:** Stale nginx configuration or nginx not pointing at `/opt/app-root/src`. Force reconcile:

```bash
oc delete configmap pulse-nginx -n openshiftpulse
oc delete secret pulse-nginx -n openshiftpulse
# Operator recreates it within seconds
```

Expand Down Expand Up @@ -874,8 +878,8 @@ oc get clusterrolebinding | grep monitoring-view
**Cause:** The Alerts view reads firing alerts/rules from `/api/prometheus/` (Thanos-querier) but silences from a separate `/api/alertmanager/` proxy — a pre-existing gap where that location didn't exist in nginx at all, so requests fell through to the SPA's own `index.html` (200 OK, `text/html`) instead of reaching Alertmanager. The UI correctly detected the non-JSON response and reported the backend as unreachable, even though Prometheus itself was fine.

```bash
# Confirm the proxy exists in the live ConfigMap
oc get configmap {name}-nginx -n <namespace> -o jsonpath='{.data.nginx\.conf}' | grep -A5 'location /api/alertmanager/'
# Confirm the proxy exists in the live configuration Secret
oc get secret <name>-nginx -n <namespace> -o jsonpath='{.data.nginx\.conf}' | base64 --decode | grep -A5 'location /api/alertmanager/'
```

If it's missing, the operator image predates this fix — upgrade and restart the UI pods. Also confirm the logged-in user holds `monitoring-alertmanager-view` (or `-edit`) in `openshift-monitoring`; `cluster-monitoring-view` alone (which covers the Thanos path) is not sufficient for silences.
Expand Down Expand Up @@ -912,7 +916,7 @@ See [SECURITY.md](SECURITY.md) to report a vulnerability.

- All managed pods run as non-root with `AllowPrivilegeEscalation=false`, `Capabilities.Drop=ALL`, and `SeccompProfile=RuntimeDefault`.
- PostgreSQL sets `ReadOnlyRootFilesystem=false` (PG requires writable socket and temp paths).
- The operator's own ClusterRole ([`config/rbac/role.yaml`](config/rbac/role.yaml)) does **not** include `escalate`/`bind` on RBAC resources — every rule it ever writes into a generated agent/UI/MCP ClusterRole is already a permission it holds itself, so Kubernetes' RBAC "you already have this" rule lets `create`/`update` succeed without those verbs. It's still a privilege-concentration point (it *creates* ClusterRoles/ClusterRoleBindings for every managed instance): restrict exec access to `pulse-operator-system` via NetworkPolicy.
- The operator's own ClusterRole ([`config/rbac/role.yaml`](config/rbac/role.yaml)) does **not** include `escalate`/`bind` on RBAC resources — every rule it ever writes into a generated agent/UI/MCP ClusterRole is already a permission it holds itself, so Kubernetes' RBAC "you already have this" rule lets `create`/`update` succeed without those verbs. It's still a privilege-concentration point (it *creates* ClusterRoles/ClusterRoleBindings for every managed instance): restrict `pods/exec` access in `pulse-operator-system` with RBAC. NetworkPolicy controls pod traffic and does not authorize Kubernetes exec requests.
- The agent, UI, PostgreSQL, and MCP server pods each get their own NetworkPolicy restricting ingress to only the pods/namespaces that legitimately call them (e.g. only the UI pod may reach the agent on :8080; only the agent pod may reach the MCP server on :8081) — no pod is reachable cluster-wide by default.
- Every cluster-scoped resource the operator creates — the agent/UI/MCP ClusterRoles and ClusterRoleBindings, the agent's `-monitoring-view` binding, and the OAuthClient — is named `{namespace}-{name}-…` to prevent collision when multiple CRs coexist on the same cluster. Namespaced resources keep plain `{name}-…` names; Kubernetes already scopes those.
- The agent's ServiceAccount is bound to OpenShift's built-in `cluster-monitoring-view` ClusterRole (read-only) so its own alert-scanning/trend-monitoring features can query `thanos-querier` — this is separate from, and in addition to, the agent's own scoped-down ClusterRole.
Expand Down
24 changes: 24 additions & 0 deletions SECURITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,30 @@ following are especially welcome:
- Secrets (ws-token, oauth cookie/client secret, PostgreSQL password) leaking
via logs, events, or status fields

## Refreshing container dependencies after a CVE

The `container-scan` job builds the operator image and runs Trivy with
`severity: CRITICAL,HIGH`, `ignore-unfixed: true`, and `exit-code: 1`.
The runtime base is digest-pinned, and `microdnf update` applies available
package errata during each build. A new fixable vulnerability can therefore
change CI results even when the source tree has not changed.

Dependabot checks Docker digests daily. Review its existing PRs before opening
another update. To validate a candidate, build the complete operator image and
scan that image with the same Trivy settings; scanning only the base does not
exercise the package update or inspect the operator binary. Preserve the
`microdnf update` step when resolving an older digest PR against current main.

The builder contributes only the statically linked manager executable
(`CGO_ENABLED=0`) to the runtime stage. Updating builder RPMs alone does not
repair a runtime RPM vulnerability. Updating Go modules or the Go toolchain
may still be needed for vulnerabilities in the executable.

Kubernetes Go modules and controller-runtime are grouped in Dependabot because
independent minor updates can produce incompatible generated validation code.
Resolve them together, run `go mod tidy`, and run the full envtest suite before
landing the update.

## Supported versions

This project is `alpha` maturity (see the CSV's `maturity` field) with a
Expand Down
Loading
Loading