Protocol Version: 2
Defines the REST and WebSocket protocol between the Pulse UI and Pulse Agent. Both repos must implement the same protocol version for compatibility.
UI integration reference, not an exhaustive backend endpoint catalog. Authoritative current schemas are in
engine/agentClient.ts,engine/monitorClient.ts,engine/agentComponents.ts(undersrc/kubeview/) and the agent API contract. Update both code and documentation together. Examples use illustrative values, not current release/tool counts.
Browser paths are /api/agent/...; production nginx and the development proxy strip that prefix and inject shared credentials. Browser requests do not append the shared token. Backend paths below omit that proxy prefix. A shared token authenticates the proxy connection; administrator endpoints also require a validated caller identity. Deployment configuration is owned by the operator.
| Method | Path | Auth | Description |
|---|---|---|---|
GET |
/healthz |
public | Liveness probe. Returns {"status": "ok"} |
GET |
/version |
public | Protocol version, agent version (dynamic from package), tool count, feature flags |
GET |
/health |
token | Circuit breaker state, error summary, investigation stats, autofix_paused status |
GET |
/tools |
token | All tools grouped by mode (sre, security) with requires_confirmation flags |
GET |
/fix-history |
token | Paginated fix history with filters (status, category, since, search) |
GET |
/fix-history/{id} |
token | Single action detail with before/after state |
POST |
/fix-history/{id}/rollback |
admin + caller access token | Roll back supported snapshot/revision actions under the caller's Kubernetes authority; refusal/error is surfaced |
POST |
/fix-history/{id}/approve |
admin | Re-plan and approve a proposal; 409 if no longer applicable |
GET |
/eval/status |
token | Cached quality gate snapshot (release, safety, integration, outcomes) |
GET |
/predictions |
token | Returns empty — predictions are WebSocket-only (/ws/monitor) |
GET |
/memory/export |
token | Export learned runbooks and patterns as JSON |
POST |
/memory/import |
token | Import runbooks and patterns from another pod's export |
GET |
/monitor/capabilities |
token | Max trust level and supported auto-fix categories |
POST |
/monitor/pause |
token | Emergency kill switch — pause all auto-fix actions |
POST |
/monitor/resume |
token | Resume auto-fix actions after a pause |
GET |
/context |
token | View recent shared context bus entries across all agents |
GET |
/skills |
token | List all skills with routing rules and metadata |
GET |
/skills/{name} |
token | Get skill detail (prompt, tools, routing, versions) |
PUT |
/admin/skills/{name} |
admin | Edit skill (prompt, tools, routing rules) |
DELETE |
/admin/skills/{name} |
admin | Delete a skill |
POST |
/admin/skills/{name}/clone |
admin | Clone a skill with a new name |
POST |
/admin/skills/test |
token | Test routing — returns which skill matches a given query |
GET |
/admin/skills/{name}/versions |
token | Version history for a skill |
GET |
/admin/skills/{name}/diff |
token | Diff between two skill versions |
POST |
/admin/mcp/toolsets |
admin | Toggle MCP toolsets on/off |
GET |
/components |
token | Component registry — list supported component kinds with schemas |
Authentication: Token-authenticated backend endpoints accept Authorization: Bearer <token> header or ?token=<token> query parameter. The configured token is PULSE_AGENT_WS_TOKEN. REST query tokens are deprecated; use the Authorization header at the proxy/backend. Missing/invalid tokens return 401; an unconfigured server may return 503. Admin endpoints additionally validate user identity/allowlist.
{
"protocol": "2",
"agent": "2.7.1",
"tools": 138,
"features": ["component_specs", "ws_token_auth", "rate_limiting", "monitor", "fix_history", "predictions"]
}The agent version is read dynamically from the installed package metadata. The tools value is reported by the backend; do not use a release-era hardcoded count.
{
"status": "ok",
"circuit_breaker": {
"state": "closed",
"failure_count": 0,
"recovery_timeout": 60
},
"errors": {
"total": 0,
"by_category": {},
"recent": []
},
"investigations": {},
"autofix_paused": false
}{
"sre": [
{"name": "list_pods", "description": "...", "requires_confirmation": false},
{"name": "delete_pod", "description": "...", "requires_confirmation": true}
],
"security": [
{"name": "scan_pod_security", "description": "...", "requires_confirmation": false}
],
"write_tools": ["apply_yaml", "cordon_node", "delete_pod", "..."]
}Only two WebSocket routes are registered by the agent today — /ws/sre and
/ws/security as separate top-level routes no longer exist. All chat
traffic goes through /ws/agent, which classifies each message and routes
it to the appropriate skill internally (7 skills, not just SRE/Security —
see pulse-agent's API_CONTRACT.md and CLAUDE.md for the full ORCA
routing design).
| Path | Auth | Description |
|---|---|---|
/ws/agent?token=... |
token | Auto-routing orchestrated agent — classifies intent per message and routes to the matching skill (sre, security, view_designer, capacity_planner, plan_builder, postmortem, slo_management) |
/ws/monitor?token=... |
token | Autonomous cluster monitoring (Protocol v2) |
At the backend, WebSocket endpoints validate the shared token in the token query parameter and reject failures with code 4001. Browser clients connect to /api/agent/ws/agent or /api/agent/ws/monitor without a query token; nginx/rspack inject it. Never copy shared tokens into browser code.
{
"type": "message",
"content": "Why are pods crash-looping in production?",
"context": {
"kind": "Deployment",
"name": "api-server",
"namespace": "production",
"gvr": "apps~v1~deployments"
},
"fleet": false
}| Field | Type | Required | Description |
|---|---|---|---|
type |
"message" |
yes | |
content |
string |
yes | User's message text |
context |
ResourceContext |
no | Resource the user is viewing |
fleet |
boolean |
no | Enable fleet/multi-cluster mode |
| Field | Type | Required | Description |
|---|---|---|---|
kind |
string |
yes | K8s resource kind (e.g., "Deployment") |
name |
string |
yes | Resource name |
namespace |
string |
no | Resource namespace (omit for cluster-scoped) |
gvr |
string |
no | GVR key (group~version~plural) |
{
"type": "confirm_response",
"approved": true,
"nonce": "abc123..."
}| Field | Type | Required | Description |
|---|---|---|---|
type |
"confirm_response" |
yes | |
approved |
boolean |
yes | Whether the user approved the action |
nonce |
string |
yes | Must match the nonce from confirm_request (replay prevention) |
{
"type": "clear"
}{
"type": "text_delta",
"text": "The pods are crash-looping because"
}{
"type": "thinking_delta",
"thinking": "Let me check the pod logs first..."
}{
"type": "tool_use",
"tool": "get_pod_logs"
}{
"type": "component",
"tool": "list_pods",
"spec": {
"kind": "data_table",
"title": "Pods in production",
"columns": [
{"id": "name", "header": "Name"},
{"id": "status", "header": "Status"}
],
"rows": [
{"name": "api-server-abc", "status": "Running"}
]
}
}See Component Specs for all spec.kind values.
{
"type": "confirm_request",
"tool": "delete_resource",
"input": {"kind": "Pod", "name": "my-pod", "namespace": "default"},
"nonce": "abc123..."
}| Field | Type | Description |
|---|---|---|
tool |
string |
Tool name requiring confirmation |
input |
object |
Tool input parameters (shown to user) |
nonce |
string |
JIT nonce for replay prevention — client must echo this back |
{
"type": "done",
"full_response": "The pods are crash-looping because..."
}{
"type": "error",
"message": "Rate limited. Max 10 messages per minute."
}{
"type": "cleared"
}Sent as the first message after connecting to /ws/monitor. Configures the monitoring session.
{
"type": "subscribe_monitor",
"trustLevel": 1,
"autoFixCategories": ["crash_loop", "resource_pressure"]
}| Field | Type | Required | Description |
|---|---|---|---|
type |
"subscribe_monitor" |
yes | |
trustLevel |
integer |
no | Browser preference (0-4), clamped by the server. Current monitor effective trust uses configured server trust as a floor; this field does not lower it. Default: 1 |
autoFixCategories |
string[] |
no | Browser category preference. Current backend seeds its effective set with all handlers; this field is not a restrictive server allowlist |
{
"type": "trigger_scan"
}Triggers an immediate cluster scan. If a scan is already in progress, returns an error. Results are pushed as finding and monitor_status events.
{
"type": "action_response",
"actionId": "abc123",
"approved": true
}| Field | Type | Required | Description |
|---|---|---|---|
type |
"action_response" |
yes | |
actionId |
string |
yes | ID of the proposed action |
approved |
boolean |
yes | Whether the user approved the action |
{
"type": "get_fix_history",
"page": 1,
"filters": {"status": "applied", "category": "crash_loop"}
}| Field | Type | Required | Description |
|---|---|---|---|
type |
"get_fix_history" |
yes | |
page |
integer |
no | Page number (default: 1) |
filters |
object |
no | Optional filters (status, category, since, search) |
{
"type": "finding",
"id": "f-abc123",
"severity": "warning",
"category": "crash_loop",
"resource": {"kind": "Pod", "name": "api-server-xyz", "namespace": "production"},
"summary": "Pod crash-looping: CrashLoopBackOff (5 restarts in 10m)",
"details": "...",
"timestamp": 1711540800
}{
"type": "prediction",
"id": "p-abc123",
"category": "resource_pressure",
"resource": {"kind": "Node", "name": "worker-03"},
"summary": "Node memory predicted to exceed 90% within 2 hours",
"confidence": 0.87,
"horizon": "2h",
"timestamp": 1711540800
}{
"type": "action_report",
"actionId": "a-abc123",
"findingId": "f-abc123",
"action": "restart_pod",
"status": "applied",
"summary": "Restarted pod api-server-xyz",
"before": {},
"after": {},
"timestamp": 1711540800
}action_report may include optional verification fields once post-fix verification completes:
verificationStatus:"verified"|"still_failing"verificationEvidence:stringverificationTimestamp:number
{
"type": "investigation_report",
"id": "i-abc123",
"findingId": "f-abc123",
"category": "crashloop",
"status": "completed",
"summary": "Crashloop due to missing ConfigMap key",
"suspectedCause": "ConfigMap key removed in recent rollout",
"recommendedFix": "Restore key and restart deployment",
"confidence": 0.82,
"timestamp": 1711540800
}{
"type": "verification_report",
"id": "v-abc123",
"actionId": "a-abc123",
"findingId": "f-abc123",
"status": "verified",
"evidence": "No active crashloop findings for affected resources",
"timestamp": 1711540800
}Sent after each scan cycle. Contains the IDs of all currently active findings. The UI removes any locally-held findings whose IDs are not in activeIds, preventing stale entries from accumulating after issues are resolved.
{
"type": "findings_snapshot",
"activeIds": ["f-abc123", "f-def456"],
"timestamp": 1711540800
}| Field | Type | Description |
|---|---|---|
activeIds |
string[] |
IDs of all findings that are still active |
timestamp |
number |
Unix timestamp of the snapshot |
{
"type": "monitor_status",
"activeWatches": ["crashloop", "pending", "workloads", "nodes", "cert_expiry", "alerts", "oom", "image_pull", "operators", "daemonsets", "hpa"],
"lastScan": 1711540800,
"findingsCount": 3,
"nextScan": 1711540860
}{
"type": "fix_history",
"items": [],
"total": 0,
"page": 1,
"pageSize": 20
}{
"type": "error",
"message": "Rate limited. Max 10 messages per minute."
}/ws/agent is the only chat endpoint (see note above — /ws/sre//ws/security were removed). Each incoming message is classified by the ORCA skill selector and routed internally to one of 7 skills (sre, security, view_designer, capacity_planner, plan_builder, postmortem, slo_management) with that skill's own system prompt and tool set. This routing is entirely server-side and transparent to the message/event schema documented above — the UI does not need to know which skill handled a given turn.
Structured UI components returned by agent tools via the component event. The UI renders these inline in the chat.
There are 25 component kinds total (GET /components returns the full
registry with schemas). The 9 below were the original set; the rest were
added since and are only summarized here — see the live /components
response or pulse-agent's component_registry.py for authoritative field
lists.
kind |
Description | Key Fields |
|---|---|---|
data_table |
Sortable table | columns[], rows[] |
info_card_grid |
Metric cards | cards[]{label, value, sub?} |
badge_list |
Colored badges | badges[]{text, variant} |
status_list |
Health status items | items[]{name, status, detail?} |
key_value |
Key-value pairs | pairs[]{key, value} |
chart |
Time-series chart | series[]{label, data[][], color?} |
tabs |
Tabbed content | tabs[]{label, content: ComponentSpec} |
grid |
Grid layout | columns, items: ComponentSpec[] |
section |
Titled section | title, content: ComponentSpec |
relationship_tree |
Resource ownership tree | nodes[], rootId |
log_viewer |
Searchable log lines | lines[]{message, level?, timestamp?} |
yaml_viewer |
Read-only YAML display | content, language? |
metric_card |
Single KPI with optional sparkline | title, value, query?, thresholds? |
node_map |
Cluster node health map | nodes[], pods? |
bar_list |
Ranked horizontal bars | items[]{label, value} |
progress_list |
Utilization progress bars | items[]{label, value, max} |
stat_card |
Single big stat with trend | title, value, trend? |
timeline |
Correlated event lanes | lanes[], correlations? |
resource_counts |
Clickable namespace resource summary | items[]{resource, count, gvr?} |
topology |
Dependency graph visualization | nodes[], edges[], perspective |
action_button |
Executes a tool from a component | label, tool, input |
confidence_badge |
Confidence score indicator | score, label? |
resolution_tracker |
Fix verification progress | status, steps[] |
blast_radius |
Impact-scope visualization | resources[], severity |
status_pipeline |
Multi-stage status indicator | stages[]{label, status} |
success | warning | error | info | default
healthy | warning | error | pending | unknown
| Constraint | Value | Enforced By |
|---|---|---|
| Max message size | 1 MB | Agent |
| Rate limit | 10 messages/minute per connection | Agent |
| Confirmation timeout | 120 seconds | Agent |
| Pending confirmation TTL | 5 minutes | Agent |
| Context field validation | ^[a-zA-Z0-9\-._/: ]{0,253}$ |
Agent |
| Chat reconnect intent | five retries with linear delay + jitter; explicit disconnect cancels retries | UI |
| Monitor reconnect | indefinite exponential backoff; hidden-tab delay | UI |
The UI sends a GET /version request before connecting. If the agent's protocol field is outside the UI's SUPPORTED_PROTOCOLS (1, 2), the UI shows a warning but still connects (graceful degradation).
| Version | Changes | UI Version | Agent Version |
|---|---|---|---|
2 |
/ws/monitor for autonomous scanning, /ws/agent for auto-routing orchestration, subscribe_monitor / trigger_scan / action_response / get_fix_history client messages, finding / prediction / action_report / investigation_report / verification_report / findings_snapshot / monitor_status server events, fix history / predictions / memory / context REST endpoints, monitor pause/resume, nonce-based confirmation replay prevention |
v5.12.0+ | v1.4.0+ |
1 |
Initial protocol: text/thinking streaming, tool use, components, confirmations | v5.0.0+ | v1.0.0+ |
These entries record older releases; they are not current compatibility guarantees. Check /version, capabilities, and end-to-end contract tests for the revisions being deployed.
Starting with the org move to PulseSRE and the rename to
pulse-ui, the UI's version numbering was reset from thev6.xline to a freshv2.xline. Both repos now share the same version number for each release.
| UI Version | Agent Version | Protocol | Status |
|---|---|---|---|
| v2.7.1 | v2.7.1 | 2 | Historical |
| v6.2.0 | v2.3.0 | 2 | Compatible (pre versioning reset) |
| v6.1.0 | v2.2.0 | 2 | Compatible |
| v6.0.0 | v2.0.0 | 2 | Compatible |
| v5.21.0 | v1.16.0 | 2 | Compatible |
| v5.13.0 | v1.5.0 | 2 | Compatible |
| v5.12.0 | v1.4.0 | 2 | Compatible |
| v5.11.0 | v1.3.0 | 1 | Compatible |
| v5.10.0 | v1.3.0 | 1 | Compatible |
| v5.8.0 | v1.2.0 | 1 | Compatible |
| v5.0.0-v5.7.0 | v1.0.0-v1.1.0 | 1 | Compatible |
Both repos should tag releases together when protocol changes occur. A shared protocol number alone does not guarantee endpoint, capability, or semantic compatibility; validate the pair being deployed.
REST helpers also cover inbox/episode lifecycle, topology/blast radius, analytics, plans, custom view persistence/sharing, skill governance, MCP connections, memory, and SLOs. See their modules under src/kubeview/engine/ and the backend route inventory rather than assuming the overview table above is exhaustive. Monitor events include resolution, scan reports, investigation progress, inbox lifecycle and skill activity; chat includes feedback acknowledgements, view updates, session expiry, and multi-skill events. Canonical discriminated unions are in the two client modules.
Action status can include expired; verification can include pending and verified_then_recurred. Do not render an expired proposal as executed or a recurred fix as lasting success. REST rollback and approve helpers parse backend refusal messages, including conditions no longer applicable.
GET /api/agent/readiness proxies the authenticated agent /readiness endpoint. The onboarding/readiness page presents its read-only installation observations separately from the existing cluster checklist. HTTP 401/403, unavailable endpoints and malformed/incomplete reports establish no readiness; the panel shows an actionable error and a re-check button.
The response has status (healthy, degraded, unknown), checked_at (UTC ISO timestamp), scope, checks, and limitations. Each check has id, status (healthy, unhealthy, unknown), message, remediation (possibly empty), and source. Required IDs are provider_configuration, provider_connectivity, database, kubernetes_pods, kubernetes_deployments, kubernetes_nodes, kubernetes_events, kubernetes_logs, and monitor. Overall degraded means at least one unhealthy check; otherwise any unknown check makes overall status unknown.
The agent caches observations for 60 seconds, retaining the original timestamp. Provider connectivity is currently unknown because this diagnostic sends no model request. Kubernetes checks use installation credentials and do not establish a browser user's authority. List/permission probes, database health and a running monitor do not establish full scanner, inference, schema or incident-recovery readiness. A failed refresh labels retained data as a previous report.