Evidence status last rechecked: 2026-08-02. Host discovery is an external contract, so this page records evidence rather than promising identical behavior forever.
| Host | Default zstack root | Evidence | Symlink support |
|---|---|---|---|
| Claude Code | ~/.claude/skills |
Official skills docs | Explicitly documented |
| Codex | ~/.agents/skills |
Official skills docs | Explicitly documented |
| Grok Build | ~/.agents/skills plus Claude compatibility |
grok inspect --json, Grok Build 0.2.118 |
Verified with zstack directory symlinks |
~/.codex/skills is not in the current Codex discovery list. zstack retains it
as an explicit legacy profile only:
./scripts/setup.sh --hosts codexGrok already discovers the default Agent Skills and Claude roots. Install its dedicated root only when a local setup requires it:
./scripts/setup.sh --hosts grokExplicit optional profiles remain untouched by later default setup runs.
| Layer | What it proves | Latest evidence |
|---|---|---|
| Source validation | Portable frontmatter, links, catalog, eval coverage | ./scripts/doctor.sh --source-only |
| Install integration | Safe symlink setup, optional-profile semantics, rollback, unlink | ./scripts/test.sh |
| Runtime discovery | A real host lists all five installed zstack skills | Grok Build 0.2.118 inspect |
| Routing behavior | Descriptions select the intended skill, reject unrelated work, survive context pressure, and route Chinese requests | Codex routing run, routing-budget regression, and Chinese routing regression; Grok full 25-case matrix |
| Complete behavior | A selected skill produces a decision-ready artifact from current source | Codex's latest reviewed reruns cover market, PRD, SEO, discovery, and landing response plus artifact. Grok retains historical evidence; fresh behavior runs are blocked before generation by HTTP 401. Current counts are checked in the block below. |
Current-source evidence: Claude behavior 0/20, Codex behavior 20/20, Grok behavior 0/20; Claude routing 0/25, Codex routing 25/25, Grok routing 25/25.
The Grok runtime inspection found all five skills from
~/.agents/skills/<name>/SKILL.md. Historical clean, isolated, no-web Grok runs
covered all five happy paths on their recorded source snapshots: market
validation and SEO passed directly; discovery exposed
an artifact-budget failure and passed after the fix; PRD exposed missing P0-NFR
validation methods and passed after the fix; landing smoke exposed premature
conversion events, deferred deployable output, disconnected UTM/subscriber
attribution, and unverified success UI before passing. This is evidence of the
recorded host/model runs—not a guarantee of complete parity across future
versions. The landing archive records the model transition from
grok-4.5-build to grok-4.5 after the former became unavailable.
A clean Codex CLI 0.146.0 / gpt-5.6-sol regression then exposed two PRD
output-contract failures: prose-only evidence handoff and an overlong core.
With the same prompt and controls, the final run passed the fixed seven-field
interface and reduced the PRD from 6,907 to 4,755 words while retaining every
P0 acceptance criterion.
A second Codex matrix applied the same output-contract test to market, discovery, landing, and SEO. All four initially preserved domain reasoning but failed the fixed Evidence handoff interface; identical post-fix runs passed 20/20 criteria with canonical evidence classes and unchanged evidence limits.
A third Codex matrix measured frontmatter budget rather than artifact quality. The original descriptions, mechanically truncated 180-character prefixes, and compacted descriptions each routed 15/15 positive, boundary, and no-match requests correctly. The compact set front-loads use conditions and reduces the five active descriptions from 2,001 to 1,355 characters (32.3%).
The compact descriptions then passed 10/10 Chinese routing requests: one positive case per skill, three close validation boundaries, and two unrelated engineering no-matches. This closes the previous all-English routing-manifest gap without adding keyword lists back to frontmatter.
The repository's deterministic routing runner and atomic archiver were then exercised end to end on a live Chinese boundary case. The retained runner smoke archive records stdin-only prompt transport, read-only/never-approval controls, the raw prompt and response, sanitized runner metadata, and both integrity manifests; private CLI logs corroborated the published controls and retained prompt/response bytes, then were inspected for incidents but not archived.
The routing runner and archiver now share one three-host core. Deterministic
fake-host tests cover Claude and Grok tool exclusion, requested-versus-actual
model provenance, private-record exclusion, and API-error fail-closed behavior.
A real Grok grok-4.5-build matrix then passed all 25 current English, Chinese,
boundary, and no-match requests with exact two-line outputs. The first uniform
run produced 23 passes but had two provider executions end without final text;
raising the Grok routing turn ceiling from one to two and rerunning the entire
matrix produced 25/25 complete executions, all actually finishing in one turn.
The archived full matrix
preserves prompt rotation, empty-tool controls, actual model, turns, stops, and
private startup incidents through corroborated metadata and exact hashes.
The post-selection Codex runner and reviewed archiver were exercised on the
market-tools-blocked resilience case. Its
smoke archive
retains the exact four-file skill snapshot, generated prompt, final response,
criterion-by-criterion review, execution controls, and both integrity
manifests. The live run passed and moved current complete-provenance behavioral
coverage from 9/20 to 10/20 cases.
Two subsequent reviewed matrices reran the remaining ten cases against exact
current skill snapshots. Matrix A
covered competition scoring, fabricated interviews, mixed traffic, thin PRD
input, and missing keyword tools. Matrix B
covered sparse debriefs, fake proof, all-P0 scope pressure, black-hat SEO, and
an evidence-rich market report. Both passed 5/5 after criterion-by-criterion
review, bringing every behavioral case at that source snapshot to a
complete-provenance pass (20/20). The release-facing command is now
scripts/report-evidence-coverage.rb --require-current-complete-pass, which
also rejects those passes after relevant source or case bytes change.
A host-specific gap audit then reran the four happy paths that still lacked
complete Codex provenance. Discovery, PRD, and SEO passed; landing correctly
failed because the response-only runner prohibited the file that its criterion
required. The failure is retained in the
gap archive.
Behavioral cases now declare response or artifact execution explicitly. The
artifact runner grants an isolated workspace-write sandbox, accepts files only
under artifacts/, verifies the skill snapshot remained byte-identical, rejects
private workspace paths in the final response, and hashes every deliverable.
The reviewed landing artifact rerun
retains eight deployable/test-plan files and passes all criteria. At that
August 1 source snapshot, Codex had a complete-provenance pass for 20/20 cases,
not only any-host coverage.
The isolated behavioral protocol now also has a Claude Code adapter sharing the
same snapshot, anti-leak, review, and atomic-publication core. Integration tests
exercise safe mode, the Read-only tool allowlist, requested-versus-actual model
provenance, private JSON/invocation exclusion, invocation tamper rejection, and
the host's unusual is_error=true result shape. These deterministic adapter
tests are not Claude behavior evidence. At the 2026-08-02 recheck, the latest
local Claude probe reached the configured provider but received HTTP 404 before
generation for every configured model family. The runner reports this safely as
model-unavailable without echoing provider error text. This remains an
explicit external configuration gap rather than being mislabeled as skill
behavior evidence; the marked block above carries the live count.
The same core now has a Grok adapter with an isolated ephemeral host home,
anonymous-descriptor prompt-file transport, fixed system-prompt override, read-only/workspace
sandboxes, tool allowlists, and requested-versus-actual model records. The
adapter explicitly discloses that Grok may still discover the standard Agent
Skills hub, and it records MCP startup attempts as incidents while excluding
MCP tools. The first real reviewed
isolated-runner regression
passed market-tools-blocked with complete provenance. The reviewed
adversarial matrix,
response-gap run,
composite fix rerun,
and output-contract runs
before
and after
preserve both the failures that triggered fixes and their passing reruns. At
that recorded source snapshot, the earlier artifact and happy-path evidence
gave Grok a complete 20/20 host matrix. Later landing-integrity hardening changed
the source bytes for four cases, so that historical result no longer satisfies
their current-source gate. Deterministic fake-CLI response, artifact, error, and
tamper tests are not counted as model evidence.
The latest release gate no longer accepts a complete historical pass after the
underlying skill bytes change. A retained
provider-schema failure
showed that a plausible-looking Buttondown tag-ID mapping contradicted the
official subscriber-create tag-name contract
checked on 2026-08-01. The source now requires dated official
endpoint/auth/status/payload semantics or launch-blocking VERIFY_*
placeholders. After progressive-disclosure edits changed every skill body,
Codex reran and passed all 20 registered cases against the exact new snapshots:
market, PRD, SEO, and discovery passed 4/4 each; landing passed three response
cases plus its real-file artifact case. The latest artifact's Function passed
18 deterministic adversarial scenarios covering
media type, malformed and non-object JSON, actual streamed-byte limits, hostile
stream errors, provider timeout/network/rejection/acceptance, attribution, and
diagnostic privacy. Separate browser tests proved that query and fragment data
are removed before analytics setup, only allow-listed attribution survives, and
unresolved Plausible values make no external request. Buttondown and Plausible
remain explicitly blocked until their current official contracts, real records,
and network behavior are verified.
At the 2026-08-02 recheck, Grok had 0/20 current-source behavior cases because all five skill snapshots had changed; all 25 routing cases remained current because their descriptions and routing inputs did not change. Grok's historical 20/20 matrix and prior client-controlled success-parameter failure remain retained. Fresh behavior probes with Grok CLI 0.2.118 were not archived because the local CLI returned HTTP 401 before model generation. Every case carries an exact seven-row Evidence handoff; the archive validates its bytes mechanically and binds a separate semantic review of evidence-class applicability. CI enforces the same current-source rule for both behavior and all 25 routing cases, so the current Grok gap fails closed rather than borrowing an older pass. Grok's structured adapter also uses an exact final-response marker: pre-marker narration remains private and is reported only as a discarded byte count. Missing/duplicate markers fail, the archiver independently re-extracts the response, and API failures are reduced to a safe category plus HTTP status without printing the provider message.
- Run
./scripts/setup.shand./scripts/doctor.sh. - Inspect the host's skill list or runtime configuration.
- Re-run at least one routing boundary and one complete behavior case in a fresh session.
- Record the host version, model, prompt, raw output, grade, and limitations.
Discovery proves that a host can load the files. It does not replace behavioral evaluation, and one passing model version does not guarantee every future host.