M6 status: complete on 2026-07-04.
CommandAgent is generalized within scope S when the same profile, evidence, and terminal-state mechanisms handle the covered task families without depending on a single prompt, corpus case, or Next.js game shape.
Scope S is explicit:
| dimension | in scope |
|---|---|
| Profiles | nextjs, python-cli, and generic, including the no-profile start path, static/reduced markers, and known-manifest promotion |
| Scenarios | GAME, TOOL, CONTENT, CLI, and AMBIGUOUS from uat/scenarios.md |
| Models | UAT evidence from qwen3.6:27b-coding-nvfp4 main execution with the configured planner model used by the runbook |
| Languages | Japanese scenario prompts, TypeScript/TSX Next.js output, Python CLI output, and English/Japanese diagnostic text |
| OS | macOS/Darwin UAT hosts only |
Within S, "generalized" means:
- Contract inference maps task intent to capabilities and evidence without requiring scenario IDs or one prompt string.
- Runtime acceptance records full success only when the required profile, evidence, and release gate pass.
- Generic app-intent goals bind a minimal static contract
(
user_input_handler_evidence,stateful_update_evidence,visible_interactive_surface_evidence) and render static-assurance markers only when that contract is verified from source evidence. Generic goals without app intent keep the reduced-assurance empty-contract path. - No-profile app-intent runs start as
generic, bind the generic static contract, then may promote to a known profile after a recognized workspace manifest appears. Assurance has three honest tiers: reduced for empty or unsupported generic contracts, static for verified generic source evidence, and full only after a known promoted profile passes its profile and behavioral release gates. - Profile promotion is intentionally table-limited. The current promotion table
covers known manifests only, such as Next.js package manifests and
python-clipyproject.tomlworkspaces. Unknown stacks terminate at the generic static tier honestly; that is a correct terminal state, not a hidden failure. - Missing probe evidence, unsupported profile confidence, incomplete behavior evidence, or generic goals outside the minimal static contract render reduced-assurance markers instead of full success.
- Generic static evidence is source-only. For files outside the comment
stripper's supported extension set, keyword-tier evidence is accepted as
weak_accepted_generic; co-signal absence on those unsupported languages is not a hard failure. This is a stated limit, not behavioral verification. - Generic contract binding emits
generic_contract_boundwith the inferred static evidence keys and the matched application-intent token. - Every observed UAT anomaly is either explained as out of scope or harvested into the corpus before probe, evidence, or profile logic changes.
- The scenario suite is rerun for any probe, evidence, or profile change.
- Any new
DomainProfileor execution pathway must add a row totests/conformance/and passcargo test --test conformance. The conformance matrix is the definition of done for shared runner/profile interface contracts; new rows reuse the named assertions instead of adding pathway-specific copies.
Named guarantees produced by the Generic Assurance Track:
- Monotonic promotion rebind: when a generic app-intent run promotes to a known profile, the promoted contract is a union. Generic interactive requirements remain bound and known-profile requirements are added; requirements never decrease during promotion.
- Earned assurance:
fullassurance is computed from executed gate statuses. A promoted interactive web run cannot earn full assurance from disconnected ornot_applicablebrowser readiness / interaction gates; those gaps fail loudly asacceptance_gates_disconnected. - Authority symmetry: dependency needs created by runtime manifests,
repairs, or promotion are paired with runtime-sanctioned install authority.
Without that authority, the terminal state is an explicit
dependency_setup_authority_requiredfailure, not a repair loop or full success.
The verified nextjs profile scenario matrix is Space Invaders / Breakout / Quiz. Quiz is the non-game row: it contains no game vocabulary while still exercising a browser-interactive Next.js app with start, answer/input, score/progress state, and retry behavior. Any change that injects or broadens evidence vocabulary, contract wording, preset planning knowledge, or similar profile-specific guidance must be admitted only after regression measurement across this three-scenario matrix. The current basis is test0710_bs_006, where Breakout exposed vocabulary-dependent arbitration, and test0711_bs_003, where the first Quiz round reached 2/2 full.
challenge_or_adversary is the general category for opposing elements, time pressure, or other sources of difficulty, while failure_or_collision is the general category for failure conditions; neither category is intrinsically game-specific. Contract inference scopes whether each category is required from the goal, and scenario-specific repair guidance must follow that inferred contract rather than the category name alone.
出自注記: 本バンドの入力12セットは移行前計測に由来し、現リポジトリからの再生成は現在未対応(analysis.md参照)。
The Phase A band measurement covers the post-hardening window from
uat-test0711-bs-003 through uat-test0713-g-001, including
uat-test0713-28-001. The target planner is
qwen3.6:27b-coding-nvfp4; measured executors are
qwen3.6:35b-a3b-coding-nvfp4 and gemma4:31b-cloud. The all-window
denominator is 78 Next.js create records. Of those records, 74 used the target
qwen27 planner and four Space control rows used gemma4:31b-cloud as planner;
the controls are retained because the band rule counts every run in the
measurement window rather than selecting favorable rows.
| scenario family | executor | full | n | full rate | measured characteristic |
|---|---|---|---|---|---|
| Quiz | gemma4:31b-cloud |
12 | 14 | 86% | Most stable row; validates non-game browser-interactive contracts. |
| Quiz | qwen3.6:35b-a3b-coding-nvfp4 |
11 | 12 | 92% | Fully local qwen35 executor is effectively tied with gemma31 here. |
| Breakout | gemma4:31b-cloud |
3 | 6 | 50% n<10 | Middle band; cloud executor improves full recurrence. |
| Breakout | qwen3.6:35b-a3b-coding-nvfp4 |
2 | 11 | 18% | Fully local qwen35 retains a high honest non-full frontier. |
| Space | gemma4:31b-cloud |
2 | 8 | 25% n<10 | Hardest row; full is not zero and cloud executor is the practical follow-up path. |
| Space | qwen3.6:35b-a3b-coding-nvfp4 |
1 | 27 | 4% | Fully local 20GB-class row; restart/recoverable-state evidence is the accepted capability wall. |
gemma4:31b-cloud is a cloud-delivered model. For a fully local 20GB-class
configuration, use the qwen3.6:35b-a3b-coding-nvfp4 rows.
The executor full-rate gradient widens with scenario complexity: Quiz is nearly
tied, Breakout is about 2.7x, and Space is about 6x.
Operational recommendation for complex state-machine scenarios at Space scale:
use the gemma31 executor and recover failed runs with recovery YAML as cloud
follow-up. If a fully local requirement is strict, plan around the 4% Space
band.
Full-run elapsed time across the measured window was 5m02s minimum, 7m02s
median, and 12m53s maximum. Scenario medians were Space 6m25s, Breakout 7m49s,
and Quiz 7m45s. The false-full count across the window is zero: every
full_success row had browser-interaction pass evidence.
These bands are measured facts for this model tier, host class, and
Space/Breakout/Quiz scenario matrix. Future model or mechanism changes must
rerun the same aggregation rather than editing the bands by hand. The source
summary is generated by
band_aggregate.py
and recorded at
band_summary.md.
The data band fixes the planner at qwen3.6:27b-coding-nvfp4, measures the
qwen35 and gemma31 executor families, and uses the deterministic
data/sales.csv fixture with SHA-256
2f6c04e42b0ebdff85a7eb6b52a342610155be6796bd89e5729075d87c78d873.
Recorded goal text assigns each run to aggregation, timeseries, or
unknown; an unknown goal remains in the all-history denominator rather than
being silently excluded.
Window B is the primary, family-specific fixed-code declaration. Aggregation
starts at UAT #7 (B-2i code HEAD 7b177fe), while timeseries starts at UAT #9
(B-2k code HEAD 2028eb4). Each measured family cell has n=6, so all rates
remain explicitly n<10.
| family | Window B start | full | partial+static | failed | n | full rate | median full duration |
|---|---|---|---|---|---|---|---|
| aggregation | uat-test0715-data-007 |
2 | 1 | 3 | 6 | 33% n<10 | 1747s (n=2) |
| timeseries | uat-test0716-data-009 |
0 | 3 | 3 | 6 | 0% n<10 | N/A |
| unknown | no stable threshold | 0 | 0 | 0 | 0 | N/A | N/A |
Window A is the honest all-history reference from data UAT #1 through #9. It contains 60 observed rows: ten invalid or discarded rows stay listed with reasons but outside the denominator, leaving 50 valid runs.
| family | full | partial+static | failed | n | full rate | median full duration |
|---|---|---|---|---|---|---|
| aggregation | 2 | 15 | 21 | 38 | 5% | 1747s (n=2) |
| timeseries | 0 | 7 | 5 | 12 | 0% | N/A |
| unknown | 0 | 0 | 0 | 0 | N/A | N/A |
The two full runs both belong to Window B: UAT #7 Run 3
(data7_gemma31_profile_001) took 2030s, and Run 4
(data7_qwen35_none_001) took 1464s (n=2). In this band, full means only
mechanical integrity backed by E1–E4: pipeline execution, reconciliation,
numeric claim binding, schema checks, and rerun consistency. It does not claim
that the analysis or recommendations are substantively correct; see
data-profile-contract.md §2.
時系列族では、移動平均・比率導出を含むパイプラインの完遂が現行ローカル
ティアの能力の壁である。機械偽陽性はDATA-13 / DATA-7bで除去済みで、
uat-test0716-data-009では両クラスとも再発0を確認した。運用はrecovery
YAMLによるクラウド後詰めを推奨する。時系列族は全史0/12、B-2k後の固定
コード窓でも0/6である。
E2のpercent claim照合は、時系列族がE2 evidenceへ未到達のため未実戦で ある。timeseries初の完走時に実戦確認されるまで、%正規化の能力をこの バンドから推定しない。
The disclosed residual class is inspection write non-follow-through. Across
UAT #7–#9 it recurred six times, including gemma31, and is recorded as
executor-varying model dispersion after the mechanical countermeasures were
exhausted. It is accepted as a band characteristic, not hidden from the
denominator. See the
Data profile first fulls
ledger entry.
The two-window rule uses the family-specific mechanism-stable states, not a
rate-selected cutoff. Window A remains beside it so both the defect-era history
and the initial timeseries measurement stay visible. Honest failed runs remain
in their applicable denominator. uat-test0714-m4-002 is excluded for operator
model-ID substitution, and uat-test0714-m4-004 is excluded because cargo-test
preflight was not green and the campaign was interrupted before four of five
data rows completed. No interrupted outcome is inferred.
Re-measure only by running
python3 workspace/management/scripts/band_aggregate.py --profile data.
Do not hand-edit the generated numbers. The complete per-run ledger,
failure-class distributions, exclusions, duration rows, and false-full
cross-check are in
band_summary_data.md.
The initial fix/nextjs band covers D-1 UAT #1 through #4
(uat-test0717-fix-001–004) with the fix contract v0 fixed, the planner at
qwen3.6:27b-coding-nvfp4, and the qwen35 and gemma31 executor families.
Window A retains all 24 observed runs as the raw history. The declared
denominator is 22: fix2_hook_qwen35_001 and
fix2_hook_qwen35_002 from #1 are held out because inherited
NODE_ENV=production skipped devDependencies before FIX-1 normalized the
bounded child environment. Their failures remain visible in the raw table;
no model outcome is inferred in the declared band.
The numbers below are transcribed from the machine-generated summary. Family
classification is compile_error_fix or contract_hook_fix; an unclassified
goal remains a separate unknown row rather than being reassigned.
| family | executor | raw full | raw failed | raw n | raw full rate | declared full | declared failed | declared n | declared full rate |
|---|---|---|---|---|---|---|---|---|---|
| compile_error_fix | gemma4:31b |
1 | 3 | 4 | 25% | 1 | 3 | 4 | 25% |
| compile_error_fix | qwen3.6:35b-a3b-coding-nvfp4 |
0 | 5 | 5 | 0% | 0 | 5 | 5 | 0% |
| contract_hook_fix | gemma4:31b |
0 | 7 | 7 | 0% | 0 | 7 | 7 | 0% |
| contract_hook_fix | qwen3.6:35b-a3b-coding-nvfp4 |
0 | 8 | 8 | 0% | 0 | 6 | 6 | 0% |
The first fix full is
uat-test0717-fix-001 / fix1_compile_gemma31_001.
Its adjudication evidence records F1 npm run build failure at epoch 1, the
same-lineage F2 success at epoch 2, and both frozen F3 regressions succeeding
at epochs 3 and 4; the complete embedded and standalone chain is anchored by
fix-…-adjudication.json.
Full makes no design-quality claim beyond the fixed
fix intent contract: the initially failing R passed
after repair and every bound regression passed.
For the compile family, diagnostic facts reached executed repair-context prompts after FIX-4a removed verify-step contamination, improving repair reachability without weakening planner lint. For the hook family, the route-bound R suggestion and contract-attribute repair target are wired, but Phase 2 edit completion remains the current local-tier capability wall. Use a cloud executor as the operational follow-up path for these honest non-full terminals.
FIX-5 closed the disclosed profile-invariant repair-target gap: generic
required_path had selected package.json first in 1 of the 22 declared runs,
and the unified fix precedence now selects the diagnosed definition source
first. The historical run remains in the denominator. Window B is based at
FIX-5 HEAD 6decdce; its cells remain empty until the first post-FIX-5 campaign.
The fix spoof-resistance gate was exercised in live runs: #2 rejected two
initially successful or task-irrelevant reproducers as
baseline_not_reproduced, with zero false full. Contract-derived R guidance
removed the observed relevance deviations from #3 onward (0 recurrence).
Lineage mismatch, regression-set shrink, and epoch inversion rejection were
not exercised by these campaigns and are not inferred from the live band.
Re-measure only by running
python3 workspace/management/scripts/band_aggregate.py --profile fix.
Do not hand-edit generated band values. The complete per-run intent ledger,
raw and declared windows, exclusions, F1–F3 false-full cross-check, and the
post-FIX-5 Window B are in
band_summary_fix.md; the
immutable source records are
uat-test0717-fix-001,
002,
003, and
004.
The investigation/data cell uses the profile_synthesis arm by default. Its
fixed investigation contract was measured over two six-run campaigns with no
exclusions: Window A is uat-test0718-inv-001 plus
uat-test0718-inv-002 (12 runs), and Window B is the post-INV-1 campaign
uat-test0718-inv-002 at baseline HEAD 3302dd9 (6 runs).
| cell | arm | window | full | failed | n | full rate |
|---|---|---|---|---|---|---|
| investigate × data | profile_synthesis (default) |
A: inv-001 + inv-002 | 0 | 12 | 12 | 0% |
| investigate × data | profile_synthesis (default) |
B: post-INV-1 inv-002 | 0 | 6 | 6 | 0% |
Family and executor cells are likewise all 0% full: Window A has pipe
gemma4:31b 0/2, pipe qwen3.6:35b-a3b-coding-nvfp4 0/4, schema
gemma4:31b 0/2, and schema qwen3.6:35b-a3b-coding-nvfp4 0/4;
Window B denominators are respectively 1, 2, 1, and 2.
I1, the existence of an executed failing reproducer, was established live in
12/12 runs. The remaining walls were diagnosis completion (8/12 stopped before
I2) and diagnosis honesty (the other 4/12 reached I2 but were rejected as
diagnosis_unbound for quoting non-existent errors or code). The local tier
therefore did not complete a verifiably bound diagnosis; use a cloud executor
as the operational follow-up tier.
The live spoof-resistance record contains 14 rejected code-snippet violations, the requested “14 violations” count. Three rejected error-quote violations bring the all-kind total to 17; every affected run was failed and false success was 0. This distinction is retained so the evidence is not undercounted.
The run-6-style format deviation remains open. Contract-derived guidance reduced proposal/example code blocks from five in inv-001 run 6 to two in inv-002 run 6, but did not eliminate them. Distinguishing quoted existing code from proposal blocks in the binding extractor remains WATCH. Following the E2 lesson, do not relax the binder until more violation originals establish a recurrent rule rather than a single prompt dialect.
Re-measure only by running
python3 workspace/management/scripts/band_aggregate.py --profile investigation.
Do not hand-edit generated band values. The complete family × executor tables,
I1/I2 invariants, per-run ledger, and Window A/B definitions are in
band_summary_investigation.md.
The workflow-circle local arm is declared from uat-test0722-circle-003; the
elevated arm is declared from uat-test0722-circle-elev-008 only.
Its three runs used the complete node configuration and the same UltraPlan
execution mode as the historical intent bands. circle-001 is retained as an
excluded measurement because profile propagation was missing (P1-a failed),
and circle-002 is retained as an excluded measurement because the node
execution mode was missing. Neither invalid campaign contributes a model
outcome to the denominator.
| arm | circle_full | circle_failed | n | full rate |
|---|---|---|---|---|
| local | 0 | 3 | 3 | 0% |
| elevated | 1 | 2 | 3 | 33% |
円環の機械は完全動作(伝播・封じ込め・earned辺・正直終端)。壁は
investigateノードのモデル能力(契約形R構築不能)。elevatedアーム
(ノード別上位モデル、schema v0.1)は1/3の circle_full を実測した。
円環の全区間(earned辺・lineage連続・検証済み診断の照準化・verify_origin
閉門)が実測で機能し、現在の壁はfixノードのモデル修復力である。初circle_fullは
2026-07-22、18秒、全証拠連鎖つき(uat-test0722-circle-elev-008 run1)。
ノードモデル構成が異なるworkflowは同じ分母へ混ぜず、
elevatedアームを別計測列として宣言する。
This 0/3 is an honest local-tier capability price, not a workflow-layer
failure rate: all three runs reached workflow_adjudicated, fired only earned
edges, remained confined to their origin workspace, and terminated
circle_failed(node_failed:investigate) when the investigate contract was not
earned. The machine-generated denominator, exclusions, and evidence paths are
in
band_summary_circle.md.
The verdict meaning remains fixed by the
workflow circle contract; a fix-node full
cannot be projected to circle_full without verify_origin.
The CLI/create band fixes the planner at
qwen3.6:27b-coding-nvfp4. The local reference arm combines
qwen3.6:35b-a3b-coding-nvfp4 and gemma4:31b; the formal elevated Window B
uses gemma4:31b-cloud. The local and cloud prices remain separate because
they are different executor tiers.
| cell | arm / window | executor | C checks reached | C2 | C3 | C4 | full | n | full rate | measured characteristic |
|---|---|---|---|---|---|---|---|---|---|---|
| cli × create | local reference | qwen35 + gemma31 | 0 | not measured | not measured | not measured | 0 | 6 | 0% | 6/6 honest static terminals under the admission-off cap |
| stats | elevated Window B | gemma4:31b-cloud |
0/3 | not reached | not reached | not reached | 0 | 3 | 0% | Three honest pre-C failures; one machine-owned verify canonicalization loss is disclosed below |
| filter | elevated Window B | gemma4:31b-cloud |
2/3 | 2/2 pass | 6/6 false claims rejected | 2/2 pass | 0 | 3 | 0% | Both reached runs built executable CLIs, then failed on README output honesty |
モデルはCLI本体を作れる(formal Window Bの到達2runでC1正常実行 exit 0)が、ドキュメントの誠実性で落ちる(C3がREADME捏造6/6を拒否)。 C2のhelp双方向照合とC4の再実行一致は到達2/2で実戦成立した。現在の 能力壁はモデルのドキュメント忠実性であり、値札はlocal 0/6、cloud Window B 0/6、いずれもfull 0%である。この主張はCへ到達した2runの 観測に限定し、未到達runへ外挿しない。
F-2a settlementは、このGemma基準線からモデル因子単独観測へ到達するまでの
Luna 8窓(n=48)を固定した。001–002はmachine BLOCKED、003–005はtext bridge、
006–008はResponses/nativeである。native 3窓はC到達10/18、C3 pass 6 /
violation 1 / claims_absent 3、full 2/18。007と008はそれぞれfull 1/6・
C3 pass 2を再現した。Gemmaのformal/pack/directive全援助段ではC3 pass 0であり、
証言壁がモデル階級で動く比較が成立した。全窓・費用・protocol別の正準対照表は
mechanism-ledger.mdのF-2a-8
に置き、F-1スコアとT2F設計の第一検算材料とする。
The elevated history preserves the defect-era campaigns without charging them
to Window B. uat-test0725-cli-elev-001 is excluded because C1-not-run was
projected as partial instead of static; elev-002 is excluded because final
acceptance did not dispatch the C runtime. elev-003 is retained as the
post-wiring, pre-calibration predecessor with one reached run. Formal Window B
is elev-004 only: six honest failed terminals, two reached C runs, and four
recorded C evidence files for each reached run.
Window B contains one machine-attributed terminal,
stats_cloud_003: verify canonicalization dropped the positional CSV input
from a command that had already succeeded and reported
verify_command_false_negative. It remains visible in the six-run observed
price; it is not evidence for the documentation-fidelity model wall. The
wall statement comes from the two reached filter runs, both of which passed
C1, C2, and C4 before C3 rejected the generated README claims.
CLI C3 and ingest N2/N3 extend the standing spoof-rejection lineage rather than introducing weaker, profile-specific truth rules:
| gate family | claim under test | observed rejection record |
|---|---|---|
| E2 | data/output numeric claim | claim must bind to executed pipeline output; unsupported normalization is calibrated from originals |
| I2 | diagnosis error/code quotation | 17 fabricated or unbound claims rejected in the investigation band |
| F1–F3 | failing baseline, repaired result, frozen regressions | successful or irrelevant reproducers rejected; lineage and regression set remain frozen |
| C3 | README execution-output claim | 6/6 fabricated CLI output claims rejected against live stdout in elev-004 |
| N2/N3 | ingest record value and frozen candidate accounting | source-absent/altered fields are rejected by N2; N3 rejected two fabricated accepted candidates against a detected set of zero in elev-005 |
Re-measure only by running
python3 workspace/management/scripts/band_aggregate.py --profile cli.
Do not hand-edit the generated numbers. The campaign dispositions, per-run
ledger, reached-run evidence invariant, and Window B denominator are in
band_summary_cli.md.
The fixed verdict meaning remains
docs/cli-profile-contract.md §4.
The ingest/create band fixes the planner at
qwen3.6:27b-coding-nvfp4. The local reference arm uses qwen35 and gemma31;
formal elevated Window B is uat-test0726-ingest-elev-008 with
gemma4:31b-cloud. Historical elev-001 through elev-007 remain visible as
the calibration arc but do not enter either formal denominator.
| cell | arm / window | executor | N checks reached | N2 | N3 | full-equivalent | n | full rate | measured characteristic |
|---|---|---|---|---|---|---|---|---|---|
| ingest × create | local reference | qwen35 + gemma31 | 0/6 | not measured | not measured | 0 | 6 | 0% | Six honest pre-N terminals; no artifact presence was promoted to N evidence |
| list | elevated Window B | gemma4:31b-cloud |
3/3 | 3/3 pass | 3/3 pass | 3 | 3 | 100% | Candidate/date fragments plus the positioned document-year fragment bind 8/3(月) to 2026-08-03 |
| table | elevated Window B | gemma4:31b-cloud |
3/3 | 1/3 pass | 3/3 pass | 1 | 3 | 33.3% | Two runs adopted an empty required date and were honestly rejected by N2 |
| ingest × create | elevated Window B total | gemma4:31b-cloud |
6/6 | 4/6 pass | 6/6 pass | 4 | 6 | 66.7% | Machine-attributed terminals 0/6; all N1–N5 evidence sets present |
自治体イベント整形という実ユースケースで、第1段
(保存済みsnapshot→抽出・整形)が成立した。出力fieldは候補内断片へ束縛され、
部分日付は候補断片と文書共有文脈の両位置を記録した場合だけ補完される。
日本語日付・和暦からISO日付への宣言済み値保存正規化、凍結candidate集合の
detected = accepted + reasoned excluded勘定、sourceにないevent/valueの
拒否は、conformanceだけでなく較正弧とformal Window Bの実入力で作動した。
N3の実弾にはelev-005でdetected 0に対する架空accepted 2件の拒否があり、
N2のformal実弾にはelev-008で空日付2件の拒否がある。
値札はlocal 0/6、elevated 4/6(66.7%)。現在の能力壁はモデルの欠落値規律
であり、table 2/6では日付欠落candidateを理由付き除外せず、required dateを
空値のまま採用した。N3の候補勘定は6/6で成立したが、N2が空値を
source_binding_violationとして拒否したため偽成功は0である。
Re-measure only by running
python3 workspace/management/scripts/band_aggregate.py --profile ingest.
Do not hand-edit the generated values. The complete 54-run ledger, seven
calibration exclusions, N-evidence invariant, and formal denominators are in
band_summary_ingest.md.
The stage-1 verdict meaning remains fixed by
docs/ingest-profile-contract.md; network
fetch and freshness remain stage 2 and are not implied by admission.
CommandAgent intentionally does not adopt declarative adjudication. YAML and TOML describe configuration, manifests, bindings, and registry metadata; verification and adjudication behavior remains executable Rust. Turning the verification implementation itself into declarative data would let a semantic gate change without passing the same code-review, focused-test, fixture, snapshot, clippy, and CI checkpoints as production logic.
This is an intentional divergence from the external static-analysis
recommendation received on 2026-07-27. E-5a centralizes the strings that cross
the protocol boundary and checks them against classes.toml, including
cross-language producers, but it does not move comparator or verdict semantics
into YAML/TOML. In short: YAML is configuration; verification is Rust.
Registry metadata may describe an adjudication class, but cannot implement or
alter the verifier that earns it.
Production-quality implementation outcomes require an implementation model in
the gemini-3.5-flash class or above. Lower-tier models remain safe: the
honest-degradation guarantees still require partial, reduced, failed, or
handoff terminal states instead of false full success. They do not, however,
produce the same full-pass rate.
Measured evidence:
- M5 harvest runs used
gemini-3.1-flash-litefor implementation withgemini-3.5-flashplanning. The harvested distribution was 0 full / 2 partial / 2 failed / 1 not-checked early harvest:test0702_008was an early/not-checked harvest,test0703_002failed on missing completion evidence,test0703_005_4was partial because the probe was unavailable,test0704_001was partial on interaction state-change evidence, andtest0704_003failed on start-transition behavior. - Q1-final used
gemini-3.5-flashas the implementation and planner model on build0604a76band produced 8/8 honest terminal states: 6 full, 1 reasoned partial, and 1 behavioral failed. - Final local-model compatibility round
test0704-999-Q1-62646566676869707172_001usedqwen3.6:27b-coding-nvfp4as the local planner andornith:35bas the local executor. It produced 8/8 honest terminal states and 2/8 full passes: CLI 1/2, web 1/6. The web full pass was CONTENT b, including browser readiness, interaction evidence, persistence, release-gate pass, and full earned assurance. The other 6/8 runs failed with concrete terminal reasons: CLI b failed the Python behavior smoke probe, TOOL a hit a dangerous-command policy boundary, TOOL b hit thetsconfigalias/baseUrl route-closure gap, CONTENT a failed dependency setup lifecycle, GAME a failed multi-file TypeScript coherence, and GAME b failed compile repair follow-through. Verdict: the 27b/35b local pair is below the recommended tier for reliable web delivery. CLI is moderately viable; web full-pass delivery is possible but low-probability on this distribution. The recommended implementation tier remains thegemini-3.5-flashclass or above. - The local-model compatibility track left permanent, model-agnostic assets:
tool-argument path normalization and corrupted-prefix salvage; verifier
command transforms for shell-control splitting,
cd/cwd normalization, and output-pipe stripping; deterministic substitution after repeated verifier policy rejection; deterministic scaffold completion; per-provider timeout calibration; planner-call chokepoint instrumentation; and precise exhaustion classification. The honest scope limit is that dialect coverage was measured against this 27b/35b model pair. New model families may require a new dialect-absorption round; that is expected bounded cost, not a portability claim for unseen output dialects. - Local single-model GAME quality track, instructions 81-87, culminated in
test0707_009on builda9fea8ee, usingqwen3.6:27b-coding-nvfp4for the single local model configuration. The six-run series moved the failure frontier from phase 0, to phase 1, to the final acceptance gate, then to a full pass with behavioral verification. The full-pass run recorded browser readinesspassed, interaction evidencepassed, state changes inaliensRemainingandscore, a primary start transition, and an in-play recovery/restart transition. This is a distinct local data point from the 27b/35b mixed-pair verdict above: single 27b can complete the GAME path after the 81-87 guard sequence, but the observation is narrow to that track and does not revise the broader recommended tier. Residualfits_viewport=false/ bottom overflow 136px is informational presentation quality guidance, not a release gate.
Model-tier observations:
New model-family entries follow the standard order in model-probe.md: model-probe card review, CLI and TOOL smoke checks, then the full scenario round with pre-committed landing criteria. The tier table cites the probe profile as dialect evidence, not as a capability benchmark or an automatic runtime configuration source. For local speed measurements, the recommended zero-risk default is a single Ollama model used for both planner and executor with the model kept resident by Ollama-side keep-alive settings; cloud-hosted providers benefit from prefix stability only when the server exposes compatible prompt caching.
| configuration | sample | distribution / frontier | verdict |
|---|---|---|---|
gemini-3.5-flash implementation/planner |
Q1-final test0704-999-Q1-62_001, 8 runs |
6 full / 1 reasoned partial / 1 behavioral failed, 8/8 honest terminal states | Recommended implementation tier baseline. |
Local mixed pair: qwen3.6:27b-coding-nvfp4 planner + ornith:35b executor |
test0704-999-Q1-62646566676869707172_001, 8 runs |
2/8 full, 8/8 honest terminal states; web 1/6 | Below recommended tier for reliable web delivery; CLI moderately viable. |
Local single model: qwen3.6:27b-coding-nvfp4 |
GAME quality track instructions 81-87, six-run series ending at test0707_009 |
frontier advanced phase0 -> phase1 -> final gate -> full pass with behavioral verification, in-play restart/recovery, score and enemy-state mutation | Golden local single-model GAME reference; narrow positive data point, not a broad tier upgrade. |
Local single model: gemma4:31b-local |
test0708_007 GAME run |
Reached 4/5 phases with a build-passing, playable-looking app, then exposed restart-hook coverage and pending-evidence exhaustion honesty gaps. | Useful local stress datum for observability coverage; not recommended for reliable web delivery until restart/input evidence follow-through recurs as full pass. |
Local MoE single model: qwen3.6:35b-a3b |
test0707_010 / test0708_001 |
Empty planner responses, missing descriptive StepPlan fields, manifest drift, and dual dependency/compile blockers exercised planner and deterministic-repair ladders. | Dialect-heavy model: viable only with the empty-response ladder, schema defaulting, manifest reconciliation, and compile repair ordering guards engaged. |
Cloud single model: gemma4:31b-cloud |
test0708_009 GAME run |
Cross-file hook contract mismatch reached compile repair, but compact zero-edit repair exhausted without regeneration; no last-known-good page snapshot existed for rollback. | Model-family datum for zero-edit repair behavior; motivates regeneration rung and does not change tier recommendation. |
Cloud single model: gemma4:31b-cloud post-97 round |
test0708_016 / test0708_017 GAME runs |
Runtime Bash shell-control at the Bash-tool boundary killed one run before normalization; a later run terminated honestly with pending restart/input capability evidence. | Model-family datum for boundary dialect and honest exhaustion behavior; no tier upgrade. |
Cloud single model: gemma4:31b-cloud Q1 CLI |
cli_b |
First full pass in the CLI domain for this model family, with the generated command-line workflow satisfying the process-profile gates. | Positive domain-specific datum; keep separate from GAME/web reliability because dialect and repair failure modes still recur in app runs. |
Cloud single model: gemma4:31b-cloud post-101 Q1 round |
8-run Q1 round, post-101 | 3/8 full, 8/8 honest terminal states; tool_b repeated model_stagnation:read_only_loop, while game_a exposed camelCase collision evidence and content_b exposed invalid semver feedback. |
Mixed datum: honesty held, but app reliability remains below recommended web tier; install-substitution, semver remedy, and case-insensitive evidence guards are corpus-backed. |
Cloud single model: gemma4:31b-cloud final multi-model round |
final gemma4-cloud round, 2026-07-08, fixture-verified | CLI capable; TOOL unstable; web not recommended. Residuals were multi-grep verify normalization and python-cli scaffold fallback, not live re-round failures. Cloud-hosted models have no identity pinning, so this verdict is valid only with the campaign date and probe card. | Track-closing verdict: use for CLI experiments with review, not as a recommended TOOL/web implementation tier. |
Observability contracts may require only invisible instrumentation: data attributes, state snapshots, and other non-visible probe hooks. Demands on visible design, including control placement, reachability, or UX layout, are preferences offered as tradeoffs. When declined or absent, the runner degrades honestly to unverified or partial evidence instead of forcing design choices.
Two recent applications make this boundary permanent. Instruction 82A' treats restart reachability as evidence honesty: an in-play restart/recovery path can earn behavioral verification, while an overlay-only or unreachable restart is reported as unverified/partial rather than rejected as a required visible layout. Instruction 86D' treats styling as presence-conditional coherence: Tailwind artifacts must be coherent when present, but plain CSS with no Tailwind artifacts is a valid Next.js path and must not be overwritten by forced stack injection.
Q1 final is the standing quality baseline as of 2026-07-05. The judgment rule has two parts:
- Mandatory stranding elimination: every sampled run must terminate honestly with a closed terminal state, concrete status/reason fields, and a recovery handoff when the run is not full. Max-iteration, human-interrupt, absent terminal status, and false-full exits disqualify the round. A run is stranded when its primary termination reason is an iteration or budget label instead of the concrete blocker, such as missing artifacts, policy rejection, compile failure, probe infrastructure failure, or failed behavioral evidence.
- Distribution over samples: after the mandatory condition passes, quality is judged by the distribution over at least two samples per scenario.
Q1-final matrix (test0704-999-Q1-62_001, implementation/planner model
gemini-3.5-flash, build 0604a76b):
| scenario/run | run id | terminal status | primary reason | key telemetry |
|---|---|---|---|---|
| CLI a | 019f318c-9931-7350-a720-c3840f640968 |
full | none | python-cli, contract_origin=initial, external contract OK, browser/interaction not applicable for non-web, assurance full |
| CLI b | 019f3193-227e-7251-ba45-d32d61762737 |
full | none | python-cli, contract_origin=initial, external contract OK, browser/interaction not applicable for non-web, assurance full |
| TOOL a | 019f319b-aef6-7961-8ff6-15a4ad7a253e |
reasoned partial | interaction_unverified:not_evaluated:no_mutation_observed |
browser readiness performed/passed, interaction performed/passed, state dimension todos, persistence_after_reload_reason=no_mutation_observed, recovery prompt/YAML recorded |
| TOOL b | 019f31a1-ff5f-7040-866d-3e405fc43e39 |
full | none | browser readiness performed/passed, interaction performed/passed, state dimension todos, release gate pass, assurance full |
| CONTENT a | 019f31a9-7700-7c12-b947-9597eca40d16 |
full | none | browser readiness performed/passed, interaction performed/passed, state dimension currentContent, token echoed, assurance full |
| CONTENT b | 019f31af-8988-7042-bea6-19f7aae0fccf |
full | none | browser readiness performed/passed, interaction performed/passed, state dimension contentLength, token echoed, assurance full |
| GAME a | 019f31b4-632c-7413-b322-34f0f8db8e91 |
full | none | browser readiness performed/passed, interaction performed/passed, state dimension bulletsCount, primary/restart hooks, assurance full |
| GAME b | 019f31c0-14f5-7c53-898f-4612c883da3f |
behavioral failed | missing_required_evidence:interactive_ui_source_evidence; browser_interaction_failed:probe_script_error |
browser readiness HTTP 200, interaction performed_failed at surface_wait, recovery prompt/YAML recorded; served-DOM inspection showed visible canvas and primary action after a hidden first state instrumentation hook, so this is a harvested probe-calibration limit, not a false full |
Result: Q1-final concludes the quality track for the current scoped host/model pair with 8/8 honest termination and a 6 full / 1 reasoned partial / 1 behavioral failed distribution.
| clause | run evidence | harvested corpus |
|---|---|---|
| Web profile behavior is not Space-Invaders-only; generic interaction, persistence, and content-editing obligations are separately checked. | M2: test0704-464748_001, test0704-464748_002, and test0704-48.1 CONTENT re-run |
Not harvested in this source corpus; the regression target is the scenario suite in uat/scenarios.md. |
| Contract inference survives renamed or opaque scenario IDs and English/Japanese prompt variation. | M3: test0704-49_001, test0704-50_001 |
Covered by required golden tests tests/eval/test_acceptance_contract.py::AcceptanceContractTest and tests/eval/test_completion_contract_snapshots.py::CompletionContractSnapshotTest. |
| No-profile app-intent runs use generic contract binding and terminate honestly at the generic static tier when the scaffolded manifest is unknown. | Static-tier fallback, live-proven: test0704-4030444542434647484814950515354_001 scaffolded a Vite/React manifest, emitted no profile_reinferred, and ended without pretending to run promoted web gates; report: workspace/management/runs/uat-test0704-4030444542434647484814950515354-001/uat-report.md. |
test0704-4030444542434647484814950515354_001 |
| No-profile app-intent runs promote only when a known manifest appears, preserving generic obligations and earning full only through executed web gates. | Promotion path, live-proven: final G3 revalidation test0704-403044454243464748481495051535455565758_000, run id 019f30c7-da99-7d83-a715-1db6b6a6a3b6, recorded profile_reinferred, contract_origin=promoted_union, dependency reconciliation, browser readiness HTTP 200, interaction probe execution, and assurance_level=full; report: workspace/management/runs/uat-test0704-403044454243464748481495051535455565758-000/uat-report.md. |
test0704-403044454243464748481495051535455565758_000 |
| The runner lifecycle supports a non-web process profile without Next.js probe or port assumptions. | M4: test0704-51_001 web run and test0704-51_001 CLI run |
The CLI contract is specified in uat/scenarios.md#cli-python-cli-profile; no app corpus case is expected for the process-only run. |
| App evidence detectors and interaction probe selection are fixture-backed. | M5 Round A four runs | test0702_008, test0703_002, test0703_005_4, test0704_001 |
| The second corpus round confirms the same detectors over a new harvested app snapshot. | M5 Round B | test0704_003 |
Q1-final residuals are corpus-backed: persistence not_evaluated must render a reason, and the GAME b HTTP-200 hidden-first surface failure is recorded as a probe calibration limit. |
Q1-final: test0704-999-Q1-62_001 TOOL a and GAME b |
q1-final-tool-a-persistence-not-evaluated, q1-final-game-b-rendered-hidden-probe-limit |
Local-tier residuals and the first local web full pass are corpus-backed: earlier GAME b records artifact stagnation, earlier TOOL b records output-pipe verifier normalization, final CONTENT b records the complete local web behavioral-gate path, final TOOL b records the tsconfig alias/baseUrl gap, and final GAME a records multi-file TypeScript coherence failure. |
Local rounds: test0704-999-Q1-6264656667686970_001 GAME b / TOOL b, and test0704-999-Q1-62646566676869707172_001 CONTENT b / TOOL b / GAME a |
local-q1-game-b-artifact-stagnation, local-q1-tool-b-output-pipe-verify, local-q1-final-content-b-web-full-pass, local-q1-final-tool-b-tsconfig-alias-gap, local-q1-final-game-a-typescript-coherence |
Local single-model GAME full pass is corpus-backed, including behavioral start, input state mutation, in-play recovery/restart, and informational fits_viewport overflow. |
test0707_009, single-model qwen3.6:27b-coding-nvfp4, run id 019f3c60-f7d3-7f21-8f8c-6b3b0626370a |
local-single-qwen36-game-full-pass |
| Post-97 residuals are corpus-backed: root-anchor path salvage miss, Bash-tool shell-control normalization, and gemma4-cloud honest capability exhaustion. | test0708_013, test0708_016, test0708_017 |
test0708_013, test0708_016, test0708_017 |
| Q1 boundedness residuals are corpus-backed: repeated broad Bash timeout loops and same-run dev-server port conflicts. | Q1-full content_b, Q1-full tool_a |
q1-full-content-b-timeout-loop, q1-full-tool-a-orphan-port-in-use |
| Post-101 residuals are corpus-backed: camelCase/PascalCase gameplay evidence no longer false-negatives, and invalid semver feedback names the manifest entry and corrected example. | Q1 post-101 game_a, content_b |
q1-post101-game-a-camelcase-collision, q1-post101-content-b-invalid-semver |
| Final multi-model residuals are corpus-backed: multi-grep verify shell-control normalization and python-cli setup fallback substitution. | Final gemma4-cloud fixture round | q1-final-multi-grep-shell-control, q1-final-cli-scaffold-fallback |
| Guardrails are permanent and cheap enough to run in CI. | M6 docs and tests | cargo test --test corpus_regression, cargo test --test generality_guardrails, cargo test --test conformance, plus the scenario contract golden tests named below. |
These gates are required for M6 branch protection or equivalent release approval:
cargo test --test corpus_regressioncargo test --test generality_guardrailscargo test --test conformancecargo test generic_ultra_promotes_to_nextjs_after_workspace_manifest --libcargo test generic_ultra_promotes_to_python_cli_after_pyproject_manifest --libcargo test generic_ultra_without_manifest_keeps_static_tier --libpython3 -m unittest tests/eval/test_acceptance_contract.pypython3 -m unittest tests/eval/test_completion_contract_snapshots.pypython3 -m unittest tests/eval/test_false_positive_regression.py
The corpus harness guards harvested probe/evidence/profile behavior. The conformance matrix guards shared interface contracts for profiles and execution pathways. The scenario contract tests are the golden suite for contract inference. The false-positive regression protects the "static screen is not a game success" boundary. The generality guardrails protect static- and reduced-assurance rendering and the Next.js boundary-erosion tripwire.
| milestone | completion date | evidence |
|---|---|---|
| M0 Baseline scope and runbook | 2026-07-04 | Scope S and UAT runbook fixed in this declaration and uat/scenarios.md. |
| M1 Acceptance clauses | 2026-07-04 | Scenario final-acceptance clauses in uat/scenarios.md. |
| M2 Web scenario variation | 2026-07-04 | test0704-464748_001, test0704-464748_002, test0704-48.1 CONTENT re-run. |
| M3 Contract inference | 2026-07-04 | test0704-49_001, test0704-50_001 plus required golden tests. |
| M4 Non-web process profile | 2026-07-04 | test0704-51_001 web and CLI runs. |
| M5 Corpus harvest | 2026-07-04 | Round A four corpus cases plus Round B corpus case. |
| M6 Declaration and guardrails | 2026-07-04 | This document, cross-references, and generality_guardrails tests. |
Status: complete on 2026-07-05.
| milestone | completion date | evidence |
|---|---|---|
| G0 Scope | 2026-07-04 | Generic assurance scope, limits, and named guarantees recorded in Scope S. |
| G1 Generic contract binding | 2026-07-04 | generic_contract_bound event and generic static contract tests. |
| G2 Known-manifest promotion | 2026-07-04 | generic_ultra_promotes_to_nextjs_after_workspace_manifest, generic_ultra_promotes_to_python_cli_after_pyproject_manifest, and generic_ultra_without_manifest_keeps_static_tier. |
| G3 Ambiguous UAT evidence | 2026-07-05 | Two live evidence entries: static-tier fallback test0704-4030444542434647484814950515354_001 and final promoted full-assurance run test0704-403044454243464748481495051535455565758_000. |
| G4 Codification | 2026-07-05 | Default Next.js no-port policy, AMBIGUOUS scenario runbook hardening, G3 corpus harvest, and mechanism-ledger.md#generic-assurance-track-cross-reference. |
Status: Q1 concluded on 2026-07-05.
| milestone | completion date | evidence |
|---|---|---|
| Q1 model-tier baseline | 2026-07-05 | Recommended model tier and M5/Q1-final distributions recorded in Recommended Model Tier. |
| Q1 final round | 2026-07-05 | test0704-999-Q1-62_001: 8 runs, 8/8 honest termination, 6 full / 1 reasoned partial / 1 behavioral failed on gemini-3.5-flash. |
| Q1 residual corpus harvest | 2026-07-05 | TOOL a persistence reason and GAME b rendered-hidden probe-limit fixtures in Clause Evidence. |
| Q1 boundedness closure | 2026-07-05 | Provider-turn and verify-command wall-clock invariants recorded in mechanism-ledger.md#boundedness-guarantees. |
| Local single-model GAME closure | 2026-07-07 | Instructions 81-87 concluded with test0707_009, a full behavioral GAME pass on single-model qwen3.6:27b-coding-nvfp4; corpus fixture local-single-qwen36-game-full-pass. |
| Multi-model generalization track | 2026-07-08 | Instructions 89-102 closed the family-adoption loop: cost curve 9 -> 7 -> ~2 -> ~2 instructions as fixes shifted from family dialects to model-agnostic assets; standing protocol is probe -> smoke x2 (CLI + TOOL) -> full round with pre-committed landing criteria -> tier entry citing the probe card; boundedness is five-dimensional: transport, spawn, planner, step wall clock, and server reaping. |
Optional backlog after Q1:
| optional track | trigger |
|---|---|
tsconfig paths alias deterministic invariant repair |
Start on recurrence of a route-bound @/* import gap or tsconfig baseUrl/paths missing @/* alias terminal reason, using local-q1-final-tool-b-tsconfig-alias-gap as the fixture. |
| Dangerous-command rejection feedback categories | Start after recurrence analysis shows repeated local-model failures at the same policy category instead of isolated blocked-command attempts. |
| CONTENT a dependency lifecycle variance | Start on recurrence of dependency setup lifecycle failure; first check offline/network/package-registry state before treating it as a CommandAgent behavior defect. |
fits_viewport responsive guidance |
Start on recurrence across scenarios. The test0707_009 GAME pass records bottom canvas overflow as informational presentation quality, not a gate. Repeated overflow should become responsive guidance and visual QA, not a hidden release blocker. |
| UX track resumption | Start when instruction 80 is resumed with a visual acceptance protocol. Keep it separate from invisible observability contracts so visible design preferences remain reviewable trades. |
| Data-profile/workflow track | Start when a non-web workflow needs first-class contracts beyond python-cli, including fixture/data lifecycle and acceptance probes. |
| T2 model-tier expansion | Start only after a probe card plus smoke/full-round evidence defines the target tier criteria in advance. |
| Second web profile / web-family layer | Start when a non-Next.js web stack needs first-class promotion, shared web obligations, or release-gate parity. Until then, unknown manifests remain generic/static by design. |
| Linux run | Start before claiming Linux host parity, adding Linux-specific release support, or treating Darwin UAT results as portable. |
| English-goal run | Start before claiming English prompt-distribution parity or using English-goal quality as release/marketing evidence. |
The declaration does not claim:
- A second web framework. The web-family shared layer has not yet been extracted from the Next.js profile boundary.
- Promotion for arbitrary web frameworks. Unknown manifests remain generic and must stop at the static tier unless a known profile boundary is added with tests and corpus evidence.
- Non-process domains outside web apps and CLI/process tools.
- Linux behavior. Current UAT evidence is macOS/Darwin-only.
- Output depth beyond the observed model pair. Depth and polish remain model-bound even when the contract, probe, and terminal-state mechanisms hold.
The admitted profile-synthesis arm is recorded separately from the none arm: none Window A: 24 runs, full 0; profile synthesis Window A/B: 6 runs, full 0. The synthesis arm eliminated the prior mechanical failure classes (0/6); all six runs reached F1 and stopped in local-tier repair read-only stagnation. Cloud-tier execution is recommended. F1 live evidence exists for 30 runs; F2/F3 live evidence exists only for one fix×nextjs run, not for fix×data. FIX-6b reproducer_defect remains WATCH.