Skip to content

Latest commit

 

History

History
714 lines (606 loc) · 55.4 KB

File metadata and controls

714 lines (606 loc) · 55.4 KB

Generality Declaration

M6 status: complete on 2026-07-04.

Scope S

CommandAgent is generalized within scope S when the same profile, evidence, and terminal-state mechanisms handle the covered task families without depending on a single prompt, corpus case, or Next.js game shape.

Scope S is explicit:

dimension in scope
Profiles nextjs, python-cli, and generic, including the no-profile start path, static/reduced markers, and known-manifest promotion
Scenarios GAME, TOOL, CONTENT, CLI, and AMBIGUOUS from uat/scenarios.md
Models UAT evidence from qwen3.6:27b-coding-nvfp4 main execution with the configured planner model used by the runbook
Languages Japanese scenario prompts, TypeScript/TSX Next.js output, Python CLI output, and English/Japanese diagnostic text
OS macOS/Darwin UAT hosts only

Within S, "generalized" means:

  • Contract inference maps task intent to capabilities and evidence without requiring scenario IDs or one prompt string.
  • Runtime acceptance records full success only when the required profile, evidence, and release gate pass.
  • Generic app-intent goals bind a minimal static contract (user_input_handler_evidence, stateful_update_evidence, visible_interactive_surface_evidence) and render static-assurance markers only when that contract is verified from source evidence. Generic goals without app intent keep the reduced-assurance empty-contract path.
  • No-profile app-intent runs start as generic, bind the generic static contract, then may promote to a known profile after a recognized workspace manifest appears. Assurance has three honest tiers: reduced for empty or unsupported generic contracts, static for verified generic source evidence, and full only after a known promoted profile passes its profile and behavioral release gates.
  • Profile promotion is intentionally table-limited. The current promotion table covers known manifests only, such as Next.js package manifests and python-cli pyproject.toml workspaces. Unknown stacks terminate at the generic static tier honestly; that is a correct terminal state, not a hidden failure.
  • Missing probe evidence, unsupported profile confidence, incomplete behavior evidence, or generic goals outside the minimal static contract render reduced-assurance markers instead of full success.
  • Generic static evidence is source-only. For files outside the comment stripper's supported extension set, keyword-tier evidence is accepted as weak_accepted_generic; co-signal absence on those unsupported languages is not a hard failure. This is a stated limit, not behavioral verification.
  • Generic contract binding emits generic_contract_bound with the inferred static evidence keys and the matched application-intent token.
  • Every observed UAT anomaly is either explained as out of scope or harvested into the corpus before probe, evidence, or profile logic changes.
  • The scenario suite is rerun for any probe, evidence, or profile change.
  • Any new DomainProfile or execution pathway must add a row to tests/conformance/ and pass cargo test --test conformance. The conformance matrix is the definition of done for shared runner/profile interface contracts; new rows reuse the named assertions instead of adding pathway-specific copies.

Named guarantees produced by the Generic Assurance Track:

  • Monotonic promotion rebind: when a generic app-intent run promotes to a known profile, the promoted contract is a union. Generic interactive requirements remain bound and known-profile requirements are added; requirements never decrease during promotion.
  • Earned assurance: full assurance is computed from executed gate statuses. A promoted interactive web run cannot earn full assurance from disconnected or not_applicable browser readiness / interaction gates; those gaps fail loudly as acceptance_gates_disconnected.
  • Authority symmetry: dependency needs created by runtime manifests, repairs, or promotion are paired with runtime-sanctioned install authority. Without that authority, the terminal state is an explicit dependency_setup_authority_required failure, not a repair loop or full success.

Scenario matrix (nextjs)

The verified nextjs profile scenario matrix is Space Invaders / Breakout / Quiz. Quiz is the non-game row: it contains no game vocabulary while still exercising a browser-interactive Next.js app with start, answer/input, score/progress state, and retry behavior. Any change that injects or broadens evidence vocabulary, contract wording, preset planning knowledge, or similar profile-specific guidance must be admitted only after regression measurement across this three-scenario matrix. The current basis is test0710_bs_006, where Breakout exposed vocabulary-dependent arbitration, and test0711_bs_003, where the first Quiz round reached 2/2 full.

challenge_or_adversary is the general category for opposing elements, time pressure, or other sources of difficulty, while failure_or_collision is the general category for failure conditions; neither category is intrinsically game-specific. Contract inference scopes whether each category is required from the goal, and scenario-specific repair guidance must follow that inferred contract rather than the category name alone.

Measured capability bands (nextjs × create)

出自注記: 本バンドの入力12セットは移行前計測に由来し、現リポジトリからの再生成は現在未対応(analysis.md参照)。

The Phase A band measurement covers the post-hardening window from uat-test0711-bs-003 through uat-test0713-g-001, including uat-test0713-28-001. The target planner is qwen3.6:27b-coding-nvfp4; measured executors are qwen3.6:35b-a3b-coding-nvfp4 and gemma4:31b-cloud. The all-window denominator is 78 Next.js create records. Of those records, 74 used the target qwen27 planner and four Space control rows used gemma4:31b-cloud as planner; the controls are retained because the band rule counts every run in the measurement window rather than selecting favorable rows.

scenario family executor full n full rate measured characteristic
Quiz gemma4:31b-cloud 12 14 86% Most stable row; validates non-game browser-interactive contracts.
Quiz qwen3.6:35b-a3b-coding-nvfp4 11 12 92% Fully local qwen35 executor is effectively tied with gemma31 here.
Breakout gemma4:31b-cloud 3 6 50% n<10 Middle band; cloud executor improves full recurrence.
Breakout qwen3.6:35b-a3b-coding-nvfp4 2 11 18% Fully local qwen35 retains a high honest non-full frontier.
Space gemma4:31b-cloud 2 8 25% n<10 Hardest row; full is not zero and cloud executor is the practical follow-up path.
Space qwen3.6:35b-a3b-coding-nvfp4 1 27 4% Fully local 20GB-class row; restart/recoverable-state evidence is the accepted capability wall.

gemma4:31b-cloud is a cloud-delivered model. For a fully local 20GB-class configuration, use the qwen3.6:35b-a3b-coding-nvfp4 rows. The executor full-rate gradient widens with scenario complexity: Quiz is nearly tied, Breakout is about 2.7x, and Space is about 6x. Operational recommendation for complex state-machine scenarios at Space scale: use the gemma31 executor and recover failed runs with recovery YAML as cloud follow-up. If a fully local requirement is strict, plan around the 4% Space band.

Full-run elapsed time across the measured window was 5m02s minimum, 7m02s median, and 12m53s maximum. Scenario medians were Space 6m25s, Breakout 7m49s, and Quiz 7m45s. The false-full count across the window is zero: every full_success row had browser-interaction pass evidence.

These bands are measured facts for this model tier, host class, and Space/Breakout/Quiz scenario matrix. Future model or mechanism changes must rerun the same aggregation rather than editing the bands by hand. The source summary is generated by band_aggregate.py and recorded at band_summary.md.

Measured capability bands (data × create)

The data band fixes the planner at qwen3.6:27b-coding-nvfp4, measures the qwen35 and gemma31 executor families, and uses the deterministic data/sales.csv fixture with SHA-256 2f6c04e42b0ebdff85a7eb6b52a342610155be6796bd89e5729075d87c78d873. Recorded goal text assigns each run to aggregation, timeseries, or unknown; an unknown goal remains in the all-history denominator rather than being silently excluded.

Window B is the primary, family-specific fixed-code declaration. Aggregation starts at UAT #7 (B-2i code HEAD 7b177fe), while timeseries starts at UAT #9 (B-2k code HEAD 2028eb4). Each measured family cell has n=6, so all rates remain explicitly n<10.

family Window B start full partial+static failed n full rate median full duration
aggregation uat-test0715-data-007 2 1 3 6 33% n<10 1747s (n=2)
timeseries uat-test0716-data-009 0 3 3 6 0% n<10 N/A
unknown no stable threshold 0 0 0 0 N/A N/A

Window A is the honest all-history reference from data UAT #1 through #9. It contains 60 observed rows: ten invalid or discarded rows stay listed with reasons but outside the denominator, leaving 50 valid runs.

family full partial+static failed n full rate median full duration
aggregation 2 15 21 38 5% 1747s (n=2)
timeseries 0 7 5 12 0% N/A
unknown 0 0 0 0 N/A N/A

The two full runs both belong to Window B: UAT #7 Run 3 (data7_gemma31_profile_001) took 2030s, and Run 4 (data7_qwen35_none_001) took 1464s (n=2). In this band, full means only mechanical integrity backed by E1–E4: pipeline execution, reconciliation, numeric claim binding, schema checks, and rerun consistency. It does not claim that the analysis or recommendations are substantively correct; see data-profile-contract.md §2.

時系列族では、移動平均・比率導出を含むパイプラインの完遂が現行ローカル ティアの能力の壁である。機械偽陽性はDATA-13 / DATA-7bで除去済みで、 uat-test0716-data-009では両クラスとも再発0を確認した。運用はrecovery YAMLによるクラウド後詰めを推奨する。時系列族は全史0/12、B-2k後の固定 コード窓でも0/6である。

E2のpercent claim照合は、時系列族がE2 evidenceへ未到達のため未実戦で ある。timeseries初の完走時に実戦確認されるまで、%正規化の能力をこの バンドから推定しない。

The disclosed residual class is inspection write non-follow-through. Across UAT #7–#9 it recurred six times, including gemma31, and is recorded as executor-varying model dispersion after the mechanical countermeasures were exhausted. It is accepted as a band characteristic, not hidden from the denominator. See the Data profile first fulls ledger entry.

The two-window rule uses the family-specific mechanism-stable states, not a rate-selected cutoff. Window A remains beside it so both the defect-era history and the initial timeseries measurement stay visible. Honest failed runs remain in their applicable denominator. uat-test0714-m4-002 is excluded for operator model-ID substitution, and uat-test0714-m4-004 is excluded because cargo-test preflight was not green and the campaign was interrupted before four of five data rows completed. No interrupted outcome is inferred.

Re-measure only by running python3 workspace/management/scripts/band_aggregate.py --profile data. Do not hand-edit the generated numbers. The complete per-run ledger, failure-class distributions, exclusions, duration rows, and false-full cross-check are in band_summary_data.md.

Measured capability bands (fix × nextjs)

The initial fix/nextjs band covers D-1 UAT #1 through #4 (uat-test0717-fix-001–004) with the fix contract v0 fixed, the planner at qwen3.6:27b-coding-nvfp4, and the qwen35 and gemma31 executor families. Window A retains all 24 observed runs as the raw history. The declared denominator is 22: fix2_hook_qwen35_001 and fix2_hook_qwen35_002 from #1 are held out because inherited NODE_ENV=production skipped devDependencies before FIX-1 normalized the bounded child environment. Their failures remain visible in the raw table; no model outcome is inferred in the declared band.

The numbers below are transcribed from the machine-generated summary. Family classification is compile_error_fix or contract_hook_fix; an unclassified goal remains a separate unknown row rather than being reassigned.

family executor raw full raw failed raw n raw full rate declared full declared failed declared n declared full rate
compile_error_fix gemma4:31b 1 3 4 25% 1 3 4 25%
compile_error_fix qwen3.6:35b-a3b-coding-nvfp4 0 5 5 0% 0 5 5 0%
contract_hook_fix gemma4:31b 0 7 7 0% 0 7 7 0%
contract_hook_fix qwen3.6:35b-a3b-coding-nvfp4 0 8 8 0% 0 6 6 0%

The first fix full is uat-test0717-fix-001 / fix1_compile_gemma31_001. Its adjudication evidence records F1 npm run build failure at epoch 1, the same-lineage F2 success at epoch 2, and both frozen F3 regressions succeeding at epochs 3 and 4; the complete embedded and standalone chain is anchored by fix-…-adjudication.json. Full makes no design-quality claim beyond the fixed fix intent contract: the initially failing R passed after repair and every bound regression passed.

For the compile family, diagnostic facts reached executed repair-context prompts after FIX-4a removed verify-step contamination, improving repair reachability without weakening planner lint. For the hook family, the route-bound R suggestion and contract-attribute repair target are wired, but Phase 2 edit completion remains the current local-tier capability wall. Use a cloud executor as the operational follow-up path for these honest non-full terminals.

FIX-5 closed the disclosed profile-invariant repair-target gap: generic required_path had selected package.json first in 1 of the 22 declared runs, and the unified fix precedence now selects the diagnosed definition source first. The historical run remains in the denominator. Window B is based at FIX-5 HEAD 6decdce; its cells remain empty until the first post-FIX-5 campaign.

The fix spoof-resistance gate was exercised in live runs: #2 rejected two initially successful or task-irrelevant reproducers as baseline_not_reproduced, with zero false full. Contract-derived R guidance removed the observed relevance deviations from #3 onward (0 recurrence). Lineage mismatch, regression-set shrink, and epoch inversion rejection were not exercised by these campaigns and are not inferred from the live band.

Re-measure only by running python3 workspace/management/scripts/band_aggregate.py --profile fix. Do not hand-edit generated band values. The complete per-run intent ledger, raw and declared windows, exclusions, F1–F3 false-full cross-check, and the post-FIX-5 Window B are in band_summary_fix.md; the immutable source records are uat-test0717-fix-001, 002, 003, and 004.

Measured capability bands (investigate × data)

The investigation/data cell uses the profile_synthesis arm by default. Its fixed investigation contract was measured over two six-run campaigns with no exclusions: Window A is uat-test0718-inv-001 plus uat-test0718-inv-002 (12 runs), and Window B is the post-INV-1 campaign uat-test0718-inv-002 at baseline HEAD 3302dd9 (6 runs).

cell arm window full failed n full rate
investigate × data profile_synthesis (default) A: inv-001 + inv-002 0 12 12 0%
investigate × data profile_synthesis (default) B: post-INV-1 inv-002 0 6 6 0%

Family and executor cells are likewise all 0% full: Window A has pipe gemma4:31b 0/2, pipe qwen3.6:35b-a3b-coding-nvfp4 0/4, schema gemma4:31b 0/2, and schema qwen3.6:35b-a3b-coding-nvfp4 0/4; Window B denominators are respectively 1, 2, 1, and 2.

I1, the existence of an executed failing reproducer, was established live in 12/12 runs. The remaining walls were diagnosis completion (8/12 stopped before I2) and diagnosis honesty (the other 4/12 reached I2 but were rejected as diagnosis_unbound for quoting non-existent errors or code). The local tier therefore did not complete a verifiably bound diagnosis; use a cloud executor as the operational follow-up tier.

The live spoof-resistance record contains 14 rejected code-snippet violations, the requested “14 violations” count. Three rejected error-quote violations bring the all-kind total to 17; every affected run was failed and false success was 0. This distinction is retained so the evidence is not undercounted.

The run-6-style format deviation remains open. Contract-derived guidance reduced proposal/example code blocks from five in inv-001 run 6 to two in inv-002 run 6, but did not eliminate them. Distinguishing quoted existing code from proposal blocks in the binding extractor remains WATCH. Following the E2 lesson, do not relax the binder until more violation originals establish a recurrent rule rather than a single prompt dialect.

Re-measure only by running python3 workspace/management/scripts/band_aggregate.py --profile investigation. Do not hand-edit generated band values. The complete family × executor tables, I1/I2 invariants, per-run ledger, and Window A/B definitions are in band_summary_investigation.md.

Measured capability bands (workflow circle)

The workflow-circle local arm is declared from uat-test0722-circle-003; the elevated arm is declared from uat-test0722-circle-elev-008 only. Its three runs used the complete node configuration and the same UltraPlan execution mode as the historical intent bands. circle-001 is retained as an excluded measurement because profile propagation was missing (P1-a failed), and circle-002 is retained as an excluded measurement because the node execution mode was missing. Neither invalid campaign contributes a model outcome to the denominator.

arm circle_full circle_failed n full rate
local 0 3 3 0%
elevated 1 2 3 33%

円環の機械は完全動作(伝播・封じ込め・earned辺・正直終端)。壁は investigateノードのモデル能力(契約形R構築不能)。elevatedアーム (ノード別上位モデル、schema v0.1)は1/3の circle_full を実測した。 円環の全区間(earned辺・lineage連続・検証済み診断の照準化・verify_origin 閉門)が実測で機能し、現在の壁はfixノードのモデル修復力である。初circle_fullは 2026-07-22、18秒、全証拠連鎖つき(uat-test0722-circle-elev-008 run1)。 ノードモデル構成が異なるworkflowは同じ分母へ混ぜず、 elevatedアームを別計測列として宣言する。

This 0/3 is an honest local-tier capability price, not a workflow-layer failure rate: all three runs reached workflow_adjudicated, fired only earned edges, remained confined to their origin workspace, and terminated circle_failed(node_failed:investigate) when the investigate contract was not earned. The machine-generated denominator, exclusions, and evidence paths are in band_summary_circle.md. The verdict meaning remains fixed by the workflow circle contract; a fix-node full cannot be projected to circle_full without verify_origin.

Measured capability bands (cli × create)

The CLI/create band fixes the planner at qwen3.6:27b-coding-nvfp4. The local reference arm combines qwen3.6:35b-a3b-coding-nvfp4 and gemma4:31b; the formal elevated Window B uses gemma4:31b-cloud. The local and cloud prices remain separate because they are different executor tiers.

cell arm / window executor C checks reached C2 C3 C4 full n full rate measured characteristic
cli × create local reference qwen35 + gemma31 0 not measured not measured not measured 0 6 0% 6/6 honest static terminals under the admission-off cap
stats elevated Window B gemma4:31b-cloud 0/3 not reached not reached not reached 0 3 0% Three honest pre-C failures; one machine-owned verify canonicalization loss is disclosed below
filter elevated Window B gemma4:31b-cloud 2/3 2/2 pass 6/6 false claims rejected 2/2 pass 0 3 0% Both reached runs built executable CLIs, then failed on README output honesty

モデルはCLI本体を作れる(formal Window Bの到達2runでC1正常実行 exit 0)が、ドキュメントの誠実性で落ちる(C3がREADME捏造6/6を拒否)。 C2のhelp双方向照合とC4の再実行一致は到達2/2で実戦成立した。現在の 能力壁はモデルのドキュメント忠実性であり、値札はlocal 0/6、cloud Window B 0/6、いずれもfull 0%である。この主張はCへ到達した2runの 観測に限定し、未到達runへ外挿しない。

F-2a settlementは、このGemma基準線からモデル因子単独観測へ到達するまでの Luna 8窓(n=48)を固定した。001–002はmachine BLOCKED、003–005はtext bridge、 006–008はResponses/nativeである。native 3窓はC到達10/18、C3 pass 6 / violation 1 / claims_absent 3、full 2/18。007と008はそれぞれfull 1/6・ C3 pass 2を再現した。Gemmaのformal/pack/directive全援助段ではC3 pass 0であり、 証言壁がモデル階級で動く比較が成立した。全窓・費用・protocol別の正準対照表は mechanism-ledger.mdのF-2a-8 に置き、F-1スコアとT2F設計の第一検算材料とする。

The elevated history preserves the defect-era campaigns without charging them to Window B. uat-test0725-cli-elev-001 is excluded because C1-not-run was projected as partial instead of static; elev-002 is excluded because final acceptance did not dispatch the C runtime. elev-003 is retained as the post-wiring, pre-calibration predecessor with one reached run. Formal Window B is elev-004 only: six honest failed terminals, two reached C runs, and four recorded C evidence files for each reached run.

Window B contains one machine-attributed terminal, stats_cloud_003: verify canonicalization dropped the positional CSV input from a command that had already succeeded and reported verify_command_false_negative. It remains visible in the six-run observed price; it is not evidence for the documentation-fidelity model wall. The wall statement comes from the two reached filter runs, both of which passed C1, C2, and C4 before C3 rejected the generated README claims.

CLI C3 and ingest N2/N3 extend the standing spoof-rejection lineage rather than introducing weaker, profile-specific truth rules:

gate family claim under test observed rejection record
E2 data/output numeric claim claim must bind to executed pipeline output; unsupported normalization is calibrated from originals
I2 diagnosis error/code quotation 17 fabricated or unbound claims rejected in the investigation band
F1–F3 failing baseline, repaired result, frozen regressions successful or irrelevant reproducers rejected; lineage and regression set remain frozen
C3 README execution-output claim 6/6 fabricated CLI output claims rejected against live stdout in elev-004
N2/N3 ingest record value and frozen candidate accounting source-absent/altered fields are rejected by N2; N3 rejected two fabricated accepted candidates against a detected set of zero in elev-005

Re-measure only by running python3 workspace/management/scripts/band_aggregate.py --profile cli. Do not hand-edit the generated numbers. The campaign dispositions, per-run ledger, reached-run evidence invariant, and Window B denominator are in band_summary_cli.md. The fixed verdict meaning remains docs/cli-profile-contract.md §4.

Measured capability bands (ingest × create, stage 1)

The ingest/create band fixes the planner at qwen3.6:27b-coding-nvfp4. The local reference arm uses qwen35 and gemma31; formal elevated Window B is uat-test0726-ingest-elev-008 with gemma4:31b-cloud. Historical elev-001 through elev-007 remain visible as the calibration arc but do not enter either formal denominator.

cell arm / window executor N checks reached N2 N3 full-equivalent n full rate measured characteristic
ingest × create local reference qwen35 + gemma31 0/6 not measured not measured 0 6 0% Six honest pre-N terminals; no artifact presence was promoted to N evidence
list elevated Window B gemma4:31b-cloud 3/3 3/3 pass 3/3 pass 3 3 100% Candidate/date fragments plus the positioned document-year fragment bind 8/3(月) to 2026-08-03
table elevated Window B gemma4:31b-cloud 3/3 1/3 pass 3/3 pass 1 3 33.3% Two runs adopted an empty required date and were honestly rejected by N2
ingest × create elevated Window B total gemma4:31b-cloud 6/6 4/6 pass 6/6 pass 4 6 66.7% Machine-attributed terminals 0/6; all N1–N5 evidence sets present

自治体イベント整形という実ユースケースで、第1段 (保存済みsnapshot→抽出・整形)が成立した。出力fieldは候補内断片へ束縛され、 部分日付は候補断片と文書共有文脈の両位置を記録した場合だけ補完される。 日本語日付・和暦からISO日付への宣言済み値保存正規化、凍結candidate集合の detected = accepted + reasoned excluded勘定、sourceにないevent/valueの 拒否は、conformanceだけでなく較正弧とformal Window Bの実入力で作動した。 N3の実弾にはelev-005でdetected 0に対する架空accepted 2件の拒否があり、 N2のformal実弾にはelev-008で空日付2件の拒否がある。

値札はlocal 0/6、elevated 4/6(66.7%)。現在の能力壁はモデルの欠落値規律 であり、table 2/6では日付欠落candidateを理由付き除外せず、required dateを 空値のまま採用した。N3の候補勘定は6/6で成立したが、N2が空値を source_binding_violationとして拒否したため偽成功は0である。

Re-measure only by running python3 workspace/management/scripts/band_aggregate.py --profile ingest. Do not hand-edit the generated values. The complete 54-run ledger, seven calibration exclusions, N-evidence invariant, and formal denominators are in band_summary_ingest.md. The stage-1 verdict meaning remains fixed by docs/ingest-profile-contract.md; network fetch and freshness remain stage 2 and are not implied by admission.

Why adjudication logic is not declarative

CommandAgent intentionally does not adopt declarative adjudication. YAML and TOML describe configuration, manifests, bindings, and registry metadata; verification and adjudication behavior remains executable Rust. Turning the verification implementation itself into declarative data would let a semantic gate change without passing the same code-review, focused-test, fixture, snapshot, clippy, and CI checkpoints as production logic.

This is an intentional divergence from the external static-analysis recommendation received on 2026-07-27. E-5a centralizes the strings that cross the protocol boundary and checks them against classes.toml, including cross-language producers, but it does not move comparator or verdict semantics into YAML/TOML. In short: YAML is configuration; verification is Rust. Registry metadata may describe an adjudication class, but cannot implement or alter the verifier that earns it.

Recommended Model Tier

Production-quality implementation outcomes require an implementation model in the gemini-3.5-flash class or above. Lower-tier models remain safe: the honest-degradation guarantees still require partial, reduced, failed, or handoff terminal states instead of false full success. They do not, however, produce the same full-pass rate.

Measured evidence:

  • M5 harvest runs used gemini-3.1-flash-lite for implementation with gemini-3.5-flash planning. The harvested distribution was 0 full / 2 partial / 2 failed / 1 not-checked early harvest: test0702_008 was an early/not-checked harvest, test0703_002 failed on missing completion evidence, test0703_005_4 was partial because the probe was unavailable, test0704_001 was partial on interaction state-change evidence, and test0704_003 failed on start-transition behavior.
  • Q1-final used gemini-3.5-flash as the implementation and planner model on build 0604a76b and produced 8/8 honest terminal states: 6 full, 1 reasoned partial, and 1 behavioral failed.
  • Final local-model compatibility round test0704-999-Q1-62646566676869707172_001 used qwen3.6:27b-coding-nvfp4 as the local planner and ornith:35b as the local executor. It produced 8/8 honest terminal states and 2/8 full passes: CLI 1/2, web 1/6. The web full pass was CONTENT b, including browser readiness, interaction evidence, persistence, release-gate pass, and full earned assurance. The other 6/8 runs failed with concrete terminal reasons: CLI b failed the Python behavior smoke probe, TOOL a hit a dangerous-command policy boundary, TOOL b hit the tsconfig alias/baseUrl route-closure gap, CONTENT a failed dependency setup lifecycle, GAME a failed multi-file TypeScript coherence, and GAME b failed compile repair follow-through. Verdict: the 27b/35b local pair is below the recommended tier for reliable web delivery. CLI is moderately viable; web full-pass delivery is possible but low-probability on this distribution. The recommended implementation tier remains the gemini-3.5-flash class or above.
  • The local-model compatibility track left permanent, model-agnostic assets: tool-argument path normalization and corrupted-prefix salvage; verifier command transforms for shell-control splitting, cd/cwd normalization, and output-pipe stripping; deterministic substitution after repeated verifier policy rejection; deterministic scaffold completion; per-provider timeout calibration; planner-call chokepoint instrumentation; and precise exhaustion classification. The honest scope limit is that dialect coverage was measured against this 27b/35b model pair. New model families may require a new dialect-absorption round; that is expected bounded cost, not a portability claim for unseen output dialects.
  • Local single-model GAME quality track, instructions 81-87, culminated in test0707_009 on build a9fea8ee, using qwen3.6:27b-coding-nvfp4 for the single local model configuration. The six-run series moved the failure frontier from phase 0, to phase 1, to the final acceptance gate, then to a full pass with behavioral verification. The full-pass run recorded browser readiness passed, interaction evidence passed, state changes in aliensRemaining and score, a primary start transition, and an in-play recovery/restart transition. This is a distinct local data point from the 27b/35b mixed-pair verdict above: single 27b can complete the GAME path after the 81-87 guard sequence, but the observation is narrow to that track and does not revise the broader recommended tier. Residual fits_viewport=false / bottom overflow 136px is informational presentation quality guidance, not a release gate.

Model-tier observations:

New model-family entries follow the standard order in model-probe.md: model-probe card review, CLI and TOOL smoke checks, then the full scenario round with pre-committed landing criteria. The tier table cites the probe profile as dialect evidence, not as a capability benchmark or an automatic runtime configuration source. For local speed measurements, the recommended zero-risk default is a single Ollama model used for both planner and executor with the model kept resident by Ollama-side keep-alive settings; cloud-hosted providers benefit from prefix stability only when the server exposes compatible prompt caching.

configuration sample distribution / frontier verdict
gemini-3.5-flash implementation/planner Q1-final test0704-999-Q1-62_001, 8 runs 6 full / 1 reasoned partial / 1 behavioral failed, 8/8 honest terminal states Recommended implementation tier baseline.
Local mixed pair: qwen3.6:27b-coding-nvfp4 planner + ornith:35b executor test0704-999-Q1-62646566676869707172_001, 8 runs 2/8 full, 8/8 honest terminal states; web 1/6 Below recommended tier for reliable web delivery; CLI moderately viable.
Local single model: qwen3.6:27b-coding-nvfp4 GAME quality track instructions 81-87, six-run series ending at test0707_009 frontier advanced phase0 -> phase1 -> final gate -> full pass with behavioral verification, in-play restart/recovery, score and enemy-state mutation Golden local single-model GAME reference; narrow positive data point, not a broad tier upgrade.
Local single model: gemma4:31b-local test0708_007 GAME run Reached 4/5 phases with a build-passing, playable-looking app, then exposed restart-hook coverage and pending-evidence exhaustion honesty gaps. Useful local stress datum for observability coverage; not recommended for reliable web delivery until restart/input evidence follow-through recurs as full pass.
Local MoE single model: qwen3.6:35b-a3b test0707_010 / test0708_001 Empty planner responses, missing descriptive StepPlan fields, manifest drift, and dual dependency/compile blockers exercised planner and deterministic-repair ladders. Dialect-heavy model: viable only with the empty-response ladder, schema defaulting, manifest reconciliation, and compile repair ordering guards engaged.
Cloud single model: gemma4:31b-cloud test0708_009 GAME run Cross-file hook contract mismatch reached compile repair, but compact zero-edit repair exhausted without regeneration; no last-known-good page snapshot existed for rollback. Model-family datum for zero-edit repair behavior; motivates regeneration rung and does not change tier recommendation.
Cloud single model: gemma4:31b-cloud post-97 round test0708_016 / test0708_017 GAME runs Runtime Bash shell-control at the Bash-tool boundary killed one run before normalization; a later run terminated honestly with pending restart/input capability evidence. Model-family datum for boundary dialect and honest exhaustion behavior; no tier upgrade.
Cloud single model: gemma4:31b-cloud Q1 CLI cli_b First full pass in the CLI domain for this model family, with the generated command-line workflow satisfying the process-profile gates. Positive domain-specific datum; keep separate from GAME/web reliability because dialect and repair failure modes still recur in app runs.
Cloud single model: gemma4:31b-cloud post-101 Q1 round 8-run Q1 round, post-101 3/8 full, 8/8 honest terminal states; tool_b repeated model_stagnation:read_only_loop, while game_a exposed camelCase collision evidence and content_b exposed invalid semver feedback. Mixed datum: honesty held, but app reliability remains below recommended web tier; install-substitution, semver remedy, and case-insensitive evidence guards are corpus-backed.
Cloud single model: gemma4:31b-cloud final multi-model round final gemma4-cloud round, 2026-07-08, fixture-verified CLI capable; TOOL unstable; web not recommended. Residuals were multi-grep verify normalization and python-cli scaffold fallback, not live re-round failures. Cloud-hosted models have no identity pinning, so this verdict is valid only with the campaign date and probe card. Track-closing verdict: use for CLI experiments with review, not as a recommended TOOL/web implementation tier.

Contract-Design Principle

Observability contracts may require only invisible instrumentation: data attributes, state snapshots, and other non-visible probe hooks. Demands on visible design, including control placement, reachability, or UX layout, are preferences offered as tradeoffs. When declined or absent, the runner degrades honestly to unverified or partial evidence instead of forcing design choices.

Two recent applications make this boundary permanent. Instruction 82A' treats restart reachability as evidence honesty: an in-play restart/recovery path can earn behavioral verification, while an overlay-only or unreachable restart is reported as unverified/partial rather than rejected as a required visible layout. Instruction 86D' treats styling as presence-conditional coherence: Tailwind artifacts must be coherent when present, but plain CSS with no Tailwind artifacts is a valid Next.js path and must not be overwritten by forced stack injection.

Q1 Final Quality Baseline

Q1 final is the standing quality baseline as of 2026-07-05. The judgment rule has two parts:

  • Mandatory stranding elimination: every sampled run must terminate honestly with a closed terminal state, concrete status/reason fields, and a recovery handoff when the run is not full. Max-iteration, human-interrupt, absent terminal status, and false-full exits disqualify the round. A run is stranded when its primary termination reason is an iteration or budget label instead of the concrete blocker, such as missing artifacts, policy rejection, compile failure, probe infrastructure failure, or failed behavioral evidence.
  • Distribution over samples: after the mandatory condition passes, quality is judged by the distribution over at least two samples per scenario.

Q1-final matrix (test0704-999-Q1-62_001, implementation/planner model gemini-3.5-flash, build 0604a76b):

scenario/run run id terminal status primary reason key telemetry
CLI a 019f318c-9931-7350-a720-c3840f640968 full none python-cli, contract_origin=initial, external contract OK, browser/interaction not applicable for non-web, assurance full
CLI b 019f3193-227e-7251-ba45-d32d61762737 full none python-cli, contract_origin=initial, external contract OK, browser/interaction not applicable for non-web, assurance full
TOOL a 019f319b-aef6-7961-8ff6-15a4ad7a253e reasoned partial interaction_unverified:not_evaluated:no_mutation_observed browser readiness performed/passed, interaction performed/passed, state dimension todos, persistence_after_reload_reason=no_mutation_observed, recovery prompt/YAML recorded
TOOL b 019f31a1-ff5f-7040-866d-3e405fc43e39 full none browser readiness performed/passed, interaction performed/passed, state dimension todos, release gate pass, assurance full
CONTENT a 019f31a9-7700-7c12-b947-9597eca40d16 full none browser readiness performed/passed, interaction performed/passed, state dimension currentContent, token echoed, assurance full
CONTENT b 019f31af-8988-7042-bea6-19f7aae0fccf full none browser readiness performed/passed, interaction performed/passed, state dimension contentLength, token echoed, assurance full
GAME a 019f31b4-632c-7413-b322-34f0f8db8e91 full none browser readiness performed/passed, interaction performed/passed, state dimension bulletsCount, primary/restart hooks, assurance full
GAME b 019f31c0-14f5-7c53-898f-4612c883da3f behavioral failed missing_required_evidence:interactive_ui_source_evidence; browser_interaction_failed:probe_script_error browser readiness HTTP 200, interaction performed_failed at surface_wait, recovery prompt/YAML recorded; served-DOM inspection showed visible canvas and primary action after a hidden first state instrumentation hook, so this is a harvested probe-calibration limit, not a false full

Result: Q1-final concludes the quality track for the current scoped host/model pair with 8/8 honest termination and a 6 full / 1 reasoned partial / 1 behavioral failed distribution.

Clause Evidence

clause run evidence harvested corpus
Web profile behavior is not Space-Invaders-only; generic interaction, persistence, and content-editing obligations are separately checked. M2: test0704-464748_001, test0704-464748_002, and test0704-48.1 CONTENT re-run Not harvested in this source corpus; the regression target is the scenario suite in uat/scenarios.md.
Contract inference survives renamed or opaque scenario IDs and English/Japanese prompt variation. M3: test0704-49_001, test0704-50_001 Covered by required golden tests tests/eval/test_acceptance_contract.py::AcceptanceContractTest and tests/eval/test_completion_contract_snapshots.py::CompletionContractSnapshotTest.
No-profile app-intent runs use generic contract binding and terminate honestly at the generic static tier when the scaffolded manifest is unknown. Static-tier fallback, live-proven: test0704-4030444542434647484814950515354_001 scaffolded a Vite/React manifest, emitted no profile_reinferred, and ended without pretending to run promoted web gates; report: workspace/management/runs/uat-test0704-4030444542434647484814950515354-001/uat-report.md. test0704-4030444542434647484814950515354_001
No-profile app-intent runs promote only when a known manifest appears, preserving generic obligations and earning full only through executed web gates. Promotion path, live-proven: final G3 revalidation test0704-403044454243464748481495051535455565758_000, run id 019f30c7-da99-7d83-a715-1db6b6a6a3b6, recorded profile_reinferred, contract_origin=promoted_union, dependency reconciliation, browser readiness HTTP 200, interaction probe execution, and assurance_level=full; report: workspace/management/runs/uat-test0704-403044454243464748481495051535455565758-000/uat-report.md. test0704-403044454243464748481495051535455565758_000
The runner lifecycle supports a non-web process profile without Next.js probe or port assumptions. M4: test0704-51_001 web run and test0704-51_001 CLI run The CLI contract is specified in uat/scenarios.md#cli-python-cli-profile; no app corpus case is expected for the process-only run.
App evidence detectors and interaction probe selection are fixture-backed. M5 Round A four runs test0702_008, test0703_002, test0703_005_4, test0704_001
The second corpus round confirms the same detectors over a new harvested app snapshot. M5 Round B test0704_003
Q1-final residuals are corpus-backed: persistence not_evaluated must render a reason, and the GAME b HTTP-200 hidden-first surface failure is recorded as a probe calibration limit. Q1-final: test0704-999-Q1-62_001 TOOL a and GAME b q1-final-tool-a-persistence-not-evaluated, q1-final-game-b-rendered-hidden-probe-limit
Local-tier residuals and the first local web full pass are corpus-backed: earlier GAME b records artifact stagnation, earlier TOOL b records output-pipe verifier normalization, final CONTENT b records the complete local web behavioral-gate path, final TOOL b records the tsconfig alias/baseUrl gap, and final GAME a records multi-file TypeScript coherence failure. Local rounds: test0704-999-Q1-6264656667686970_001 GAME b / TOOL b, and test0704-999-Q1-62646566676869707172_001 CONTENT b / TOOL b / GAME a local-q1-game-b-artifact-stagnation, local-q1-tool-b-output-pipe-verify, local-q1-final-content-b-web-full-pass, local-q1-final-tool-b-tsconfig-alias-gap, local-q1-final-game-a-typescript-coherence
Local single-model GAME full pass is corpus-backed, including behavioral start, input state mutation, in-play recovery/restart, and informational fits_viewport overflow. test0707_009, single-model qwen3.6:27b-coding-nvfp4, run id 019f3c60-f7d3-7f21-8f8c-6b3b0626370a local-single-qwen36-game-full-pass
Post-97 residuals are corpus-backed: root-anchor path salvage miss, Bash-tool shell-control normalization, and gemma4-cloud honest capability exhaustion. test0708_013, test0708_016, test0708_017 test0708_013, test0708_016, test0708_017
Q1 boundedness residuals are corpus-backed: repeated broad Bash timeout loops and same-run dev-server port conflicts. Q1-full content_b, Q1-full tool_a q1-full-content-b-timeout-loop, q1-full-tool-a-orphan-port-in-use
Post-101 residuals are corpus-backed: camelCase/PascalCase gameplay evidence no longer false-negatives, and invalid semver feedback names the manifest entry and corrected example. Q1 post-101 game_a, content_b q1-post101-game-a-camelcase-collision, q1-post101-content-b-invalid-semver
Final multi-model residuals are corpus-backed: multi-grep verify shell-control normalization and python-cli setup fallback substitution. Final gemma4-cloud fixture round q1-final-multi-grep-shell-control, q1-final-cli-scaffold-fallback
Guardrails are permanent and cheap enough to run in CI. M6 docs and tests cargo test --test corpus_regression, cargo test --test generality_guardrails, cargo test --test conformance, plus the scenario contract golden tests named below.

Required Gates

These gates are required for M6 branch protection or equivalent release approval:

  • cargo test --test corpus_regression
  • cargo test --test generality_guardrails
  • cargo test --test conformance
  • cargo test generic_ultra_promotes_to_nextjs_after_workspace_manifest --lib
  • cargo test generic_ultra_promotes_to_python_cli_after_pyproject_manifest --lib
  • cargo test generic_ultra_without_manifest_keeps_static_tier --lib
  • python3 -m unittest tests/eval/test_acceptance_contract.py
  • python3 -m unittest tests/eval/test_completion_contract_snapshots.py
  • python3 -m unittest tests/eval/test_false_positive_regression.py

The corpus harness guards harvested probe/evidence/profile behavior. The conformance matrix guards shared interface contracts for profiles and execution pathways. The scenario contract tests are the golden suite for contract inference. The false-positive regression protects the "static screen is not a game success" boundary. The generality guardrails protect static- and reduced-assurance rendering and the Next.js boundary-erosion tripwire.

Roadmap Completion

milestone completion date evidence
M0 Baseline scope and runbook 2026-07-04 Scope S and UAT runbook fixed in this declaration and uat/scenarios.md.
M1 Acceptance clauses 2026-07-04 Scenario final-acceptance clauses in uat/scenarios.md.
M2 Web scenario variation 2026-07-04 test0704-464748_001, test0704-464748_002, test0704-48.1 CONTENT re-run.
M3 Contract inference 2026-07-04 test0704-49_001, test0704-50_001 plus required golden tests.
M4 Non-web process profile 2026-07-04 test0704-51_001 web and CLI runs.
M5 Corpus harvest 2026-07-04 Round A four corpus cases plus Round B corpus case.
M6 Declaration and guardrails 2026-07-04 This document, cross-references, and generality_guardrails tests.

Generic Assurance Track Completion

Status: complete on 2026-07-05.

milestone completion date evidence
G0 Scope 2026-07-04 Generic assurance scope, limits, and named guarantees recorded in Scope S.
G1 Generic contract binding 2026-07-04 generic_contract_bound event and generic static contract tests.
G2 Known-manifest promotion 2026-07-04 generic_ultra_promotes_to_nextjs_after_workspace_manifest, generic_ultra_promotes_to_python_cli_after_pyproject_manifest, and generic_ultra_without_manifest_keeps_static_tier.
G3 Ambiguous UAT evidence 2026-07-05 Two live evidence entries: static-tier fallback test0704-4030444542434647484814950515354_001 and final promoted full-assurance run test0704-403044454243464748481495051535455565758_000.
G4 Codification 2026-07-05 Default Next.js no-port policy, AMBIGUOUS scenario runbook hardening, G3 corpus harvest, and mechanism-ledger.md#generic-assurance-track-cross-reference.

Quality Track Completion

Status: Q1 concluded on 2026-07-05.

milestone completion date evidence
Q1 model-tier baseline 2026-07-05 Recommended model tier and M5/Q1-final distributions recorded in Recommended Model Tier.
Q1 final round 2026-07-05 test0704-999-Q1-62_001: 8 runs, 8/8 honest termination, 6 full / 1 reasoned partial / 1 behavioral failed on gemini-3.5-flash.
Q1 residual corpus harvest 2026-07-05 TOOL a persistence reason and GAME b rendered-hidden probe-limit fixtures in Clause Evidence.
Q1 boundedness closure 2026-07-05 Provider-turn and verify-command wall-clock invariants recorded in mechanism-ledger.md#boundedness-guarantees.
Local single-model GAME closure 2026-07-07 Instructions 81-87 concluded with test0707_009, a full behavioral GAME pass on single-model qwen3.6:27b-coding-nvfp4; corpus fixture local-single-qwen36-game-full-pass.
Multi-model generalization track 2026-07-08 Instructions 89-102 closed the family-adoption loop: cost curve 9 -> 7 -> ~2 -> ~2 instructions as fixes shifted from family dialects to model-agnostic assets; standing protocol is probe -> smoke x2 (CLI + TOOL) -> full round with pre-committed landing criteria -> tier entry citing the probe card; boundedness is five-dimensional: transport, spawn, planner, step wall clock, and server reaping.

Optional backlog after Q1:

optional track trigger
tsconfig paths alias deterministic invariant repair Start on recurrence of a route-bound @/* import gap or tsconfig baseUrl/paths missing @/* alias terminal reason, using local-q1-final-tool-b-tsconfig-alias-gap as the fixture.
Dangerous-command rejection feedback categories Start after recurrence analysis shows repeated local-model failures at the same policy category instead of isolated blocked-command attempts.
CONTENT a dependency lifecycle variance Start on recurrence of dependency setup lifecycle failure; first check offline/network/package-registry state before treating it as a CommandAgent behavior defect.
fits_viewport responsive guidance Start on recurrence across scenarios. The test0707_009 GAME pass records bottom canvas overflow as informational presentation quality, not a gate. Repeated overflow should become responsive guidance and visual QA, not a hidden release blocker.
UX track resumption Start when instruction 80 is resumed with a visual acceptance protocol. Keep it separate from invisible observability contracts so visible design preferences remain reviewable trades.
Data-profile/workflow track Start when a non-web workflow needs first-class contracts beyond python-cli, including fixture/data lifecycle and acceptance probes.
T2 model-tier expansion Start only after a probe card plus smoke/full-round evidence defines the target tier criteria in advance.
Second web profile / web-family layer Start when a non-Next.js web stack needs first-class promotion, shared web obligations, or release-gate parity. Until then, unknown manifests remain generic/static by design.
Linux run Start before claiming Linux host parity, adding Linux-specific release support, or treating Darwin UAT results as portable.
English-goal run Start before claiming English prompt-distribution parity or using English-goal quality as release/marketing evidence.

Out Of Scope

The declaration does not claim:

  • A second web framework. The web-family shared layer has not yet been extracted from the Next.js profile boundary.
  • Promotion for arbitrary web frameworks. Unknown manifests remain generic and must stop at the static tier unless a known profile boundary is added with tests and corpus evidence.
  • Non-process domains outside web apps and CLI/process tools.
  • Linux behavior. Current UAT evidence is macOS/Darwin-only.
  • Output depth beyond the observed model pair. Depth and polish remain model-bound even when the contract, probe, and terminal-state mechanisms hold.

D-2 close: fix × data

The admitted profile-synthesis arm is recorded separately from the none arm: none Window A: 24 runs, full 0; profile synthesis Window A/B: 6 runs, full 0. The synthesis arm eliminated the prior mechanical failure classes (0/6); all six runs reached F1 and stopped in local-tier repair read-only stagnation. Cloud-tier execution is recommended. F1 live evidence exists for 30 runs; F2/F3 live evidence exists only for one fix×nextjs run, not for fix×data. FIX-6b reproducer_defect remains WATCH.