test(port): continue Python parity at stateful boundaries - #44
Conversation
|
Important Review skippedDraft detected. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Plus Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Integration checkpoint from the Rust implementation branch:
The snapshot had two test-infrastructure blockers during validation:
PR #44 advanced again after the pin. Future updates will merge the pinned old head to the new head as an increment, without rewriting this PR's history. |
|
Final checkpoint after the test-port agent paused:
The first integration compile-only workspace build also found a remaining unrelated imported compile blocker: |
|
Integration branch compile status
Additional test-infrastructure fixes after the previous checkpoint:
The snapshot still emits unused/dead-code warnings in several imported tests, so full Clippy with warnings denied is not yet green. The active parity tests are also intentionally expected to fail until PR #45 closes their implementation gaps. |
|
Downstream integration note (paused head Two more reviewed slices have been promoted to draft PR #45 without modifying this branch:
The complete PR #44 snapshot and ancestry remain preserved on |
|
PR #44 can resume from |
|
Checkpoint: packaging parity wave (oracle
Validation status for this checkpoint:
Reason: this execution container has no |
|
Checkpoint: i18n parity wave (oracle
Validation status:
The container still has no Rust toolchain and DNS blocks Rust bootstrap. No PR-triggered GitHub Actions workflow run exists for this head. I therefore classify execution as unclassified / not run, not passing. Static source/API/provenance/strength review completed; no production source changed. |
|
Checkpoint: responsive TUI parity wave (oracle
Validation status for this checkpoint:
Reason remains the same: this execution container has no Rust toolchain and DNS blocks Rust bootstrap. GitHub reports no PR workflow run for this head. Execution is therefore unclassified / not run, not passing. |
|
uvman integration audit (tested against PR commit
Please keep this as a test-only harness correction. It also prevents the suite from making an unintended real network download. |
|
Checkpoint — test-only Python v0.4 parity migration; no Rust production/source changes. Newly accounted since the previous progress wave:
Formal behavior inventory is now 62/84 modules, 1,209/3,018 frozen Python test functions accounted. The master inventory remains intentionally red until all 84 behavior modules have audited completeness guards. Validation status for this environment: execution remains unclassified / not run. The local container has no Required validation commands when an executable Rust environment is available:
PR stays draft. Red behavioral assertions remain implementation work for the separate Rust-fix agent; they are not weakened here. |
|
Interpreter migration checkpoint (test-only):
No Rust implementation fixes were made. No local cargo execution is being claimed here; this checkpoint is static/provenance-completeness work plus the hermetic test-fixture correction. Red behavior tests remain valid parity findings for the implementation agent. |
|
Delta after the interpreter checkpoint: |
|
Integration checkpoint:
|
|
Blocking accounting finding in 6168183: the frozen oracle at main@206f9ef has 78 def test_ declarations in tests/test_store.py, not 60. The command Please restore test_store.py to 78 and the total/message to 3,018/3018 on this PR branch. I did not modify the PR branch. The integration branch carries the correction in follow-up commit 17e9ab0. |
|
Audit pinned to Test-quality blockers:
Small candidates after formatting: Exact integration gates: workspace/all-targets/all-features |
|
Audit pinned to Blocking findings in this increment:
Do not count or import these files as complete as-is. Main extracted only three faithful owners (TS suffix, Unix 0600, bad-value prelaunch), removed/relabelled their duplicate language owners, and excluded the weak/duplicate/red rows. |
|
Test-port checkpoint (head
|
|
Checkpoint correction/update for the previous test-port note:
|
# Conflicts: # crates/skit-cli/tests/port_test_zz_behavior_inventory_gate.rs
|
Pinned-head integration/audit update (2026-08-13)
This is a synchronization result, not approval to move the new tests to main. The fixed PR batch still has blocking audit findings:
Observed manifest status on the fixed tree: config green, Fish green, Shell analyzer green, JS injection red. Because the green manifests still have duplicate/closure/static-gate defects, none of these clusters should be copied to main wholesale. Any mainline batch should be reduced to individually faithful, non-duplicate green owners and rerun through the hard gates. |
|
Pinned review of This is a linear 18-commit, 13-path test-only increment from the previously reviewed pin Blocking accounting findings:
Honest small batches are still available: the two JS singletons plus corrected manifest includes; shim staging as a two-contract batch after runtime availability is verified; and small shim core/runtime behavior groups without the manifest or inventory. Atomic and Store need owner deduplication/reclassification first. |
|
Pause checkpoint recorded.
The latest raw increment remains under review rather than green: Store accounting hides 12 duplicate owners; Atomic lock wrappers do not distinguish their claimed seams; two Shim names are strict semantic duplicates; and the two JavaScript gate contracts require a production No further PR #44 integration work will run until the paused goal is resumed. |
|
Benchmark-tooling checkpoint:live occurrence audit 現在 83 executable + 15 narrowly closed = 98/156;master 仍保持 83/84 / 2,862/3,018,尚未 attach benchmark module。 Front door 本輪新增 6 個真 binary owners: 我刻意保留 frozen 新增的 2 個 closures 只限 Python wrapper/injection seam: 另外已修 live accounting 自己的一個重要 bug:frozen 156 是 function occurrences,不是 156 unique bare names;例如 |
|
Benchmark-tooling checkpoint:live occurrence audit 已到 126 executable + 23 narrowly closed = 149/156;master 仍未 attach,仍是 83/84 / 2,862/3,018。 這一輪有一個重要 test-oracle 修正:我重新讀 frozen 新增 executable evidence:env/reuse/search 8;source integrity 6(精確 bytes + parser/analyzer);repo contract-sync 5(budgets canonical、analyzer filename shared registry、3 個 benchmark workflows 統一走 Hyperfine action、legacy compare 的 pyperf pin、BENCH_CI_RUNNER == runs-on)。 新增 closures 只限已消失/被型別消除的 Python seams:pyperf worker inheritance、duplicated Hyperfine/BROKEN_LINES constants、Python console-script census、argparse private parser hash、Python subprocess.run AST timeout keyword(Rust 現在只剩 7 個 exact frozen occurrences,不再有其他暗漏:
前 6 個我正用 public |
|
接手 checkpoint(固定 head
完成審查後才 attach |
|
Final-7 strength audit found test-side defects that must be fixed before master attach (these are not acceptable deliberate-red implementation findings):
I am correcting these fixtures/keys and strengthening the two silent-undercount owners. I will also replace the existing architecture closure for ordinary |
Replace invalid benchmark-suite fixtures and invented keys with public-suite owners that preserve the frozen Python raw/statistical contracts. Keep the runover silent-undercount oracle deliberately red, promote the ordinary generator undercount defense to executable evidence, and attach the completed 156-occurrence audit to the master behavior inventory. Test-only: no production source, manifests, docs, or workflows changed.
|
本輪 final-7 強度修復已完成並推送: Scope
最終 accounting
本輪修正的 test-side 問題
預期紅燈
Validation
後續 Rust 環境若揭露 test compilation/fixture 錯誤,只修測試;若揭露 behavior mismatch,保留紅燈交由 implementation agent 修 production。 |
|
Implementation-agent handoff note — the PR description now has a dedicated Implementation-agent handoff section so the next local Rust implementation pass does not need to reconstruct intent from the full comment history. Key instructions for the implementation branch:
Authoritative disputed-behavior source remains |
|
Final fixed-head diff audit against This increment cannot be merged or cherry-picked as one batch:
Duplicate-name adjudication is body-based. Every candidate is compared against both the current Rust owner and the frozen Python body. One concrete rejected rewrite is Three corrected, uniquely-owned folds have landed on PR #45:
The combined implementation checkpoint is green at 3,207 passed, 0 failed, and 929 ignored, with fmt, workspace Clippy Further useful bodies remain, but they must be folded into existing owners in small batches. Raw split files, module manifests, and the master inventory should not be imported. |
|
Second selective fold checkpoint on PR #45: Three more corrected batches landed after the final-head audit:
The selected-runner preflight failure contract was not imported: current production sets a status but does not route back to Library, and the PR's loose ANSI-output check cannot prove final screen state. It remains a production TDD blocker. Current combined workspace: 3,214 passed, 0 failed, 922 ignored. Fmt, workspace Clippy with warnings denied, and Rustdoc with warnings denied pass. Exact owner count is one for every folded name. |
|
Third selective fold checkpoint on PR #45: Four more production-backed batches landed after three-way PR/main/Python review:
The combined workspace is green at 3,229 passed, 0 failed, and 913 ignored. Explicit The next reviewed candidate is not a raw PR file import: remaining general CLI/add/params contracts continue to require body-by-body adjudication against the frozen oracle. |
|
Final-head selective-fold update at Rust checkpoint
Both accepted owners were unique, ran genuine TestBackend/key/mouse paths, produced pre-fix REDs, passed controlled reversals, and passed workspace/fmt/Clippy/Rustdoc. No PR split file, support target, manifest, or inventory guard was imported. |
|
Selective-fold update at Rust checkpoint Accepted only after PR/current/frozen body comparison and independent RED→GREEN evidence:
Rejected or reclassified instead of imported:
No split target, manifest, or inventory guard from PR #44 was imported. Every landed fold passed workspace/fmt/Clippy/Rustdoc on its corrected owner. |
Continue the independent Python-to-Rust behavioral-oracle port from
origin/main@206f9ef946fc45835cb2479593794431f2620c32onto the Rust rewrite base8687325591dd9a5463dafa5535d01dfb5bc91585.Scope: test-only
This draft intentionally changes tests and test support only. Production
src/, product Cargo manifests/dependencies, docs, and durable workflow changes are out of scope. Frozen Python behavior is authoritative; Rust implementation fixes belong in another branch/agent.A ported test that fails is a parity finding and may stay red. Red behavior is not a reason to weaken assertions, accept alternate wording, substitute dry-run for real execution, add
#[ignore]/#[should_panic], invent Rust behavior, architecture-close an implementation gap, or patch production code here.rust_additive_*tests do not count toward Python parity.Frozen inventory
Source of truth:
main@206f9ef946fc45835cb2479593794431f2620c32.def test_function occurrences.test_shell_inject.py= 87,test_flows.py= 102. Earlier transient77 / 62 / 2,968accounting is superseded.Current accounting — 2026-08-18
84/84 behavior modules and 3,018/3,018 frozen test-function occurrences are audited/accounted.
Every accounted module points to an executable Rust completeness guard. Each guard parses or otherwise pins the preserved Python denominator, rejects missing/extra/duplicate owners, and fixes any narrowly allowed architecture-closure set so it cannot expand silently.
Final module:
tests/test_benchmarks_tooling.py134 executable + 22 narrowly architecture-closed = 156/156.
crates/skit-benchmarks/tests/port_test_benchmarks_tooling_manifest.rstreats frozen names as a multiset, including the cross-class duplicatetest_rejects_bad_inputs, scans 15 canonical Rust owner files, rejects non-frozen/duplicate/missing owners, and fixes the closure partition at exactly 22.The final strength repair replaced invalid test fixtures and invented output keys rather than accommodating Rust behavior:
suites::runpath and assert the actual frozen nested raw-sample schema plus full median/p95/sample-standard-deviation metadata.total = store + stateandbytes_per_entry = total / nthrough a real generated library and public Footprint suite.uvvenv creation plus a failed and successful install attempt, the real two-second retry delay, bounded spawn sites, isolated cwd, and the complete constructed child environment. It checks the realfootprint.closure_bytes,footprint.skit_installed_bytes, and distribution-count metrics.test_generate_refuses_silent_store_undercountis executable evidence: it performs real generation/store scanning and pins the post-write real-store count plus the frozengenerated {found} entries, expected {expected}diagnostic.test_runover_refuses_silent_store_undercountis deliberately red until production counts the real runover store after the final commit and preservesrunover library has {found} entries, expected {expected}. This missing public defense is not architecture-closed.The 22 closures are limited to Python-only parser, private monkeypatch/injection, representation, or removed-duplicate-constant seams for which no meaningful deterministic Rust public seam exists. Their concrete reasons are pinned in the live manifest. Public behavior surrounding those seams remains executable.
Other completed high-risk audits
tests/test_flows.py: 102 executable + 0 closed = 102/102. Exact wording, drift guidance, runner warnings/help, launch/staging/transparency, process, filesystem, and locking contracts remain strict.tests/test_js_deps.py: 136 executable + 15 narrowly closed = 151/151. Closures are fixed Python-only scanner/tempfile/clock/syscall/Textual/gettext parser seams, not implementation-gap waivers.tests/test_path_tui.py: 56 executable + 5 narrowly closed = 61/61. Public Ratatui traversal/rendering/events/reducer/token/picker/insertion/workdir behavior remains executable.tests/test_cli.py: 135 executable + 5 narrowly closed = 140/140. Owners cross real CLI/process/PTY/filesystem/editor/state boundaries; closures are private Python parser/monkeypatch seams only.tests/test_prompt_cli.py: 149 executable + 1 narrowly closed = 150/150. Real-run snapshots use actual lock ordering as a deterministic post-validation barrier; the single closure is the private consecutive_read_bodycall-count seam.tests/test_prompt_kind.py: 110 executable + 5 narrowly closed = 115/115. Invalid-Unicode representation and private loader call-count/override seams are closed; public grammar, frozen bytes, real child injection, raw CAS/type sensitivity, process locks, limits, persistence, and error contracts remain executable.tests/test_prompt_tui.py: 83 executable + 0 closed = 83/83. Textual-to-Ratatui differences are not closures by themselves. Real PTY/child/editor/store behavior remains executable, and known Settings/Library-edit implementation gaps intentionally stay red.Implementation-agent handoff
This PR is now an oracle bundle, not a request to make the tests green by changing the tests. When the Rust implementation branch consumes it:
005bc9b7365fca1cfa7173acb61a2e8629f03bc9. If an older PR test(port): continue Python parity at stateful boundaries #44 snapshot is already merged locally, merge this head as an increment; do not cherry-pick only green-looking tests or drop the completeness/master gates.#[ignore],#[should_panic], dry-run substitution, looser substring checks, alternate accepted outputs, and expanding architecture closures are not acceptable ways to get green.generate_runovermust verify the real post-commit store count and preserverunover library has {found} entries, expected {expected};resyncguidance;cargo fmt --all --check;cargo test --locked --workspace --all-targets --all-features --no-run; then the exact manifest/master-accounting tests; then targeted failing parity tests; finally the broader workspace suite. The test-port environment did not have a usable Rust toolchain, so current head has no compile-green claim.The authoritative source for disputed behavior remains
main@206f9ef946fc45835cb2479593794431f2620c32. If a Rust behavior seems more convenient but conflicts with that frozen oracle, assume the test is intentional until the test itself is shown to be defective.Guardrails
Validation status
Current head:
005bc9b7365fca1cfa7173acb61a2e8629f03bc9(test(benchmarks): repair final frozen owners).No current-head compile/pass claim is made. This environment has no usable Rust toolchain, and GitHub Actions has returned no runs for the current head. CodeRabbit reports success and there are no unresolved inline review threads.
Performed validation:
Noneguard;/bin/shsmoke checks for the generated fake TUI/benchmark and fake-uvendpoints;cd5733cf..005bc9b: 32 commits / 17 changed files, all undercrates/skit-benchmarks/tests/*except the single test-only master inventory gate.If a Rust environment exposes a test compilation or fixture mistake, fix the test only. If it exposes a behavior mismatch, keep the parity test red for the implementation branch.