Skip to content

Show three instant and three generated predictions - #974

Merged
enaboapps merged 3 commits into
mainfrom
codex/973-six-predictions
Oct 4, 2026
Merged

enaboapps merged 3 commits into
mainfrom
codex/973-six-predictions

Conversation

@enaboapps

@enaboapps enaboapps commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Six fixed optional slots replace the refinement row. Slots 1–3 always remain the statistical predictor's first three words; slots 4–6 receive additional generated words. Empty slots are skipped in both scan directions, and the entire row remains stable while scanned. Exact batch-token acceptance and existing input-safety checks are retained.

Closes #973. Depends on switchifyapp/switchify-prediction#24.

Both SDKs and worker provenance pin release commit 2a2ed0c2b0013663ffa9317dba3389ec96b6f14b from published companion v0.2.1. All four permanent worker URLs were downloaded anonymously and verified against archive and selected-file hashes. CI no longer needs a GitHub token for asset fetching. Model/tokenizer/database hashes are unchanged. The two-second reply and 30-second startup bounds remain.

Validation

Independent read-only review of latest desktop head 07ff11e found no actionable findings. Latest-head CI and CodeQL passed. All required local checks also passed with these release pins. Node 24 lint, 253 UI tests, 12 script tests and frontend build passed. Rust 1.97.1 formatting, warning-denied Clippy, 634 library tests and 7 config tests passed locally and platform CI passed. Fake-adapter tests cover slot gaps, equal-width rendering, stable scanning, stale results and original/completed batch acceptance. Automated tests never inject desktop input.

Final Windows installer and macOS application contents passed resource verification and offline model loading. A final unsigned Windows development installer was also built and checked locally. Signing/notarization and interactive Accessibility testing are not established by these unsigned CI builds.

Final release-pin integration measurements:

Metric Windows x64 macOS ARM
Attempts / successful refinements 1001 / 1000 1000 / 1000
Failures 1 timeout 0
Immediate median / p95 / maximum 11.20 / 14.35 / 45.23 ms 7.81 / 15.28 / 31.90 ms
Successful generation median / p95 / maximum 825.36 / 861.35 / 1268.07 ms 721.65 / 1348.93 / 2029.22 ms
Peak process-tree RSS 691,376,128 bytes 861,257,728 bytes

Windows recorded one timeout after 82 successful refinements. The benchmark explicitly simulated reopening the keyboard and completed the remaining 918 successes in a second session. Production does not retry automatically, and statistical suggestions remain available after neural failure. Both runs filled all 3,000 neural slots across successful refinements. The timeout is included in failure counts, not in successful-generation latency statistics.

These fixtures use five repeated synthetic contexts, the production engine/model adapter and child IPC with 20 ms polling. Timing starts before immediate statistical prediction and includes dispatch, reset handling and scheduling. The 2.03-second macOS end-to-end maximum is reported openly; the separate per-reply deadline remains two seconds. Measurements exclude the outer desktop pipe and rendering and do not establish corpus accuracy. Machine-readable reports are CI artifacts prediction-benchmark-Windows and prediction-benchmark-macOS in run 37219775614.

Corpus comparison and limitations

Final companion qualification sampled 1,000 of the 1,760 frozen development/test queries per platform. Six-slot top-six accuracy was 72.6% Windows and 72.7% macOS, versus 48.2% for six statistical suggestions. Existing reranker top-six was 51.8–52.0%, with at most five results. Keeping statistical first-three ordering reduces top-three from 50.1–50.4% to 39.9%. Individual regressions remain: 16 Windows and 15 macOS queries lost against six statistical suggestions, with 260 gains on each. Both recovered 29/63 OOV target query instances.

Neural fill on the corpus was 91.6% Windows portable and 93.4% macOS ARM. Each completed 1,000 generation replies without generation failures; each comparison reranker failed once. Generation p95 was 1676.94 ms and 1466.30 ms. Search cutoff and runner speed can affect fill. Unknown training overlap prevents claims of unseen-data accuracy, and historical model-quality qualification failures remain.

Reproduction commands and earlier source-labeled machine-readable measurements are in docs/six-predictions.md and docs/measurements/six-predictions. Final companion reports are artifacts generation-windows-latest and generation-macos-latest in the linked qualification run.

Before merge or release

Companion v0.2.1 has been published and the temporary asset blocker is resolved. The desktop uses permanent checksum-pinned release URLs. Verified caches and installed applications work offline. The desktop PR remains unmerged; merging and publishing the desktop RC require separate authorization.

@enaboapps enaboapps added this to the v1.0.0-rc.20 milestone Oct 4, 2026
@enaboapps
enaboapps marked this pull request as draft October 4, 2026 17:14
@enaboapps
enaboapps marked this pull request as ready for review October 4, 2026 18:00
@enaboapps
enaboapps merged commit 43656aa into main Oct 4, 2026
6 checks passed
@enaboapps
enaboapps deleted the codex/973-six-predictions branch October 4, 2026 18:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Show three instant and three generated predictions

2 participants