Show three instant and three generated predictions - #974
Merged
Merged
Conversation
enaboapps
marked this pull request as ready for review
October 4, 2026 16:10
enaboapps
marked this pull request as draft
October 4, 2026 17:14
enaboapps
marked this pull request as ready for review
October 4, 2026 18:00
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Six fixed optional slots replace the refinement row. Slots 1–3 always remain the statistical predictor's first three words; slots 4–6 receive additional generated words. Empty slots are skipped in both scan directions, and the entire row remains stable while scanned. Exact batch-token acceptance and existing input-safety checks are retained.
Closes #973. Depends on switchifyapp/switchify-prediction#24.
Both SDKs and worker provenance pin release commit 2a2ed0c2b0013663ffa9317dba3389ec96b6f14b from published companion v0.2.1. All four permanent worker URLs were downloaded anonymously and verified against archive and selected-file hashes. CI no longer needs a GitHub token for asset fetching. Model/tokenizer/database hashes are unchanged. The two-second reply and 30-second startup bounds remain.
Validation
Independent read-only review of latest desktop head 07ff11e found no actionable findings. Latest-head CI and CodeQL passed. All required local checks also passed with these release pins. Node 24 lint, 253 UI tests, 12 script tests and frontend build passed. Rust 1.97.1 formatting, warning-denied Clippy, 634 library tests and 7 config tests passed locally and platform CI passed. Fake-adapter tests cover slot gaps, equal-width rendering, stable scanning, stale results and original/completed batch acceptance. Automated tests never inject desktop input.
Final Windows installer and macOS application contents passed resource verification and offline model loading. A final unsigned Windows development installer was also built and checked locally. Signing/notarization and interactive Accessibility testing are not established by these unsigned CI builds.
Final release-pin integration measurements:
Windows recorded one timeout after 82 successful refinements. The benchmark explicitly simulated reopening the keyboard and completed the remaining 918 successes in a second session. Production does not retry automatically, and statistical suggestions remain available after neural failure. Both runs filled all 3,000 neural slots across successful refinements. The timeout is included in failure counts, not in successful-generation latency statistics.
These fixtures use five repeated synthetic contexts, the production engine/model adapter and child IPC with 20 ms polling. Timing starts before immediate statistical prediction and includes dispatch, reset handling and scheduling. The 2.03-second macOS end-to-end maximum is reported openly; the separate per-reply deadline remains two seconds. Measurements exclude the outer desktop pipe and rendering and do not establish corpus accuracy. Machine-readable reports are CI artifacts prediction-benchmark-Windows and prediction-benchmark-macOS in run 37219775614.
Corpus comparison and limitations
Final companion qualification sampled 1,000 of the 1,760 frozen development/test queries per platform. Six-slot top-six accuracy was 72.6% Windows and 72.7% macOS, versus 48.2% for six statistical suggestions. Existing reranker top-six was 51.8–52.0%, with at most five results. Keeping statistical first-three ordering reduces top-three from 50.1–50.4% to 39.9%. Individual regressions remain: 16 Windows and 15 macOS queries lost against six statistical suggestions, with 260 gains on each. Both recovered 29/63 OOV target query instances.
Neural fill on the corpus was 91.6% Windows portable and 93.4% macOS ARM. Each completed 1,000 generation replies without generation failures; each comparison reranker failed once. Generation p95 was 1676.94 ms and 1466.30 ms. Search cutoff and runner speed can affect fill. Unknown training overlap prevents claims of unseen-data accuracy, and historical model-quality qualification failures remain.
Reproduction commands and earlier source-labeled machine-readable measurements are in docs/six-predictions.md and docs/measurements/six-predictions. Final companion reports are artifacts generation-windows-latest and generation-macos-latest in the linked qualification run.
Before merge or release
Companion v0.2.1 has been published and the temporary asset blocker is resolved. The desktop uses permanent checksum-pinned release URLs. Verified caches and installed applications work offline. The desktop PR remains unmerged; merging and publishing the desktop RC require separate authorization.