Repository navigation
Add optional on-device language-model word prediction - #879
Merged
Merged
Conversation
Enhanced word prediction runs SmolLM2-135M (int8 ONNX) inside the existing prediction worker. The model spells words from subword pieces, so it can suggest current words the 2017 lookup lacks. Filters keep out fragments, blocked words and improbable unknown words. The lookup answers while the model loads and whenever it fails. The setting is off by default because it needs about 300-450 MB more memory. The model is fetched from a pinned revision with SHA-256 verification, because it exceeds GitHub's file size limit. Refs #878 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Stop the model's word search after 400 ms, and switch the worker to the lookup for good if one call exceeds 1 s, so a slow machine cannot hit the two-second deadline and lose all suggestions. - Treat curly apostrophes as part of a word, and keep a typed curly apostrophe in completions. - Use the lookup when clipping leaves no complete earlier word. - Prune only with words that can pass the final filter, and check the context vocabulary length. - Release the keyboard when the setting changes mid-acceptance. - Disable ONNX Runtime telemetry explicitly. - Run the real-model test in CI and cover the model/lookup merge path with a fake predictor. - Document the fetch step in AGENTS.md validation and the build-time ONNX Runtime download in PROVENANCE, and back off fetch retries. Refs #878 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The context pass is the one model step without a deadline. Reading at most the last 256 buffered characters, starting at a word boundary, bounds its cost on slow machines. Buffers with no complete word in that window use the lookup. Refs #878 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
enaboapps
marked this pull request as ready for review
September 24, 2026 21:10
enaboapps
marked this pull request as draft
September 25, 2026 08:22
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #878
Summary
Adds Enhanced word prediction, an opt-in setting under Keyboard → Suggestions, off by default. It runs a small on-device language model (SmolLM2-135M, int8 ONNX, Apache-2.0) inside the existing prediction worker. The model spells words from subword pieces, so it can suggest current words the 2017 lookup never will (WhatsApp, covid, emoji, spotify, physio…).
prediction/model.rs): beam search over word pieces, constrained to the typed prefix. Each word is scored with P(pieces) × P(word boundary), and case variants are merged. The search is written against aScorertrait, so unit tests use a fake model.prediction/database.rs): the model loads on a background thread, and the lookup answers until it is ready and whenever the model fails. Lookup suggestions fill any remaining positions.Context::before), with the same privacy boundary as today. Nothing is logged, persisted or sent, and the model never learns.enhancedWordPredictioninConfig/PointSettings, with a serde default offalsefor existing saved settings. Changing it restarts the worker.npm run prediction-modelfetches it from a pinned Hugging Face revision and verifies size and SHA-256. It runs from TauribeforeDevCommand/beforeBuildCommandand from a CI step and both release signing jobs. Provenance is inPROVENANCE.md.ort2.0.0-rc.13 (default-features = false, nocopy-dylibs). I verified that inference runs with noDirectML.dllpresent.Evidence (local experiment, same typing simulation, 5 suggestions)
Validation (Windows 11, Node 24.13.0, Rust 1.97.1)
npm run lint✅npm test: 224 vitest + 5 node tests ✅. New UI test covers the toggle: default off, saves, disabled when Word prediction is off.npm run build✅cargo fmt --check✅;cargo clippy --locked --all-targets -D warnings✅cargo test --locked: 558 passed, 1 ignored ✅. The real-model test (bundled_model_suggests_current_words) now runs in CI; it asserts "tea" and "kettle" appear. Fake-scorer and fake-predictor tests cover:bundled_model_suggests_current_wordsagainst the real model: startup 650 ms, predictions 40–93 ms.Independent review
A separate agent that did not write the code reviewed the PR, then reviewed each new head again.
a7e56cb): no blockers. One major finding: no time budget on slow machines, now fixed. It also confirmed the tensor maths, KV-cache tiling, privacy and settings compatibility.cf83e10): every fix verified, no regressions, two nits. The context cap is fixed ina7e6c59. The’./’,boundary edge case has negligible quality impact and is left as is.a7e6c59): no issues.a2a329d, latest head): independent reviewer found no actionable findings in the Windows startup-timeout fix.Windows VM follow-up (2026-09-25)
a2a329dallows 30 s for the first reply, then retains the 2 s deadline for later queries. Unit tests cover cold-start and established-worker deadlines.BUILD_EXIT=0), opened the keyboard with a local switch, and observed the prediction worker running beyond the prior timeout with no Predictions unavailable message. The keyboard paused normally after scan passes. The VM switch settings were restored to the original Space=Select binding after testing.Not yet validated
npm run macos:run), following the checklist indocs/word-prediction.md, with Enhanced word prediction on. The Windows check above verifies worker startup and the reported error, but did not exercise the full typing checklist.Known limitations
ortdownloads prebuilt ONNX Runtime fromcdn.pyke.ioat build time, hash-checked by the crate; this is recorded in PROVENANCE.tokenizerspulls inpaste, which has an unmaintained-crate advisory (RUSTSEC-2024-0436; a warning, and the audit passes).🤖 Generated with Claude Code