Skip to content

Add optional on-device language-model word prediction - #879

Merged
enaboapps merged 4 commits into
mainfrom
878-language-model-prediction
Sep 25, 2026
Merged

enaboapps merged 4 commits into
mainfrom
878-language-model-prediction

Conversation

@enaboapps

@enaboapps enaboapps commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Closes #878

Summary

Adds Enhanced word prediction, an opt-in setting under Keyboard → Suggestions, off by default. It runs a small on-device language model (SmolLM2-135M, int8 ONNX, Apache-2.0) inside the existing prediction worker. The model spells words from subword pieces, so it can suggest current words the 2017 lookup never will (WhatsApp, covid, emoji, spotify, physio…).

  • Engine (prediction/model.rs): beam search over word pieces, constrained to the typed prefix. Each word is scored with P(pieces) × P(word boundary), and case variants are merged. The search is written against a Scorer trait, so unit tests use a fake model.
  • Quality gate:
    • ASCII letters and apostrophes only;
    • a capital is required at the start of text (rejects fragments);
    • no single letters unless the lookup knows them;
    • LDNOOBW blocklist (CC BY 4.0);
    • words the lookup doesn't know need log-probability ≥ −10;
    • known words use the lookup's spelling ("who", not "WHO"), and unknown words keep the model's spelling ("WhatsApp").
  • Fallback and time limits (prediction/database.rs): the model loads on a background thread, and the lookup answers until it is ready and whenever the model fails. Lookup suggestions fill any remaining positions.
    • The word search stops after 400 ms.
    • If one model call exceeds 1 s, that worker uses the lookup for good, so a slow machine can never hit the 2 s reply deadline and lose all suggestions.
    • The model reads at most the last 256 buffered characters, starting at a word boundary, which bounds the one model pass that has no deadline.
  • Context: the model gets the whole tracked buffer before the prefix (Context::before), with the same privacy boundary as today. Nothing is logged, persisted or sent, and the model never learns.
  • Setting plumbing: enhancedWordPrediction in Config/PointSettings, with a serde default of false for existing saved settings. Changing it restarts the worker.
  • Distribution: the model (135.7 MB) exceeds GitHub's file limit, so npm run prediction-model fetches it from a pinned Hugging Face revision and verifies size and SHA-256. It runs from Tauri beforeDevCommand/beforeBuildCommand and from a CI step and both release signing jobs. Provenance is in PROVENANCE.md.
  • ONNX Runtime: linked statically via ort 2.0.0-rc.13 (default-features = false, no copy-dylibs). I verified that inference runs with no DirectML.dll present.

Evidence (local experiment, same typing simulation, 5 suggestions)

Engine Everyday AAC text Modern-vocabulary text
Current lookup 53.8% keystrokes saved 45.3%
This engine 55.3% 51.8%

Validation (Windows 11, Node 24.13.0, Rust 1.97.1)

  • npm run lint ✅
  • npm test: 224 vitest + 5 node tests ✅. New UI test covers the toggle: default off, saves, disabled when Word prediction is off.
  • npm run build ✅
  • cargo fmt --check ✅; cargo clippy --locked --all-targets -D warnings ✅
  • cargo test --locked: 558 passed, 1 ignored ✅. The real-model test (bundled_model_suggests_current_words) now runs in CI; it asserts "tea" and "kettle" appear. Fake-scorer and fake-predictor tests cover:
    • spelling and ranking, the blocklist, and the single-letter and unknown-word filters;
    • the start-of-text capital rule, casing, and curly apostrophes;
    • attached-prefix fallback and the expired deadline;
    • the context cap and clipped buffers;
    • model words leading, lookup fill-in, de-duplication and the 5-word cap;
    • load failure, a missing model, and the slow-call switch-off;
    • the settings default and round trip.
  • Release run of bundled_model_suggests_current_words against the real model: startup 650 ms, predictions 40–93 ms.
    • "Can you send a wh" → white, WhatsApp, whole, wheelchair, who
    • "I would like a cup of" → coffee, tea, hot, water, lemonade
    • "Please put the ket" → kettle, ketch, ketchup
  • Fetch script: fresh download, tampered file replaced, verified files skipped.
  • Peak worker memory in the lab harness: ~425 MB with the model vs ~120 MB lookup-only.

Independent review

A separate agent that did not write the code reviewed the PR, then reviewed each new head again.

  • First round (a7e56cb): no blockers. One major finding: no time budget on slow machines, now fixed. It also confirmed the tensor maths, KV-cache tiling, privacy and settings compatibility.
  • Minor findings, all fixed:
    • curly-apostrophe word ends;
    • clipped-buffer start of text;
    • the pruning bound counting filtered words;
    • the vocabulary length check;
    • releasing the keyboard when the setting changes mid-accept;
    • explicitly disabling ONNX Runtime telemetry;
    • the real-model test running in CI, and merge-path tests;
    • the AGENTS.md fetch step, the ONNX Runtime provenance, and fetch backoff.
  • Second round (cf83e10): every fix verified, no regressions, two nits. The context cap is fixed in a7e6c59. The ’./’, boundary edge case has negligible quality impact and is left as is.
  • Third round (a7e6c59): no issues.
  • Follow-up review (a2a329d, latest head): independent reviewer found no actionable findings in the Windows startup-timeout fix.

Windows VM follow-up (2026-09-25)

  • Reproduced Predictions unavailable on the Windows 11 Parallels VM: the bundled 50.8 MB lookup took 7.1 s to load on a cold worker, beyond the original 2 s first-reply deadline. The lookup and model resources were present and the worker returned a valid reply.
  • Commit a2a329d allows 30 s for the first reply, then retains the 2 s deadline for later queries. Unit tests cover cold-start and established-worker deadlines.
  • Built the corrected native Windows ARM64 debug executable in the VM (BUILD_EXIT=0), opened the keyboard with a local switch, and observed the prediction worker running beyond the prior timeout with no Predictions unavailable message. The keyboard paused normally after scan passes. The VM switch settings were restored to the original Space=Select binding after testing.
  • Re-ran Node 24.19.0 lint, test and build; Rust 1.97.1 fmt, clippy and tests: all passed (559 Rust unit tests, 7 config tests, 1 ignored).

Not yet validated

  • Full native manual validation of Enhanced word prediction on Windows and macOS (npm run macos:run), following the checklist in docs/word-prediction.md, with Enhanced word prediction on. The Windows check above verifies worker startup and the reported error, but did not exercise the full typing checklist.
  • Packaged installer size: the model adds about 136 MB to the installer and to every updater download.
  • macOS aarch64 build and notarization with statically linked ONNX Runtime; CI will exercise the build.

Known limitations

  • A 135M model misses some modern words; "facetime" scores below invented words.
  • Some technical or odd words can still pass the unknown-word threshold, e.g. "php" and "backend".
  • Supply chain: ort downloads prebuilt ONNX Runtime from cdn.pyke.io at build time, hash-checked by the crate; this is recorded in PROVENANCE. tokenizers pulls in paste, which has an unmaintained-crate advisory (RUSTSEC-2024-0436; a warning, and the audit passes).
  • Personal learning is out of scope (separate privacy-reviewed issue).

🤖 Generated with Claude Code

Enhanced word prediction runs SmolLM2-135M (int8 ONNX) inside the
existing prediction worker. The model spells words from subword pieces,
so it can suggest current words the 2017 lookup lacks. Filters keep out
fragments, blocked words and improbable unknown words. The lookup
answers while the model loads and whenever it fails.

The setting is off by default because it needs about 300-450 MB more
memory. The model is fetched from a pinned revision with SHA-256
verification, because it exceeds GitHub's file size limit.

Refs #878

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@enaboapps enaboapps added this to the v1.0.0-rc.16 milestone Sep 24, 2026
@enaboapps enaboapps added the enhancement New feature or request label Sep 24, 2026
OwenMcGirr and others added 2 commits September 24, 2026 21:39
- Stop the model's word search after 400 ms, and switch the worker to the
  lookup for good if one call exceeds 1 s, so a slow machine cannot hit
  the two-second deadline and lose all suggestions.
- Treat curly apostrophes as part of a word, and keep a typed curly
  apostrophe in completions.
- Use the lookup when clipping leaves no complete earlier word.
- Prune only with words that can pass the final filter, and check the
  context vocabulary length.
- Release the keyboard when the setting changes mid-acceptance.
- Disable ONNX Runtime telemetry explicitly.
- Run the real-model test in CI and cover the model/lookup merge path
  with a fake predictor.
- Document the fetch step in AGENTS.md validation and the build-time
  ONNX Runtime download in PROVENANCE, and back off fetch retries.

Refs #878

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The context pass is the one model step without a deadline. Reading at
most the last 256 buffered characters, starting at a word boundary,
bounds its cost on slow machines. Buffers with no complete word in that
window use the lookup.

Refs #878

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@enaboapps
enaboapps marked this pull request as ready for review September 24, 2026 21:10
@enaboapps
enaboapps marked this pull request as draft September 25, 2026 08:22
@enaboapps
enaboapps marked this pull request as ready for review September 25, 2026 08:54
@enaboapps
enaboapps merged commit b18deeb into main Sep 25, 2026
6 checks passed
@enaboapps
enaboapps deleted the 878-language-model-prediction branch September 25, 2026 09:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add optional on-device language-model word prediction

2 participants