Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 9 additions & 3 deletions docs/word-prediction.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,15 +4,15 @@ Word prediction is enabled by default in Scanning settings. Its five-position ro

Predictions use only a temporary buffer of successful Switchify keyboard input, starting with the first letter. Existing text, pasted text and hardware keyboard typing are never read into it. Switchify does not inspect fields, selections, passwords or caret positions. Suggestions may therefore appear anywhere the keyboard is open, including password fields or applications without a text field. The buffer records successful input injection; it cannot verify what an application actually accepted.

The buffer holds at most 512 characters, retaining complete Unicode graphemes at its leading boundary. It tracks ordinary characters, spaces, Backspace and accepted completions across keyboard pages and top/bottom docking. Navigation, Delete, Enter, Tab, shortcuts, failed input, external typing/clicks/scrolling and foreground changes clear context. Opening or closing the keyboard, ending scanning, disconnecting and exiting also discard it. The keyboard remains open across foreground changes, resets modifiers and suggestions, and sends subsequent input to the new foreground application.
The buffer holds at most 512 characters, retaining complete Unicode graphemes at its leading boundary. It tracks ordinary characters, spaces, Backspace and accepted completions across keyboard pages and top/bottom docking. Navigation, Delete, Enter, Tab, shortcuts, failed input, external typing/clicks/scrolling and foreground changes clear context. Opening or closing the keyboard, ending scanning, disconnecting and exiting also discard it, by killing the process that held it. The keyboard remains open across foreground changes, resets modifiers and suggestions, and sends subsequent input to the new foreground application.

A passive observer records only an activity counter and timestamp, never external text. Prediction is unavailable if this observer cannot start or loses access. Edits made before the observer is ready are discarded because intervening activity cannot be verified. Each queued edit is scoped to its foreground target and the time before injection, so edits preceding an observed external change are discarded. Changes within an application that produce no observed input cannot be detected without inspecting its fields.

## On-device model

One Word prediction setting controls an offline SmolLM2-135M int8 ONNX model. The saved `enhancedWordPrediction` field is retained for compatibility but does not choose an engine. The model spells candidates from subword pieces and reads only the last 256 characters of the temporary buffer, beginning at a word boundary. It never learns from typing.

For a first typed prefix, a fixed local context keeps the model from favoring website names at the start of a document; it adds no user text. The model loads in the worker while keyboard input remains available. Suggestions are blank until loading finishes. A passive badge beside the scan prompt distinguishes loading, a ready keyboard awaiting typed context, no matching suggestions, available suggestions, paused activity tracking, and prediction failure. It never shows typed text and is not a scan target. If loading or inference fails, or a call exceeds 1.5 seconds, the keyboard continues accepting input and offers **Retry predictions** in its toolbar. Retry restarts only the prediction worker, clears its private text context, and leaves the keyboard open; type a new prefix afterward. A 400 ms search budget bounds candidate exploration. A clipped buffer with no complete earlier word yields no suggestions.
For a first typed prefix, a fixed local context keeps the model from favoring website names at the start of a document; it adds no user text. The model loads in the worker while keyboard input remains available, which takes about a second. Suggestions are blank until loading finishes. Later keyboard opens normally skip this wait by using the spare worker described under Worker process. A passive badge beside the scan prompt distinguishes loading, a ready keyboard awaiting typed context, no matching suggestions, available suggestions, paused activity tracking, and prediction failure. It never shows typed text and is not a scan target. If loading or inference fails, or a call exceeds 1.5 seconds, the keyboard continues accepting input and offers **Retry predictions** in its toolbar. Retry restarts only the prediction worker, clears its private text context, and leaves the keyboard open; type a new prefix afterward. A 400 ms search budget bounds candidate exploration. A clipped buffer with no complete earlier word yields no suggestions.

Candidates contain ASCII letters and apostrophes and must have sufficient model probability. There is no vocabulary filter: any word the model finds likely can be suggested, including swearing, because the person typing chose it. Single-letter candidates are limited to “a” and “I”. Up to five suggestions are shown: the first, third and fifth are the most likely single words, and the second and fourth are the two most likely two-word phrases that begin with one of the top three words, ranked by the probability of the pair. When fewer phrases are found within an extra 200 ms, single words fill the remaining slots, and the other way round. Accepting a phrase inserts the rest of its first word, a space, the second word and a trailing space.

Expand Down Expand Up @@ -64,6 +64,12 @@ The model files are too large to commit. `npm run prediction-model` downloads th

A separate process owns the buffer, activity observer and model. Private bounded inherited pipes carry successful edits and results to native rendering. Text and suggestions are not sent to the React UI, diagnostic history or telemetry. One request is outstanding at a time with a two-second deadline. Timeout stops predictions until the keyboard is reopened; ordinary keyboard operation remains available. Closing the keyboard, ending scanning, or exiting kills and reaps the worker.

### Spare worker

Loading the model takes about a second, so a keyboard that had a working worker leaves a spare behind when it closes. The worker that held the typed text is killed and reaped first. A new process is then started, which loads the model and waits. It has received no request, so it holds no text, and it starts its activity observer only on its first request, so it observes nothing while it waits. The next keyboard open adopts it and suggestions are ready without a reload.

At most one spare exists, and only between keyboard opens. It is killed and reaped after two minutes without a keyboard open, whenever scanning ends or restarts, such as after saving settings, when Word prediction is turned off, on Retry predictions and on exit. A spare started for different switch keys, or one that has exited, is discarded and a new worker is started instead. A keyboard whose worker failed leaves no spare. In each of these cases, and on the first open of a scanning session, the next open loads the model again. A keyboard reopened within about a second adopts a spare that is still loading and shows the loading badge until it finishes. While it waits, the spare holds the loaded model in memory, roughly 250 MB.

Before accepting a suggestion, the worker checks its token, edit revision, foreground identity and external activity. The main process checks its generation and foreground again before injection. Verification and native insertion cannot be atomic across applications; the target can still change in that short interval.

## Validation
Expand All @@ -81,6 +87,6 @@ For native validation, use disposable synthetic text in Notepad and a browser on
3. Change pages and docking, type punctuation and numbers, and return to Letters. Verify context survives these layout changes and Backspace edits the tracked buffer.
4. Use navigation, shortcuts, failed edits and external keyboard/mouse activity. Verify suggestions clear and a new first letter starts fresh context.
5. Change foreground apps with locked modifiers selected. Verify the keyboard stays open, modifiers and suggestions reset, and later keys go to the new app.
6. Close the keyboard, stop scanning, disconnect and exit. Verify input releases and the prediction worker exits. Test observer failure separately; typing should remain usable without predictions.
6. Close the keyboard, stop scanning, disconnect and exit. Verify input releases and the prediction worker exits. After closing the keyboard one spare worker remains; reopen after a few seconds and within two minutes, and verify the loading badge clears almost immediately, then verify the spare exits after two minutes idle and when scanning stops. Test observer failure separately; typing should remain usable without predictions.

Compilation and fake-adapter tests do not establish native application compatibility. Record live results separately.
2 changes: 1 addition & 1 deletion src-tauri/src/point_scan_runtime.rs
Original file line number Diff line number Diff line change
Expand Up @@ -151,7 +151,7 @@ impl Adapter for PointScan {
return result.map(|()| None);
}
Request::OpenKeyboard | Request::OpenMouse | Request::OpenPoint => {
crate::prediction::stop();
crate::prediction::close();
crate::scan_executor::activate(request)
}
Request::MouseDrag => {
Expand Down
Loading
Loading