Repository navigation
Add bounded asynchronous word generation - #24
Merged
Merged
Conversation
enaboapps
marked this pull request as ready for review
October 4, 2026 16:07
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds asynchronous SmolLM2 whole-word generation for three additional keyboard slots while preserving statistical and reranking APIs. Protocol 2 uses bounded byte-level beam search with normalized prefix constraints and exclusions, including incomplete UTF-8 handling. Existing model/tokenizer hashes remain unchanged. All package versions are aligned to v0.2.1 for a separately authorized release.
Closes #23.
Final head c25f068 passed independent read-only review with no actionable findings. Local Rust formatting, warning-denied Clippy, tests and 16 Python checks passed. Root CI passed on Windows, macOS and Linux. Neural tests, audits and all four platform worker packages passed; final Windows/macOS qualification also passed in run 37212539426.
The reviewed generation implementation completed 1,000 warmed queries on Windows portable and macOS ARM with no generation failures and zero generation context-cache hits. Generation p95 was 1388 ms and 767 ms respectively, below the unchanged two-second reply deadline. Neural slot fill was 93.5% and 93.4%. On the sampled frozen corpus, six-slot top-six accuracy was 72.6–72.7%, versus 48.2% for six statistical suggestions and 51.8–52.0% for the existing reranker with at most five results. Preserving statistical first-three ordering reduces top-three from 50.1–50.3% to 39.9%. Individual regressions and historical quality qualification failures remain. Unknown pretraining overlap prevents unseen-data accuracy claims.
These are 1,000-query samples of the 1,760-query workload; omit --samples for the full comparison. Reports cover normalized duplicates, OOV query instances, gains/regressions, failures, latency and process-tree RSS. The Windows local run overlapped development checks; final CI repeats measurement. Machine-readable reports and a concise comparison are included in dependent switchifyapp/switchify-pc#974 under docs/measurements/six-predictions and docs/six-predictions.md. Reproduce with scripts/generation_evaluate.py and explicit --cli, --baseline, --bundle, --worker, --samples 1000 and --output paths. Automated tests use fake adapters and never inject desktop input.
Release notes and generation-capable worker archives are prepared. This PR does not authorize merging, public artifact publication, a desktop RC release or promotion of model quality. Desktop final pins use checksum-verified prepared archives until separately authorized publication provides permanent URLs.
Final-head evidence: Neural CI and qualification and root CI passed. Final 1,000-query generation runs had zero generation/warmup failures on both platforms; each comparison reranker failed once. Windows portable generation median/p95/max was 1643.66/1676.94/1712.09 ms, with 690,319,360 bytes peak process-tree RSS. macOS ARM was 689.82/1466.30/1868.47 ms with 909,574,144 bytes peak RSS. Cold-load times were 4.02 and 5.04 seconds.
Final top-six accuracy remained 72.6% Windows and 72.7% macOS. Neural fill was 2749/3000 and 2803/3000. Generated OOV occurrences were 654 and 701; each run recovered 29/63 OOV target query instances. The bounded search can return fewer words on slower runners, and timings vary by machine. The final JSON reports are attached to that CI run as generation-windows-latest and generation-macos-latest. Earlier checked-in reports remain labeled with their measured source and are not substituted for this final evidence.