Skip to content

Add bounded asynchronous word generation - #24

Merged
enaboapps merged 5 commits into
mainfrom
codex/neural-generation
Oct 4, 2026
Merged

enaboapps merged 5 commits into
mainfrom
codex/neural-generation

Conversation

@enaboapps

@enaboapps enaboapps commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Adds asynchronous SmolLM2 whole-word generation for three additional keyboard slots while preserving statistical and reranking APIs. Protocol 2 uses bounded byte-level beam search with normalized prefix constraints and exclusions, including incomplete UTF-8 handling. Existing model/tokenizer hashes remain unchanged. All package versions are aligned to v0.2.1 for a separately authorized release.

Closes #23.

Final head c25f068 passed independent read-only review with no actionable findings. Local Rust formatting, warning-denied Clippy, tests and 16 Python checks passed. Root CI passed on Windows, macOS and Linux. Neural tests, audits and all four platform worker packages passed; final Windows/macOS qualification also passed in run 37212539426.

The reviewed generation implementation completed 1,000 warmed queries on Windows portable and macOS ARM with no generation failures and zero generation context-cache hits. Generation p95 was 1388 ms and 767 ms respectively, below the unchanged two-second reply deadline. Neural slot fill was 93.5% and 93.4%. On the sampled frozen corpus, six-slot top-six accuracy was 72.6–72.7%, versus 48.2% for six statistical suggestions and 51.8–52.0% for the existing reranker with at most five results. Preserving statistical first-three ordering reduces top-three from 50.1–50.3% to 39.9%. Individual regressions and historical quality qualification failures remain. Unknown pretraining overlap prevents unseen-data accuracy claims.

These are 1,000-query samples of the 1,760-query workload; omit --samples for the full comparison. Reports cover normalized duplicates, OOV query instances, gains/regressions, failures, latency and process-tree RSS. The Windows local run overlapped development checks; final CI repeats measurement. Machine-readable reports and a concise comparison are included in dependent switchifyapp/switchify-pc#974 under docs/measurements/six-predictions and docs/six-predictions.md. Reproduce with scripts/generation_evaluate.py and explicit --cli, --baseline, --bundle, --worker, --samples 1000 and --output paths. Automated tests use fake adapters and never inject desktop input.

Release notes and generation-capable worker archives are prepared. This PR does not authorize merging, public artifact publication, a desktop RC release or promotion of model quality. Desktop final pins use checksum-verified prepared archives until separately authorized publication provides permanent URLs.

Final-head evidence: Neural CI and qualification and root CI passed. Final 1,000-query generation runs had zero generation/warmup failures on both platforms; each comparison reranker failed once. Windows portable generation median/p95/max was 1643.66/1676.94/1712.09 ms, with 690,319,360 bytes peak process-tree RSS. macOS ARM was 689.82/1466.30/1868.47 ms with 909,574,144 bytes peak RSS. Cold-load times were 4.02 and 5.04 seconds.

Final top-six accuracy remained 72.6% Windows and 72.7% macOS. Neural fill was 2749/3000 and 2803/3000. Generated OOV occurrences were 654 and 701; each run recovered 29/63 OOV target query instances. The bounded search can return fewer words on slower runners, and timings vary by machine. The final JSON reports are attached to that CI run as generation-windows-latest and generation-macos-latest. Earlier checked-in reports remain labeled with their measured source and are not substituted for this final evidence.

@enaboapps enaboapps added this to the v0.2.1 milestone Oct 4, 2026
@enaboapps
enaboapps marked this pull request as ready for review October 4, 2026 16:07
@enaboapps
enaboapps merged commit 2a2ed0c into main Oct 4, 2026
15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add bounded neural whole-word generation

2 participants