Repository navigation
Add opt-in SmolLM2 companion library and CLI - #18
Merged
Merged
Conversation
enaboapps
marked this pull request as ready for review
October 3, 2026 20:07
enaboapps
marked this pull request as draft
October 3, 2026 20:24
enaboapps
marked this pull request as ready for review
October 3, 2026 20:37
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Add an opt-in SmolLM2-135M Q8 companion library and CLI. Applications receive existing statistical suggestions immediately, followed by asynchronous refinement from a persistent isolated worker. The root Predictor/Suggestion APIs, personal learning and SQLite formats remain unchanged.
The companion bounds requests and queues, suppresses stale results, enforces a 500 ms worker deadline, supports explicit recovery and verifies model bytes against compiled pins. A portable parent checks AVX2/FMA/F16C support before dispatching to the separately built optimized worker. Runtime prediction is offline and in memory. Model assets remain separate and have not been published.
Qualification result
This candidate is not production-qualified and remains opt-in. The fixed policy improved aggregate top-five accuracy but failed the agreed one-percentage-point regression limit in two small development cells. No policy tuning, automatic promotion, desktop integration or release was performed.
Optimized top-five accuracy increased from 47.39% to 52.33%. Two development cells each lost one hit out of eight. Linux and both macOS architectures have build/test/package coverage; actual model performance on those platforms is unmeasured.
Qualification report and reproducible commands link to machine-readable results and the frozen protocol. The model converter reproduced the pinned Q8 SHA-256 exactly. Binary packages include qualification status and license notices, without model assets.
The converter supports only the pinned Q8 format; the unused worker failure reply and experimental F32 conversion option have been removed.
Validation
8b201d342940d2233a6df71f1175d2a457366815. All 13 GitHub checks passed on this cleanup head. Local formatting, Clippy, companion tests, 15 Python tests and exact Q8 conversion hash verification also passed. Findings on stale results, worker paths and warm-up failure reporting were addressed.Closes #17