Skip to content

Add opt-in SmolLM2 companion library and CLI - #18

Merged
enaboapps merged 8 commits into
mainfrom
codex/smollm-production
Oct 4, 2026
Merged

enaboapps merged 8 commits into
mainfrom
codex/smollm-production

Conversation

@enaboapps

@enaboapps enaboapps commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Add an opt-in SmolLM2-135M Q8 companion library and CLI. Applications receive existing statistical suggestions immediately, followed by asynchronous refinement from a persistent isolated worker. The root Predictor/Suggestion APIs, personal learning and SQLite formats remain unchanged.

The companion bounds requests and queues, suppresses stale results, enforces a 500 ms worker deadline, supports explicit recovery and verifies model bytes against compiled pins. A portable parent checks AVX2/FMA/F16C support before dispatching to the separately built optimized worker. Runtime prediction is offline and in memory. Model assets remain separate and have not been published.

Qualification result

This candidate is not production-qualified and remains opt-in. The fixed policy improved aggregate top-five accuracy but failed the agreed one-percentage-point regression limit in two small development cells. No policy tuning, automatic promotion, desktop integration or release was performed.

Reference Windows measurement Optimized Portable
Queries 1,760 1,760
Successful neural refinements 1,759 1,757
Failed requests 0 2
Immediate p95 7.95 ms 7.89 ms
Refinement p95 including IPC 122.44 ms 327.69 ms

Optimized top-five accuracy increased from 47.39% to 52.33%. Two development cells each lost one hit out of eight. Linux and both macOS architectures have build/test/package coverage; actual model performance on those platforms is unmeasured.

Qualification report and reproducible commands link to machine-readable results and the frozen protocol. The model converter reproduced the pinned Q8 SHA-256 exactly. Binary packages include qualification status and license notices, without model assets.

The converter supports only the pinned Q8 format; the unused worker failure reply and experimental F32 conversion option have been removed.

Validation

  • Root and companion formatting, Clippy with warnings denied, Rust tests and Python tests passed locally.
  • Fake processes cover stalls, blocked input pipes, crashes, bad/oversized replies, queue replacement, session changes, rejected-input invalidation, explicit retry and shutdown. No keyboard or pointer injection.
  • Explicit local-model fixtures passed for both portable and optimized workers, including context reuse and reset. Kernel-specific order fixtures reflect observed Q8 differences.
  • The root audit remains strict. The companion has one documented maintenance-only exception for RUSTSEC-2024-0436, paste 1.0.15.
  • Independent agent review is clear on head 8b201d342940d2233a6df71f1175d2a457366815. All 13 GitHub checks passed on this cleanup head. Local formatting, Clippy, companion tests, 15 Python tests and exact Q8 conversion hash verification also passed. Findings on stale results, worker paths and warm-up failure reporting were addressed.
  • The downloaded Windows CI package passed archive and 13-file checksum verification, model-bundle validation, and actual portable/optimized prediction smoke tests on the reference machine.

Closes #17

@enaboapps enaboapps added this to the v0.2.0 milestone Oct 3, 2026
@enaboapps
enaboapps marked this pull request as ready for review October 3, 2026 20:07
@enaboapps
enaboapps marked this pull request as draft October 3, 2026 20:24
@enaboapps
enaboapps marked this pull request as ready for review October 3, 2026 20:37
@enaboapps
enaboapps merged commit 2595d66 into main Oct 4, 2026
13 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add opt-in production SmolLM2 companion library and CLI

2 participants