Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
36 changes: 36 additions & 0 deletions .github/workflows/neural.yml
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,42 @@ jobs:
name: neural-${{ matrix.os }}
path: artifacts/neural-cli/*
if-no-files-found: error
generation-qualification:
strategy:
fail-fast: false
matrix:
os: [windows-latest, macos-latest]
runs-on: ${{ matrix.os }}
timeout-minutes: 120
steps:
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
- uses: dtolnay/rust-toolchain@4716b85f2fac3e324e64fa2810f6b5c3905760a5
- uses: Swatinem/rust-cache@6323deb102c322ba6fcbdcafc7e3dddab59af2b6
with:
workspaces: neural -> ../target/qualification
- uses: actions/setup-python@a26af69be951a213d495a4c3e4e4022e16d87065
with:
python-version: '3.13'
- run: python -m pip install psutil==7.0.0 regex==2025.11.3
- run: cargo build --manifest-path neural/Cargo.toml --workspace --release --locked --target-dir target/qualification
- name: Prepare pinned offline bundle
shell: bash
run: |
ext=""
if [ "$RUNNER_OS" = "Windows" ]; then ext=".exe"; fi
python scripts/prepare_generation_qualification.py --quantize "target/qualification/release/quantize$ext" --output target/generation-inputs
- name: Measure 1000 warmed queries and compare frozen fixtures
shell: bash
run: |
ext=""
if [ "$RUNNER_OS" = "Windows" ]; then ext=".exe"; fi
python scripts/generation_evaluate.py --cli "target/qualification/release/switchify-prediction-neural$ext" --worker "target/qualification/release/switchify-smol-worker$ext" --baseline target/generation-inputs/english.sqlite --bundle target/generation-inputs/bundle --samples 1000 --output "artifacts/generation-${{ runner.os }}.json"
- uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02
if: always()
with:
name: generation-${{ matrix.os }}
path: artifacts/generation-*.json
if-no-files-found: error
neural-audit:
runs-on: ubuntu-latest
steps:
Expand Down
2 changes: 1 addition & 1 deletion Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

2 changes: 1 addition & 1 deletion Cargo.toml
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
[package]
name = "switchify-prediction"
version = "0.2.0"
version = "0.2.1"
edition = "2024"
rust-version = "1.97.1"
license = "MIT"
Expand Down
23 changes: 23 additions & 0 deletions docs/releases/v0.2.1.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,23 @@
# v0.2.1

Adds asynchronous whole-word generation to the offline SmolLM2 companion while
preserving the statistical and reranking APIs and SQLite formats. Generation
uses protocol 2, so the matching v0.2.1 workers are required. Model and tokenizer
bytes are unchanged.

The generation request carries context, prefix, session identity, up to three
excluded instant words and a result limit of up to three. Search uses eight
beams, up to eight tokens per word and 64 forward evaluations. Results are
normalized whole words with a probable following boundary. Words need not
belong to the statistical vocabulary. Empty results are valid.

The inference/reset reply deadline increases from 500 ms in v0.2.0 to two
seconds. Startup remains bounded to 30 seconds. Failure requires explicit retry. No network access is needed at runtime.

Generation is not automatically production-qualified. Corpus comparisons are
regression tests with unknown pretraining overlap. Existing reranking quality
regressions remain documented in the historical qualification reports.
Prepared CI bundles are unsigned; platform applications sign their embedded
workers through their own packaging workflows.

Publication of this release and the Switchify PC RC require separate approval.
13 changes: 13 additions & 0 deletions docs/six-slot-generation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
# Six-slot generation qualification

Issue #23 adds generation while preserving statistical and reranking APIs.
The first three slots belong to the unchanged statistical predictor. The next
three are generated words, excluding the normalized instant words. Empty positions
stay empty.

The worker uses protocol 2 and the unchanged pinned Q8 model/tokenizer. See
neural/model-bundle.json for the frozen search policy. Reports are produced by
scripts/generation_evaluate.py on the existing frozen fixtures. Unknown overlap
with model pretraining prevents claims of unseen-data accuracy.

Validation and platform measurements are in progress. No release is published.
8 changes: 5 additions & 3 deletions neural/Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

2 changes: 1 addition & 1 deletion neural/Cargo.toml
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
[package]
name = "switchify-prediction-neural"
version = "0.2.0"
version = "0.2.1"
edition = "2024"
rust-version = "1.97.1"
license = "MIT"
Expand Down
32 changes: 30 additions & 2 deletions neural/README.md
Original file line number Diff line number Diff line change
@@ -1,10 +1,10 @@
# SmolLM2 companion

An optional library and CLI for immediate statistical suggestions followed by offline SmolLM2 refinement. The root predictor, learning APIs and SQLite formats are unchanged. This package has its own workspace and lockfile. It is not integrated with Switchify PC.
An optional library and CLI for immediate statistical suggestions followed by offline SmolLM2 refinement. The root predictor, learning APIs and SQLite formats are unchanged. This package has its own workspace and lockfile. Switchify PC integration is maintained in its separate repository.

This candidate is not yet production-qualified. The optimized Windows worker passed the latency target and improved overall top-five accuracy, but two development cells failed the frozen quality gate. See `docs/smol-production-results.md` in the repository for the full comparison and platform limits. `Ready` means the worker is available, not that the quality gate has passed.

The fixed policy reranks eight statistical candidates using SmolLM2-135M Q8. It scores every token of a word plus the probability of a following word boundary, uses the current sentence capped at 64 tokens and returns at most five ordered words. Scores from different models are never combined or exposed as shared probabilities. Limits above five are errors. Minimum grapheme settings are honored; unigram-only requests stay statistical. The same caller-owned predictor supplies immediate suggestions and the shortlist, including its current personal snapshot.
The preserved reranking policy reranks eight statistical candidates using SmolLM2-135M Q8. It scores every token of a word plus the probability of a following word boundary, uses the current sentence capped at 64 tokens and returns at most five ordered words. Scores from different models are never combined or exposed as shared probabilities. Limits above five are errors. Minimum grapheme settings are honored; unigram-only requests stay statistical. The same caller-owned predictor supplies immediate suggestions and the shortlist, including its current personal snapshot.

The controller returns immediate words synchronously. A persistent child loads the model once and refines through private bounded pipes. There is one active request and one replaceable pending request. `poll()` yields only the latest request; callers should also match its request ID before displaying it. Session changes and `reset()` invalidate old results and clear the worker's context. Loading is separate from the 2 second reply deadline for inference and reset. Startup has a 30 second bound. Failure stops the worker and leaves immediate suggestions available. Recovery requires `retry()`. `shutdown()` kills and joins the worker and clears pending text.

Expand Down Expand Up @@ -64,3 +64,31 @@ python scripts/neural_evaluate.py --cli target/smol-portable/release/switchify-p
Add `--accelerated-worker target/smol-avx2/release/switchify-smol-worker` for the optimized comparison. Each run contains 1,760 warmed queries with personal learning disabled. Reports include quality cells, immediate and IPC-inclusive refinement timings, cache hits/misses, cold load and sampled process-tree RSS. Training overlap is checked; unknown neural pretraining overlap remains possible. These are regression comparisons and do not prove unseen-data accuracy. A failed quality or latency gate prevents a production-quality claim; it does not cause test-set tuning or automatic promotion.

`python scripts/package_neural.py --portable target/smol-portable/release --accelerated target/smol-avx2/release` packages binaries, hashes and dependency notices without model files. Omit the accelerated path for ARM. CI builds/tests Windows x64, Linux x64, macOS ARM64 and macOS x64. Build/test success is not a claim of measured model latency on those platforms. See `SECURITY.md` for the scoped dependency advisory exception and deployment boundaries.


## Asynchronous generation

The generation API returns up to three additional normalized whole words without
requiring statistical-vocabulary membership. Call
Refiner::generate(before, prefix, session, instant_words, 3), then poll using the
returned request ID. A subsequent submit, generate or reset invalidates old results.
The existing submit/reranking API is unchanged. The stream CLI accepts
{"command":"generate","before":"please send the","prefix":"","session":1} on stdin.

Worker protocol 2 is required; protocol 1 workers fail cleanly. Model weights and
tokenizer hashes are unchanged. The manifest records beam width 8, at most 8 tokens
per word, at most 64 forward evaluations including uncached context, and a 1600 ms
search cutoff within the unchanged 2000 ms parent reply deadline. Results are
ranked by whole-word probability including following boundary mass. A boundary
probability of at least 0.5 excludes likely unfinished fragments. Search can return
fewer than three words, including none. This is English-focused and not a spelling
dictionary; plausible but incorrect words remain possible.

Run scripts/generation_evaluate.py with explicit --cli, --worker, --baseline,
--bundle and --output paths. Optional --samples 1000 selects a reproducible spread
across the existing frozen corpus partitions and zero through four graphemes.
Omit --samples for the full comparison. Install psutil==7.0.0 and regex==2025.11.3.
The report includes top-three/top-six counts, fill, OOV, regressions, latency and
process-tree RSS. This is a regression comparison, not unseen-data qualification.
Existing model-quality failures remain; generation has not been promoted to a
qualified model. CI artifacts are prepared without publishing a release.
11 changes: 11 additions & 0 deletions neural/fixtures/baseline-pin.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,11 @@
{
"archive": {
"url": "https://github.com/switchifyapp/switchify-prediction/releases/download/v0.1.0/switchify-english-en-aac-oanc-v1.zip",
"sha256": "4537c70b44f553b31e618378a5cc500cd83ffa2dff940263ecabf88e18945b17",
"bytes": 10393776
},
"database": {
"sha256": "222253417d0a7a705823ffb7e599a3bcf5d5d3daf4a9d76161ac6b3e555aeaad",
"bytes": 29802496
}
}
14 changes: 13 additions & 1 deletion neural/model-bundle.json
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,19 @@
"context_tokens": 64,
"shortlist": 8,
"max_results": 5,
"scoring": "whole-word log probability plus boundary probability, sequential"
"scoring": "whole-word log probability plus boundary probability, sequential",
"generation": {
"protocol_version": 2,
"beam_width": 8,
"max_tokens_per_word": 8,
"max_forward_evaluations": 64,
"max_results": 3,
"max_word_bytes": 128,
"reply_deadline_ms": 2000,
"scoring": "whole-word probability including following boundary; normalized duplicates excluded",
"minimum_boundary_probability": 0.5,
"search_cutoff_ms": 1600
}
},
"files": {
"model.gguf": {
Expand Down
5 changes: 4 additions & 1 deletion neural/qualification.json
Original file line number Diff line number Diff line change
Expand Up @@ -9,5 +9,8 @@
],
"optimized_windows_latency_passed": true,
"desktop_integration": false,
"automatic_promotion": false
"automatic_promotion": false,
"scope": "Historical v0.2.0 reranking qualification; retained unchanged reasons are not generation qualification claims.",
"generation_report_command": "python scripts/generation_evaluate.py",
"generation_production_qualified": false
}
50 changes: 46 additions & 4 deletions neural/src/lib.rs
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ pub mod bundle;
mod process;
pub mod protocol;

use protocol::Query;
use protocol::{Command, GenerationQuery, Query};
use serde::Serialize;
use std::{
path::PathBuf,
Expand Down Expand Up @@ -74,7 +74,7 @@ pub struct Refined {
struct Shared {
latest: u64,
session: u64,
pending: Option<Query>,
pending: Option<Command>,
result: Option<Refined>,
status: Status,
reset: bool,
Expand Down Expand Up @@ -224,13 +224,13 @@ impl Refiner {
&& !candidates.is_empty()
&& matches!(shared.status, Status::Loading | Status::Ready);
if refinement_requested {
shared.pending = Some(Query {
shared.pending = Some(Command::Predict(Query {
id: shared.latest,
session,
before: effective_context(before),
candidates,
limit: options.limit,
});
}));
}
let result = Immediate {
request_id: shared.latest,
Expand All @@ -242,6 +242,48 @@ impl Refiner {
Ok(result)
}

/// Generate additional whole words without a statistical vocabulary restriction.
/// Results use the same request-ID-based poll method as reranking.
/// None means generation is unavailable until an explicit retry.
pub fn generate(
&mut self,
before: &str,
prefix: &str,
session: u64,
exclude: &[String],
limit: usize,
) -> Result<Option<u64>> {
let mut query = GenerationQuery {
id: 0,
session,
before: before.to_owned(),
prefix: normalize(prefix),
exclude: exclude.iter().map(|s| normalize(s)).collect(),
limit,
};
if before.len() > 16_384 || prefix.len() > 256 || !query.valid() {
self.reset();
return Err(Error::Input);
}
query.before = effective_context(before);
let mut shared = self.state.0.lock().unwrap();
shared.latest = shared.latest.checked_add(1).ok_or(Error::Input)?;
shared.pending = None;
shared.result = None;
if shared.session != session {
shared.session = session;
shared.reset = true;
}
let id = shared.latest;
let requested = limit > 0 && matches!(shared.status, Status::Loading | Status::Ready);
if requested {
query.id = id;
shared.pending = Some(Command::Generate(query));
}
self.state.1.notify_one();
Ok(requested.then_some(id))
}

/// A subsequent submit/reset invalidates any previously unconsumed result.
pub fn poll(&mut self) -> Option<Refined> {
self.state.0.lock().unwrap().result.take()
Expand Down
38 changes: 38 additions & 0 deletions neural/src/main.rs
Original file line number Diff line number Diff line change
Expand Up @@ -47,6 +47,11 @@ struct Run {
#[derive(Deserialize)]
#[serde(tag = "command", rename_all = "snake_case", deny_unknown_fields)]
enum Input {
Generate {
before: String,
prefix: String,
session: u64,
},
Predict {
before: String,
prefix: String,
Expand Down Expand Up @@ -140,6 +145,39 @@ fn run(args: Run, once: bool) -> Result<(), ()> {
continue;
}
match rx.recv_timeout(Duration::from_millis(1)) {
Ok(Ok(Input::Generate {
before,
prefix,
session,
})) => {
started = Instant::now();
predicted = true;
let statistical: Vec<String> = predictor
.predict(
&before,
&prefix,
Options {
limit: 6,
min_chars: 0,
unigram_only: false,
},
)
.into_iter()
.map(|s| s.word)
.collect();
let words: Vec<String> = statistical.iter().take(3).cloned().collect();
let id = engine
.generate(&before, &prefix, session, &words, 3)
.map_err(|_| ())?;
outstanding = id.is_some();
emit(json!({"type":"immediate", "result": {
"request_id":id, "words":words, "statistical_six":statistical,
"status":engine.status(), "refinement_requested":outstanding
}, "elapsed_ms":started.elapsed().as_secs_f64()*1000.}))?;
if once {
eof = true;
}
}
Ok(Ok(Input::Predict {
before,
prefix,
Expand Down
Loading
Loading