Skip to content

feat(inference/mlx): on-device self-speculative decoding + OpenAI server (Apple Silicon) - #4

Open
guks-trl wants to merge 9 commits into
mainfrom
mlx-on-device
Open

guks-trl wants to merge 9 commits into
mainfrom
mlx-on-device

Conversation

@guks-trl

Copy link
Copy Markdown
Collaborator

Summary

Adds inference/mlx/: an on-device (Apple Silicon, MLX) backend for Trida2.0-4B with lossless
self-speculative decoding, plus an OpenAI-compatible server usable from agent harnesses (tested with Hermes Agent).

  • Decoder: bd_bidir_shift b7/g4 semantics of HybridDiffusionSelfSpec, on top of mlx-lm's Qwen3.5 model.
    Attention rows 0..N-1 causal, MASK rows see the whole canvas. Gated-delta clean rows read h[t], MASK rows read h[block_end].
    Nothing persists until commit(adv). Fused canvas GDN Metal kernel (adapted from mlx-lm, MIT). Greedy + top-k/top-p speculative sampling.
  • Server: /v1/chat/completions + /v1/completions, SSE, reasoning split, tool calls → OpenAI tool_calls,
    context_length on /v1/models, 3-slot prompt cache with a system-block snapshot, single MLX worker thread.
  • Tools: trida-mlx-convert (q8/q6/q4), -verify, -bench, -profile, -agent; uv project.
  • Docs: README, CHANGELOG, NOTICE/COMPLIANCE (mlx-lm kernel, chat-template test fixture).

Results (M3 Pro 18 GB, q8, greedy, 512 tokens)

causal self-spec
decode tok/s 27.1 49.0 (≈1.8–2.0×)
tokens / forward 1.00 2.32
Self-spec step = 1.13× an AR step (44.5 vs 39.3 ms). Canvas rows == AR logits within bf16 noise.
Greedy output matches plain decoding except where the top-2 logits tie exactly in bf16.

Testing

  • cd inference/mlx && uv run pytest: 20 offline tests on a tiny random model (CPU or GPU).
  • Real checkpoint: trida-mlx-verify / trida-mlx-bench on M3 Pro. Hermes Agent end to end with the full default toolset.

Notes

  • The checkpoint is text-only (Qwen3_5ForCausalLM); vision tools need a separate model.
  • The decoder reimplements the HybridDiffusion (PolyForm-NC) self-spec algorithm independently; no upstream code is copied. Flagging for license review.
  • Single stream, no batching.

guks-trl and others added 4 commits September 30, 2026 01:29
…icon

Serve Trida2.0-4B's lossless self-spec (bd_bidir_shift, block 7 / gen 4,
strict_truncated) with MLX, on top of mlx-lm's Qwen3.5 model and Metal
gated-delta kernel:

- canvas forward: full-attention rows 0..N-1 causal, MASK rows see the whole
  canvas; GDN clean rows read h[t], MASK rows read h[block_end]; nothing is
  persisted until commit(adv) (KV trim, GDN state of the accepted prefix,
  conv window from the saved pre-conv rows)
- fused canvas GDN Metal kernel (one launch per layer; verify step 49.8 ->
  44.5 ms on M3 Pro q8, 1.13x an AR step)
- greedy and top-k/top-p speculative sampling, AR baseline, prompt cache
  with slots and snapshots (GDN layers cannot rewind)
- convert (MLX q8/q6/q4), verify, bench, profile tools; uv project
- tests on a tiny random model: canvas rows == AR logits, block-end readout,
  commit == AR state, greedy lossless, sampled distribution == AR

M3 Pro 18 GB, q8, greedy: 27 -> 49 tok/s (2.3 tokens per forward).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013o6LMxkLD6mBHGy1C6MBR7
…port

- /v1/chat/completions and /v1/completions with SSE streaming, reasoning
  split, qwen3-coder / JSON tool calls parsed into OpenAI tool_calls with
  schema-typed parameters, stop sequences, per-request thinking control
- /v1/models advertises context_length; context_length_exceeded 400s
- all MLX work on one worker thread (MLX streams are per-thread; the
  threaded HTTP server crashed on chunked prefills without it)
- Hermes Agent compatibility: template-safe message normalisation,
  prefill keep-alives, 3 cache slots + a system-block snapshot so side
  requests and rewritten user turns keep the tool-schema prefix cached
- minimal stdlib agent loop (files, search, calculator, opt-in shell)
- offline end-to-end tests on a tiny raw-HF checkpoint (loader, convert
  q8, prompt cache, server, parsing, the real Trida chat template)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013o6LMxkLD6mBHGy1C6MBR7
…ty code

- inference/mlx/README.md: quickstart (uv), design mapping to the SGLang
  reference, prompt cache, Hermes setup, measured M3 Pro numbers, limits
- root README layout + CHANGELOG entry
- NOTICE / COMPLIANCE: the mlx-lm-derived Metal kernel (MIT) and the
  bundled chat-template test fixture (Apache-2.0); MLX + mlx-lm as
  runtime dependencies

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013o6LMxkLD6mBHGy1C6MBR7
guks-trl and others added 2 commits September 30, 2026 09:21
Remove inference/mlx/tests (tiny-model decoder/server tests and the chat
template fixture) and everything that referenced it: the README Tests
section, the pytest dev group and config (uv.lock re-locked), and the
fixture's NOTICE / COMPLIANCE entries.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013o6LMxkLD6mBHGy1C6MBR7
…da-2.0-4B-1006)

- trida_mlx/vision.py: Qwen3.5 vision encoder + projector (adapted from
  mlx-vlm, MIT), Qwen2-VL preprocessing (bit-identical to the HF PIL
  processor), interleaved multimodal RoPE
- runtime: image tokens are per-image hash keys, so the prompt cache never
  confuses two images and follow-up turns don't re-encode them; prompts
  with images prefill with MRoPE, SeqCache.pos tracks the next RoPE
  position, decode and the self-spec canvas stay plain RoPE
- loader/convert: load the vision tower when the checkpoint has one; keep
  it bf16 when quantizing; carry processor configs and transformers_version
- server: OpenAI image_url parts (data URL / http / path), 400 on bad
  images, supports_vision on /v1/models, --max-image-pixels
- verify --image; self-spec default stays N=4 (on M3 Pro q8 N=8 is 2.72 vs
  2.40 tok/fwd but 43 vs 52 tok/s); README block-size table, NOTICE /
  COMPLIANCE for the mlx-vlm-derived encoder, pillow dependency

Checked: full multimodal prefill matches mlx-vlm's Qwen3.5 to 4e-6 (fp32,
tiny model); on M3 Pro the 1006 q8 model describes a screenshot correctly
and greedy self-spec matches AR with a 975-token image prompt.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013o6LMxkLD6mBHGy1C6MBR7
usik-luke-trl added a commit that referenced this pull request Oct 6, 2026
…wered alert

Prepares the chat channel without opening one. The guideline's first instruction is
"첫 주에는 서버부터 만들지 않는다" -- purpose, channel map, roles and response
expectations are settled before any server exists, and this is that work.

Public-facing
- COMMUNITY.md: where to ask what, and what to expect back. The chat row is left
  empty on purpose rather than carrying a placeholder link.
- README gains a help/contribute/report section. The repository checklist is judged
  from a logged-out browser, and until now the README linked to neither
  CONTRIBUTING.md nor SECURITY.md -- the first checklist item, failing.

Operations (docs/community/)
- CHANNEL_MATRIX: which system each kind of enquiry COMPLETES in, the order to decide
  by, what a transition comment must carry, and alarm design.
- CHANNEL_POLICY: minimum channel structure, roles with least privilege, enforcement.
- MODERATOR_RUNBOOK: intake, S1-S4, incident records, the 8-step promotion procedure,
  and the four drills to run before opening.
- LABELS: what each label means AND when it comes off, which the checklist requires.

Automation
- community-check.yml: required files present, and the reporting routes still stated.
  It fails today's README, which is how the gap above was found rather than assumed.
- stale-unanswered.yml + tools/notify_stale.py: weekday mornings, items with no human
  response for 48h. Of everything the guideline says is worth sending to chat, this is
  the one that cannot be a GitHub notification -- it is the absence of an event.
  Dry run found a real case: PR #4 open 138 hours, unanswered.
- community_to_issue.yml: the chat-to-issue form, so promotion is a form rather than
  something remembered.

Discord or Slack is deliberately not decided here. Only the webhook differs, and
notify_stale posts to whichever secret is set; with neither it prints and does not fail.

Why promotion is written into the runbook rather than left as etiquette: M-001
measures response from GitHub comment timestamps. An answer given only in chat makes
the metric silently wrong AND loses the knowledge, which is the named risk of chat
channels -- 지식 소실, 검색·보존 제한. The 48h alert backstops it: an item answered
only in chat keeps reappearing until someone records it.

Also enabled Discussions and added the triage/community labels.

Not done, and listed as launch-hold conditions in docs/community/README.md: nobody
has verified that security@ and conduct@ actually deliver to two people each. CI can
check an address is stated; it cannot check mail arrives. The guideline says do not
open public invitations until that is confirmed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K599fUWZZYdiL4ekUq4Vhs
usik-luke-trl added a commit that referenced this pull request Oct 6, 2026
…wered alert

Prepares the chat channel without opening one. The guideline's first instruction is
"첫 주에는 서버부터 만들지 않는다" -- purpose, channel map, roles and response
expectations are settled before any server exists, and this is that work.

Public-facing
- COMMUNITY.md: where to ask what, and what to expect back. The chat row is left
  empty on purpose rather than carrying a placeholder link.
- README gains a help/contribute/report section. The repository checklist is judged
  from a logged-out browser, and until now the README linked to neither
  CONTRIBUTING.md nor SECURITY.md -- the first checklist item, failing.

Operations (docs/community/)
- CHANNEL_MATRIX: which system each kind of enquiry COMPLETES in, the order to decide
  by, what a transition comment must carry, and alarm design.
- CHANNEL_POLICY: minimum channel structure, roles with least privilege, enforcement.
- MODERATOR_RUNBOOK: intake, S1-S4, incident records, the 8-step promotion procedure,
  and the four drills to run before opening.
- LABELS: what each label means AND when it comes off, which the checklist requires.

Automation
- community-check.yml: required files present, and the reporting routes still stated.
  It fails today's README, which is how the gap above was found rather than assumed.
- stale-unanswered.yml + tools/notify_stale.py: weekday mornings, items with no human
  response for 48h. Of everything the guideline says is worth sending to chat, this is
  the one that cannot be a GitHub notification -- it is the absence of an event.
  Dry run found a real case: PR #4 open 138 hours, unanswered.
- community_to_issue.yml: the chat-to-issue form, so promotion is a form rather than
  something remembered.

Discord or Slack is deliberately not decided here. Only the webhook differs, and
notify_stale posts to whichever secret is set; with neither it prints and does not fail.

Why promotion is written into the runbook rather than left as etiquette: M-001
measures response from GitHub comment timestamps. An answer given only in chat makes
the metric silently wrong AND loses the knowledge, which is the named risk of chat
channels -- 지식 소실, 검색·보존 제한. The 48h alert backstops it: an item answered
only in chat keeps reappearing until someone records it.

Also enabled Discussions and added the triage/community labels.

Not done, and listed as launch-hold conditions in docs/community/README.md: nobody
has verified that security@ and conduct@ actually deliver to two people each. CI can
check an address is stated; it cannot check mail arrives. The guideline says do not
open public invitations until that is confirmed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K599fUWZZYdiL4ekUq4Vhs
usik-luke-trl added a commit that referenced this pull request Oct 6, 2026
…wered alert (#14)

* community: channel map, chat policy, moderator runbook, and the unanswered alert

Prepares the chat channel without opening one. The guideline's first instruction is
"첫 주에는 서버부터 만들지 않는다" -- purpose, channel map, roles and response
expectations are settled before any server exists, and this is that work.

Public-facing
- COMMUNITY.md: where to ask what, and what to expect back. The chat row is left
  empty on purpose rather than carrying a placeholder link.
- README gains a help/contribute/report section. The repository checklist is judged
  from a logged-out browser, and until now the README linked to neither
  CONTRIBUTING.md nor SECURITY.md -- the first checklist item, failing.

Operations (docs/community/)
- CHANNEL_MATRIX: which system each kind of enquiry COMPLETES in, the order to decide
  by, what a transition comment must carry, and alarm design.
- CHANNEL_POLICY: minimum channel structure, roles with least privilege, enforcement.
- MODERATOR_RUNBOOK: intake, S1-S4, incident records, the 8-step promotion procedure,
  and the four drills to run before opening.
- LABELS: what each label means AND when it comes off, which the checklist requires.

Automation
- community-check.yml: required files present, and the reporting routes still stated.
  It fails today's README, which is how the gap above was found rather than assumed.
- stale-unanswered.yml + tools/notify_stale.py: weekday mornings, items with no human
  response for 48h. Of everything the guideline says is worth sending to chat, this is
  the one that cannot be a GitHub notification -- it is the absence of an event.
  Dry run found a real case: PR #4 open 138 hours, unanswered.
- community_to_issue.yml: the chat-to-issue form, so promotion is a form rather than
  something remembered.

Discord or Slack is deliberately not decided here. Only the webhook differs, and
notify_stale posts to whichever secret is set; with neither it prints and does not fail.

Why promotion is written into the runbook rather than left as etiquette: M-001
measures response from GitHub comment timestamps. An answer given only in chat makes
the metric silently wrong AND loses the knowledge, which is the named risk of chat
channels -- 지식 소실, 검색·보존 제한. The 48h alert backstops it: an item answered
only in chat keeps reappearing until someone records it.

Also enabled Discussions and added the triage/community labels.

Not done, and listed as launch-hold conditions in docs/community/README.md: nobody
has verified that security@ and conduct@ actually deliver to two people each. CI can
check an address is stated; it cannot check mail arrives. The guideline says do not
open public invitations until that is confirmed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K599fUWZZYdiL4ekUq4Vhs

* docs(community): Discord launch plan

Turns the channel decision into an executable plan. Discord is settled -- public
invitation is the default, and the guideline warns against Slack's free tier as a
permanent knowledge store.

The structure that matters is the launch-hold section. Three conditions must hold
before public invitations open, and the first is unverified: nobody has confirmed
that security@ and conduct@ actually deliver, to two people each. A code of conduct
whose report address is dead is not operational, and community-check.yml can only
test that an address is stated.

Phases: merge #14 and assign roles -> build the server -> wire notifications ->
seed content -> limited release. Phase 3 is the largest, because FAQs and
first-issue candidates have to be real ones rather than filler.

Two platform constraints are written down rather than discovered later: GitHub
repository webhooks subscribe by event type and cannot filter on labels, so
"help wanted" needs an Actions workflow; and replying from Discord into GitHub is
not practically possible, which is why promotion is a documented duty.

Promotion is also what keeps M-001 honest -- it reads response time from GitHub
comment timestamps, so an answer given only in chat makes the metric wrong and
loses the knowledge. stale-unanswered.yml is the backstop: such an item keeps
reappearing every weekday until someone records it.

12 acceptance criteria, each checkable. Notably #12 requires one real promotion to
have happened, because a rule nobody has exercised is not yet a practice.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K599fUWZZYdiL4ekUq4Vhs

* docs(metrics): keep the definitions, drop the collection

#13 proposed collecting metrics in this repository -- a collector, tests, a monthly
workflow committing snapshots to main, and a baseline. That was wrong on three
counts. The team said they would handle collection themselves; a scheduled workflow
with contents: write pushing to the default branch is a real permission on a public
repo for data no outside reader consumes; and the current snapshot is mostly no_data
because there are no external contributors yet.

What stays is one page. It exists because three community documents cite it, and
because someone filing an issue should be able to see how their question is measured
and that they are not the thing being measured.

Definitions, the external-contributor scope and why it matters, why response time
and no-response rate are read together, the observation period with no targets yet,
why answering only in chat makes the numbers wrong, the explicit no-individual-
measurement stance, and the limitations.

No numbers, no collector, no workflow. Collection and reporting happen internally.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K599fUWZZYdiL4ekUq4Vhs

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
usik-luke-trl

This comment was marked as outdated.

@usik-luke-trl usik-luke-trl left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: tests claimed in the description are not in the tree. Two line comments below.

Comment thread inference/mlx/pyproject.toml
Comment thread inference/mlx/trida_mlx/vision.py Outdated
guks-trl and others added 3 commits October 6, 2026 06:39
- profile: --context N profiles at agent-length prefixes (14K-33K); adds
  timing-only bounds for a cheaper canvas attention mask (no mask, built-in
  causal) to size any mask-path optimization.
- bench: --prompts takes a JSONL of OpenAI chat requests (e.g. logged Hermes
  requests), normalized like the server, images included.

GQA head folding for canvas attention was tried and measured no gain on M3 Pro
at 28K context (68.29 vs 67.83 ms), so it is not included.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013o6LMxkLD6mBHGy1C6MBR7
mlx-vlm is by Prince Canuma (Blaizzy), not Apple; matches NOTICE.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013o6LMxkLD6mBHGy1C6MBR7
Brings back inference/mlx/tests (reverts d344fdc): tiny random Qwen3.5-hybrid
decoder invariants (test_tiny.py), tiny raw-HF checkpoint end to end
(test_e2e.py, updated for load_model's vision return and
_normalize_messages' image list), the chat-template fixture with its NOTICE /
COMPLIANCE entries, the pytest dev group and config, and the README Tests
section.

Adds test_vision.py on a tiny multimodal checkpoint (tiny_vl.py): image
expansion, greedy self-spec = greedy AR with an image, chunked prefill across
an image, MRoPE position advance, prompt cache never reusing across different
images, convert keeping the vision tower bit-identical, and the server
image_url path. 27 tests, ~30 s on CPU (uv run pytest).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013o6LMxkLD6mBHGy1C6MBR7
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants