Repository navigation
Conversation
…icon Serve Trida2.0-4B's lossless self-spec (bd_bidir_shift, block 7 / gen 4, strict_truncated) with MLX, on top of mlx-lm's Qwen3.5 model and Metal gated-delta kernel: - canvas forward: full-attention rows 0..N-1 causal, MASK rows see the whole canvas; GDN clean rows read h[t], MASK rows read h[block_end]; nothing is persisted until commit(adv) (KV trim, GDN state of the accepted prefix, conv window from the saved pre-conv rows) - fused canvas GDN Metal kernel (one launch per layer; verify step 49.8 -> 44.5 ms on M3 Pro q8, 1.13x an AR step) - greedy and top-k/top-p speculative sampling, AR baseline, prompt cache with slots and snapshots (GDN layers cannot rewind) - convert (MLX q8/q6/q4), verify, bench, profile tools; uv project - tests on a tiny random model: canvas rows == AR logits, block-end readout, commit == AR state, greedy lossless, sampled distribution == AR M3 Pro 18 GB, q8, greedy: 27 -> 49 tok/s (2.3 tokens per forward). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013o6LMxkLD6mBHGy1C6MBR7
…port - /v1/chat/completions and /v1/completions with SSE streaming, reasoning split, qwen3-coder / JSON tool calls parsed into OpenAI tool_calls with schema-typed parameters, stop sequences, per-request thinking control - /v1/models advertises context_length; context_length_exceeded 400s - all MLX work on one worker thread (MLX streams are per-thread; the threaded HTTP server crashed on chunked prefills without it) - Hermes Agent compatibility: template-safe message normalisation, prefill keep-alives, 3 cache slots + a system-block snapshot so side requests and rewritten user turns keep the tool-schema prefix cached - minimal stdlib agent loop (files, search, calculator, opt-in shell) - offline end-to-end tests on a tiny raw-HF checkpoint (loader, convert q8, prompt cache, server, parsing, the real Trida chat template) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013o6LMxkLD6mBHGy1C6MBR7
…ty code - inference/mlx/README.md: quickstart (uv), design mapping to the SGLang reference, prompt cache, Hermes setup, measured M3 Pro numbers, limits - root README layout + CHANGELOG entry - NOTICE / COMPLIANCE: the mlx-lm-derived Metal kernel (MIT) and the bundled chat-template test fixture (Apache-2.0); MLX + mlx-lm as runtime dependencies Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013o6LMxkLD6mBHGy1C6MBR7
# Conflicts: # CHANGELOG.md
Remove inference/mlx/tests (tiny-model decoder/server tests and the chat template fixture) and everything that referenced it: the README Tests section, the pytest dev group and config (uv.lock re-locked), and the fixture's NOTICE / COMPLIANCE entries. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013o6LMxkLD6mBHGy1C6MBR7
…da-2.0-4B-1006) - trida_mlx/vision.py: Qwen3.5 vision encoder + projector (adapted from mlx-vlm, MIT), Qwen2-VL preprocessing (bit-identical to the HF PIL processor), interleaved multimodal RoPE - runtime: image tokens are per-image hash keys, so the prompt cache never confuses two images and follow-up turns don't re-encode them; prompts with images prefill with MRoPE, SeqCache.pos tracks the next RoPE position, decode and the self-spec canvas stay plain RoPE - loader/convert: load the vision tower when the checkpoint has one; keep it bf16 when quantizing; carry processor configs and transformers_version - server: OpenAI image_url parts (data URL / http / path), 400 on bad images, supports_vision on /v1/models, --max-image-pixels - verify --image; self-spec default stays N=4 (on M3 Pro q8 N=8 is 2.72 vs 2.40 tok/fwd but 43 vs 52 tok/s); README block-size table, NOTICE / COMPLIANCE for the mlx-vlm-derived encoder, pillow dependency Checked: full multimodal prefill matches mlx-vlm's Qwen3.5 to 4e-6 (fp32, tiny model); on M3 Pro the 1006 q8 model describes a screenshot correctly and greedy self-spec matches AR with a 975-token image prompt. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013o6LMxkLD6mBHGy1C6MBR7
usik-luke-trl
added a commit
that referenced
this pull request
Oct 6, 2026
…wered alert Prepares the chat channel without opening one. The guideline's first instruction is "첫 주에는 서버부터 만들지 않는다" -- purpose, channel map, roles and response expectations are settled before any server exists, and this is that work. Public-facing - COMMUNITY.md: where to ask what, and what to expect back. The chat row is left empty on purpose rather than carrying a placeholder link. - README gains a help/contribute/report section. The repository checklist is judged from a logged-out browser, and until now the README linked to neither CONTRIBUTING.md nor SECURITY.md -- the first checklist item, failing. Operations (docs/community/) - CHANNEL_MATRIX: which system each kind of enquiry COMPLETES in, the order to decide by, what a transition comment must carry, and alarm design. - CHANNEL_POLICY: minimum channel structure, roles with least privilege, enforcement. - MODERATOR_RUNBOOK: intake, S1-S4, incident records, the 8-step promotion procedure, and the four drills to run before opening. - LABELS: what each label means AND when it comes off, which the checklist requires. Automation - community-check.yml: required files present, and the reporting routes still stated. It fails today's README, which is how the gap above was found rather than assumed. - stale-unanswered.yml + tools/notify_stale.py: weekday mornings, items with no human response for 48h. Of everything the guideline says is worth sending to chat, this is the one that cannot be a GitHub notification -- it is the absence of an event. Dry run found a real case: PR #4 open 138 hours, unanswered. - community_to_issue.yml: the chat-to-issue form, so promotion is a form rather than something remembered. Discord or Slack is deliberately not decided here. Only the webhook differs, and notify_stale posts to whichever secret is set; with neither it prints and does not fail. Why promotion is written into the runbook rather than left as etiquette: M-001 measures response from GitHub comment timestamps. An answer given only in chat makes the metric silently wrong AND loses the knowledge, which is the named risk of chat channels -- 지식 소실, 검색·보존 제한. The 48h alert backstops it: an item answered only in chat keeps reappearing until someone records it. Also enabled Discussions and added the triage/community labels. Not done, and listed as launch-hold conditions in docs/community/README.md: nobody has verified that security@ and conduct@ actually deliver to two people each. CI can check an address is stated; it cannot check mail arrives. The guideline says do not open public invitations until that is confirmed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01K599fUWZZYdiL4ekUq4Vhs
usik-luke-trl
added a commit
that referenced
this pull request
Oct 6, 2026
…wered alert Prepares the chat channel without opening one. The guideline's first instruction is "첫 주에는 서버부터 만들지 않는다" -- purpose, channel map, roles and response expectations are settled before any server exists, and this is that work. Public-facing - COMMUNITY.md: where to ask what, and what to expect back. The chat row is left empty on purpose rather than carrying a placeholder link. - README gains a help/contribute/report section. The repository checklist is judged from a logged-out browser, and until now the README linked to neither CONTRIBUTING.md nor SECURITY.md -- the first checklist item, failing. Operations (docs/community/) - CHANNEL_MATRIX: which system each kind of enquiry COMPLETES in, the order to decide by, what a transition comment must carry, and alarm design. - CHANNEL_POLICY: minimum channel structure, roles with least privilege, enforcement. - MODERATOR_RUNBOOK: intake, S1-S4, incident records, the 8-step promotion procedure, and the four drills to run before opening. - LABELS: what each label means AND when it comes off, which the checklist requires. Automation - community-check.yml: required files present, and the reporting routes still stated. It fails today's README, which is how the gap above was found rather than assumed. - stale-unanswered.yml + tools/notify_stale.py: weekday mornings, items with no human response for 48h. Of everything the guideline says is worth sending to chat, this is the one that cannot be a GitHub notification -- it is the absence of an event. Dry run found a real case: PR #4 open 138 hours, unanswered. - community_to_issue.yml: the chat-to-issue form, so promotion is a form rather than something remembered. Discord or Slack is deliberately not decided here. Only the webhook differs, and notify_stale posts to whichever secret is set; with neither it prints and does not fail. Why promotion is written into the runbook rather than left as etiquette: M-001 measures response from GitHub comment timestamps. An answer given only in chat makes the metric silently wrong AND loses the knowledge, which is the named risk of chat channels -- 지식 소실, 검색·보존 제한. The 48h alert backstops it: an item answered only in chat keeps reappearing until someone records it. Also enabled Discussions and added the triage/community labels. Not done, and listed as launch-hold conditions in docs/community/README.md: nobody has verified that security@ and conduct@ actually deliver to two people each. CI can check an address is stated; it cannot check mail arrives. The guideline says do not open public invitations until that is confirmed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01K599fUWZZYdiL4ekUq4Vhs
usik-luke-trl
added a commit
that referenced
this pull request
Oct 6, 2026
…wered alert (#14) * community: channel map, chat policy, moderator runbook, and the unanswered alert Prepares the chat channel without opening one. The guideline's first instruction is "첫 주에는 서버부터 만들지 않는다" -- purpose, channel map, roles and response expectations are settled before any server exists, and this is that work. Public-facing - COMMUNITY.md: where to ask what, and what to expect back. The chat row is left empty on purpose rather than carrying a placeholder link. - README gains a help/contribute/report section. The repository checklist is judged from a logged-out browser, and until now the README linked to neither CONTRIBUTING.md nor SECURITY.md -- the first checklist item, failing. Operations (docs/community/) - CHANNEL_MATRIX: which system each kind of enquiry COMPLETES in, the order to decide by, what a transition comment must carry, and alarm design. - CHANNEL_POLICY: minimum channel structure, roles with least privilege, enforcement. - MODERATOR_RUNBOOK: intake, S1-S4, incident records, the 8-step promotion procedure, and the four drills to run before opening. - LABELS: what each label means AND when it comes off, which the checklist requires. Automation - community-check.yml: required files present, and the reporting routes still stated. It fails today's README, which is how the gap above was found rather than assumed. - stale-unanswered.yml + tools/notify_stale.py: weekday mornings, items with no human response for 48h. Of everything the guideline says is worth sending to chat, this is the one that cannot be a GitHub notification -- it is the absence of an event. Dry run found a real case: PR #4 open 138 hours, unanswered. - community_to_issue.yml: the chat-to-issue form, so promotion is a form rather than something remembered. Discord or Slack is deliberately not decided here. Only the webhook differs, and notify_stale posts to whichever secret is set; with neither it prints and does not fail. Why promotion is written into the runbook rather than left as etiquette: M-001 measures response from GitHub comment timestamps. An answer given only in chat makes the metric silently wrong AND loses the knowledge, which is the named risk of chat channels -- 지식 소실, 검색·보존 제한. The 48h alert backstops it: an item answered only in chat keeps reappearing until someone records it. Also enabled Discussions and added the triage/community labels. Not done, and listed as launch-hold conditions in docs/community/README.md: nobody has verified that security@ and conduct@ actually deliver to two people each. CI can check an address is stated; it cannot check mail arrives. The guideline says do not open public invitations until that is confirmed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01K599fUWZZYdiL4ekUq4Vhs * docs(community): Discord launch plan Turns the channel decision into an executable plan. Discord is settled -- public invitation is the default, and the guideline warns against Slack's free tier as a permanent knowledge store. The structure that matters is the launch-hold section. Three conditions must hold before public invitations open, and the first is unverified: nobody has confirmed that security@ and conduct@ actually deliver, to two people each. A code of conduct whose report address is dead is not operational, and community-check.yml can only test that an address is stated. Phases: merge #14 and assign roles -> build the server -> wire notifications -> seed content -> limited release. Phase 3 is the largest, because FAQs and first-issue candidates have to be real ones rather than filler. Two platform constraints are written down rather than discovered later: GitHub repository webhooks subscribe by event type and cannot filter on labels, so "help wanted" needs an Actions workflow; and replying from Discord into GitHub is not practically possible, which is why promotion is a documented duty. Promotion is also what keeps M-001 honest -- it reads response time from GitHub comment timestamps, so an answer given only in chat makes the metric wrong and loses the knowledge. stale-unanswered.yml is the backstop: such an item keeps reappearing every weekday until someone records it. 12 acceptance criteria, each checkable. Notably #12 requires one real promotion to have happened, because a rule nobody has exercised is not yet a practice. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01K599fUWZZYdiL4ekUq4Vhs * docs(metrics): keep the definitions, drop the collection #13 proposed collecting metrics in this repository -- a collector, tests, a monthly workflow committing snapshots to main, and a baseline. That was wrong on three counts. The team said they would handle collection themselves; a scheduled workflow with contents: write pushing to the default branch is a real permission on a public repo for data no outside reader consumes; and the current snapshot is mostly no_data because there are no external contributors yet. What stays is one page. It exists because three community documents cite it, and because someone filing an issue should be able to see how their question is measured and that they are not the thing being measured. Definitions, the external-contributor scope and why it matters, why response time and no-response rate are read together, the observation period with no targets yet, why answering only in chat makes the numbers wrong, the explicit no-individual- measurement stance, and the limitations. No numbers, no collector, no workflow. Collection and reporting happen internally. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01K599fUWZZYdiL4ekUq4Vhs --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
usik-luke-trl
left a comment
Contributor
There was a problem hiding this comment.
Blocking: tests claimed in the description are not in the tree. Two line comments below.
- profile: --context N profiles at agent-length prefixes (14K-33K); adds timing-only bounds for a cheaper canvas attention mask (no mask, built-in causal) to size any mask-path optimization. - bench: --prompts takes a JSONL of OpenAI chat requests (e.g. logged Hermes requests), normalized like the server, images included. GQA head folding for canvas attention was tried and measured no gain on M3 Pro at 28K context (68.29 vs 67.83 ms), so it is not included. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013o6LMxkLD6mBHGy1C6MBR7
mlx-vlm is by Prince Canuma (Blaizzy), not Apple; matches NOTICE. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013o6LMxkLD6mBHGy1C6MBR7
Brings back inference/mlx/tests (reverts d344fdc): tiny random Qwen3.5-hybrid decoder invariants (test_tiny.py), tiny raw-HF checkpoint end to end (test_e2e.py, updated for load_model's vision return and _normalize_messages' image list), the chat-template fixture with its NOTICE / COMPLIANCE entries, the pytest dev group and config, and the README Tests section. Adds test_vision.py on a tiny multimodal checkpoint (tiny_vl.py): image expansion, greedy self-spec = greedy AR with an image, chunked prefill across an image, MRoPE position advance, prompt cache never reusing across different images, convert keeping the vision tower bit-identical, and the server image_url path. 27 tests, ~30 s on CPU (uv run pytest). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013o6LMxkLD6mBHGy1C6MBR7
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds
inference/mlx/: an on-device (Apple Silicon, MLX) backend for Trida2.0-4B with losslessself-speculative decoding, plus an OpenAI-compatible server usable from agent harnesses (tested with Hermes Agent).
bd_bidir_shiftb7/g4 semantics ofHybridDiffusionSelfSpec, on top of mlx-lm's Qwen3.5 model.Attention rows 0..N-1 causal, MASK rows see the whole canvas. Gated-delta clean rows read h[t], MASK rows read h[block_end].
Nothing persists until
commit(adv). Fused canvas GDN Metal kernel (adapted from mlx-lm, MIT). Greedy + top-k/top-p speculative sampling./v1/chat/completions+/v1/completions, SSE, reasoning split, tool calls → OpenAItool_calls,context_lengthon/v1/models, 3-slot prompt cache with a system-block snapshot, single MLX worker thread.trida-mlx-convert(q8/q6/q4),-verify,-bench,-profile,-agent; uv project.Results (M3 Pro 18 GB, q8, greedy, 512 tokens)
Testing
cd inference/mlx && uv run pytest: 20 offline tests on a tiny random model (CPU or GPU).trida-mlx-verify/trida-mlx-benchon M3 Pro. Hermes Agent end to end with the full default toolset.Notes
Qwen3_5ForCausalLM); vision tools need a separate model.