LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
-
Updated
Sep 30, 2026 - Python
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
macOS menu bar app that tells you, in plain English, what each USB-C cable plugged into your Mac can actually do
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
Run Stable Diffusion on Mac natively
macOS and Linux VMs on Apple Silicon to use in CI and other automations
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
Native MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API.
The free AI already on your Mac. CLI tool, OpenAI-compatible server, and interactive chat — all on-device via Apple Intelligence. No API keys, no cloud, no downloads.
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
Slap your MacBook, it yells back. Uses Apple Silicon accelerometer via IOKit HID.
Apple Silicon device emulator.
Rapid-MLX is an open-source (Apache 2.0) OpenAI- and Anthropic-compatible LLM inference server and Mac app for Apple Silicon, built on MLX, focused on reliable tool calling for coding agents. Release-gated with Claude Code, Codex CLI, Aider, Hermes and DeepSeek Harness.
🦾 A list of reported app support for Apple Silicon as well as Apple M4 and M3 Ultra Macs
Run Claude Code 100% on-device with local AI on Apple Silicon. MLX-native Anthropic-API server. 6 fighters incl. Muse-Glimmer 30B (now multimodal — reads images, abliterated), Gemma 4 31B, Qwen 3.5 122B (65 tok/s), DeepSeek V4 Flash (1M ctx). Private, offline, airgap-ready. Built for NDA / legal / healthcare workflows.
An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependency.
Widelands is a free, open source real-time strategy game with singleplayer campaigns and a multiplayer mode. The game was inspired by Settlers II™ (© Bluebyte) but has significantly more variety and depth to it.
The all-in-one bootable USB creator for Mac
The fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B model on 16 GB Macs. Native MTP speculative decoding on Apple Silicon, exact at any temperature. OpenAI and Anthropic compatible local server.
To associate your repository with the apple-silicon topic, visit your repo's landing page and select "manage topics."