Real-time camera image-to-image transformation using diffusion models on Apple Silicon, accelerated with CoreML.
22.7 FPS at 512x512 resolution on Apple M3 Ultra with SDXS-512.
- macOS 14+ (Sonoma or later)
- Apple Silicon (M1 / M2 / M3 / M4 series)
- Python 3.9-3.12 (coremltools does not support 3.13+)
- Camera (built-in or USB webcam)
# 1. Clone
git clone https://github.com/ochyai/streamdiffusion-mac.git
cd streamdiffusion-mac
# 2. Setup environment
chmod +x setup.sh
./setup.sh
# 3. Activate
source .venv/bin/activate
# 4. Convert models to CoreML (one-time, ~5 minutes)
python scripts/convert_models.py
# 5. Run camera
python camera.py --prompt "oil painting style, masterpiece"setup.sh installs all dependencies from requirements.txt. Key version constraints:
| Package | Default | Legacy (--legacy) |
|---|---|---|
| PyTorch | 2.1-2.5.x | 2.0.x |
| CoreML Tools | 7.x-8.x | 6.x |
| diffusers | 0.21+ | 0.21.x |
| numpy | 1.x (< 2.0) | 1.x (< 2.0) |
- Default (
requirements.txt): Tested with torch 2.5.1 + coremltools 8.3 on M3 Ultra / M4 Max. - Legacy (
requirements-legacy.txt): Conservative versions for environments where the default fails.
If ./setup.sh fails or you experience issues, try the legacy profile:
rm -rf .venv
./setup.sh --legacyIf you prefer manual setup:
# Create virtual environment
python3 -m venv .venv
source .venv/bin/activate
# Install dependencies
pip install -r requirements.txt
# Convert models
python scripts/convert_models.py
# Run
python camera.py# Default (SDXS-512, best speed/quality balance)
python camera.py --prompt "oil painting style, masterpiece"
# Watercolor style
python camera.py --prompt "watercolor painting, soft brushstrokes"
# With built-in prompt gallery (10 styles, press n/p to switch)
python camera.py --prompts| Key | Action |
|---|---|
q |
Quit |
s |
Save current frame |
n / p |
Next / Previous prompt |
+ / - |
Adjust camera blend ratio |
e / d |
Adjust EMA smoothing |
# Use SD-Turbo instead of SDXS (slower but different style)
python camera.py --model sd-turbo --prompt "anime style"
# Blend camera with AI output (30% camera)
python camera.py --blend 0.3
# Adjust temporal smoothing
python camera.py --ema 0.9 --feedback 0.4
# Lower resolution for faster inference on smaller Macs
python camera.py --render-size 384
# Select camera device
python camera.py --camera 1# Convert SDXS-512 (default, recommended)
python scripts/convert_models.py
# Convert SD-Turbo
python scripts/convert_models.py --model sd-turbo
# Custom output directory
python scripts/convert_models.py --output-dir ./my_models
python camera.py --coreml-dir ./my_modelsCamera Thread ──→ latest_frame ──→ Inference Thread ──→ latest_ai_result
(30 FPS) │ (CoreML pipeline) │
│ │
└──────→ Display Thread ←──────────────┘
(blends camera + AI, 30+ FPS)
The pipeline decouples camera capture, AI inference, and display into three independent threads. The inference thread runs the full CoreML pipeline on every frame:
- Preprocess: Center crop, resize to 512x512, normalize
- VAE Encode: Image → Latent space (CoreML TAESD, ~5ms)
- Noise Addition: Fixed noise for temporal coherence
- UNet Inference: Single-step denoising (CoreML, ~24ms for SDXS)
- VAE Decode: Latent → Image (CoreML TAESD, ~5ms)
- Postprocess: Denormalize, resize, display
Temporal coherence is maintained through:
- Fixed noise seed: Same noise pattern every frame eliminates flickering
- Latent feedback: 30% of previous frame's denoised latent blended into current input
- EMA smoothing: Exponential moving average on display output
Benchmarked on Apple M3 Ultra (60-core GPU, 512GB unified memory):
| Model | Parameters | UNet Latency | Camera FPS | Quality |
|---|---|---|---|---|
| SDXS-512 | 328M | 24.4ms | 22.7 | Good |
| SD-Turbo | 866M | 53.2ms | 13.8 | Good |
| Tiny-SD | 323M | 31.3ms | ~20 | Fair |
Performance scales with GPU core count. Expected approximate FPS:
- M1/M2: ~5-8 FPS
- M1/M2 Pro: ~8-12 FPS
- M1/M2/M3 Max: ~12-18 FPS
- M4 Max (40-core GPU, 128GB): ~15.4 FPS (SDXS-512, measured)
- M3 Ultra: ~22 FPS
This project is the result of a systematic 10-phase optimization study on real-time diffusion model inference on Apple Silicon. Below is a summary of key findings.
| Technique | Effect | Notes |
|---|---|---|
| CoreML conversion | +64% | Only effective UNet acceleration method |
| Distilled models (SDXS) | +118% | Best speed/quality trade-off |
| 3-thread pipeline | Smooth display | Decouples inference from rendering |
| Technique | Effect | Why |
|---|---|---|
| Quantization (INT8 to 2-bit) | 0% | M3 Ultra is compute-bound, not memory-bandwidth-bound |
| Token Merging (ToMe) | -10% | MPS overhead exceeds attention savings |
| Parallel CoreML inference | 0% | Metal serializes GPU commands |
| Neural Engine for UNet | -19% to -520% | ANE unsuitable for large (866M) models |
| torch.compile | Crash | MPS backend not supported |
| Attention Slicing | -40% | MPS memory management overhead |
The most important finding is that optimization techniques established for NVIDIA GPUs and the CUDA ecosystem largely do not transfer to Apple Silicon's unified memory architecture:
- Quantization is ineffective because Apple Silicon is compute-bound (not memory-bandwidth-bound). The 800 GB/s unified memory bandwidth is sufficient for model weights, so reducing precision doesn't help.
- Parallel inference is impossible because CoreML serializes Metal GPU commands, unlike CUDA Streams which allow fine-grained kernel-level parallelism.
- The software ecosystem is immature compared to CUDA's decades of optimization (cuDNN, TensorRT, xformers, Flash Attention). torch.compile doesn't work on MPS, and many PyTorch operations have suboptimal Metal implementations.
Several creative approaches were also tested and yielded negative results:
- kNN search-based synthesis (Phase 7): 512GB memory enables searching 100M vectors in 0.5ms, but kNN retrieval fundamentally cannot replace the continuous nonlinear function approximation of a UNet.
- pix2pix-turbo (Phase 8): Skip-connection VAE design prevents CoreML conversion, creating a 160ms VAE bottleneck (vs 53ms UNet). Result: 4 FPS.
- Optical flow frame skipping (Phase 9): Warping between UNet frames produces jelly-like distortion artifacts. 17.4 FPS, worse than SDXS baseline.
- Knowledge distillation (Phase 10): 875K-parameter feedforward CNN trained with L1 loss produces blank output. The capacity gap vs 328M-parameter diffusion model is too large.
A detailed academic paper covering all experiments is available:
- paper.tex (Japanese)
- paper_en.tex (English)
MIT License
@article{ochiai2025streamdiffusion_mac,
title={Systematic Optimization of Real-Time Diffusion Model Inference on Apple M3 Ultra},
author={Ochiai, Yoichi},
journal={arXiv preprint},
year={2025}
}- StreamDiffusion — Original pipeline architecture
- SDXS — Distillation-specialized model
- SD-Turbo — One-step diffusion baseline
- TAESD — Tiny Autoencoder for Stable Diffusion