Sparse-upcycle a dense Gemma-3-270m into a Mixture-of-Experts (gemma3moe) and serve it on llama.cpp (CPU). Upcycle, train, inspect routing, GGUF.
transformers moe gemma upcycling mixture-of-experts cpu-inference llama-cpp small-language-model gemma3 sparse-upcycling
-
Updated
Jul 27, 2026 - Python