I quantize frontier models down into EXL3 and GGUF to fit on smaller GPUs, while also creating the most optimal recipes possible for the latest local LLMs!
Popular repositories Loading
-
Qwen3.8-Flash-Next-EXL3-DGX-Spark-recipe
Qwen3.8-Flash-Next-EXL3-DGX-Spark-recipe PublicServe Qwen3.8-Flash-Next (turboderp ExLlamaV3 pack) on one NVIDIA DGX Spark with exllamav3
-
GLM-5.3-Flash-EXL3-K2-DGX-Spark-recipe
GLM-5.3-Flash-EXL3-K2-DGX-Spark-recipe PublicvLLM recipe: GLM-5.3-Flash EXL3 K2 on one DGX Spark GB10. Native MTP k=2. Measured tok/s.
-
-
DeepSeek-V4.1-Flash-EXL3-DGX-Spark-recipe
DeepSeek-V4.1-Flash-EXL3-DGX-Spark-recipe Public -
hermes-agentic-bench
hermes-agentic-bench PublicAgentic test batteries for local models via Hermes Agent - scripted tools + real Hermes CLI sessions
Python 15
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.




