e-accelerate
Popular repositories Loading
-
pd-bridge
pd-bridge PublicForked from chadhurley25075-png/pd-bridge
Heterogeneous prefill/decode for DeepSeek-V4-Flash: CUDA prefill (DGX Spark, vLLM) -> Metal decode (Mac Studio, oMLX) over plain 10GbE
Python 1
-
Qwen3.8-27B-DFlash2-Triton-Blackwell
Qwen3.8-27B-DFlash2-Triton-Blackwell PublicQwen3.8-27B on SGLang for Blackwell SM120: DFlash2 block-16 + triton. 268 tok/s on one RTX PRO 5000 (48GB), full 262K context. Measured, with honest sample sizes.
Shell
-
GLM-5.3-Flash-EXL3-2x-DGX-Sparks
GLM-5.3-Flash-EXL3-2x-DGX-Sparks PublicForked from MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks
GLM-5.3 Flash EXL3 for 2-4x DGX Sparks
Python
-
glm53-flash-2x-recipe
glm53-flash-2x-recipe PublicGLM-5.3-Flash EXL3 on 2x GB10 (ASUS GX10 + DGX Spark), dual-rail: tuned, measured recipe on MiaAI-Lab's kit
Python
Repositories
- GLM-5.3-Flash-EXL3-2x-DGX-Sparks Public Forked from MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks
GLM-5.3 Flash EXL3 for 2-4x DGX Sparks
- glm53-flash-2x-recipe Public
GLM-5.3-Flash EXL3 on 2x GB10 (ASUS GX10 + DGX Spark), dual-rail: tuned, measured recipe on MiaAI-Lab's kit
- pd-bridge Public Forked from chadhurley25075-png/pd-bridge
Heterogeneous prefill/decode for DeepSeek-V4-Flash: CUDA prefill (DGX Spark, vLLM) -> Metal decode (Mac Studio, oMLX) over plain 10GbE
- Qwen3.8-27B-DFlash2-Triton-Blackwell Public
Qwen3.8-27B on SGLang for Blackwell SM120: DFlash2 block-16 + triton. 268 tok/s on one RTX PRO 5000 (48GB), full 262K context. Measured, with honest sample sizes.
People
This organization has no public members. You must be a member to see who’s a part of this organization.
Top languages
Loading…
Most used topics
Loading…