Skip to content
@e-accelerate

e-accelerate

Popular repositories Loading

  1. pd-bridge pd-bridge Public

    Forked from chadhurley25075-png/pd-bridge

    Heterogeneous prefill/decode for DeepSeek-V4-Flash: CUDA prefill (DGX Spark, vLLM) -> Metal decode (Mac Studio, oMLX) over plain 10GbE

    Python 1

  2. Qwen3.8-27B-DFlash2-Triton-Blackwell Qwen3.8-27B-DFlash2-Triton-Blackwell Public

    Qwen3.8-27B on SGLang for Blackwell SM120: DFlash2 block-16 + triton. 268 tok/s on one RTX PRO 5000 (48GB), full 262K context. Measured, with honest sample sizes.

    Shell

  3. GLM-5.3-Flash-EXL3-2x-DGX-Sparks GLM-5.3-Flash-EXL3-2x-DGX-Sparks Public

    Forked from MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks

    GLM-5.3 Flash EXL3 for 2-4x DGX Sparks

    Python

  4. glm53-flash-2x-recipe glm53-flash-2x-recipe Public

    GLM-5.3-Flash EXL3 on 2x GB10 (ASUS GX10 + DGX Spark), dual-rail: tuned, measured recipe on MiaAI-Lab's kit

    Python

Repositories

Showing 4 of 4 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…