Popular repositories Loading
-
qwen38-flash-next-2x3090
qwen38-flash-next-2x3090 PublicQwen3.8-Flash-Next on 2× RTX 3090: with 128 GB RAM up to 4,191 tok/s prefill · 111.5 tok/s decode (131K prompt; prefill and agent 128K profiles), full 256K window at 2,865 tok/s; with 64 GB RAM 3,4…
-
qwen38-flash-next-3090
qwen38-flash-next-3090 PublicRun Qwen3.8-Flash-Next on ONE RTX 3090 (24 GB) + 64 GB RAM: 128K context, up to 2,100 tok/s prefill, 43–51 tok/s decode. vLLM runtime with hot MoE experts on the GPU and cold experts computed on th…
Python 1
-
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.

