Skip to content
#

qwen38

Here are 16 public repositories matching this topic...

Qwen3.8-Flash-Next on 2× RTX 3090: with 128 GB RAM up to 4,191 tok/s prefill · 111.5 tok/s decode (131K prompt; prefill and agent 128K profiles), full 256K window at 2,865 tok/s; with 64 GB RAM 3,410 tok/s prefill · 84 tok/s decode (64 GB profile).

  • Updated Sep 30, 2026
  • Python

Run Qwen3.8-Flash-Next on ONE RTX 3090 (24 GB) + 64 GB RAM: 128K context, up to 2,100 tok/s prefill, 43–51 tok/s decode. vLLM runtime with hot MoE experts on the GPU and cold experts computed on the CPU, INT8 KV cache, Docker, OpenAI-compatible API.

  • Updated Sep 30, 2026
  • Python

Add this topic to your repo

To associate your repository with the qwen38 topic, visit your repo's landing page and select "manage topics."

Learn more