RapidLLM is an LLM inference framework with continuous batching, tensor and data parallelism, quantized weights, and with pluggable Triton/CUDA kernels.
python3 attention continues llm llm-inference llama3 continuous-batching flash-attention-3 triton-kernels qwen3-moe qwen3-vl tensor-parallel
-
Updated
Sep 30, 2026 - Python