Benchmark workbench for MAX (Mojo) LLM decode kernels vs llama.cpp / cuBLAS / FlashInfer and the memory roofline on consumer NVIDIA GPUs (sm_86/sm_89). A public record, not a competing kernel library.
benchmark mojo cuda max nvidia quantization gpu-kernels roofline gemv llama-cpp llm-inference flash-attention gguf flashinfer q4-0
-
Updated
Sep 3, 2026 - Python