Skip to content

feat: add FP8 blockwise dequantize/GEMM operators and FP8(E4M3) KV cache support in paged attention - #1568

Open
shsaihdsaiudh wants to merge 11 commits into
InfiniTensor:mainfrom
shsaihdsaiudh:feat/fp8-blockwise-dequantize
Open

shsaihdsaiudh wants to merge 11 commits into
InfiniTensor:mainfrom
shsaihdsaiudh:feat/fp8-blockwise-dequantize

fix: route F8 caches to the infiniop fallback in paged_caching dispatch

264ebc1
Select commit
Loading
Failed to load commit list.

Workflow runs completed with no jobs