Pinned Loading
Repositories
Showing 10 of 44 repositories
- speculators Public
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
- semantic-router Public
Intelligent Mixture-of-Models Router for Efficient Heterogeneous LLMs Inference
- compressed-tensors Public
A safetensors extension to efficiently store sparse quantized tensors on disk
- llm-compressor Public
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Top languages
Loading…
Most used topics
Loading…