Skip to content
@local-inference-lab

Local Inference Lab

Popular repositories Loading

  1. rtx6kpro rtx6kpro Public

    RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink

    Python 1.1k 87

  2. b12x b12x Public

    Python 272 71

  3. llm-inference-bench llm-inference-bench Public

    LLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context lengths. Supports SGLang and vLLM engines.

    Python 113 17

  4. blackwell-llm-docker blackwell-llm-docker Public

    Docker images for LLM inference (SGLang + vLLM) on NVIDIA Blackwell GPUs (SM120, CUDA 13.2)

    Python 83 23

  5. vllm vllm Public

    Forked from vllm-project/vllm

    A high-throughput and memory-efficient inference and serving engine for LLMs

    Python 41 34

  6. quant-toolkit quant-toolkit Public

    Python 17 9

Repositories

Showing 10 of 19 repositories
  • lilmon Public

    Local Inference Lab Monitor

    local-inference-lab/lilmon's past year of commit activity
    Rust 0 Apache-2.0 0 0 0 Updated Oct 6, 2026
  • llm-inference-bench Public

    LLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context lengths. Supports SGLang and vLLM engines.

    local-inference-lab/llm-inference-bench's past year of commit activity
    Python 113 17 5 5 Updated Oct 6, 2026
  • blackwell-llm-docker Public

    Docker images for LLM inference (SGLang + vLLM) on NVIDIA Blackwell GPUs (SM120, CUDA 13.2)

    local-inference-lab/blackwell-llm-docker's past year of commit activity
    Python 83 23 3 8 Updated Oct 6, 2026
  • vllm Public Forked from vllm-project/vllm

    A high-throughput and memory-efficient inference and serving engine for LLMs

    local-inference-lab/vllm's past year of commit activity
    Python 41 Apache-2.0 23,285 87 236 Updated Oct 6, 2026
  • LMCache Public Forked from LMCache/LMCache

    LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

    local-inference-lab/LMCache's past year of commit activity
    Python 0 Apache-2.0 2,013 0 31 Updated Oct 6, 2026
  • flashinfer Public Forked from flashinfer-ai/flashinfer

    FlashInfer: Kernel Library for LLM Serving

    local-inference-lab/flashinfer's past year of commit activity
    Cuda 0 Apache-2.0 1,541 0 2 Updated Oct 5, 2026
  • nccl-canonical Public Forked from NVIDIA/nccl

    Optimized primitives for collective multi-GPU communication

    local-inference-lab/nccl-canonical's past year of commit activity
    C++ 0 1,451 0 2 Updated Oct 5, 2026
  • InstantTensor Public Forked from voipmonitor/InstantTensor

    An ultra-fast, distributed Safetensors loader

    local-inference-lab/InstantTensor's past year of commit activity
    C++ 0 Apache-2.0 19 0 1 Updated Oct 5, 2026
  • rtx6kpro Public

    RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink

    local-inference-lab/rtx6kpro's past year of commit activity
    Python 1,121 87 60 19 Updated Oct 5, 2026
  • b12x Public
    local-inference-lab/b12x's past year of commit activity
    Python 272 Apache-2.0 71 60 84 Updated Oct 4, 2026

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Most used topics

Loading…