Skip to content
@foundry-org

foundry-org

Pinned Loading

  1. foundry foundry Public

    Foundry materializes CUDA graphs along with its execution context to disk to support fast cold start of serving engines.

    C++ 62 11

  2. sglang sglang Public

    Forked from sgl-project/sglang

    SGLang is a high-performance serving framework for large language models and multimodal models.

    Python

  3. vllm vllm Public

    Forked from vllm-project/vllm

    Adapted fork of vLLM+Foundry integration. vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs

    Python 1

  4. TensorRT-LLM TensorRT-LLM Public

    Forked from NVIDIA/TensorRT-LLM

    TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. Tensor…

    Python

Repositories

Showing 4 of 4 repositories
  • sglang Public Forked from sgl-project/sglang

    SGLang is a high-performance serving framework for large language models and multimodal models.

    foundry-org/sglang's past year of commit activity
    Python 0 Apache-2.0 9,087 0 4 Updated Sep 20, 2026
  • foundry Public

    Foundry materializes CUDA graphs along with its execution context to disk to support fast cold start of serving engines.

    foundry-org/foundry's past year of commit activity
    C++ 62 Apache-2.0 11 2 2 Updated Sep 18, 2026
  • vllm Public Forked from vllm-project/vllm

    Adapted fork of vLLM+Foundry integration. vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs

    foundry-org/vllm's past year of commit activity
    Python 1 Apache-2.0 22,720 0 0 Updated Sep 1, 2026
  • TensorRT-LLM Public Forked from NVIDIA/TensorRT-LLM

    TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.

    foundry-org/TensorRT-LLM's past year of commit activity
    Python 0 2,797 0 0 Updated May 26, 2026

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…