This project aims to provide a high effective KV cache manage framework for llm inference and improve memory utilization and inference speed.
-
Updated
Apr 24, 2026 - Python
This project aims to provide a high effective KV cache manage framework for llm inference and improve memory utilization and inference speed.
NPU powered On-device AI Mobile applications using Melange
The Buildfunctions SDK for Agents examples: Hardware-isolated CPU and GPU Sandboxes for untrusted AI actions 🪐 Now in beta, with $100 in free credits!
The Buildfunctions SDK for interacting with serverless functions, hardware-isolated sandboxes, and runtime controls for AI agents 🪐 Now in beta, with $100 in free credits!
160+ pages of handwritten notes from Stanford CS336: Language Modeling from Scratch, covering language modeling, LLM systems, and full-stack AI engineering from data to deployment and evaluation and RL alignment
The Buildfunctions SDK for interacting with serverless functions, hardware-isolated sandboxes, and runtime controls for AI agents 🪐 Now in beta, with $100 in free credits!
A .NET MCP control-gate project that discovers live tools, validates contracts, and blocks risky tool changes before promotion.
Secure Agent Governed Execution OS
This repositry contains the practice files necessary for the MLOps and AI infrastructure project
GPU performance optimization labs covering Roofline Analysis, LLM Decode Optimization, and CUDA Graphs using PyTorch
EvoNano = Evolve + Nano-VLLM Evolving hands-on capabilities to build and polish lightweight AI inference engines on top of the concise Nano-VLLM framework via continuous iteration.
UniGPU is a peer-to-peer GPU marketplace that enables students to rent out idle GPUs while allowing clients to execute secure, containerized GPU workloads remotely.
Add a description, image, and links to the aiinfrastructure topic page so that developers can more easily learn about it.
To associate your repository with the aiinfrastructure topic, visit your repo's landing page and select "manage topics."