🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.
-
Updated
Sep 13, 2026 - TypeScript
🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
🤘 awesome-semantic-segmentation
Easily fine-tune, evaluate and deploy Qwen, Gemma, or any open weight LLM!
OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.
🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓
Next-generation AI Agent Optimization Platform: Cozeloop addresses challenges in AI agent development by providing full-lifecycle management capabilities from development, debugging, and evaluation to monitoring.
AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.
The platform for LLM evaluations and AI agent testing
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
Python package for the evaluation of odometry and SLAM
Building a modern functional compiler from first principles. (http://dev.stephendiehl.com/fun/)
MTEB: State-of-the-art evaluation of embeddings across languages and modalities
To associate your repository with the evaluation topic, visit your repo's landing page and select "manage topics."