Field guide to building production-grade evaluation infrastructure for LLM agent systems. Based on real production work at Airbnb (BPI Virtual Analyst), Shell (NLP), and contributions to LangChain, LiveKit, and Ragas.
python airbnb rag langchain llm-eval llm-evaluation agent-infrastructure agent-eval eval-driven-deployment llm-production
-
Updated
Sep 7, 2026 - Jupyter Notebook