An open-source agentic experience evaluation platform.
Evaluate how AI agents use your product, approach real tasks, and respond to feedback. Define success with reusable criteria, follow each run live, and inspect the evidence behind every result through the Portal or CLI.
Users and contributors, visit the official documentation website for setup instructions, evaluation workflows, architecture, and reference guides.
Start with Getting started or Local development. To help improve Scope, read the contributor guide. For help or vulnerability reporting, see Support and security.