LLM-as-judge evaluator for scoring AI agent trajectories on task completion, tool correctness, and efficiency
-
Updated
Jul 24, 2026 - Python
LLM-as-judge evaluator for scoring AI agent trajectories on task completion, tool correctness, and efficiency
Multi-agent routing system built on Amazon Bedrock AgentCore and LangChain, utilizing deterministic prompt engineering to handle bug reports and FAQs.
Independent RAG evaluation framework: a fictional knowledge base stress-tested for hallucination, scored by a blinded, cross-family LLM judge validated on a held-out set.
Microservices infrastructure for web parsing evaluation using Llama 3.2 as an LLM-as-a-Judge and traditional and advanced metrics. Computer Engineering Thesis, Sapienza University of Rome.
"Benchmarking LLMs on structured argument extraction from dense, low-resource halakhic texts using an LLM-as-a-Judge pipeline."
A modular framework for evaluating Retrieval-Augmented Generation (RAG) systems with novel pedagogical metrics (PES) for E-Learning environments. Master's Thesis @ Politecnico di Torino.
Lightweight Rust engine for model-vs-model and adversarial LLM evaluation.
A tool that scores an LLM's response to a prompt against a five-part rubric, using Claude as the judge
To associate your repository with the llmasajudge topic, visit your repo's landing page and select "manage topics."