Documented, reproducible finding: VulcanBench declarative grader mis-scored all functional tasks as 0.0 due to a repo-root pytest-cov addopts leak. Filed upstream issue #79.
benchmark coverage forensics pytest ai-evaluation llm-agents llm-benchmarking agent-eval vulcanbench
-
Updated
Aug 28, 2026