Lynn GitHub 镜像仓 · Primary repository: https://github.com/MerkyorLynn/Lynn · Downloads: https://download.merkyorlynn.com/download.html
-
Updated
Sep 11, 2026 - TypeScript
Lynn GitHub 镜像仓 · Primary repository: https://github.com/MerkyorLynn/Lynn · Downloads: https://download.merkyorlynn.com/download.html
Meta-harness optimization loop wired onto Islo sandboxes. POC: 0/5→5/5 in four proposer steps. Built on islo.dev.
Observe an agent run and GroundEval drafts the policy and diagram for you, no hand-written policy required, then scores what it checked, what it skipped, and what it wasn't allowed to touch.
Live, open-source benchmark for comparing AI coding agents on real GitHub issues
Scenario Testing for AI Agents
AI-operated company. Building agent-friend: universal tool adapter for AI agents. @tool → OpenAI, Claude, Gemini, MCP. Live 24/7 on Twitch.
Framework-agnostic evaluation harness for Go — test your MCP servers and AI agents with scored, CI-ready checks.
Project page for Meta-harness on Islo (POC). https://zozo123.github.io/meta-harness-on-islo-page/
A reasoning benchmark runner for comparing LLMs as OpenClaw agents use them. 52 prompts, 3 eval sets, 11 traps, LLM-as-judge, tier-based leaderboard.
Transcript-first evaluation tool for comparing coding-agent sessions across Codex, Claude Code, and Pi.
A durable, long-running agent that improves and evaluates other agents.
Vendor-neutral research umbrella for measuring AI plugin, agent, and MCP server quality across CLI runtimes (Claude Code, Gemini CLI, Copilot CLI, Codex CLI).
Binary-criteria evaluation harness for Claude skills with planned extension to plugins, agents, and MCP servers. Score every change yes/no across 7 layers — package integrity, trigger quality, functional quality, regression protection, baseline value, model variance, rollout safety. Never gradients.
PandaProbe harness turns agent failures into fixes
开源通用 AI Agent 真实任务评测 · 同 Prompt、客观开奖、评分细则全公开 | Open-source evaluation of general-purpose AI Agents on real-world tasks with verifiable outcomes — by PingWest / 硅星人
Documented, reproducible finding: VulcanBench declarative grader mis-scored all functional tasks as 0.0 due to a repo-root pytest-cov addopts leak. Filed upstream issue #79.
Field guide to building production-grade evaluation infrastructure for LLM agent systems. Based on real production work at Airbnb (BPI Virtual Analyst), Shell (NLP), and contributions to LangChain, LiveKit, and Ragas.
An AI coding benchmark task evaluating automated rebase conflict resolution, MLflow tracking security, secret redaction, and Hugging Face model evaluation safety.
Deterministic agent-eval / benchmark-grading hygiene audit CLI. Detects the silent mis-scoring bug class (pytest cov-gate leakage). Free lead-magnet for the paid harness-audit service.
Auto-generate evaluation rubrics from agent audit-log trajectories (PhoneWorld pattern applied to action logs)
To associate your repository with the agent-eval topic, visit your repo's landing page and select "manage topics."