Skip to content

feat(bench): BugBench — detection-quality benchmark + CI gate (/bench) - #155

Open
shivsin25 wants to merge 3 commits into
awarexone:mainfrom
shivsin25:feat/bugbench
Open

shivsin25 wants to merge 3 commits into
awarexone:mainfrom
shivsin25:feat/bugbench

Conversation

@shivsin25

Copy link
Copy Markdown
Contributor

The claim this unlocks

"Agentic Bug Hunter measures its own detection quality — 100% precision, 0% false-positive rate on the benchmark corpus — and CI blocks any change that regresses it."

Almost no AI security tool can say that. Reputation in this space is trust, and trust is measured false positives + reproducibility. BugBench is the infrastructure that lets the project make — and keep — that claim.

Why it's needed

The repo generates findings (scanners, agent), captures them (/poc), re-checks them (/replay), and prioritizes them (/oracle). What it has never had is a way to answer: "how often are the detectors actually right, and how often do they cry wolf?" Without that number, every accuracy claim is a vibe. BugBench makes it a measurement.

What's in this PR

  • A ground-truth corpus — bughunter/bench/cases/*.json, with known-vulnerable AND known-safe cases. The safe cases are the point: they're what measure false positives. Each case carries a recorded input (the exact response a detector sees), never a live URL, so the score is byte-identical on every machine and in CI.
  • A harness — bughunter/tools/bench.py: load → run detector → score. Detectors register via a one-line @detector adapter (open/closed — adding SQLi/JWT/CRLF is a one-function change, the core never grows an if/elif tree). The cors detector is wired first against its real, pure classifier, offline.
  • A scoreboard (Markdown + JSON) that leads with precision + false-positive rate. Accuracy is deliberately not the headline — a test (test_accuracy_is_misleading_under_imbalance) proves a do-nothing detector scores 95% accuracy while catching zero bugs. That's why trust metrics lead.
  • A CI gate — a new bugbench job runs bench.py run --min-precision 1.0 --max-fp-rate 0.0, which exits non-zero on regression. A change can never quietly make a detector noisier. "We measure quality" becomes enforced, not aspirational.

The scoreboard today

# BugBench — detection quality
**100% precision · 0% false-positive rate · 100% recall · F1 1.00**  (n=2)

| detector | cases | precision | recall | FP-rate | F1 |
|---|---|---|---|---|---|
| cors     |   2   |   100%    |  100%  |   0%    | 1.00 |

(Starts with the CORS detector; the corpus is designed to grow — every new case and detector strengthens the guarantee.)

Quality

  • 28 tests (tests/test_bench.py), each layer isolated: data-contract validation, metric math (incl. the "accuracy lies" lesson), the runner against the real cors classifier, the scoreboard, and a gate test proving a wolf-crying detector fails CI.
  • Pure-stdlib, deterministic, additive: 7 files changed, 803 insertions(+), nothing removed. A malformed corpus fails loud (never silently skews a score).

How to extend (built for contribution)

  • Add a case: drop a JSON file in bughunter/bench/cases/ — no code.
  • Add a detector: one @detector("name") adapter in bench.py.

Self-contained and valuable on its own; also the measuring stick for a follow-up prove-or-suppress verification layer.

Try it

tools/bench.py run --min-precision 1.0 --max-fp-rate 0.0

🤖 Generated with Claude Code

Reputation in security tooling is trust, and trust is *measured* false positives
and reproducibility. Nothing in the repo currently measures whether the toolkit's
detectors are accurate — so it can't make the one claim that earns credibility:
"here is our false-positive rate, reproduce it yourself." BugBench adds that.

What it is:
- A ground-truth corpus (bughunter/bench/cases/*.json) of known-vulnerable AND
  known-safe cases. Each case carries a RECORDED input (the exact response a
  detector sees), never a live URL — so scores are identical on every machine
  and in CI. Cases are data: anyone can contribute one without touching Python.
- A harness (bughunter/tools/bench.py) that runs each case's detector and scores
  precision / recall / false-positive rate. Detectors register a one-function
  adapter via @detector — adding one never touches the harness core (open/closed).
  The cors detector is wired first, against its real, pure classifier (offline).
- A Markdown/JSON scoreboard leading with precision + FP-rate. Accuracy is
  deliberately NOT the headline: on a mostly-safe corpus a do-nothing detector
  scores high accuracy while catching nothing (a test asserts exactly this).

Why the CI gate matters:
- `bench.py run --min-precision 1.0 --max-fp-rate 0.0` exits non-zero on a
  regression, and a new `bugbench` CI job runs it. So a change can never quietly
  make a detector noisier or less accurate — "we measure quality" becomes an
  enforced guarantee, not a one-time claim.

Design: the whole thing is pure-stdlib and deterministic. Data model, metrics,
runner, scoreboard, and the gate are each independently unit-tested (28 tests in
tests/test_bench.py), including the end-to-end run of the real cors classifier
over the shipped corpus (100% precision, 0 false positives) and a gate test that
proves a wolf-crying detector fails CI.

This is self-contained and valuable on its own; it's also the measuring stick for
a follow-up verification layer (prove-or-suppress) that will route findings
through confirmation before they're ever reported.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

@shuvonsec shuvonsec left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good design, and the workflow is clean (pull_request not pull_request_target, permissions contents:read, no secrets, static run command). Verified 28 tests pass, and I ran the exact CI gate command against current main - it exits 0 and the gate passes, so no CI breakage. Two asks:

  1. expected.min_severity / expected.class are silently ignored - scoring only checks the boolean vulnerable (severity != INFO), so a detector that downgrades CRITICAL->LOW still passes the gate. Either enforce min_severity in scoring, or drop it from the schema + docs so it doesn't imply a guarantee it doesn't give.
  2. The gate runs --min-precision 1.0 --max-fp-rate 0.0 on a 2-case corpus, on every PR. The first detector regression or a newly-added hard case will red-gate every unrelated open PR. Consider continue-on-error: true (or not a required check) until the corpus matures, plus a doc note that adding a currently-failing case reds the whole repo.

Nice touch: test_accuracy_is_misleading_under_imbalance. Later, worth adding null-origin and wildcard+creds as corpus cases (they're unit-tested but not in the shipped corpus).

Shivendra-Coherent and others added 2 commits September 28, 2026 13:59
Addresses @shuvonsec's review on /bench.

1. min_severity was silently ignored — scoring only checked the boolean
   `vulnerable` (severity != INFO), so a detector that downgrades CRITICAL->LOW
   still passed. Scoring now enforces each case's min_severity: a below-severity
   detection on a vulnerable case is counted as a MISS (failure kind
   "severity_downgrade"), not a catch. `class` is documented as descriptive
   metadata (detector routing already fixes it).

2. The CI gate ran strict thresholds on a tiny corpus on every PR, which could
   red-gate unrelated PRs on the first regression or a newly-added hard case. The
   `bugbench` job is now continue-on-error (non-blocking) while the corpus
   matures — it reports quality without blocking — with a comment + doc note that
   adding a currently-failing case shows red but does not block merges, and how to
   flip it to a required check later.

3. Grew the corpus with the two cases the reviewer suggested: null-origin +
   credentials (HIGH) and wildcard ACAO + credentials (MEDIUM), each with a
   min_severity so they also exercise the new enforcement. Corpus is now 4 cases,
   still 100% precision / 0 FP.

Tests: +4 in tests/test_bench.py — severity ranking, a downgrade counted as a
miss (regression guard), at/above min_severity counted as a catch, and the
expanded shipped corpus scoring clean. 32 pass. (This commit also merges main and
resolves the CLAUDE.md conflict, keeping both /dashboard and /bench entries.)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@shivsin25

Copy link
Copy Markdown
Contributor Author

good catches, both fixed.

on min_severity: yeah that was a real gap, it was only checking the boolean so a downgrade would sail through. scoring now enforces it — if a case says min_severity high and the detector only flags it low/medium it's counted as a miss (i tagged that failure kind severity_downgrade so it's obvious in the report). left class as descriptive metadata since the detector routing already pins it, so it's not pretending to guarantee anything now.

on the gate: fair point, strict thresholds on a 2-case corpus would've nuked unrelated PRs the moment anything regressed. made the bugbench job continue-on-error for now so it still reports quality on each PR but doesn't block, and dropped a comment + doc note saying a failing case shows red without blocking, and how to make it a required check later once the corpus is bigger.

also went ahead and added the null-origin and wildcard+creds cases you mentioned (both were unit-tested but not in the shipped corpus). corpus is 4 now, still 100% precision / 0 FP, and both carry a min_severity so they exercise the new check too.

rebased on main and fixed the CLAUDE.md conflict while i was at it. 32 tests pass.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants