Conversation
Reputation in security tooling is trust, and trust is *measured* false positives and reproducibility. Nothing in the repo currently measures whether the toolkit's detectors are accurate — so it can't make the one claim that earns credibility: "here is our false-positive rate, reproduce it yourself." BugBench adds that. What it is: - A ground-truth corpus (bughunter/bench/cases/*.json) of known-vulnerable AND known-safe cases. Each case carries a RECORDED input (the exact response a detector sees), never a live URL — so scores are identical on every machine and in CI. Cases are data: anyone can contribute one without touching Python. - A harness (bughunter/tools/bench.py) that runs each case's detector and scores precision / recall / false-positive rate. Detectors register a one-function adapter via @detector — adding one never touches the harness core (open/closed). The cors detector is wired first, against its real, pure classifier (offline). - A Markdown/JSON scoreboard leading with precision + FP-rate. Accuracy is deliberately NOT the headline: on a mostly-safe corpus a do-nothing detector scores high accuracy while catching nothing (a test asserts exactly this). Why the CI gate matters: - `bench.py run --min-precision 1.0 --max-fp-rate 0.0` exits non-zero on a regression, and a new `bugbench` CI job runs it. So a change can never quietly make a detector noisier or less accurate — "we measure quality" becomes an enforced guarantee, not a one-time claim. Design: the whole thing is pure-stdlib and deterministic. Data model, metrics, runner, scoreboard, and the gate are each independently unit-tested (28 tests in tests/test_bench.py), including the end-to-end run of the real cors classifier over the shipped corpus (100% precision, 0 false positives) and a gate test that proves a wolf-crying detector fails CI. This is self-contained and valuable on its own; it's also the measuring stick for a follow-up verification layer (prove-or-suppress) that will route findings through confirmation before they're ever reported. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
shuvonsec
left a comment
There was a problem hiding this comment.
Good design, and the workflow is clean (pull_request not pull_request_target, permissions contents:read, no secrets, static run command). Verified 28 tests pass, and I ran the exact CI gate command against current main - it exits 0 and the gate passes, so no CI breakage. Two asks:
expected.min_severity/expected.classare silently ignored - scoring only checks the booleanvulnerable(severity != INFO), so a detector that downgrades CRITICAL->LOW still passes the gate. Either enforcemin_severityin scoring, or drop it from the schema + docs so it doesn't imply a guarantee it doesn't give.- The gate runs
--min-precision 1.0 --max-fp-rate 0.0on a 2-case corpus, on every PR. The first detector regression or a newly-added hard case will red-gate every unrelated open PR. Considercontinue-on-error: true(or not a required check) until the corpus matures, plus a doc note that adding a currently-failing case reds the whole repo.
Nice touch: test_accuracy_is_misleading_under_imbalance. Later, worth adding null-origin and wildcard+creds as corpus cases (they're unit-tested but not in the shipped corpus).
# Conflicts: # CLAUDE.md
Addresses @shuvonsec's review on /bench. 1. min_severity was silently ignored — scoring only checked the boolean `vulnerable` (severity != INFO), so a detector that downgrades CRITICAL->LOW still passed. Scoring now enforces each case's min_severity: a below-severity detection on a vulnerable case is counted as a MISS (failure kind "severity_downgrade"), not a catch. `class` is documented as descriptive metadata (detector routing already fixes it). 2. The CI gate ran strict thresholds on a tiny corpus on every PR, which could red-gate unrelated PRs on the first regression or a newly-added hard case. The `bugbench` job is now continue-on-error (non-blocking) while the corpus matures — it reports quality without blocking — with a comment + doc note that adding a currently-failing case shows red but does not block merges, and how to flip it to a required check later. 3. Grew the corpus with the two cases the reviewer suggested: null-origin + credentials (HIGH) and wildcard ACAO + credentials (MEDIUM), each with a min_severity so they also exercise the new enforcement. Corpus is now 4 cases, still 100% precision / 0 FP. Tests: +4 in tests/test_bench.py — severity ranking, a downgrade counted as a miss (regression guard), at/above min_severity counted as a catch, and the expanded shipped corpus scoring clean. 32 pass. (This commit also merges main and resolves the CLAUDE.md conflict, keeping both /dashboard and /bench entries.) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
good catches, both fixed. on min_severity: yeah that was a real gap, it was only checking the boolean so a downgrade would sail through. scoring now enforces it — if a case says min_severity high and the detector only flags it low/medium it's counted as a miss (i tagged that failure kind severity_downgrade so it's obvious in the report). left class as descriptive metadata since the detector routing already pins it, so it's not pretending to guarantee anything now. on the gate: fair point, strict thresholds on a 2-case corpus would've nuked unrelated PRs the moment anything regressed. made the bugbench job continue-on-error for now so it still reports quality on each PR but doesn't block, and dropped a comment + doc note saying a failing case shows red without blocking, and how to make it a required check later once the corpus is bigger. also went ahead and added the null-origin and wildcard+creds cases you mentioned (both were unit-tested but not in the shipped corpus). corpus is 4 now, still 100% precision / 0 FP, and both carry a min_severity so they exercise the new check too. rebased on main and fixed the CLAUDE.md conflict while i was at it. 32 tests pass. |
The claim this unlocks
Almost no AI security tool can say that. Reputation in this space is trust, and trust is measured false positives + reproducibility. BugBench is the infrastructure that lets the project make — and keep — that claim.
Why it's needed
The repo generates findings (scanners, agent), captures them (
/poc), re-checks them (/replay), and prioritizes them (/oracle). What it has never had is a way to answer: "how often are the detectors actually right, and how often do they cry wolf?" Without that number, every accuracy claim is a vibe. BugBench makes it a measurement.What's in this PR
bughunter/bench/cases/*.json, with known-vulnerable AND known-safe cases. The safe cases are the point: they're what measure false positives. Each case carries a recorded input (the exact response a detector sees), never a live URL, so the score is byte-identical on every machine and in CI.bughunter/tools/bench.py: load → run detector → score. Detectors register via a one-line@detectoradapter (open/closed — adding SQLi/JWT/CRLF is a one-function change, the core never grows an if/elif tree). Thecorsdetector is wired first against its real, pure classifier, offline.test_accuracy_is_misleading_under_imbalance) proves a do-nothing detector scores 95% accuracy while catching zero bugs. That's why trust metrics lead.bugbenchjob runsbench.py run --min-precision 1.0 --max-fp-rate 0.0, which exits non-zero on regression. A change can never quietly make a detector noisier. "We measure quality" becomes enforced, not aspirational.The scoreboard today
(Starts with the CORS detector; the corpus is designed to grow — every new case and detector strengthens the guarantee.)
Quality
tests/test_bench.py), each layer isolated: data-contract validation, metric math (incl. the "accuracy lies" lesson), the runner against the realcorsclassifier, the scoreboard, and a gate test proving a wolf-crying detector fails CI.7 files changed, 803 insertions(+), nothing removed. A malformed corpus fails loud (never silently skews a score).How to extend (built for contribution)
bughunter/bench/cases/— no code.@detector("name")adapter inbench.py.Self-contained and valuable on its own; also the measuring stick for a follow-up prove-or-suppress verification layer.
Try it
🤖 Generated with Claude Code