feat: deterministic-verification + parallel-swarm architecture - #170
Merged
Merged
Conversation
Separate discovery from validation so no finding reaches a report without a
deterministic, non-AI oracle confirming it against the live target.
- verifiers: tools/verifiers/ library (redirect, sensitive-file, auth-bypass,
sqli error+timing, idor two-identity, oob ssrf/xxe/rce, dom-xss) each
returning VerifyResult(confirmed, trace); REGISTRY + verify_finding dispatcher
- validator gate: tools/validate_core.py non-tty verify_finding_programmatic
encoding kill-signal / chain-required / Q8-identity tables as code;
validate.py --auto wrapper
- report gate: brain.write_report + hunt.py refuse any finding that is not a
verified validated_finding with linked evidence; agent._classify_obs emits
candidate leads, never direct findings
- swarm: fcntl-locked + adjudicated lead_board; tools/swarm.py coordinator fans
leads out to ephemeral per-lead workers (ThreadPoolExecutor, per-worker
session+scope); hunt.py --swarm flag
- llm skills: skills/llm-redteam + skills/agentic-app-audit; expanded
llm_redteam corpus (multi-turn, cross-lingual, cipher, Sneaky Bits
token-smuggling); unified ASI01-ASI10 canonical mapping; commands/llm-app-audit
- benchmark: tests/benchmark/ FLAG{} canary target + precision/recall/FP
scorer; tests/test_benchmark.py smoke test (recall=1.0, fp_rate=0.0)
Resolve conflicts from the bughunter/ package restructure: - Relocate new verifier/swarm/validate_core/evidence/redaction modules under bughunter/ - hunt.py: keep main's shell=False argv (command-injection fix) and pass SCOPE_ENFORCED via the environment instead of the shell string - adapters.py: keep both scope-filter helpers and path-escape guards - README: take main (supersedes the older sponsor block; keeps #169 contributors) - Union .gitignore; keep plugin.json version 6.0.1; keep unified ASI table - Fix test paths to bughunter/ layout and assert SCOPE_ENFORCED via env Co-authored-by: Cursor <cursoragent@cursor.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Reworks the hunter around the XBOW architecture thesis: discovery and validation are separate systems, and no finding reaches a report without a deterministic, non-AI oracle confirming it against the live target. Benchmark on a canary target shows precision=1.0, recall=1.0, fp_rate=0.0, 0 unverified findings reaching report.
Pillars
tools/verifiers/) — one oracle per bug class (redirect, sensitive-file, auth-bypass, sqli error+timing, idor two-identity, oob ssrf/xxe/rce, dom-xss), each returningVerifyResult(confirmed, trace). No LLM, no confidence scores.REGISTRY+verify_findingdispatcher + CLI.tools/validate_core.py) — non-ttyverify_finding_programmaticencoding the kill-signal / chain-required / Q8-identity tables as code.brain.write_reportandhunt.pyrefuse any finding that isn't a verifiedvalidated_findingwith linked evidence;agent._classify_obsnow emits candidate leads, never direct findings.tools/swarm.py) —fcntl-locked + adjudicated lead board; coordinator fans leads out to ephemeral per-lead workers (ThreadPoolExecutor) with per-worker session + scope. Gated behindhunt.py --swarm(default off).skills/llm-redteam/(canonical ASI01–ASI10 source of truth) andskills/agentic-app-audit/;llm_redteamcorpus expanded (multi-turn, cross-lingual, cipher, Sneaky Bits token-smuggling); unified the previously-conflicting ASI tables;commands/llm-app-audit.md.tests/benchmark/) —FLAG{}canary vuln target for recall, hardened demo for precision, precision/recall/FP scorer, andtests/test_benchmark.pysmoke test.Testing
885 passedacross the full suite.tests/test_benchmark.py: recall=1.0, fp_rate=0.0, precision=1.0 (TP=6, FN=0, FP=0, TN=5).