Skip to content

feat: deterministic-verification + parallel-swarm architecture - #170

Merged
shuvonsec merged 2 commits into
mainfrom
feat/xbow-killer-deterministic-verification
Oct 2, 2026
Merged

shuvonsec merged 2 commits into
mainfrom
feat/xbow-killer-deterministic-verification

Conversation

@shuvonsec

@shuvonsec shuvonsec commented Oct 2, 2026 •

Copy link
Copy Markdown
Member

Summary

Reworks the hunter around the XBOW architecture thesis: discovery and validation are separate systems, and no finding reaches a report without a deterministic, non-AI oracle confirming it against the live target. Benchmark on a canary target shows precision=1.0, recall=1.0, fp_rate=0.0, 0 unverified findings reaching report.

Note: this branch also carries other previously-uncommitted project work (phase docs, evidence/redaction store, scope hardening, MCP changes) because those changes are commingled in shared files (agent.py, brain.py, lead_board.py, scope_checker.py) and could not be cleanly separated at file level. The XBOW-killer work is the headline below.

Pillars

  1. Deterministic verifier library (tools/verifiers/) — one oracle per bug class (redirect, sensitive-file, auth-bypass, sqli error+timing, idor two-identity, oob ssrf/xxe/rce, dom-xss), each returning VerifyResult(confirmed, trace). No LLM, no confidence scores. REGISTRY + verify_finding dispatcher + CLI.
  2. Programmatic validator + hard report gate (tools/validate_core.py) — non-tty verify_finding_programmatic encoding the kill-signal / chain-required / Q8-identity tables as code. brain.write_report and hunt.py refuse any finding that isn't a verified validated_finding with linked evidence; agent._classify_obs now emits candidate leads, never direct findings.
  3. Parallel swarm (tools/swarm.py) — fcntl-locked + adjudicated lead board; coordinator fans leads out to ephemeral per-lead workers (ThreadPoolExecutor) with per-worker session + scope. Gated behind hunt.py --swarm (default off).
  4. LLM-attack skills — new skills/llm-redteam/ (canonical ASI01–ASI10 source of truth) and skills/agentic-app-audit/; llm_redteam corpus expanded (multi-turn, cross-lingual, cipher, Sneaky Bits token-smuggling); unified the previously-conflicting ASI tables; commands/llm-app-audit.md.
  5. Benchmark harness (tests/benchmark/) — FLAG{} canary vuln target for recall, hardened demo for precision, precision/recall/FP scorer, and tests/test_benchmark.py smoke test.

Testing

  • 885 passed across the full suite.
  • tests/test_benchmark.py: recall=1.0, fp_rate=0.0, precision=1.0 (TP=6, FN=0, FP=0, TN=5).

shuvonsec and others added 2 commits October 2, 2026 14:27
Separate discovery from validation so no finding reaches a report without a
deterministic, non-AI oracle confirming it against the live target.

- verifiers: tools/verifiers/ library (redirect, sensitive-file, auth-bypass,
  sqli error+timing, idor two-identity, oob ssrf/xxe/rce, dom-xss) each
  returning VerifyResult(confirmed, trace); REGISTRY + verify_finding dispatcher
- validator gate: tools/validate_core.py non-tty verify_finding_programmatic
  encoding kill-signal / chain-required / Q8-identity tables as code;
  validate.py --auto wrapper
- report gate: brain.write_report + hunt.py refuse any finding that is not a
  verified validated_finding with linked evidence; agent._classify_obs emits
  candidate leads, never direct findings
- swarm: fcntl-locked + adjudicated lead_board; tools/swarm.py coordinator fans
  leads out to ephemeral per-lead workers (ThreadPoolExecutor, per-worker
  session+scope); hunt.py --swarm flag
- llm skills: skills/llm-redteam + skills/agentic-app-audit; expanded
  llm_redteam corpus (multi-turn, cross-lingual, cipher, Sneaky Bits
  token-smuggling); unified ASI01-ASI10 canonical mapping; commands/llm-app-audit
- benchmark: tests/benchmark/ FLAG{} canary target + precision/recall/FP
  scorer; tests/test_benchmark.py smoke test (recall=1.0, fp_rate=0.0)
Resolve conflicts from the bughunter/ package restructure:
- Relocate new verifier/swarm/validate_core/evidence/redaction modules under bughunter/
- hunt.py: keep main's shell=False argv (command-injection fix) and pass
  SCOPE_ENFORCED via the environment instead of the shell string
- adapters.py: keep both scope-filter helpers and path-escape guards
- README: take main (supersedes the older sponsor block; keeps #169 contributors)
- Union .gitignore; keep plugin.json version 6.0.1; keep unified ASI table
- Fix test paths to bughunter/ layout and assert SCOPE_ENFORCED via env

Co-authored-by: Cursor <cursoragent@cursor.com>
@shuvonsec
shuvonsec merged commit 647a9ea into main Oct 2, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant