Repository navigation
Conversation
The hunt-memory flywheel records every finding's outcome (confirmed/rejected/ false_positive, payout, severity, vuln_class, endpoint, tech) in journal.jsonl, but nothing learns from it — the lead board still ranks by hardcoded regex priority. docs/TODOS.md TODO-6 flags exactly this: the memory→hunt loop "never spins up in practice." Oracle closes the loop: it trains on the hunter's own history and scores any lead by P(this becomes a real, valuable finding). Model choice is deliberate. Bug-bounty history is small (tens–hundreds of labelled findings), so a heavy ML stack would overfit and add a dependency this repo avoids. Oracle uses a smoothed Naive-Bayes log-odds model — right for the data regime: calibrated, fully explainable (every score carries its per-feature log-odds), pure-stdlib, deterministic. A cold-start PRIOR (the same high-value intuition the regex scorer encodes) is blended in and decays as evidence grows (prior → blended → learned), so it's useful on day one and data-driven later. Why it beats the static prioritizer: the regex priority is identical for every program; Oracle learns program-specific reality. In the demo history where IDOR is always a duplicate and open-redirect chains to ATO, leave-one-out CV scores Oracle 1.0 accuracy vs the prior heuristic's 0.471 (+0.529) — it correctly down-ranks the duped IDOR and up-ranks the paying open-redirect. `oracle eval` runs that CV against prior + majority baselines on the user's own data so the lift is provable before it's trusted. Commands: train (learn + save + report), eval (LOOCV vs baselines), rank <target> (scores lead_board leads, reuses load_ledger), score (one hypothetical lead). Read-only — never mutates the lead board. Compounds with /poc + /replay, whose outcomes feed the journal Oracle learns from. 22 tests in tests/test_oracle.py (features, learning, calibration, explanations, serialization, cold-start blend, LOOCV-beats-baseline, data loading, all CLI). /oracle command + CLAUDE.md entries. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
shuvonsec
left a comment
There was a problem hiding this comment.
Really like this - JSON persistence (no pickle), div-by-zero guarded throughout, cold-start handled and tested, and the tests actually exercise the learning signal + LOOCV-beats-baseline. Verified 22 tests pass. One accuracy ask (non-blocking):
- LOOCV in
train/evalscores folds with_learned_score(pure learned, w=1), but productionrank/scorecallmodel.score()which blends learned with the cold-start prior weightedw = n/(n+20). On the small histories Oracle targets (e.g.34 samples -> w0.63) the deployed scorer is only ~63% learned, so the reported CV lift overstates what users actually get. Either run LOOCV through the deployedscore(), or label the number as "learned-only upper bound."
|
Great catch, @shuvonsec — you're exactly right that the CV number was measuring
The eval table and the headline lift now use the deployed number: Added 2 tests: one asserts both metrics are reported (and the UB ≥ deployed), and a regression guard that LOOCV actually calls Also merged the latest |
# Conflicts: # CLAUDE.md
…only Addresses @shuvonsec's accuracy note on /oracle. train/eval computed LOOCV with _learned_score (pure learned, w=1), but production rank/score use the blended score() (w = n/(n+SHRINKAGE_K)) — so on small histories (e.g. n=34 → ~63% learned) the reported CV lift overstated what users actually get. loocv() now scores each held-out fold with the DEPLOYED score() — the same blend users receive — as the primary "oracle" metric, and reports the pure-learned result separately as "oracle_learned_only" (the upper bound the blend converges to as history grows). The eval table and the headline lift now use the deployed number; the upper bound is labelled as such. Implements both of the reviewer's suggested options: run LOOCV through the deployed scorer AND label the learned-only number as an upper bound. Tests: +2 in tests/test_oracle.py — asserts both metrics are reported and that LOOCV exercises the deployed score() path (regression guard against reverting to _learned_score). 24 pass. (This commit also merges the updated main and resolves the CLAUDE.md conflict, keeping both the /dashboard and /oracle entries.) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The one-line case
The repo already stores every finding's outcome — and never learns from it. Oracle is the model that finally makes the hunt-memory flywheel decide what to hunt next.
Why this needs to exist (no hand-waving)
Every hunt writes outcomes to
journal.jsonl—confirmed/rejected/false_positive, plus payout, severity, vuln class, endpoint, tech stack. That is a labelled dataset of what actually pays for this hunter on these surfaces. Yet the lead board still ranks purely by a hardcoded regex priority (high/med/low), identical for every program and every hunter. The maintainers already flagged this exact hole:Oracle closes that loop. It trains on the hunter's own journal and scores any lead by P(this becomes a real, valuable finding) — turning a dormant data asset into the prioritization brain.
Why it's provably better than what exists (the part that removes doubt)
The static regex scorer is context-blind: it thinks IDOR is always high-value and open-redirect is always low. Real programs disagree — one program dupes every IDOR; another pays big for an open-redirect→ATO chain. Oracle learns that program-specific reality.
oracle evalproves it on your own data with leave-one-out cross-validation against two baselines. On a history where IDOR is always duped and open-redirect always pays:Oracle correctly down-ranks the duped IDOR (
30%) and up-ranks the paying open-redirect (81%) — the static scorer gets both backwards. The lift is measured, not asserted, and it's shown before you trust the model.Why this model, specifically (engineering judgement)
Bug-bounty history is small (tens–hundreds of labelled findings). A gradient-boosted / neural stack would overfit that and drag in a dependency this repo deliberately avoids (
requirements: justrequests+mcp). So Oracle uses a smoothed Naive-Bayes log-odds model — the correct tool for the data regime:class:idor -2.84 (seen 0✓/18✗)), so a hunter can see why. No black box.prior→blended→learned). Useful on day one, data-driven by day thirty.What it does
/oracle trainoracle_model.json, print CV report/oracle eval/oracle rank <target>lead_board.load_ledger)/oracle score --class idor --url …Expected value =
P(worth pursuing) × mean historical payout for that class, so ranking optimizes for money, not just probability. Read-only — Oracle never mutates the lead board.It compounds with what's already here
/poc(merged, #144) and/replay(#152) produce the very outcome labels Oracle learns from: capture → confirm/replay → journal → Oracle gets smarter → better leads. This is the flywheel actually turning.Safety & quality
tests/test_oracle.py): feature extraction, that the model learns a real signal and stays a probability, explanations point at the right feature, serialization round-trips, cold-start falls back to prior, LOOCV beats the baselines, data loading, and all four CLI commands. Additive only:4 files changed, 860 insertions(+).Try it