Living burn-down tracker. Done today: the core loop — admin → scan → verify → triage →
generate → apply band-aid (service_policy create+attach+exploit-validate+rollback;
malicious_user LB-enable+config-validate+rollback) → open code-fix PR — model-
independent, with a localhost console and a 9/9 benchmark. This file tracks what's left.
Effort: S ≈ <1 session · M ≈ 1–2 · L ≈ multi. Priority: P0 foundational · P1 high-value · P2 later/bigger. Check items off as we land them.
Every control generate can emit should also be apply-able + validated, behind a dispatcher.
- A0 Unified apply dispatcher — DONE:
apply_control(control, lb, **kw)routes malicious_user / rate_limit / bot_defense to their handlers. (M, P1) - A1
bot_defenseapply — DONE:apply_bot_defenseflipsdisable_bot_defense→bot_defensewith a valid default flag-only policy (add-on IS present on the tenant), config validation + rollback + guardrails. Live round-trip validated on nimbus-www (default policy accepted by XC: enable→readback→rollback). CLIapply-bot. (S, P1) - A2
rate_limitapply — DONE:apply_rate_limitenables LB rate limiting (requests/unit/burst), config-validate + rollback; live round-trip validated on nimbus-www. CLIapply-ratelimit. (M, P1) - A3
waf/waf_data_guardapply — DONE:apply_wafcreates a Blocking app_firewall (cloned from a template), attaches it via a fully-qualified ref (name+namespace+tenant, popping thedisable_wafoneof), fires a SQLi and confirms the block (XC serves a 200Request Rejectedpage — the prober matches on the body), then rolls back.apply_data_guardensures the WAF is on (Data Guard requires it) then adds maskingdata_guard_rules; config-readback validated. Both live-validated onvpcopilot-lab. (M, P2) - A4
api_schemaapply — DONE:apply_api_schemauploads an OpenAPI to the XC object store (put_swagger, PUT to the stored-objects/swagger endpoint), creates anapi_definitionreferencing it, then attachesapi_specification.validation_all_spec_endpointswithvalidation_mode_active+request_validation_properties:[PROPERTY_HTTP_BODY]+enforcement_block+fall_through_mode_allow. Validated live onvpcopilot-lab: a-1payment (OpenAPIamount: exclusiveMinimum 0) returns 403 as a schema violation while+1passes; then rolls back. The schema-preferred positive-security band-aid. (L, P2)
- B1 Finding-correlation step — DONE:
correlate.pycoverage_key(LB-wide controls collapse to one instance;service_policykeyed per endpoint); pipeline skips generating a band-aid an earlier finding already covers, writescorrelations.json+ a summary line. Live: 4 redundant band-aids deduped. (M, P1) - B2 Verify confidence threshold — DONE:
--min-confidence(default 0.5) drops verified findings below it (logged, no silent cap); wired through scan/bench/console +run_pipeline. (S, P1) - B3 Behavioral validation — DONE:
apply-ratelimit --behavioralenables the limit, thenprobe_rate_limitdrives a burst above it and confirms the excess is rate-limited (429), proving mitigation vs config-only. Live-validated onvpcopilot-lab: 10/MINUTE + a 30-burst → 10 pass / 20 × 429. Surfaced in the report's Band-aid impact panel. (malicious-user/bot stay config-level — behavioral proof there needs sustained abuse + telemetry over minutes, tier- dependent; noted honestly.) (L, P2)
- C1 Remediation ledger — DONE:
ledger.pypersists per-findingfound→mitigated→remediated→retired(forward-only) inledger.json; pipeline seedsfound(+ apolicies.jsonpolicy→finding index),applymarksmitigated,prmarksremediated.vpcopilot ledgerCLI +/api/ledgerconsole endpoint. Tests added. (M, P0) - C2 Auto-retire band-aid on cure-merge — DONE:
retire.py+vpcopilot retire [--finding <id> | --all] [--force]: checks the finding's cure PR is merged (GitHub API, parsed from the ledgercure.pr_url), then detaches its control from the LB (_detach_control, the inverse of each apply_*) andledger.mark_retired, closing found→mitigated→remediated→retired. Same protected-LB guardrail. Live-validated: apply-ratelimit --keep → retire --force → rate_limit removed + ledger retired. 5 tests. (M, P2) - C3 PR tracking — DONE: "Open all code-fix PRs" batch button; PR links surfaced
inline (dashboard actions) and in the Ledger tab (from the ledger
cure). (S, P1)
- D1 Bonus-vuln scoring — DONE:
bonus:section inanswer_key.yaml; scorer credits real extra findings (bonus_found) and reports only genuinenoise. (S, P1) - D2 Per-stage metrics — DONE:
run_pipelinetimes each stage and counts outcomes →metrics.json(+summary.metrics): timing_s {discover, verify, synthesize, total}; verify {candidates, verified, refuted, dropped_low_confidence, confirm_rate, avg_confidence}; synthesize {policies, dupe_bandaids_collapsed, code_fix_prs}. Rendered as a Pipeline metrics panel in the HTML report. Live run: discover 7.4s / verify 3.1s / synth 75.9s, verify 12→11 (0.92 confirm), 4 dupes collapsed. (Also fixed a latent E3 bug — the auto-report call usedlogoutside its scope in_write_out; moved torun_pipeline.) (M, P2) - D3 Multi-provider proof run — DONE (see MODELS.md): config-only swap ran the
full pipeline on
gpt-4o(Claude 9/9, gpt-4o ~8/9 real, triage 100% on both). Surfaced + fixed the "trust intentional/demo comments" reviewer weakness for all models. (S, P0)
- E1 Per-finding action buttons — DONE: dashboard rows have inline Apply {control} (routes service_policy→/api/apply, malicious_user/rate_limit/bot_defense→their endpoints) + Open PR, driven by an action-settings bar; per-row result inline. (M, P1)
- E2 Ledger view — DONE: Ledger tab renders
/api/ledger(found→mitigated→remediated→retired) with mitigation control + cure PR links. (S, P1) - E3 Standalone shareable HTML export — DONE:
report.pyreads the out/ artifacts and writes a single self-containedreport.html(inline CSS, native<details>, no server/external assets; model content HTML-escaped): run-summary chips, per-finding cards (severity, class, band-aid chips, code-cure badge, expandable exploit/snippet), grouped XC policies, and the ledger. Every scan auto-writesout/report.html;vpcopilot report [--open]+ a console Open HTML report button (/api/report) rebuild it. 3 tests. (M, P2) - E4 Richer before/after panel — DONE: the exploit-validated applies (service_policy,
waf, api_schema) fire a baseline exploit BEFORE mutating and return a
before_after{before/after → exploit_status, exploit_blocked, legit_ok} (normalized across probes viaprobe.normalize), persisted to the audit log. The HTML report renders a Band-aid impact table (exploit200 allowed → 403 blocked, legit ok, PASS). Live-validated. (M, P2) - E5 Workflow tab — DONE: visual agent pipeline (discover→verify→triage→generate→
remediate) with each agent's configured model (from
/api/agents) + roles, the deterministic spine (correlate / human gate / apply / PR), and last-run counts. (S, P1)
- F1 Packaging — DONE:
--version; wheelforce-includeships the console HTML;vpcopilotconsole-script entrypoint. (S, P1) - F2 Test coverage — DONE: unit tests for the service-policy normalizer, the protected-LB + protected-policy guardrails (fake XC env, no network), the ledger, correlate, audit, and schemas. 16 tests. (M, P0)
- F3 Pipeline concurrency — DONE: discover + verify run in a
ThreadPoolExecutor(--concurrency, default 8); the first discover call runs solo to warm instructor's (non-thread-safe) mode registry, then the rest parallelize. Validated live. (M, P1) - F4 Audit log — DONE:
audit.pyappends every mutating action (create/attach/enable/rollback/PR) to<out>/audit.log(UTC ts + details);vpcopilot auditCLI +/api/audit. (S, P1) - F5 Customer docs — DONE:
docs/USAGE.md(install, config/model-independence, all commands, console, safety model + guardrails, worked Nimbus example). (M, P1)
- C1 ledger + F2 tests — foundations everything else leans on.
- D3 multi-provider proof — cheap, and it substantiates the headline claim.
- A0 → A1 → A2 — finish the easy apply toolbox behind the dispatcher.
- B2 → B1 — triage quality (confidence gate, correlation).
- C3 → E1 → E2 — cure tracking + console UX.
- D1, F1, F4, F5 — eval polish + hardening + docs.
- A3, A4, B3, C2, D2, E3, E4 — bigger / optional.
(BACKLOG.md holds looser "someday" ideas; this file is the committed plan.)