## Problem Many attacks test 'echo this token' with substring detector = tautological. Proves obedience, not vulnerability. ## Tasks - [ ] Tag attacks as obedience vs policy-bypass in attack metadata - [ ] Add attack_type field to schema - [ ] Report obedience and policy-bypass rates separately - [ ] Update METHODOLOGY.md to explain the distinction ## Acceptance Criteria - Results clearly distinguish 'model followed instruction' from 'model violated safety policy'
Problem
Many attacks test 'echo this token' with substring detector = tautological. Proves obedience, not vulnerability.
Tasks
Acceptance Criteria