Synthetic-document fine-tuning on Qwen2.5-7B: a controlled study of whether SDF installs sandbagging, finding a layered recognition/generation/behavior dissociation.
-
Updated
May 25, 2026 - Python
Synthetic-document fine-tuning on Qwen2.5-7B: a controlled study of whether SDF installs sandbagging, finding a layered recognition/generation/behavior dissociation.
Reproducible red-team findings for openai/gpt-oss-20b: five minimal harnesses with checks, zips & manifest (v0.9.3).
Preregistered AI-safety study of sandbagging model organisms: trigger type sets the sign of cross-capability alignment (task-local locks dismantle it, situational locks amplify it) and cue-sharing sets its size. All five predictions failed, four reversed.
Proving that the neural network is honest about its lack of capabilities
To associate your repository with the sandbagging topic, visit your repo's landing page and select "manage topics."