Skip to content
#

password-locked-models

Here is 1 public repository matching this topic...

Preregistered AI-safety study of sandbagging model organisms: trigger type sets the sign of cross-capability alignment (task-local locks dismantle it, situational locks amplify it) and cue-sharing sets its size. All five predictions failed, four reversed.

  • Updated Sep 5, 2026
  • Python

Add this topic to your repo

To associate your repository with the password-locked-models topic, visit your repo's landing page and select "manage topics."

Learn more