Skip to content
#

specification-gaming

Here are 6 public repositories matching this topic...

Language: All
Filter by language

98% of the Alpha Was a Gate — a documented reward hack in a self-optimizing LLM research agent. Full search trajectory, frozen evaluation harness, both strategies, and re-runnable ablations. The agent found a second exploit 18 minutes after the first was patched.

  • Updated Aug 22, 2026
  • TeX

A reinforcement learning workbench that imports nothing. Every algorithm written out, every run reproducible from its seed, and three environments where the best possible policy under the stated reward is the wrong behaviour.

  • Updated Sep 7, 2026
  • Python

Add this topic to your repo

To associate your repository with the specification-gaming topic, visit your repo's landing page and select "manage topics."

Learn more