EECS E6895 final project measuring reward-gaming behavior in Gemma 2B with shell-game evals, LoRA SFT, and leakage-aware probes.
-
Updated
May 12, 2026 - Python
EECS E6895 final project measuring reward-gaming behavior in Gemma 2B with shell-game evals, LoRA SFT, and leakage-aware probes.
Predict how a policy games a reward. RL environments scored by an executable-exploit verifier.
Reward hacking is largely a property of the scaffold: same model, same 103 tasks, 63% cheating in ImpossibleBench's harness vs 2% in an ordinary coding agent
98% of the Alpha Was a Gate — a documented reward hack in a self-optimizing LLM research agent. Full search trajectory, frozen evaluation harness, both strategies, and re-runnable ablations. The agent found a second exploit 18 minutes after the first was patched.
Multi-agent specification-gaming benchmark on Google ADK — N LLM workers compete on LeetCode under an exploitable reward function
A reinforcement learning workbench that imports nothing. Every algorithm written out, every run reproducible from its seed, and three environments where the best possible policy under the stated reward is the wrong behaviour.
To associate your repository with the specification-gaming topic, visit your repo's landing page and select "manage topics."