I study whether post-training methods that reach the same reward leave the same model behind.
In DuraSeed, I compare skills acquired through supervised fine-tuning and on-policy RL, then test how those models retain the skill and respond to identical later training. In Kintsugi, I study whether repeated on-policy repair restores a model’s ability to keep learning, not only its visible behavior.


