Track benchmark-agnostic fixes for long MiniSWE and Harbor runs.
Acceptance criteria:
- expose the OpenAI-compatible aliases expected by MiniSWE
- support bounded, tool-safe message history
- treat authoritative verifier rewards correctly after agent failures
- rebuild converted output from raw Harbor results during rescoring
Implementation: #759
Track benchmark-agnostic fixes for long MiniSWE and Harbor runs.
Acceptance criteria:
Implementation: #759