Skip to content

Improve generic Harbor and MiniSWE evaluation resilience #761

Description

@neubig

Track benchmark-agnostic fixes for long MiniSWE and Harbor runs.

Acceptance criteria:

  • expose the OpenAI-compatible aliases expected by MiniSWE
  • support bounded, tool-safe message history
  • treat authoritative verifier rewards correctly after agent failures
  • rebuild converted output from raw Harbor results during rescoring

Implementation: #759

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions