Skip to content

Port ProgramBench benchmark to Harbor #723

Description

@neubig

Context

benchmarks/programbench currently contains substantial benchmark-specific inference and evaluation logic for cleanroom task images and ProgramBench submission/evaluation behavior. A Harbor adapter is needed before this benchmark can be removed from the bespoke OpenHands path.

Proposed scope

  • Create a Harbor adapter/dataset for ProgramBench tasks.
  • Preserve cleanroom task image usage, workspace submission format, and upstream evaluation semantics.
  • Support resource settings appropriate for ProgramBench long rebuild tasks.
  • Run parity against the current programbench-infer/programbench-eval workflow.

Acceptance criteria

  • ProgramBench tasks can be run through Harbor.
  • Evaluation output matches upstream ProgramBench expectations.
  • Parity results and resource caveats are documented.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions