Skip to content

Port SWE-Bench Multimodal benchmark to Harbor #726

Description

@neubig

Context

benchmarks/swebenchmultimodal does not appear to have a direct Harbor registry dataset/adapter yet. Harbor supports multimodal trajectories, but this benchmark needs a task adapter that preserves SWE-Bench Multimodal task setup, resolved-instance selection, and evaluation behavior.

Proposed scope

  • Create a Harbor adapter/dataset for SWE-Bench Multimodal.
  • Preserve the current resolved-instance selection behavior and multimodal assets.
  • Validate the verifier against the current swebenchmultimodal-infer/swebenchmultimodal-eval workflow.
  • Document any differences in image handling or multimodal output/trajectory handling.

Acceptance criteria

  • SWE-Bench Multimodal tasks can be run through Harbor.
  • Parity results are recorded against the existing OpenHands benchmark workflow.
  • The OpenHands wrapper can delegate SWE-Bench Multimodal execution to Harbor.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions