Skip to content

Port OpenAgentSafety benchmark to Harbor #722

Description

@neubig

Context

benchmarks/openagentsafety does not appear to have a corresponding Harbor registry dataset/adapter yet. To remove benchmark-specific execution code from this repository, OpenAgentSafety needs a Harbor-native task representation and verifier.

Proposed scope

  • Model OpenAgentSafety scenarios as Harbor tasks.
  • Preserve NPC/workplace interaction setup and safety scoring semantics.
  • Validate Oracle/reference behavior where available.
  • Compare aggregate results with the current OpenHands benchmark harness.

Acceptance criteria

  • OpenAgentSafety tasks can be run through Harbor.
  • Verifier/scoring parity is documented.
  • The OpenHands wrapper can delegate OpenAgentSafety execution to Harbor.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions