Skip to content

Explicit Edit Benchmark: testing how harness helps an agent to edit #223

Description

@alexshpunt

Hi! I’m maintaining the Explicit Edit Benchmark, a small public benchmark for measuring how much agent tooling and harness design affect a model’s ability to make precise text edits.

Text editing is one of the most common operations in coding-agent workflows. In my experience, it is also one of the operations agents struggle with the most. Encodings, text blocks, line endings, invisible characters, unicode, and file formats with strict formatting or structural requirements create a huge number of variations. I have seen an agent fail to edit a single line because it could not reproduce an invisible character in that line. There are many cases like this.

The benchmark currently contains 226 deterministic tasks across different file formats and edit types: replacements, insertions, deletions, moves, copies, unicode cases, large files, and other edge cases. The tasks are intentionally simple, so they can be run even with low reasoning and keep the focus on the model and tooling combination rather than deep problem solving.

I’ve been running as many agents, models, configurations, and Pi editing extensions as I can and publishing every accepted observation in the public dataset. There are too many combinations for one person to cover, and model behavior is stochastic, so a single run cannot give us a reliable picture. The benchmark was designed from the start as a public, community-driven dataset that anyone interested and able to run it can contribute to. We can get reliable information only by collecting enough observations across different setups.

I stumbled across your agent and realized it's pretty similar to what I'm doing with my pi-agent-ide, hence decided to leave a note here in case you are interested.

Viewer: https://huggingface.co/spaces/alexshpunt/benchmark-explorer
Repo: https://github.com/alexshpunt/explicit-edit-benchmark

If you’re interested, you’re very welcome to run the benchmark for your own configuration of the agent and a model of your choice.

No action required. I wanted to share the info with you because your project is super related to my sphere of interests and I wondered if you were willing to test it as well.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions