Skip to content

Evaluation roadmap v2 #96

Description

@dangng2004

This supersedes the v1 roadmap #51 as it has led to this paper: acl_latex.pdf.

That benchmark measures systems' error detection ability, which is good for a start, but we also need a metric for the quality of each comment and the overall feedback.

Some desiderata of the overall feedback:

  • high-level feedback that is less standard and includes more substance.
  • Prioritization of the issues: do more serious issues get foregrounded?
  • Clarity: is the communication clear?
  • Balance: does it balance the strengths and weaknesses of a paper?
  • Conciseness: does it communicate in a concise manner?

Some desiderata for the comments:

  • Relevance
  • Accuracy

We will likely use LLM as a judge for this phase of the benchmark.

Metadata

Metadata

Assignees

No one assigned

    Labels

    evaluationEvaluation related issueshelp wantedExtra attention is needed

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions