This supersedes the v1 roadmap #51 as it has led to this paper: acl_latex.pdf.
That benchmark measures systems' error detection ability, which is good for a start, but we also need a metric for the quality of each comment and the overall feedback.
Some desiderata of the overall feedback:
- high-level feedback that is less standard and includes more substance.
- Prioritization of the issues: do more serious issues get foregrounded?
- Clarity: is the communication clear?
- Balance: does it balance the strengths and weaknesses of a paper?
- Conciseness: does it communicate in a concise manner?
Some desiderata for the comments:
We will likely use LLM as a judge for this phase of the benchmark.
This supersedes the v1 roadmap #51 as it has led to this paper: acl_latex.pdf.
That benchmark measures systems' error detection ability, which is good for a start, but we also need a metric for the quality of each comment and the overall feedback.
Some desiderata of the overall feedback:
Some desiderata for the comments:
We will likely use LLM as a judge for this phase of the benchmark.