You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Evaluation
One of the big problem in this space is that there is no public benchmark for what thorough reviews should look like. We should have a scalable way to collect this benchmark.
Algorithm design
The current approach uses incremental summarization. It will have trouble for long-term dependency.
Test recursive language modeling for this task.
Turn the key idea into a claude skill
Integrate with https://github.com/ChicagoHAI/MechEvalAgent/ to have execution-grounded evaluation
Interaction
This tool would be much more useful if the users can address the comments interactively.
Improve comment quality
We need a way to evaluate the quality of the overall feedback and individual comments and improve them.
Evaluation
One of the big problem in this space is that there is no public benchmark for what thorough reviews should look like. We should have a scalable way to collect this benchmark.
Algorithm design
The current approach uses incremental summarization. It will have trouble for long-term dependency.
Test recursive language modeling for this task.
Turn the key idea into a claude skill
Integrate with https://github.com/ChicagoHAI/MechEvalAgent/ to have execution-grounded evaluation
Interaction
This tool would be much more useful if the users can address the comments interactively.
Improve comment quality
We need a way to evaluate the quality of the overall feedback and individual comments and improve them.