diff --git a/offline/README.md b/offline/README.md index 3aa9456..8351558 100644 --- a/offline/README.md +++ b/offline/README.md @@ -17,6 +17,7 @@ Open replication of the code review benchmark used by companies like [Augment](h | [Greptile](https://www.greptile.com/) | AI code review | | [Propel](https://propelauth.com/) | AI code review | | [Qodo](https://www.qodo.ai/) | AI code review | +| [Tusk](https://www.usetusk.ai/) | AI code review & testing | Adding a new tool requires forking the benchmark PRs and collecting the tool's reviews — see Steps 0 and 1 below. diff --git a/offline/analysis/benchmark_dashboard.html b/offline/analysis/benchmark_dashboard.html index 4905b9e..ae0f6db 100644 --- a/offline/analysis/benchmark_dashboard.html +++ b/offline/analysis/benchmark_dashboard.html @@ -193,22 +193,23 @@
Highest Precision
Best for Concurrency (Precision)
Best for Complex Code (Precision)
-
Best for Bug Fixes (Recall)
Best for Bug Fixes (Recall)
Best for Bug Fixes (Recall)
+
Best for Bug Fixes (Recall)
Best for Performance Optimization
+
Best for Ui (Recall)
Highest Recall
Java + Authentication
-
Moderate Bugs + File Context (Recall)
Java + High Risk (Recall)
+
Moderate Bugs + File Context (Recall)
Best for Medium Ruby PRs
Typescript + Correctness
Best for Critical Risk
Java + Authentication
Bug Fixes + Cross-File
Best for Ui
-
Java + High Risk
File Context + Correctness
+
Java + High Risk
Medium PRs + File Context
Ruby + Features (Precision)
Best for Complex Code
@@ -413,10 +414,10 @@