Vetoo ranks among the top
in an independent AI code review benchmark
Martian’s open benchmark scores 22 AI code reviewers against comments written by human reviewers on 50 pull requests from cal.com, Sentry, Grafana, Keycloak and Discourse. Two independent judges scored every review. Vetoo placed top 5 on large PRs, complex code, cross-file changes and critical-risk PRs, and first on Ruby.
Precision and recall
Precision
Precision is the share of a tool’s comments that flag a genuine problem. The rest is noise reviewers have to read and dismiss.
Recall
Recall is the share of a pull request’s real issues that the tool catches. Every issue it misses is left for a human reviewer, or for production.
What F1 measures
F1 folds precision and recall into one balanced score, so a tool that leans hard on one side of the tradeoff is penalized. Two strategies score poorly:
- High recall, low precision
- Flagging everything and burying real issues in noise.
- High precision, low recall
- Commenting rarely and letting bugs through.
A high F1 only comes from reviews that are both accurate and thorough.
Head to head, slice by slice.
Besides the overall result, the benchmark also compares tools across 36 comparison groups. Each card shows how many of those groups Vetoo won against the given tool.
36 groups categorized:
- OverallAll 50 PRs
- LanguageGo, Java, Python, Ruby, TypeScript
- PR sizeLarge PRs, Medium PRs, Small PRs
- DomainUI, Authentication, Caching, Concurrency, Scheduling
- ComplexityComplex Code, Moderate Code
- DifficultyModerate Bugs, Subtle Bugs
- RiskCritical Risk, High Risk, Medium Risk
- ContextCross-File, File Context
- ConcernCorrectness, Reliability, Security
- Change typeBug Fixes, Features, Performance Optimization
- CombinedSmall Go PRs, Medium Java PRs, Medium Python PRs, Medium Ruby PRs, Complex & Subtle, High Risk Auth, Security Critical
- VetoovsGraphite360
- VetoovsCodeAnt315
- VetoovsClaude315
- VetoovsSourcery297
- VetoovsCodeRabbit288
- VetoovsCopilot279
- VetoovsGemini279
- VetoovsKodus2610
- VetoovsCursor Bugbot2016
- VetoovsMacroscope2016
- VetoovsGitLab Duo2016
- VetoovsGreptile v41917
- 1Qodo Extended66% · 57%
- 2Cubic64% · 58%
- 3Augment62% · 57%
- 4Qodo58% · 51%
- 5Devin52% · 52%
- 6Vetoo52% · 50%
- 7GitLab Duo53% · 48%
- 8Macroscope54% · 46%
- 9Cursor Bugbot51% · 49%
- 10Greptile v451% · 48%
- 11Gemini50% · 44%
- 12GitHub Copilot49% · 44%
- 13Kodus47% · 45%
- 14Baz46% · 43%
- 15Sourcery47% · 42%
- 16Claude Code (CLI)47% · 39%
- 17CodeRabbit42% · 43%
- 18Gitar41% · 42%
- 19Claude40% · 37%
- 20CodeAnt41% · 34%
- 21KG24% · 22%
- 22Graphite14% · 14%
Vetoo’s strengths
Vetoo ranks in the top 4 for precision under both judges: 62% of its comments matched a real issue under Opus 4.5, and 59% under Sonnet 4.5. That precision comes without giving up coverage. Vetoo catches 44% of real issues under Opus 4.5 and 43% under Sonnet 4.5, more than Devin and Graphite, the other tools at the high-precision end. Under Sonnet 4.5, no tool with higher precision finds more issues than Vetoo.
Qodo Extended
CodeRabbit
Gemini
Greptile v4
Cursor Bugbot
GraphiteThat balance comes from how a Vetoo review runs: several models review each change and a synthesis step settles what reaches the pull request.
Methodology
50 human-curated PRs
From cal.com, Sentry, Grafana, Keycloak and Discourse. Each PR carries golden comments written by human reviewers, tagged by severity and category.
Vetoo reviews each fork
All 50 PRs were forked and Vetoo posted its standard review on each. 214 comments in total, 50 of 50 reviews completed with 0 errors.
Two judges, run separately
Claude Opus 4.5 and Claude Sonnet 4.5 each matched Vetoo's findings to the golden comments to count true positives, false positives and misses per PR.
F1 on the Core profile
Precision and recall weighted equally, over bug, security, concurrency, data, API, performance, test-gap and doc-defect findings. The benchmark's default.
For a different angle, our own benchmark replays merged pull requests that CodeRabbit, Greptile or Cursor Bugbot had already reviewed and lists the defects Vetoo found that they did not.