Cookies, session recording and ad conversion measurement improve the product. Before you accept we measure page views and anonymous interactions, with no cookies, no recording and no profile. Reject keeps it that way. Cookie policy

Vetoo
ManifestoLive reviewAgent panelMemory & trustPricing
Resources
Competitor comparisonLearn centerDocs
Sign up
Sign inSign up

28 September 2026

Vetoo ranks among the top
in an independent AI code review benchmark

Martian’s open benchmark scores 22 AI code reviewers against comments written by human reviewers on 50 pull requests from cal.com, Sentry, Grafana, Keycloak and Discourse. Two independent judges scored every review. Vetoo placed top 5 on large PRs, complex code, cross-file changes and critical-risk PRs, and first on Ruby.

Precision and recall

Precision

Precision is the share of a tool’s comments that flag a genuine problem. The rest is noise reviewers have to read and dismiss.

Comment matches a real issueNoise

Recall

Recall is the share of a pull request’s real issues that the tool catches. Every issue it misses is left for a human reviewer, or for production.

Issue caughtIssue missed

What F1 measures

F1 folds precision and recall into one balanced score, so a tool that leans hard on one side of the tradeoff is penalized. Two strategies score poorly:

High recall, low precision
Flagging everything and burying real issues in noise.
High precision, low recall
Commenting rarely and letting bugs through.

A high F1 only comes from reviews that are both accurate and thorough.

Head to head, slice by slice.

Besides the overall result, the benchmark also compares tools across 36 comparison groups. Each card shows how many of those groups Vetoo won against the given tool.

36 groups categorized:

  • OverallAll 50 PRs
  • LanguageGo, Java, Python, Ruby, TypeScript
  • PR sizeLarge PRs, Medium PRs, Small PRs
  • DomainUI, Authentication, Caching, Concurrency, Scheduling
  • ComplexityComplex Code, Moderate Code
  • DifficultyModerate Bugs, Subtle Bugs
  • RiskCritical Risk, High Risk, Medium Risk
  • ContextCross-File, File Context
  • ConcernCorrectness, Reliability, Security
  • Change typeBug Fixes, Features, Performance Optimization
  • CombinedSmall Go PRs, Medium Java PRs, Medium Python PRs, Medium Ruby PRs, Complex & Subtle, High Risk Auth, Security Critical
  • VetoovsGraphite360
  • VetoovsCodeAnt315
  • VetoovsClaude315
  • VetoovsSourcery297
  • VetoovsCodeRabbit288
  • VetoovsCopilot279
  • VetoovsGemini279
  • VetoovsKodus2610
  • VetoovsCursor Bugbot2016
  • VetoovsMacroscope2016
  • VetoovsGitLab Duo2016
  • VetoovsGreptile v41917

F1 Leaderboard

RankToolF1
  1. 1Qodo Extended66% · 57%
  2. 2Cubic64% · 58%
  3. 3Augment62% · 57%
  4. 4Qodo58% · 51%
  5. 5Devin52% · 52%
  6. 6Vetoo52% · 50%
  7. 7GitLab Duo53% · 48%
  8. 8Macroscope54% · 46%
  9. 9Cursor Bugbot51% · 49%
  10. 10Greptile v451% · 48%
  11. 11Gemini50% · 44%
  12. 12GitHub Copilot49% · 44%
  13. 13Kodus47% · 45%
  14. 14Baz46% · 43%
  15. 15Sourcery47% · 42%
  16. 16Claude Code (CLI)47% · 39%
  17. 17CodeRabbit42% · 43%
  18. 18Gitar41% · 42%
  19. 19Claude40% · 37%
  20. 20CodeAnt41% · 34%
  21. 21KG24% · 22%
  22. 22Graphite14% · 14%

Vetoo’s strengths

Vetoo ranks in the top 4 for precision under both judges: 62% of its comments matched a real issue under Opus 4.5, and 59% under Sonnet 4.5. That precision comes without giving up coverage. Vetoo catches 44% of real issues under Opus 4.5 and 43% under Sonnet 4.5, more than Devin and Graphite, the other tools at the high-precision end. Under Sonnet 4.5, no tool with higher precision finds more issues than Vetoo.

VetooP 62% · R 44% · F1 52%
Recall %01020304050607044
Cubic
GitHub Copilot
Qodo Extended
CodeRabbit
Gemini
Greptile v4
Cursor Bugbot
Devin
Claude
Graphite
Vetoo
3040506070809010062Precision %
Qodo Extended: precision 67%, recall 65%, F1 66%. Cubic: precision 62%, recall 66%, F1 64%. Devin: precision 69%, recall 41%, F1 52%. Vetoo: precision 62%, recall 44%, F1 52%. Cursor Bugbot: precision 57%, recall 47%, F1 51%. Greptile v4: precision 51%, recall 51%, F1 51%. Gemini: precision 45%, recall 55%, F1 50%. GitHub Copilot: precision 38%, recall 66%, F1 49%. CodeRabbit: precision 33%, recall 60%, F1 42%. Claude: precision 43%, recall 38%, F1 40%. Graphite: precision 100%, recall 8%, F1 14%

That balance comes from how a Vetoo review runs: several models review each change and a synthesis step settles what reaches the pull request.

Methodology

  1. 01

    50 human-curated PRs

    From cal.com, Sentry, Grafana, Keycloak and Discourse. Each PR carries golden comments written by human reviewers, tagged by severity and category.

  2. 02

    Vetoo reviews each fork

    All 50 PRs were forked and Vetoo posted its standard review on each. 214 comments in total, 50 of 50 reviews completed with 0 errors.

  3. 03

    Two judges, run separately

    Claude Opus 4.5 and Claude Sonnet 4.5 each matched Vetoo's findings to the golden comments to count true positives, false positives and misses per PR.

  4. 04

    F1 on the Core profile

    Precision and recall weighted equally, over bug, security, concurrency, data, API, performance, test-gap and doc-defect findings. The benchmark's default.

For a different angle, our own benchmark replays merged pull requests that CodeRabbit, Greptile or Cursor Bugbot had already reviewed and lists the defects Vetoo found that they did not.

On this page
Precision and recallWhat F1 measuresHead to headF1 LeaderboardVetoo's strengthsMethodology

Clear the review queue
before it forms.

Connect your first repository and let Vetoo do the rest.

Connect your repository
Vetoo

Product

ManifestoLive reviewAgent panelMemory & trustPricingFAQ

Resources

All resourcesBenchmarkWhat is AI code review?Code review checklistVetoo vs CodeRabbitVetoo vs GreptileVetoo vs Bugbot

Developers

Docs

Company

AboutContact

Legal

Legal
AI code reviewTry it out