ReviewBench: An open benchmark for AI code review

Agentic code review is becoming an essential piece of how development happens. It helps you inspect pull requests, catch issues, and decide what deserves attention before code ships.
But the quality of existing AI reviewers can be hard to measure, and you need to know the strengths of a reviewer before you know if it will help you. Some reviewers surface more issues, some produce less noise, and some are stronger at catching critical problems while others surface smaller improvements, too. You may need code review to do different things within your workflow.
That makes it important to understand how reviewers actually compare: what different systems catch, what they miss, and the tradeoffs they make. A good code review benchmark should reflect the diversity of real pull requests, capture a broad set of review findings, and support meaningful breakdowns by severity, category, and precision-recall preferences. For teams building code review agents, the benchmark should also provide an offline signal that reliably tracks whether changes are likely to improve the experience in production. Existing benchmarks often make tradeoffs between label quality, coverage, and how well they represent real-world code review, leaving a gap for a rigorous and reproducible evaluation methodology that brings these pieces together.
We built ReviewBench, a new code review offline benchmark, to address that gap, and it is available for you to use today. It follows the language, repo size, and size distribution of pull requests, modeled after over 100 million real pull requests on GitHub. It uses a multi-source golden set and a consistent evaluation rubric and has been independently validated by senior engineers. Just as important, with the help of ReviewBench, our offline evaluation of Copilot code review (CCR) has become more effective at anticipating the direction of production experiments, giving us greater confidence that measured improvements reflect meaningful gains for users.
In this post, we’ll walk through how ReviewBench is constructed, how it establishes reliable ground truth and scoring, and how to onboard your own code review system and submit results.
Definitions of terms used in this blog post
- Benchmark: A standardized evaluation that tests code reviewers on a common set of pull requests using the same scoring methodology.
- Finding: A specific issue surfaced during code review.
- Golden set: A validated collection of known findings for each pull request, used as a reference for evaluating what a reviewer catches or misses.
- Precision: Of the issues a reviewer surfaces, the proportion that are valid. Higher precision generally means less noise.
- Recall: Of the known valid issues, the proportion the reviewer finds. Higher recall means broader coverage.
- F1 score: A single score that balances precision and recall equally.
- Fβ score: A variation of F1 that lets you put more weight on either precision or recall, depending on your review preference.
ReviewBench at a glance
1
What we built
A realistic, comprehensive benchmark for AI code review agents
103.9M
GitHub pull requests
Analyze distributions by language, repository size, and change shape.
Representative benchmark corpus
219 public pull requests across 19 languages, aligned to GitHub-wide distributions while preserving substantive review cases.
Multi-source golden set
- Human reviewers
- Frontier LLMs
- Static analysis
Structured findings
Every finding is labeled for severity and category, enabling user-tailored slices.
Severity
- Critical
- Medium
- Low
Category
- Correctness
- Security
- Reliability
- Maintainability
- Testing
- ......
Evaluation metrics
Four metrics measure both known and newly discovered issues.
- Grounded precision
- Grounded recall
- Augmented precision
- Augmented recall
Objective evaluation
Measure improvement and compare across agents objectively. Help users choose the reviewer that fits their needs the best.
2
How we keep it trustworthy
An auditable chain from rubric to expert validation and production checks
Published rubric
One explicit standard for all findings.
Human-labeled dev set
Senior engineers establish ground truth.
Calibrated grader
Aligned with human judgment.
Uniform labeling
Same standard across all sources.
Published agreement
Expert audit of benchmark quality.
Auditable end to end
96.6% agreement
Senior engineers independently labeled golden true-positives before release.
Offline signals that anticipate production
Benchmark movement is checked against online experiments.
- Improvements tend to show up online
- Regressions tend to show up online too
How ReviewBench works
Our benchmark is built around five principles:
1. Representative pull requests, not a demo set
We analyzed 103.9 million GitHub pull requests to characterize the real-world distribution of code review workloads. ReviewBench contains 219 pull requests from 187 public open source licensed repositories spanning 19 languages, with its language and repository-size distributions closely matching GitHub overall. The complete benchmark dataset is publicly available.
We make one deliberate adjustment to this distribution: while language and repository size mirror GitHub directly, pull request size is weighted toward the reviewable middle and tail. This reduces the overrepresentation of tiny, single-file changes while preserving more substantive, multi-file pull requests where review quality matters most.
Quick corpus snapshot: