ReviewBench: An open benchmark for AI code review

Agentic code review is becoming an essential piece of how development happens. It helps you inspect pull requests, catch issues, and decide what deserves attention before code ships.

But the quality of existing AI reviewers can be hard to measure, and you need to know the strengths of a reviewer before you know if it will help you. Some reviewers surface more issues, some produce less noise, and some are stronger at catching critical problems while others surface smaller improvements, too. You may need code review to do different things within your workflow.

That makes it important to understand how reviewers actually compare: what different systems catch, what they miss, and the tradeoffs they make. A good code review benchmark should reflect the diversity of real pull requests, capture a broad set of review findings, and support meaningful breakdowns by severity, category, and precision-recall preferences. For teams building code review agents, the benchmark should also provide an offline signal that reliably tracks whether changes are likely to improve the experience in production. Existing benchmarks often make tradeoffs between label quality, coverage, and how well they represent real-world code review, leaving a gap for a rigorous and reproducible evaluation methodology that brings these pieces together.

We built ReviewBench, a new code review offline benchmark, to address that gap, and it is available for you to use today. It follows the language, repo size, and size distribution of pull requests, modeled after over 100 million real pull requests on GitHub. It uses a multi-source golden set and a consistent evaluation rubric and has been independently validated by senior engineers. Just as important, with the help of ReviewBench, our offline evaluation of Copilot code review (CCR) has become more effective at anticipating the direction of production experiments, giving us greater confidence that measured improvements reflect meaningful gains for users.

In this post, we’ll walk through how ReviewBench is constructed, how it establishes reliable ground truth and scoring, and how to onboard your own code review system and submit results.

Definitions of terms used in this blog post

  • Benchmark: A standardized evaluation that tests code reviewers on a common set of pull requests using the same scoring methodology.
  • Finding: A specific issue surfaced during code review.
  • Golden set: A validated collection of known findings for each pull request, used as a reference for evaluating what a reviewer catches or misses.
  • Precision: Of the issues a reviewer surfaces, the proportion that are valid. Higher precision generally means less noise.
  • Recall: Of the known valid issues, the proportion the reviewer finds. Higher recall means broader coverage.
  • F1 score: A single score that balances precision and recall equally.
  • Fβ score: A variation of F1 that lets you put more weight on either precision or recall, depending on your review preference.

ReviewBench at a glance

1

What we built

A realistic, comprehensive benchmark for AI code review agents

103.9M

GitHub pull requests

Analyze distributions by language, repository size, and change shape.

Representative benchmark corpus

219 public pull requests across 19 languages, aligned to GitHub-wide distributions while preserving substantive review cases.

Multi-source golden set

  • Human reviewers
  • Frontier LLMs
  • Static analysis

Structured findings

Every finding is labeled for severity and category, enabling user-tailored slices.

Severity

  • Critical
  • Medium
  • Low

Category

  • Correctness
  • Security
  • Reliability
  • Maintainability
  • Testing
  • ......

Evaluation metrics

Four metrics measure both known and newly discovered issues.

  • Grounded precision
  • Grounded recall
  • Augmented precision
  • Augmented recall

Objective evaluation

Measure improvement and compare across agents objectively. Help users choose the reviewer that fits their needs the best.

2

How we keep it trustworthy

An auditable chain from rubric to expert validation and production checks

Published rubric

One explicit standard for all findings.

Human-labeled dev set

Senior engineers establish ground truth.

Calibrated grader

Aligned with human judgment.

Uniform labeling

Same standard across all sources.

Published agreement

Expert audit of benchmark quality.

Auditable end to end

96.6% agreement

Senior engineers independently labeled golden true-positives before release.

Offline signals that anticipate production

Benchmark movement is checked against online experiments.

  • Improvements tend to show up online
  • Regressions tend to show up online too

How ReviewBench works

Our benchmark is built around five principles:

1. Representative pull requests, not a demo set

We analyzed 103.9 million GitHub pull requests to characterize the real-world distribution of code review workloads. ReviewBench contains 219 pull requests from 187 public open source licensed repositories spanning 19 languages, with its language and repository-size distributions closely matching GitHub overall. The complete benchmark dataset is publicly available.

We make one deliberate adjustment to this distribution: while language and repository size mirror GitHub directly, pull request size is weighted toward the reviewable middle and tail. This reduces the overrepresentation of tiny, single-file changes while preserving more substantive, multi-file pull requests where review quality matters most.

Quick corpus snapshot:

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论