GitHub发布AI代码审查基准ReviewBench
GitHub推出开放基准ReviewBench,用于评估AI代码审查智能体。该基准涵盖来自187个公共代码仓库、19种语言的219个拉取请求,并采用多来源金标准,支持按精确率、召回率、严重性和类别分析。GitHub称,ReviewBench的离线结果与GitHub Copilot代码审查的生产实验趋势相符。
为什么重要:它为团队比较AI代码审查工具的覆盖率与误报水平提供了更可复现的方法。
Agentic code review is becoming an essential piece of how development happens. It helps you inspect pull requests, catch issues, and decide what deserves attention before code ships.
But the quality of existing AI reviewers can be hard to measure, and you need to know the strengths of a reviewer before you know if it will help you. Some reviewers surface more issues, some produce less noise, and some are stronger at catching critical problems while others surface smaller improvements, too. You may need code review to do different things within your workflow.
That makes it important to understand how reviewers actually compare: what different systems catch, what they miss, and the tradeoffs they make. A good code review benchmark should reflect the diversity of real pull requests, capture a broad set of review findings, and support meaningful breakdowns by severity, category, and precision-recall preferences. For teams building code review agents, the benchmark should also provide an offline signal that reliably tracks whether changes are likely to improve the experience in production. Existing benchmarks often make tradeoffs between label quality, coverage, and how well they represent real-world code review, leaving a gap for a rigorous and reproducible evaluation methodology that brings these pieces together.
We built ReviewBench, a new code review offline benchmark, to address that gap, and it is available for you to use today. It follows the language, repo size, and size distribution of pull requests, modeled after over 100 million real pull requests on GitHub. It uses a multi-source golden set and a consistent evaluation rubric and has been independently validated by senior engineers. Just as important, with the help of ReviewBench, our offline evaluation of Copilot code review (CCR) has become more effective at anticipating the direction of production experiments, giving us greater confidence that measured improvements reflect meaningful gains for users.
In this post, we’ll walk through how ReviewBench is constructed, how it establishes reliable ground truth and scoring, and how to onboard your own code review system and submit results.
ReviewBench at a glance
1
What we built
A realistic, comprehensive benchmark for AI code review agents
103.9M
GitHub pull requests
Analyze distributions by language, repository size, and change shape.
Representative benchmark corpus
219 public pull requests across 19 languages, aligned to GitHub-wide distributions while preserving substantive review cases.
Multi-source golden set
- Human reviewers
- Frontier LLMs
- Static analysis
Structured findings
Every finding is labeled for severity and category, enabling user-tailored slices.
Severity
- Critical
- Medium
- Low
Category
- Correctness
- Security
- Reliability
- Maintainability
- Testing
- ......
Evaluation metrics
Four metrics measure both known and newly discovered issues.
- Grounded precision
- Grounded recall
- Augmented precision
- Augmented recall
Objective evaluation
Measure improvement and compare across agents objectively. Help users choose the reviewer that fits their needs the best.
2
How we keep it trustworthy
An auditable chain from rubric to expert validation and production checks
Published rubric
One explicit standard for all findings.
Human-labeled dev set
Senior engineers establish ground truth.
Calibrated grader
Aligned with human judgment.
Uniform labeling
Same standard across all sources.
Published agreement
Expert audit of benchmark quality.
Auditable end to end
96.6% agreement
Senior engineers independently labeled golden true-positives before release.
Offline signals that anticipate production
Benchmark movement is checked against online experiments.
- Improvements tend to show up online
- Regressions tend to show up online too
How ReviewBench works
Our benchmark is built around five principles:
1. Representative pull requests, not a demo set
We analyzed 103.9 million GitHub pull requests to characterize the real-world distribution of code review workloads. ReviewBench contains 219 pull requests from 187 public open source licensed repositories spanning 19 languages, with its language and repository-size distributions closely matching GitHub overall. The complete benchmark dataset is publicly available.
We make one deliberate adjustment to this distribution: while language and repository size mirror GitHub directly, pull request size is weighted toward the reviewable middle and tail. This reduces the overrepresentation of tiny, single-file changes while preserving more substantive, multi-file pull requests where review quality matters most.
Quick corpus snapshot: