Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

New here?

Story

GitHub launches ReviewBench benchmark for AI code review

AI summary

GitHub introduced ReviewBench, an open benchmark for evaluating AI code review agents. The benchmark covers 219 pull requests from 187 public repositories across 19 languages, uses a multi-source golden set, and supports precision, recall, severity, and category-based analysis. GitHub says its offline results have shown alignment with production experiments for GitHub Copilot code review.

Why it matters: It gives teams and developers a more reproducible way to compare the coverage and noise levels of AI code reviewers.

GitHubGitHub CopilotReviewBench

0
Source textGitHub Blog · AI & ML · 11 min read

Agentic code review is becoming an essential piece of how development happens. It helps you inspect pull requests, catch issues, and decide what deserves attention before code ships.

But the quality of existing AI reviewers can be hard to measure, and you need to know the strengths of a reviewer before you know if it will help you. Some reviewers surface more issues, some produce less noise, and some are stronger at catching critical problems while others surface smaller improvements, too. You may need code review to do different things within your workflow.

That makes it important to understand how reviewers actually compare: what different systems catch, what they miss, and the tradeoffs they make. A good code review benchmark should reflect the diversity of real pull requests, capture a broad set of review findings, and support meaningful breakdowns by severity, category, and precision-recall preferences. For teams building code review agents, the benchmark should also provide an offline signal that reliably tracks whether changes are likely to improve the experience in production. Existing benchmarks often make tradeoffs between label quality, coverage, and how well they represent real-world code review, leaving a gap for a rigorous and reproducible evaluation methodology that brings these pieces together.

We built ReviewBench, a new code review offline benchmark, to address that gap, and it is available for you to use today. It follows the language, repo size, and size distribution of pull requests, modeled after over 100 million real pull requests on GitHub. It uses a multi-source golden set and a consistent evaluation rubric and has been independently validated by senior engineers. Just as important, with the help of ReviewBench, our offline evaluation of Copilot code review (CCR) has become more effective at anticipating the direction of production experiments, giving us greater confidence that measured improvements reflect meaningful gains for users.

In this post, we’ll walk through how ReviewBench is constructed, how it establishes reliable ground truth and scoring, and how to onboard your own code review system and submit results.

ReviewBench at a glance

1

What we built

A realistic, comprehensive benchmark for AI code review agents

103.9M

GitHub pull requests

Analyze distributions by language, repository size, and change shape.

Representative benchmark corpus

219 public pull requests across 19 languages, aligned to GitHub-wide distributions while preserving substantive review cases.

Multi-source golden set

  • Human reviewers
  • Frontier LLMs
  • Static analysis

Structured findings

Every finding is labeled for severity and category, enabling user-tailored slices.

Severity

  • Critical
  • Medium
  • Low

Category

  • Correctness
  • Security
  • Reliability
  • Maintainability
  • Testing
  • ......

Evaluation metrics

Four metrics measure both known and newly discovered issues.

  • Grounded precision
  • Grounded recall
  • Augmented precision
  • Augmented recall

Objective evaluation

Measure improvement and compare across agents objectively. Help users choose the reviewer that fits their needs the best.

2

How we keep it trustworthy

An auditable chain from rubric to expert validation and production checks

Published rubric

One explicit standard for all findings.

Human-labeled dev set

Senior engineers establish ground truth.

Calibrated grader

Aligned with human judgment.

Uniform labeling

Same standard across all sources.

Published agreement

Expert audit of benchmark quality.

Auditable end to end

96.6% agreement

Senior engineers independently labeled golden true-positives before release.

Offline signals that anticipate production

Benchmark movement is checked against online experiments.

  • Improvements tend to show up online
  • Regressions tend to show up online too

How ReviewBench works

Our benchmark is built around five principles:

1. Representative pull requests, not a demo set

We analyzed 103.9 million GitHub pull requests to characterize the real-world distribution of code review workloads. ReviewBench contains 219 pull requests from 187 public open source licensed repositories spanning 19 languages, with its language and repository-size distributions closely matching GitHub overall. The complete benchmark dataset is publicly available.

We make one deliberate adjustment to this distribution: while language and repository size mirror GitHub directly, pull request size is weighted toward the reviewable middle and tail. This reduces the overrepresentation of tiny, single-file changes while preserving more substantive, multi-file pull requests where review quality matters most.

Quick corpus snapshot:

Read the original →

Comments

I've used this: share my experience What I think: share my view
How important is this story?No ratings yet

No comments yet. Start the conversation.