GitHub has released ReviewBench, an open benchmark designed to measure how well AI tools actually review code. The project addresses a specific gap: existing tests often fail to reflect the diversity of real pull requests or provide a reliable way to track improvements before they reach production.
Teams need to know if a reviewer catches critical errors or simply adds noise. ReviewBench uses data modeled after over 100 million real pull requests on GitHub. It includes 219 public pull requests across 19 languages, preserving the actual distribution of repository sizes and code changes. The benchmark is validated by senior engineers and uses a multi-source golden set. This set combines findings from human reviewers, frontier large language models, and static analysis tools.
How it works
The benchmark evaluates systems using a consistent rubric. It scores results on precision, recall, and F1 metrics. Precision measures how many of the flagged issues are valid, while recall measures how many known valid issues were found. The Fβ score allows teams to weight these metrics differently based on their specific needs.
“With the help of ReviewBench, our offline evaluation of Copilot code review has become more effective at anticipating the direction of production experiments.”
This approach gives developers greater confidence that measured improvements reflect meaningful gains for users rather than random fluctuations.
Developers can now onboard their own code review systems and submit results to the benchmark. The project provides a standardized way to test agents on a common set of pull requests using the same scoring methodology.
What it means
For engineering teams, this tool offers a way to compare different code review agents without needing to run live experiments. It provides an offline signal that tracks whether changes are likely to improve the experience in production. This helps teams decide which tools deserve attention before code ships.




