GitHub has launched ReviewBench, an open benchmark created with Microsoft for evaluating how well AI code review agents spot useful problems in pull requests. The company's own product, GitHub Copilot code review, sits first on the inaugural leaderboard with a 40.1% grounded F1 score, which blends precision and recall against the benchmark's reference findings. The new test joins a growing field of AI software engineering benchmarks, from SWE-bench for resolving GitHub issues to Terminal-Bench for terminal operations.

ReviewBench draws on 219 public pull requests from 187 repositories covering 19 programming languages, selected after examining 103.9 million pull requests. The corpus mirrors GitHub's overall language and repository size distribution, while deliberately weighting pull request size away from single-file changes. To set what reviewers should discover, the benchmark builds a reference set from human review comments, subsequent author changes, static analysis tools, and LLM reviewers. Claude Sonnet 5 categorizes findings, while another LLM matcher decides whether candidate findings match the same underlying issues in the reference set. The approach saw 47 findings initially marked as true positives manually corrected, with human and classifier judgments aligning 96.6% of the time on whether findings were true or false positives.

According to Alejandro Carderera de Diego, a staff applied engineer at GitHub, the benchmark has served as a useful early indicator for how Copilot Code Review changes will perform in production. "We've been using ReviewBench to improve GitHub Copilot Code Review, and its offline results have consistently anticipated the direction of later production experiments," he writes. Carderera de Diego and Michelle Zhou, an applied scientist at Microsoft, argue in a blog post that prior benchmarks "often make tradeoffs between label quality, coverage, and how well they represent real-world code review," leaving space for a more rigorous evaluation method. The team generated initial leaderboard entries by running publicly available versions of each product themselves, without vendor verification—and because testing occurred on different dates, some results are months old, with Copilot tested October 1 while Cubic and Greptile were tested in June.

The contrasting results across different benchmarks highlight how much design choices shape which tools appear to perform best. Martian's Code Review Bench, launched in February by an AI research company, combines offline testing with online tracking based on how developers respond to review comments across real open source repositories. As of October 6, Martian's online leaderboard places Cubic first at 64.9% F1 score, with GitHub Copilot ranking fourth at 60.9%, while its offline leaderboard puts Qodo Deep first and Copilot fifth at 58% F2 score—a variation that weights recall more heavily than precision. These scores aren't directly comparable with ReviewBench, which uses a different dataset and scoring method, but they illustrate the variation in vendor rankings depending on the test. GitHub publishes ReviewBench's dataset, methodology, and judging setup, and allows vendors to submit their own runs through a self-service system, with new results added as they're scored—though vendors will choose which configurations to enter. Whether enough vendors participate to transform the current GitHub-produced snapshot into a broader vendor-tested comparison will be the real measure of the benchmark's staying power. For organizations evaluating AI code review tools, the multiplicity of benchmarks suggests that no single leaderboard captures the full picture, and that real-world testing on a company's own codebase remains essential.