GitHub and Microsoft released ReviewBench, an open benchmark for grading how well AI agents catch bugs, security issues and other problems in pull requests, according to the GitHub Blog. The benchmark draws on analysis of 103.9 million GitHub pull requests to build a set of 219 real pull requests spanning 19 programming languages, the post said.

Findings are gathered from human reviewers, large language models and static analysis tools, then deduplicated and checked against a shared rubric using Claude Sonnet 5 as the grading model, according to the post. Senior engineers who independently labeled the same findings agreed with ReviewBench's own judgments 96.6% of the time, GitHub said.

GitHub said it used ReviewBench to test a multi-model ensemble approach for its Copilot Code Review product before shipping it, and the benchmark's predicted improvements lined up with what the company then measured in production: an 8% rise in the rate reviewers addressed flagged issues, a 13.6% rise in recall, and a 227% jump in critical comments surfaced, compared with a baseline, according to the post.

Most code review benchmarks score a model's output in isolation. Building one from real pull requests and then checking it against a shipped product's actual production numbers is a higher bar, and a useful one for any team trying to decide if an AI reviewer is worth turning on before, not after, it ships to their own repository.