THE CRUNCH

GitHub has released ReviewBench, an open offline benchmark for judging AI code review agents, available to use today. It is modelled on the distributions of more than 100 million real GitHub pull requests, covering language, repository size and change shape, and aims to measure what different reviewers catch, what they miss and the tradeoffs between noise and coverage. The motivation is that existing benchmarks trade off label quality, coverage and real-world representativeness, leaving no rigorous, reproducible way to compare reviewers.

The benchmark's corpus consists of 219 public pull requests across 19 languages, chosen to match GitHub-wide distributions for language, repository size and change shape while keeping cases with substantive review content. Every finding in its golden set, drawn from human reviewers, frontier LLMs and static analysis tools, is labelled with a severity (critical, medium or low) and a category such as correctness, security, reliability, maintainability or testing, so results can be sliced to match a team's priorities.

Trustworthiness rests on an auditable chain: a published rubric, a human-labelled development set, a calibrated grader and uniform labelling across sources. Before release, senior engineers independently labelled the golden set's true positives, with the published audit reporting 96.6 per cent agreement on those labels. GitHub also reports that benchmark movement has been checked against online experiments, with improvements and regressions tending to show up in production.

For teams building review agents, GitHub frames ReviewBench as an offline signal for whether changes are likely to improve the production experience, and the post explains how to onboard a system and submit results.

WHAT HAPPENS NEXT

GitHub's post walks through how to onboard your own code review system and submit results to the benchmark.