DigestAI news desk

Cut through the AI noise.

Research10 min read

GitHub releases ReviewBench to evaluate AI code reviews across 219 pull requests

GitHub released ReviewBench as a research preview on October 5, 2026, creating an evaluation framework for AI code review tools. Developed by teams from GitHub and Microsoft, the benchmark tests agents on 219 pull requests across 187 open-source repositories and 19 programming languages, derived from an initial analysis of 103.9 million pull requests. The dataset emphasizes medium and large pull…

1 source

Key points

  • GitHub launched ReviewBench to test AI code reviews across 219 pull requests, 187 repositories, and 19 languages.
  • The evaluation utilizes Claude Sonnet 3.5 as a judge model to track precision, recall, and false positives.
  • A Copilot production test showed an 8.00% response rate increase, 13.58% higher recall, and an 8.00% drop in review costs.

The benchmark measures performance using dual metrics: grounded evaluations against a fixed reference dataset and augmented evaluations that assess newly discovered issues. Claude Sonnet 3.5 serves as the automated judgment model to match and verify feedback. An internal audit of the ground truth dataset by senior engineers matched ReviewBench judgments at 96.6%, though the materials represent benchmark specifications rather than academic peer-reviewed research.

GitHub also shared a production case study evaluating a multi-model configuration against Copilot code review's single-review setup. Offline testing predicted increases of 4.45% in precision and 12.88% in recall, alongside a 23.7% cost reduction. In production A/B testing, the addressed response rate rose by 8.00%, recall rose by 13.58%, and costs dropped by 8.00%. However, source figures conflicted regarding total comment volume, noting a 25.0% increase in a data table and 61% in body text.

The story so far

6 episodes →
  1. GitHub releases ReviewBench to evaluate AI code reviews across 219 pull requeststhis story
Full story from note.com · by 野崎秀吾 · via Search: GitHubOpen source ↗

GitHub Releases 'ReviewBench': Thinking About AI Code Review Misses and False Positives Through 219 PRs

note.com · 5 October 2026

Loading the full article…

This text was published by note.com and written by 野崎秀吾. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗

Topics · follow one to build your own front page
GitHubMicrosoftClaude Sonnet 3.5CopilotMichelle ZhouAlejandro Carderera de Diego

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories