{"version":1,"type":"story","url":"https://digestai.news/story/github-releases-reviewbench-to-evaluate-ai-code-reviews-across-219-pul","json":"https://digestai.news/story/github-releases-reviewbench-to-evaluate-ai-code-reviews-across-219-pul.json","markdown":"https://digestai.news/story/github-releases-reviewbench-to-evaluate-ai-code-reviews-across-219-pul.md","slug":"github-releases-reviewbench-to-evaluate-ai-code-reviews-across-219-pul","headline":"GitHub releases ReviewBench to evaluate AI code reviews across 219 pull requests","summary":"GitHub released ReviewBench as a research preview on October 5, 2026, creating an evaluation framework for AI code review tools. Developed by teams from GitHub and Microsoft, the benchmark tests agents on 219 pull requests across 187 open-source repositories and 19 programming languages, derived from an initial analysis of 103.9 million pull requests. The dataset emphasizes medium and large pull requests, including 78 changes that exceed 1,000 lines.\n\nThe benchmark measures performance using dual metrics: grounded evaluations against a fixed reference dataset and augmented evaluations that assess newly discovered issues. Claude Sonnet 3.5 serves as the automated judgment model to match and verify feedback. An internal audit of the ground truth dataset by senior engineers matched ReviewBench judgments at 96.6%, though the materials represent benchmark specifications rather than academic peer-reviewed research.\n\nGitHub also shared a production case study evaluating a multi-model configuration against Copilot code review's single-review setup. Offline testing predicted increases of 4.45% in precision and 12.88% in recall, alongside a 23.7% cost reduction. In production A/B testing, the addressed response rate rose by 8.00%, recall rose by 13.58%, and costs dropped by 8.00%. However, source figures conflicted regarding total comment volume, noting a 25.0% increase in a data table and 61% in body text.","keyPoints":["GitHub launched ReviewBench to test AI code reviews across 219 pull requests, 187 repositories, and 19 languages.","The evaluation utilizes Claude Sonnet 3.5 as a judge model to track precision, recall, and false positives.","A Copilot production test showed an 8.00% response rate increase, 13.58% higher recall, and an 8.00% drop in review costs."],"whyItMatters":"ReviewBench provides an open framework for measuring AI code review accuracy, helping software engineering teams evaluate false positives against critical misses rather than relying on generic leaderboard scores.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":["GitHub","Microsoft"],"models":["Claude Sonnet 3.5","Copilot"],"people":["Michelle Zhou","Alejandro Carderera de Diego"]},"firstPublishedAt":"2026-10-05T20:40:00Z","updatedAt":"2026-10-05T22:21:00Z","sourceCount":2,"hasPrimarySource":false,"sources":[{"outlet":"note.com","title":"GitHub Releases 'ReviewBench': Thinking About AI Code Review Misses and False Positives Through 219 PRs","url":"https://note.com/shugo/n/n527b5ed7e5bc?hl=en","publishedAt":"2026-10-05T20:40:00Z","type":"press","primary":false,"lead":true},{"outlet":"note.com","title":"For those who get stuck when AI gives too much feedback: How to distinguish 'what needs fixing'","url":"https://note.com/rock_s_1112/n/n0a50b58d46b3?hl=en","publishedAt":"2026-10-05T22:21:00Z","type":"press","primary":false,"lead":false}],"sourceNotes":null,"discussions":[],"thread":{"title":"Navigating the New Era of Autonomous AI Agents","url":"https://digestai.news/thread/google-launches-model-context-protocol-for-ai-agents-to-control-home-devices","storyCount":6},"cite":{"text":"Digest AI, \"GitHub releases ReviewBench to evaluate AI code reviews across 219 pull requests\", 5 October 2026, https://digestai.news/story/github-releases-reviewbench-to-evaluate-ai-code-reviews-across-219-pul","publisher":"Digest AI","title":"GitHub releases ReviewBench to evaluate AI code reviews across 219 pull requests","datePublished":"2026-10-05T20:40:00Z","url":"https://digestai.news/story/github-releases-reviewbench-to-evaluate-ai-code-reviews-across-219-pul"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}