{"version":1,"type":"story","url":"https://digestai.news/story/researchers-find-multi-agent-code-judge-often-declares-both-solutions","json":"https://digestai.news/story/researchers-find-multi-agent-code-judge-often-declares-both-solutions.json","markdown":"https://digestai.news/story/researchers-find-multi-agent-code-judge-often-declares-both-solutions.md","slug":"researchers-find-multi-agent-code-judge-often-declares-both-solutions","headline":"Researchers find Multi-Agent code judge often declares both solutions equally good","summary":"A study posted to arXiv examines how language models judge code correctness without access to ground truth. The researchers ran the MARCH framework on two code judging benchmarks and found it declared both solutions equally good in 78 to 95% of comparisons, achieving only 4.4% accuracy. By contrast, the same model asked directly reached 43.7% accuracy. The paper identifies two label-free measurements from the pipeline's logs that explain the issue: when gated on one measurement, the pipeline declines uncertain comparisons and raises accuracy from 20.7 to 36.9% while still answering half of all comparisons. The contribution is a method to detect when a judge lacks basis for its answer, not to improve accuracy.","keyPoints":["MARCH framework declared both solutions equally good in 78 to 95% of comparisons on two code judging benchmarks","Accuracy reached 4.4% with MARCH versus 43.7% when the same model was asked directly","Gating on one log-based measurement raised accuracy from 20.7 to 36.9% while answering half of all comparisons"],"whyItMatters":"The work highlights a fundamental flaw in AI-assisted code verification: models often guess rather than admit uncertainty. Offering a way to detect ungrounded judgments could improve reliability in automated code review and AI-assisted programming tools.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":[],"models":[],"people":[]},"firstPublishedAt":"2026-09-28T04:00:00Z","updatedAt":"2026-09-28T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.AI","title":"When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess","url":"https://arxiv.org/abs/2609.30328","publishedAt":"2026-09-28T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Researchers find Multi-Agent code judge often declares both solutions equally good\", 28 September 2026, https://digestai.news/story/researchers-find-multi-agent-code-judge-often-declares-both-solutions","publisher":"Digest AI","title":"Researchers find Multi-Agent code judge often declares both solutions equally good","datePublished":"2026-09-28T04:00:00Z","url":"https://digestai.news/story/researchers-find-multi-agent-code-judge-often-declares-both-solutions"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}