{"version":1,"type":"story","url":"https://digestai.news/story/study-finds-llm-judge-consensus-overstates-evidence-because-judges-sha","json":"https://digestai.news/story/study-finds-llm-judge-consensus-overstates-evidence-because-judges-sha.json","markdown":"https://digestai.news/story/study-finds-llm-judge-consensus-overstates-evidence-because-judges-sha.md","slug":"study-finds-llm-judge-consensus-overstates-evidence-because-judges-sha","headline":"Study finds LLM judge consensus overstates evidence because judges share errors","summary":"Researchers examined how error dependence among large language model (LLM) judges affects the reliability of consensus judgments. Analyzing a bank of ten judges, they measured an average pairwise error correlation of 0.21, meaning the judges’ mistakes are not independent. This correlation reduces the effective information of ten judges to roughly that of 3.5 independent judges.\n\nThe authors show that, for high‑accuracy frontier judges from different providers, the dependency is even stronger. In up to 28% of their comparisons, ignoring shared errors would suggest one system is significantly better, while accounting for the correlation removes that claim. They also find that the pattern of shared errors—whether widespread or concentrated—impacts which voting method works best. The paper recommends using a small set of trusted examples to estimate judge accuracy, identify shared mistakes, and select an appropriate voting scheme before applying it to new data.","keyPoints":["Average pairwise error correlation among ten LLM judges is 0.21.","Ten judges provide statistical information equivalent to about 3.5 independent judges.","Ignoring shared errors can falsely indicate significance in up to 28% of comparisons."],"whyItMatters":"The findings reveal that common evaluation practices may overstate model performance, prompting more careful design of LLM benchmarking.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":[],"models":[],"people":[]},"firstPublishedAt":"2026-09-22T04:00:00Z","updatedAt":"2026-09-22T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.AI","title":"Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus","url":"https://arxiv.org/abs/2609.22512","publishedAt":"2026-09-22T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Study finds LLM judge consensus overstates evidence because judges share errors\", 22 September 2026, https://digestai.news/story/study-finds-llm-judge-consensus-overstates-evidence-because-judges-sha","publisher":"Digest AI","title":"Study finds LLM judge consensus overstates evidence because judges share errors","datePublished":"2026-09-22T04:00:00Z","url":"https://digestai.news/story/study-finds-llm-judge-consensus-overstates-evidence-because-judges-sha"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}