Study finds LLM judge consensus overstates evidence because judges share errors
Researchers examined how error dependence among large language model (LLM) judges affects the reliability of consensus judgments. Analyzing a bank of ten judges, they measured an average pairwise error correlation of 0.21, meaning the judges’ mistakes are not independent. This correlation reduces the effective information of ten judges to roughly that of 3.5 independent judges.
Key points
- Average pairwise error correlation among ten LLM judges is 0.21.
- Ten judges provide statistical information equivalent to about 3.5 independent judges.
- Ignoring shared errors can falsely indicate significance in up to 28% of comparisons.
The authors show that, for high‑accuracy frontier judges from different providers, the dependency is even stronger. In up to 28% of their comparisons, ignoring shared errors would suggest one system is significantly better, while accounting for the correlation removes that claim. They also find that the pattern of shared errors—whether widespread or concentrated—impacts which voting method works best. The paper recommends using a small set of trusted examples to estimate judge accuracy, identify shared mistakes, and select an appropriate voting scheme before applying it to new data.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Researchers propose multi-split boundary decision to lower LLM document segmentation cost · 1 src
- GaitVista reduces gait measurement error by 27.7% in lab tests · 1 src
- Megagon Labs releases mawile workbench for auditing LLM judges · 1 src
- Researchers propose Goal-driven variant categorization using LLM · 1 src
- Study shows AI agents select fewer papers when they see others' choices · 1 src
Comments
via GitHub Discussions