DigestAI news desk

Cut through the AI noise.

Research

Study finds LLM judge consensus overstates evidence because judges share errors

Researchers examined how error dependence among large language model (LLM) judges affects the reliability of consensus judgments. Analyzing a bank of ten judges, they measured an average pairwise error correlation of 0.21, meaning the judges’ mistakes are not independent. This correlation reduces the effective information of ten judges to roughly that of 3.5 independent judges.

1 source primary source

Key points

  • Average pairwise error correlation among ten LLM judges is 0.21.
  • Ten judges provide statistical information equivalent to about 3.5 independent judges.
  • Ignoring shared errors can falsely indicate significance in up to 28% of comparisons.

The authors show that, for high‑accuracy frontier judges from different providers, the dependency is even stronger. In up to 28% of their comparisons, ignoring shared errors would suggest one system is significantly better, while accounting for the correlation removes that claim. They also find that the pattern of shared errors—whether widespread or concentrated—impacts which voting method works best. The paper recommends using a small set of trusted examples to estimate judge accuracy, identify shared mistakes, and select an appropriate voting scheme before applying it to new data.

Read the original at arXiv cs.AI · by Elias Hossain, Niloofar Yousefi, Ser-Nam Lim primary sourceOpen source ↗

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories