arXiv study finds reading LLM judges from first token overstates position bias
A new paper on arXiv examines how reading an LLM judge’s verdict from its first generated token distorts results. Researchers found this method overstates position bias in every tested condition, as judges do not always start with a verdict token. For three Qwen3 judges, this happens in 12% to 49% of cases, while Llama-3.1-8B and Phi-3.5-mini show under 3% non-compliance. When forced to read…
Key points
- Reading LLM judges from the first token overstates position bias in all tested conditions, per arXiv study
- Three Qwen3 judges fail to start with a verdict token in 12% to 49% of cases, while Llama-3.1-8B and Phi-3.5-mini show under 3%
- Forced early reading flips 89.7% of undecided pairs when responses are swapped, compared to 47.5% after full generation
The study highlights two key failures: first, the method misrepresents position bias by 42 points while barely affecting judge accuracy. Second, even when judges do start with a verdict, they sometimes begin with a partial letter before reasoning to the final judgment. The authors recommend reporting the rate at which judges lead with a verdict token alongside position-bias figures, as this requires only one forward pass and no labeled data.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- A Holistic Assessment of the Carbon Footprint of Noor, a Very Large Arabic Language Model · 1 src
- Korean legal study finds KLUE-BERT outperforms GPT models in sexual offense text classification · 1 src
- Researchers question human-derived bias measures for LLM evaluation · 1 src
- Researchers introduce Context language models that manage their own Context · 3 src
- Researchers release Build2SPARQL benchmark for text-to-SPARQL in building KGs · 1 src
Comments
via GitHub Discussions