DigestAI news desk

Cut through the AI noise.

Research

arXiv study finds reading LLM judges from first token overstates position bias

A new paper on arXiv examines how reading an LLM judge’s verdict from its first generated token distorts results. Researchers found this method overstates position bias in every tested condition, as judges do not always start with a verdict token. For three Qwen3 judges, this happens in 12% to 49% of cases, while Llama-3.1-8B and Phi-3.5-mini show under 3% non-compliance. When forced to read…

1 source primary source

Key points

  • Reading LLM judges from the first token overstates position bias in all tested conditions, per arXiv study
  • Three Qwen3 judges fail to start with a verdict token in 12% to 49% of cases, while Llama-3.1-8B and Phi-3.5-mini show under 3%
  • Forced early reading flips 89.7% of undecided pairs when responses are swapped, compared to 47.5% after full generation

The study highlights two key failures: first, the method misrepresents position bias by 42 points while barely affecting judge accuracy. Second, even when judges do start with a verdict, they sometimes begin with a partial letter before reasoning to the final judgment. The authors recommend reporting the rate at which judges lead with a verdict token alongside position-bias figures, as this requires only one forward pass and no labeled data.

Read the original at arXiv cs.CL · by Gnaneswar Villuri, Hashmath Shaik, Alex Doboli primary sourceOpen source ↗
Topics · follow one to build your own front page

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories