Researchers find Multi-Agent code judge often declares both solutions equally good
A study posted to arXiv examines how language models judge code correctness without access to ground truth. The researchers ran the MARCH framework on two code judging benchmarks and found it declared both solutions equally good in 78 to 95% of comparisons, achieving only 4.4% accuracy. By contrast, the same model asked directly reached 43.7% accuracy. The paper identifies two label-free…
Key points
- MARCH framework declared both solutions equally good in 78 to 95% of comparisons on two code judging benchmarks
- Accuracy reached 4.4% with MARCH versus 43.7% when the same model was asked directly
- Gating on one log-based measurement raised accuracy from 20.7 to 36.9% while answering half of all comparisons
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Researchers release benchmark for AI in systematic review screening · 1 src
- SlideLab framework generates scientific presentations from research papers · 1 src
- Cartograph reduces AI agent tool discovery from O(n) to O(k) · 1 src
- Survey reviews 211 fake review detection studies from 2018 to 2026 · 1 src
- Researchers audit LLM-as-judge in text-to-SQL pipeline, find low agreement · 1 src
Comments
via GitHub Discussions