DigestAI news desk

Cut through the AI noise.

Research

Researchers find Multi-Agent code judge often declares both solutions equally good

A study posted to arXiv examines how language models judge code correctness without access to ground truth. The researchers ran the MARCH framework on two code judging benchmarks and found it declared both solutions equally good in 78 to 95% of comparisons, achieving only 4.4% accuracy. By contrast, the same model asked directly reached 43.7% accuracy. The paper identifies two label-free…

1 source primary source

Key points

  • MARCH framework declared both solutions equally good in 78 to 95% of comparisons on two code judging benchmarks
  • Accuracy reached 4.4% with MARCH versus 43.7% when the same model was asked directly
  • Gating on one log-based measurement raised accuracy from 20.7 to 36.9% while answering half of all comparisons
Read the original at arXiv cs.AI · by Salma Roshdy Aly, Hussein Assaf, Ziad Kobti primary sourceOpen source ↗

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories