DigestAI news desk

Cut through the AI noise.

Research

COMED improves multi-LLM inference with selective collaboration

Researchers introduced COMED, a post-anchor controller designed to optimize multi-LLM inference by selectively invoking peer models. The system addresses the limitations of current approaches, where routing stops after initial selection and dense collaboration risks corrupting correct answers. COMED uses anchor self-consistency, router margin, and a lightweight peer probe to decide when to…

1 source primary source

Key points

  • COMED uses anchor self-consistency and router margin to selectively escalate queries to peer models.
  • The system improved GPT-5.5 on HLE from 23.1% to 28.1%, outperforming dense collaboration.
  • COMED achieved gains up to +10.7 percentage points on MedQA across 16 open-weight settings.

The method relies on a rescue-harm decomposition, improving performance only when rescued errors outweigh collaboration-induced harms. In tests across medical, scientific, and general reasoning benchmarks, COMED outperformed fixed and routed anchors in all 16 open-weight settings. It achieved gains of up to +10.7 percentage points on MedQA while using fewer decoded tokens than dense collaboration.

On the HLE benchmark with frontier models, COMED improved GPT-5.5 performance from 23.1% to 28.1%. This result surpassed dense collaboration, marking the best performance in the study. The approach aims to balance reliability and efficiency by avoiding unnecessary model invocations.

The story so far

2 episodes →
  1. COMED improves multi-LLM inference with selective collaborationthis story
Read the original at arXiv cs.CL · by Norah Alballa, Wenxuan Zhang, Salma Kharrat, Fares Fourati, Zafar Ayyub Qazi, Mohamed Elhoseiny, Marco Canini primary sourceOpen source ↗
Topics · follow one to build your own front page

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories