{"version":1,"type":"story","url":"https://digestai.news/story/comed-improves-multi-llm-inference-with-selective-collaboration","json":"https://digestai.news/story/comed-improves-multi-llm-inference-with-selective-collaboration.json","markdown":"https://digestai.news/story/comed-improves-multi-llm-inference-with-selective-collaboration.md","slug":"comed-improves-multi-llm-inference-with-selective-collaboration","headline":"COMED improves multi-LLM inference with selective collaboration","summary":"Researchers introduced COMED, a post-anchor controller designed to optimize multi-LLM inference by selectively invoking peer models. The system addresses the limitations of current approaches, where routing stops after initial selection and dense collaboration risks corrupting correct answers. COMED uses anchor self-consistency, router margin, and a lightweight peer probe to decide when to escalate queries to other models.\n\nThe method relies on a rescue-harm decomposition, improving performance only when rescued errors outweigh collaboration-induced harms. In tests across medical, scientific, and general reasoning benchmarks, COMED outperformed fixed and routed anchors in all 16 open-weight settings. It achieved gains of up to +10.7 percentage points on MedQA while using fewer decoded tokens than dense collaboration.\n\nOn the HLE benchmark with frontier models, COMED improved GPT-5.5 performance from 23.1% to 28.1%. This result surpassed dense collaboration, marking the best performance in the study. The approach aims to balance reliability and efficiency by avoiding unnecessary model invocations.","keyPoints":["COMED uses anchor self-consistency and router margin to selectively escalate queries to peer models.","The system improved GPT-5.5 on HLE from 23.1% to 28.1%, outperforming dense collaboration.","COMED achieved gains up to +10.7 percentage points on MedQA across 16 open-weight settings."],"whyItMatters":"COMED offers a more efficient way to combine multiple LLMs, reducing token usage while improving accuracy. This selective collaboration approach could lower inference costs and enhance reliability for complex reasoning tasks in production systems.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":[],"models":["GPT-5.5"],"people":[]},"firstPublishedAt":"2026-09-24T04:00:00Z","updatedAt":"2026-09-24T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.CL","title":"COMED: The Missing Middle Between Routing and Collaboration in Multi-LLM Inference","url":"https://arxiv.org/abs/2609.26913","publishedAt":"2026-09-24T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":{"title":"AI Agent Insight and Collaboration Saga","url":"https://digestai.news/thread/continual-search-framework-boosts-ai-agent-failure-attribution-accuracy-on","storyCount":2},"cite":{"text":"Digest AI, \"COMED improves multi-LLM inference with selective collaboration\", 24 September 2026, https://digestai.news/story/comed-improves-multi-llm-inference-with-selective-collaboration","publisher":"Digest AI","title":"COMED improves multi-LLM inference with selective collaboration","datePublished":"2026-09-24T04:00:00Z","url":"https://digestai.news/story/comed-improves-multi-llm-inference-with-selective-collaboration"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}