{"version":1,"type":"story","url":"https://digestai.news/story/aibuildai-2-5-ranks-first-on-mle-bench-with-73-3-medal-rate","json":"https://digestai.news/story/aibuildai-2-5-ranks-first-on-mle-bench-with-73-3-medal-rate.json","markdown":"https://digestai.news/story/aibuildai-2-5-ranks-first-on-mle-bench-with-73-3-medal-rate.md","slug":"aibuildai-2-5-ranks-first-on-mle-bench-with-73-3-medal-rate","headline":"AIBuildAI-2.5 ranks first on MLE-Bench with 73.3% medal rate","summary":"AIBuildAI-2.5 is an autonomous agentic system designed to build AI models by treating model construction as a tree‑search problem. The authors identify three efficiency gaps in prior agents—limited candidate execution, lack of resource‑aware scheduling, and reliance on a single powerful LLM for all tasks—and propose solutions.\n\nThe new system adds an LLM‑guided tree search where a judge scores each candidate on expected improvement, grounding, and feasibility, and a selector ranks the pool using those scores and the current search state. A scheduler launches training jobs based on real‑time hardware availability, while a router assigns cheaper LLMs to low‑complexity steps, reserving the strongest model for the hardest sub‑tasks. In benchmark evaluation, AIBuildAI-2.5 achieved a 73.3% medal rate, placing first on MLE‑Bench, and outperformed a strong baseline across six autonomous AI research tasks from AIRS‑Bench.","keyPoints":["AIBuildAI-2.5 introduces LLM‑guided tree search that scores candidates on expected improvement, grounding, and feasibility","The system adds a resource‑aware scheduler and a router that delegates low‑cost LLMs to simple tasks","It achieved a 73.3% medal rate, ranking first on MLE‑Bench and beating a strong baseline on six AIRS‑Bench tasks"],"whyItMatters":"Improving the efficiency of autonomous AI research agents can lower compute costs and broaden access to AI model development for scientific and engineering teams.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":[],"models":["AIBuildAI-2.5"],"people":[]},"firstPublishedAt":"2026-09-23T04:00:00Z","updatedAt":"2026-09-23T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.CL","title":"AIBuildAI-2.5: Efficient Autonomous AI Model Development Through LLM-Guided Tree Search","url":"https://arxiv.org/abs/2609.25047","publishedAt":"2026-09-23T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"AIBuildAI-2.5 ranks first on MLE-Bench with 73.3% medal rate\", 23 September 2026, https://digestai.news/story/aibuildai-2-5-ranks-first-on-mle-bench-with-73-3-medal-rate","publisher":"Digest AI","title":"AIBuildAI-2.5 ranks first on MLE-Bench with 73.3% medal rate","datePublished":"2026-09-23T04:00:00Z","url":"https://digestai.news/story/aibuildai-2-5-ranks-first-on-mle-bench-with-73-3-medal-rate"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}