{"version":1,"type":"story","url":"https://digestai.news/story/bioeval-benchmark-evaluates-llms-on-bioengineering-tasks","json":"https://digestai.news/story/bioeval-benchmark-evaluates-llms-on-bioengineering-tasks.json","markdown":"https://digestai.news/story/bioeval-benchmark-evaluates-llms-on-bioengineering-tasks.md","slug":"bioeval-benchmark-evaluates-llms-on-bioengineering-tasks","headline":"BioEVAL benchmark evaluates LLMs on bioengineering tasks","summary":"BioEVAL is a global, multi‑institutional benchmark created to test large language and multimodal models on bioengineering tasks. The benchmark contains 608 evaluation items from 22 research groups: 380 multiple‑choice questions (359 retained after audit), 218 literature‑synthesis tasks, and 10 multimodal problems that require experimental image interpretation.\n\nThe authors evaluated cloud‑scale foundation models such as ChatGPT, Gemini, and Grok, as well as locally deployable models that can run on consumer‑grade GPUs. The best model scored 90 % accuracy on the retained MCQs, achieved a similarity score of 0.72 on literature‑synthesis tasks, and reached 80 % accuracy on a small sample of multimodal reasoning questions. Performance varied widely across the 11 bioengineering subfields.\n\nLeaderboard rankings highlight current strengths and gaps, guiding future development. BioEVAL is designed to be extensible, with standardized protocols for ongoing expert item contribution and model evaluation.","keyPoints":["608 evaluation items from 22 groups, 359 MCQs retained","ChatGPT, Gemini, Grok reach up to 90% MCQ accuracy, 0.72 similarity, 80% multimodal","Leaderboard maps strengths and gaps across 11 bioengineering subfields"],"whyItMatters":"The benchmark provides a focused assessment of AI reasoning in bioengineering, a field where accurate experimental interpretation is critical. By exposing performance gaps across subfields, BioEVAL can steer model development toward practical biomedical applications and help researchers choose the right tools for experimental design and data analysis.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":["OpenAI","Google","Anthropic"],"models":["ChatGPT","Gemini","Grok"],"people":[]},"firstPublishedAt":"2026-09-28T04:00:00Z","updatedAt":"2026-09-28T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.AI","title":"BioEVAL: A global, multi-institutional benchmark of large language and multimodal models for bioengineering","url":"https://arxiv.org/abs/2609.30489","publishedAt":"2026-09-28T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"BioEVAL benchmark evaluates LLMs on bioengineering tasks\", 28 September 2026, https://digestai.news/story/bioeval-benchmark-evaluates-llms-on-bioengineering-tasks","publisher":"Digest AI","title":"BioEVAL benchmark evaluates LLMs on bioengineering tasks","datePublished":"2026-09-28T04:00:00Z","url":"https://digestai.news/story/bioeval-benchmark-evaluates-llms-on-bioengineering-tasks"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}