BioEVAL benchmark evaluates LLMs on bioengineering tasks
BioEVAL is a global, multi‑institutional benchmark created to test large language and multimodal models on bioengineering tasks. The benchmark contains 608 evaluation items from 22 research groups: 380 multiple‑choice questions (359 retained after audit), 218 literature‑synthesis tasks, and 10 multimodal problems that require experimental image interpretation.
Key points
- 608 evaluation items from 22 groups, 359 MCQs retained
- ChatGPT, Gemini, Grok reach up to 90% MCQ accuracy, 0.72 similarity, 80% multimodal
- Leaderboard maps strengths and gaps across 11 bioengineering subfields
The authors evaluated cloud‑scale foundation models such as ChatGPT, Gemini, and Grok, as well as locally deployable models that can run on consumer‑grade GPUs. The best model scored 90 % accuracy on the retained MCQs, achieved a similarity score of 0.72 on literature‑synthesis tasks, and reached 80 % accuracy on a small sample of multimodal reasoning questions. Performance varied widely across the 11 bioengineering subfields.
Leaderboard rankings highlight current strengths and gaps, guiding future development. BioEVAL is designed to be extensible, with standardized protocols for ongoing expert item contribution and model evaluation.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Researchers release benchmark for AI in systematic review screening · 1 src
- SlideLab framework generates scientific presentations from research papers · 1 src
- Cartograph reduces AI agent tool discovery from O(n) to O(k) · 1 src
- Survey reviews 211 fake review detection studies from 2018 to 2026 · 1 src
- Researchers audit LLM-as-judge in text-to-SQL pipeline, find low agreement · 1 src
Comments
via GitHub Discussions