DigestAI news desk

Cut through the AI noise.

Research

BioEVAL benchmark evaluates LLMs on bioengineering tasks

BioEVAL is a global, multi‑institutional benchmark created to test large language and multimodal models on bioengineering tasks. The benchmark contains 608 evaluation items from 22 research groups: 380 multiple‑choice questions (359 retained after audit), 218 literature‑synthesis tasks, and 10 multimodal problems that require experimental image interpretation.

1 source primary source

Key points

  • 608 evaluation items from 22 groups, 359 MCQs retained
  • ChatGPT, Gemini, Grok reach up to 90% MCQ accuracy, 0.72 similarity, 80% multimodal
  • Leaderboard maps strengths and gaps across 11 bioengineering subfields

The authors evaluated cloud‑scale foundation models such as ChatGPT, Gemini, and Grok, as well as locally deployable models that can run on consumer‑grade GPUs. The best model scored 90 % accuracy on the retained MCQs, achieved a similarity score of 0.72 on literature‑synthesis tasks, and reached 80 % accuracy on a small sample of multimodal reasoning questions. Performance varied widely across the 11 bioengineering subfields.

Leaderboard rankings highlight current strengths and gaps, guiding future development. BioEVAL is designed to be extensible, with standardized protocols for ongoing expert item contribution and model evaluation.

Read the original at arXiv cs.AI · by Shun Ye, Vinny Chandran Suja, Chenlong Li, Chongming Jiang, Reza Zamani, Xiang Li, Christopher Bain, Yuqi Zhou, Walker Peterson, Huidong Wang, Chenglang Hu, Jongchan Park, Xiao Cheng, Benjamin Swedlund, Sandra Murillo, Anjali Sivanandan, Shiyu Sun, Liang Lanfeng, Mohammad Tariqul Islam, Baju C. Joy, primary sourceOpen source ↗
Topics · follow one to build your own front page

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories