BioEVAL: A global, multi-institutional benchmark of large language and multimodal models for bioengineering
Existing LLM benchmarking has focused heavily on factual recall, giving limited insight into how models handle frontier and multimodal tasks in biomedical science. To address this, 22 research groups launched BioEVAL (BioEngineering Validation of AI and LLMs), a global multi-institutional initiative assessing experimental reasoning across bioengineering (BE) subfields. The benchmark covers 11 major BE subfields plus uncategorized items and consists of 608 PhD-level evaluation items: 380 multiple-choice questions (359 retained after audit), 218 literature synthesis tasks, and 10 multimodal problems requiring experimental image interpretation.
All items went through authoring-group expert review and centralized quality control before evaluation. After evaluation, a blinded cross-group consensus audit of the highest- and lowest-accuracy MCQ items flagged 21 questions for revision or removal; these were withheld, so all reported MCQ results are computed on the 359 retained items.
Evaluations covered cloud-scale foundation and multimodal models (including ChatGPT, Gemini, and Grok) as well as locally deployable models suited to consumer-grade GPUs. The best models reached up to 90% accuracy on MCQs, a similarity score of 0.72 on literature synthesis, and 80% accuracy on the small multimodal sample, with substantial performance variation across subfields. The resulting leaderboard highlights current capabilities, limitations, and development priorities across BE task categories.
BioEVAL is maintained as an extensible benchmark with standardized protocols for ongoing expert item contribution and model evaluation.