OpenAI Life Science Benchmark Reveals AI Passes Only 1 in 3 Scientific Research Tasks
Summary
OpenAI released LifeSciBench, a 750-task benchmark for life-science research, and the top model passed only 36.1% of tasks. The benchmark uses expert-written rubrics and free-response prompts to test real scientific workflows such as evidence handling, analysis, design, and communication. Performance drops sharply when models must interpret attached artifacts like figures, tables, genomic files, or chemical structures. The results show that frontier AI still struggles with the kinds of multistep, multimodal work common in drug discovery and research settings. The article frames the findings as a useful signal for research teams and pharma organizations deciding where AI can assist and where human review remains necessary.
Classifications
industries
No industries detected
applications
Web and Content Management
AskAI Classifications
Labels
AI Software
SaaS
Developer Tools
Linked Companies
OpenAI
$25M to $50M