OpenAI Life Science Benchmark Reveals AI Passes Only 1 in 3 Scientific Research Tasks

General News

Summary

OpenAI released LifeSciBench, a 750-task benchmark for life-science research, and the top model passed only 36.1% of tasks. The benchmark uses expert-written rubrics and free-response prompts to test real scientific workflows such as evidence handling, analysis, design, and communication. Performance drops sharply when models must interpret attached artifacts like figures, tables, genomic files, or chemical structures. The results show that frontier AI still struggles with the kinds of multistep, multimodal work common in drug discovery and research settings. The article frames the findings as a useful signal for research teams and pharma organizations deciding where AI can assist and where human review remains necessary.

Classifications

industries
No industries detected
applications
Web and Content Management

AskAI Classifications

Labels
AI Software SaaS Developer Tools

Linked Companies

OpenAI
$25M to $50M