OpenAI tested GPT-5, Claude, and Gemini on real-world tasks - the results were surprising
Summary
OpenAI introduced GDPval, a new benchmark for measuring how AI models perform on economically valuable real-world tasks. The evaluation covers 1,320 tasks across 44 occupations and uses files, documents, and multimodal outputs to better simulate workplace work. OpenAI said top frontier models are already approaching expert-level quality on some tasks, with Claude Opus 4.1 leading in aesthetics and GPT-5 leading in accuracy. The company also plans to release an experimental autograder and a subset of tasks for researchers. The article emphasizes that current AI still struggles with context, iteration, and ambiguity in real jobs.
Classifications
industries
No industries detected
applications
Web and Content Management
AskAI Classifications
Labels
AI Software
SaaS
Developer Tools
Linked Companies
OpenAI
$25M to $50M