The overselling of AI - and how to resist it

General News

Summary

A BlueOptima study (BARE) evaluated 57 LLMs on maintainability-focused refactoring tasks across 4,276 real source files and found that even the best models succeeded less than 23% of the time, while benchmark scores often exceeded 85% and poorly reflected real-world performance. The study measured strict success criteria—compilation, no behavioral regressions, and measurable maintainability improvements—and recorded wide variance by language and task (e.g., 32% success in JavaScript, 4% in C, and as low as 1.5% on complex architectural tasks). The piece warns that AI is being oversold by vendors and pundits, and quotes David Linthicum urging leaders to adopt an evidence-driven, skeptical stance to avoid costly overspending and strategic errors. It advises companies to invest in backend work, maintainability practices, and qualified decision-makers before treating AI as a turnkey solution.

Classifications

industries
No industries detected
applications
Business Intelligence

AskAI Classifications

Labels
Software Engineering Intelligence Developer Productivity Tools Code Security Software

Linked Companies

BlueOptima
$10M to $25M