I tested Opus 4.5 to see if its really the best in the world at coding - and things got weird fast
Summary
This article tests Anthropic’s Claude Opus 4.5 on several coding tasks and finds mixed results. The model fails on file handling, a WordPress plugin build, and a JavaScript validation fix, while it performs better on a more complex cross-app workflow. The author concludes that Opus 4.5 looks unreliable in the basic chatbot interface, even though it performs strongly inside Claude Code for agentic coding. The piece also highlights a gap between Anthropic’s marketing claim and the model’s real-world behavior. It ends by noting that the model may still improve, but it does not yet feel ready for prime time in this setup.
Classifications
industries
No industries detected
applications
Accounting and Taxes
AskAI Classifications
Labels
AI Software
SaaS
Developer Tools
Linked Companies
Anthropic
$10M to $25M