I tested Opus 4.5 to see if its really the best in the world at coding - and things got weird fast

General News

Summary

This article tests Anthropic’s Claude Opus 4.5 on several coding tasks and finds mixed results. The model fails on file handling, a WordPress plugin build, and a JavaScript validation fix, while it performs better on a more complex cross-app workflow. The author concludes that Opus 4.5 looks unreliable in the basic chatbot interface, even though it performs strongly inside Claude Code for agentic coding. The piece also highlights a gap between Anthropic’s marketing claim and the model’s real-world behavior. It ends by noting that the model may still improve, but it does not yet feel ready for prime time in this setup.

Classifications

industries
No industries detected
applications
Accounting and Taxes

AskAI Classifications

Labels
AI Software SaaS Developer Tools

Linked Companies

Anthropic
$10M to $25M