Anthropic has launched Claude Opus 5, a new flagship model that the company says delivers state-of-the-art performance on coding and knowledge-work tasks while costing the same to use as its predecessor. The model is available today across Anthropic’s platforms at $5 per million input tokens and $25 per million output tokens. It is also available in a faster mode that runs about 2.5 times quicker than the default setting, although that option costs twice as much. Anthropic says Opus 5 is designed to handle complex software engineering, business automation, scientific research, and computer-use tasks more efficiently. The company positions it as a daily-use model rather than one reserved only for the most demanding workloads. On Frontier-Bench v0.1, an evaluation focused on software engineering, Opus 5 reportedly outperformed other tested models and more than doubled the performance of Opus 4.8 at a lower cost per task. At maximum effort on CursorBench 3.2, it came within 0.5% of the peak score achieved by Claude Fable 5, while costing half as much per task. New model pushes performance The performance gains extend beyond coding. Anthropic says Opus 5 scored three times higher than the next-best model on ARC-AGI 3, which tests how well models solve novel problems. On Zapier AutomationBench, which measures whether models can complete business tasks from start to finish, Opus 5 achieved about 1.5 times the pass rate of the next-best model at the same cost per task. The model also outperformed other systems on OSWorld 2.0, a benchmark for computer-use capabilities, surpassing Fable 5’s best result at just over one-third of the cost. Anthropic also reports improvements in scientific research, particularly in life sciences. On an internal organic chemistry benchmark involving molecular structures inferred from spectroscopy data, Opus 5 scored 10.2 percentage points higher than Opus 4.8. It also improved by 7.7 percentage points on tasks involving the effect of protein sequence variations on function. The company says the model is also better at checking its own work and making repeated attempts when an initial approach fails. In one test, Opus 5 was given an image of a machine part but no direct way to view it and was asked to recreate the part as a 3D FreeCAD model. Instead of stopping there, the model reportedly built its own computer vision pipeline to extract geometry from the image’s raw pixels before reconstructing the part. Anthropic said it succeeded repeatedly, while competing models failed after five attempts under the same conditions. Lower cost, stronger safeguards Opus 5 is also being introduced with safety controls designed around its growing capabilities. Anthropic says it is its most aligned model to date based on automated behavioral testing and has a lower rate of deceptive behavior than recent models. The company said Opus 5 remains behind its Mythos 5 model on offensive cybersecurity capabilities. While the two models were similarly capable of identifying software vulnerabilities in Anthropic’s OSS-Fuzz evaluation, Opus 5 was considerably less successful at developing exploits. Anthropic has also added stronger restrictions around certain cybersecurity tasks, including binary-based vulnerability scanning, penetration testing, and exploit generation. Requests flagged by the safety systems can automatically fall back to Opus 4.8. Alongside the launch, Anthropic is introducing beta features that allow developers to change available tools mid-conversation without invalidating prompt caches and automatically route flagged API requests to another model.Opus 5 is now the default model for Claude Max and the strongest model available to Claude Pro users.Recommended ArticlesGet the latest in engineering, tech, space & science - delivered daily to your inbox.With over a decade-long career in journalism, Neetika Walter has worked with The Economic Times, ANI, and Hindustan Times, covering politics, business, technology, and the clean energy sector. Passionate about contemporary culture, books, poetry, and storytelling, she brings depth and insight to her writing. When she isn’t chasing stories, she’s likely lost in a book or enjoying the company of her dogs.
Anthropic debuts Claude Opus 5 with top coding benchmarks at half the per-task cost
Full Article
Original Source
Read the full article at Interestingengineering →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.