Zendoric
← Back to the day · July 28, 2026

Anthropic launches Claude Opus 5: close to Fable 5's frontier intelligence at half the price

🕒 Published on Zendoric: July 28, 2026 · 00:38

On July 24, 2026, Anthropic unveiled Claude Opus 5, the latest in its Opus line, which the company describes as a "thoughtful and proactive" model that approaches the frontier intelligence of Claude Fable 5 but at half the price.

On July 24, 2026, Anthropic unveiled Claude Opus 5, the latest entry in its Opus line, which the company describes as a "reflective and proactive" model that approaches the frontier intelligence of Claude Fable 5 but at half the price. According to Anthropic, in coding and knowledge-work evaluations such as Frontier-Bench and GDPval-AA, Opus 5 becomes the new state of the art, although it still trails Mythos 5 on cybersecurity tasks. The model becomes the default in Claude Max and the most powerful one available in Claude Pro, and Anthropic presents it as a model designed for everyday use thanks to greater operational efficiency.

On performance and cost, Anthropic stresses that Opus 5 offers a substantial improvement over its predecessor, Opus 4.8, at the same price. On Frontier-Bench v0.1, Opus 5 beats every other model and more than doubles Opus 4.8's performance at a lower cost per task. On CursorBench 3.2, at the maximum effort setting, the model comes within less than 0.5% of Fable 5's maximum, but at half the cost per task, and achieves a better performance-to-cost ratio than the rest of the models at the high, xhigh and max effort levels.

On knowledge and problem-solving tasks the results are equally striking: on ARC-AGI 3, an evaluation focused on solving novel problems, Opus 5's score is triple that of the next best model. On Zapier AutomationBench, which measures whether models complete business tasks from start to finish, Opus 5's success rate is roughly 1.5 times that of the next best model at the same cost per task, and even at its lowest effort level it beats any other model. On OSWorld 2.0, a computer-use evaluation, Opus 5 beats all other models at any given cost, and even exceeds Fable 5's best result at little more than a third of its cost. Anthropic also points to it as its most efficient model on evaluations such as GDPval-AA v2, HLE, AutomationBench and DeepSearchQA.

In the scientific domain, Opus 5 improves on Opus 4.8 across all of Anthropic's life sciences evaluations, which cover structural biology, organic chemistry and bioinformatics. The most notable gains come in organic chemistry tasks, such as inferring molecular structures from spectroscopy data, where it scores 10.2 percentage points higher than Opus 4.8 on the company's internal benchmark, and in protein-related tasks, such as predicting how variations in a protein's sequence affect its function, where the improvement is 7.7 percentage points. Anthropic also highlights a greater capacity to produce more sophisticated visual outputs, illustrating airflow over aerodynamic objects in a simulated wind tunnel and building a simplified interactive illustration of a cell.

The article offers several concrete examples of the model's agency and thoroughness. In one Frontier-Bench task, Opus 5 was given the drawing of a mechanical part and asked to write code to rebuild it as a 3D model in FreeCAD, but deliberately without any way to view the drawing directly; the model responded by writing its own computer vision pipeline to extract the geometry from the raw pixels and managed to solve the task repeatedly, something no competing model achieved in five attempts under the same configuration. In another case, faced with a real bug in a widely used open-source package manager, Opus 5 found the root cause and fixed an edge case the community patch had missed, while a competing model only fixed the surface symptom and reported the bug as resolved. An engineer at a trading firm used it to build, in a single session, a market data feed for a new exchange, a task earlier models had been unable to complete even with detailed plans from the engineer; finding no live feed to validate against, the model itself built a test harness to check that its code was correctly interpreting the exchange's data.

The announcement includes numerous testimonials from early-access customers, illustrating the model's use across different sectors. In coding tools, representatives from Devin point to performance close to Fable's but at half the cost, with particular strength in hard debugging and root cause analysis; Cursor describes a result just below Fable 5 on CursorBench with similar behaviors; and Kiro and JetBrains highlight its ability to hold context on long, multi-step tasks and its judgment in spotting logical flaws during planning before writing code. In business automation, Zapier says Opus 5 topped its AutomationBench leaderboard without spending more tokens than previous Claude models, completing a full customer churn prevention sequence —from detecting at-risk accounts to alerting the right owner and summarizing for the retention team— with a 100% success rate, something previous models did not manage. In scientific research, a genomic analysis customer describes the model as behaving "more like a careful scientist" than any previous model, choosing the right statistical tests to rule out confounding factors and verifying its own results with independent methods.

In product development, Lovable reports a 22% improvement over Opus 4.7 on its hardest agentic coding tasks, along with greater consistency across runs, while another customer notes that in full application builds the frontend shows the best animations, games and 3D work yet seen in an Opus model. In enterprise and specialized document work, Box says Opus 5 beats Opus 4.8 by 8% overall, with gains of 11% in data analysis and 17% in due diligence workflows, processes common in technology, healthcare and the public sector. In finance, a financial modeling customer reports average accuracy 9 percentage points higher with a third fewer turns and tool calls and 60% less time, while another notes that the model hits its trading benchmark using roughly one-seventh of the reasoning tokens and less than half the latency of Opus 4.8. In the legal field, one customer highlights notable improvements in corporate governance and arbitration, with similar performance achieved while generating 26% fewer tokens on average than Opus 4.8 at maximum reasoning, and another points to the highest score of any model tested on first-pass redlines, almost double Opus 4.8. Also mentioned are cases of code review and PR management, where the model checks branches, templates and testing implications before publishing; an architecture redesign case in which the model defended its position against an engineer's proposal but offered a compromise that preserved what was valuable in the original idea; and a case of production monitoring agents able to manage part of their own memory, reviewing their own assumptions, correcting them when they turn out to be wrong and retiring monitoring queries themselves once they are no longer needed.

On alignment and safety, Anthropic says its automated pre-deployment behavioral audit found Opus 5 to be its most aligned model to date: it follows Claude's Constitution better than Opus 4.8, Sonnet 5 or Fable 5, shows the lowest rates of deceptive behavior, is the least susceptible to being manipulated for misuse and is the safest at avoiding reckless actions with hard-to-reverse effects. In that audit it scores 2.3 on general misaligned behavior, the lowest of the company's recent models.

On dual-use risks, Anthropic notes that Opus 5 does not advance the frontier in risky capabilities: in rigorous evaluations conducted with private sector and government partners, the model still trails Mythos 5 in both biological research and offensive cybersecurity. As with Opus 4.8, Anthropic deliberately avoided training Opus 5 on cyber tasks, although the model has improved substantially at them as a result of greater general capability, approaching Mythos 5 in detecting cybersecurity vulnerabilities. However, it remains far behind Mythos 5 in exploiting those vulnerabilities, that is, in turning them into real cyber threats, something Anthropic illustrates with its OSS-Fuzz evaluation, in which both models identify vulnerabilities with similar success but Opus 5 lags well behind in developing exploits.

On the safeguards applied, Anthropic explains that they are similar to those for Opus 4.8, with the exception of stricter controls on a narrow range of cyber tasks. Opus 5's cybersecurity classifiers are proportionally less restrictive than Fable 5's: they allow finding vulnerabilities in source code, but block "binary-based" vulnerability scanning (a method more associated with malicious actors), penetration testing and exploit generation. Anthropic estimates these classifiers will trigger about 85% less often than in Fable 5's case. In Claude.ai, Claude Code and Claude Cowork, flagged requests will fall back by default to Opus 4.8, and that fallback option can also be enabled in the API. Anthropic's Cyber Verification Program (CVP) gives already-enrolled companies and researchers immediate access to a version of Opus 5 with fewer safety restrictions to facilitate legitimate work that would otherwise be blocked. In biology, since Opus 5 keeps a set of safeguards similar to Opus 4.8's, it becomes Anthropic's most capable generally available model for scientific research, although it still shows significant limitations on autonomous, long-running research tasks, precisely where the company believes the greatest biological risk associated with AI models lies (Mythos 5 remains the strongest model for that kind of work). With this release, biology-related requests that are blocked on Fable 5 will now be routed to Opus 5 instead of Opus 4.8.

On availability, Claude Opus 5 is already accessible across all of Anthropic's platforms, priced at $5 per million input tokens and $25 per million output tokens, identical to Opus 4.8. Developers can use it via the claude-opus-5 identifier in the Claude API. It is also offered in Fast mode, which runs about 2.5 times faster than the default mode, available at twice Opus 5's base price on the Claude Platform and through usage credits in Claude Code. Alongside the model, Anthropic is launching two features in beta: the ability to change the tools Claude can use mid-conversation on the Claude Platform without invalidating the prompt cache, and automatic fallbacks in the API, so that requests flagged by the safety classifiers on Opus 5 or Fable 5 can be automatically rerouted to another model instead of being blocked, ensuring requests are always routed to the best available model by default. As with previous Opus models, Opus 5 imposes no data retention requirements for general access.

Taken together, the launch consolidates a strategy already visible in Anthropic's previous cycles: using the Opus line as the midpoint of its range —below Fable in cost and peak performance, but increasingly close to it— while reserving Mythos as the reference model for sensitive capabilities such as offensive cybersecurity and high-risk biology. The article's insistence on efficiency metrics (fewer tokens, fewer turns, lower latency) alongside the alignment results suggests Anthropic is positioning Opus 5 not just as smarter, but as the model designed for sustained agentic use in production, with an explicit emphasis on the model verifying its own work, iterating carefully and staying on course on long, multi-step tasks.

🔗 Related on Zendoric

Sources & references