Zendoric
← Back to the day · July 26, 2026

Opus 5: Anthropic makes its frontier intelligence cheaper and shields against dual-use risk after Mythos 5

🕒 Published on Zendoric: July 26, 2026 · 00:23

Anthropic launches Claude Opus 5: it comes close to the intelligence of its flagship model, Fable 5, at half the price, and beats its own records in coding, computer use and business automation. But it deliberately keeps it behind Mythos 5 in cybersecurity and biology, the capabilities it watches most closely.

By Zendoric · July 26, 2026. Anthropic yesterday unveiled Claude Opus 5, the new model in its general-purpose family, with a central message: performance close to that of its flagship, Claude Fable 5, at the same price its predecessor cost. According to Anthropic's own data, Opus 5 costs $5 per million input tokens and $25 per million output tokens —the same as Opus 4.8— and is already the default model in Claude Max and the most powerful available in Claude Pro.

In the coding and knowledge-work tests —Frontier-Bench and GDPval-AA—, Opus 5 posts the best result ever for an Anthropic model. On Frontier-Bench v0.1 it beats every other model on the market and more than doubles Opus 4.8's result at a lower cost per task, according to the company. On CursorBench 3.2 at maximum effort it comes within just 0.5% of Fable 5's peak, at half the cost per task. On ARC-AGI 3 —a test of solving novel problems, with no prior patterns— it triples the second-best model, and on Zapier's AutomationBench, which measures full end-to-end completion of business tasks, its success rate is around 50% higher than the next best model at the same cost. On OSWorld 2.0, the reference benchmark for computer use (handling interfaces, browsing, running applications as a person would), Opus 5 beats any other model at any cost and tops Fable 5's best result at little more than a third of its price. In fact, as PCMag also reports on the same launch, Opus 5 goes as far as beating Fable 5 on agentic search tasks —orchestrating searches and multi-step information retrieval—, one of the few areas where the cheap model not only matches but surpasses the paid frontier.

The improvement also extends to scientific research: Anthropic reports significant advances in structural biology, organic chemistry and bioinformatics, with the biggest leaps in inferring molecular structures from spectroscopy data and in predicting how variations in a protein sequence affect its function.

In its internal behavioural audit, Anthropic places Opus 5 as its best-aligned model to date: a score of 2.3 on overall misaligned behaviour, the lowest of its recent models, with the lowest rates of deceptive conduct and the least susceptibility to being manipulated for misuse. These are the company's own figures, worth remembering, but they point in a direction consistent with what Anthropic has been documenting in its latest launches.

The most interesting data point in this launch, however, is not about performance but about risk architecture. Anthropic deliberately keeps Opus 5 behind Mythos 5 —the model that, according to the company, retains the lead in those areas— in offensive cybersecurity and biological research. Opus 5 was not deliberately trained on cyber tasks, but it has improved through the sheer pull of its general capabilities: on the OSS-Fuzz benchmark it comes close to Mythos 5 at identifying software vulnerabilities, though it lags far behind when it comes to turning those vulnerabilities into working exploits. Its safety classifiers are proportionally less restrictive than Fable 5's —Anthropic expects them to intervene around 85% less— and they let through searching for vulnerabilities in source code while blocking binary scanning, penetration testing and exploit generation; flagged requests fall back by default to Opus 4.8 in Claude.ai, Claude Code and Claude Cowork. Only members of Anthropic's Cyber Verification Program —its verified access programme— can use a less restricted version of Opus 5.

This is, at bottom, a governance experiment as significant as the model itself. Anthropic no longer launches a single model with a uniform risk level: it builds a capability hierarchy —Fable 5 as the general frontier, Opus 5 as the accessible, cheap frontier, Mythos 5 as the model that retains the lead in the most sensitive dual-use capabilities— and manages access to each tier with different classifiers and verification programmes. It is the most serious answer we have seen to a real problem: how to keep raising the ceiling of general capability without handing out, along the way, a vulnerability exploitation manual to anyone. Segmentation does not eliminate the risk, it redistributes it: anyone who needs long-horizon autonomous biological research or exploit development will still depend on Mythos 5. In the short term, that means access to the most powerful —and potentially most dangerous— capabilities remains concentrated among those with a direct relationship with Anthropic; it is the same pattern of power concentration we have been flagging in other dual-use model launches.

For the rest of the market, the message is simpler: what was frontier now costs half as much and is the default model in Claude Max and the most powerful available in Claude Pro. That cheapening of frontier intelligence —also visible in OSWorld 2.0 and AutomationBench, two benchmarks that directly measure the automation of office and business tasks— is the concrete mechanism behind the abundance we champion in the long term at Zendoric: the cheaper and more capable the software that executes complete administrative tasks, the closer we are to the point where that work no longer requires routine human intervention and frees people for tasks demanding greater judgement. The short term, however, remains hard: each of these leaps —specifically the AutomationBench one, which measures end-to-end business tasks— is, in practice, a leap in which administrative profiles stop being necessary.

Anthropic accompanies the launch with two beta features that speak to real friction detected in production: switching tools mid-conversation without invalidating the prompt cache (so developers do not pay again to reprocess the context every time the available tools change) and automatic fallbacks, which redirect requests blocked by the safety classifiers to another model instead of simply rejecting them. It is a minor detail compared with the benchmarks, but it reveals where Anthropic is looking: not just making smarter models, but making the safety filter get in the way as little as possible for the developer building on top. That balance —growing capability, bounded risk, decreasing friction— is, ultimately, the whole industry's bet for the coming years.

🔗 Related on Zendoric

Sources & references