Claude Opus 5 Is Out. Almost Fable 5, at Half the Price.
Claude Opus 5 launches within half a percentage point of Fable 5 on CursorBench, at half the API price. Producer-reported benchmarks, an Opus 4.8 fallback, and a rougher CodeRabbit code-review result complicate the "half price" story.
Marcin
Java since 2007, .NET since 2010 · Azure · lately deep in LLMs
Order the tasting menu and you get thirteen courses, a wine pairing, and a bill that makes your stomach hurt before the food does. Order a few courses off the same kitchen's à la carte menu instead, and you can walk away with most of what made the tasting menu worth ordering, for about half the check. That's roughly what just happened to Anthropic's flagship lineup.
On July 24, 2026, Anthropic launched Claude Opus 5, and the single number that matters most doesn't come from Anthropic at all. It comes from Cursor's independent CursorBench 3.2: at maximum effort, Opus 5 scored 70.0% on real coding tasks. Claude Fable 5, still Anthropic's most capable model on paper, scored 70.5% on the same benchmark. The catch: the average task on Opus 5 cost $8.23. The same task on Fable 5 cost $17.32.
Half a percentage point of accuracy for 52% less money is not a footnote. It is close to the whole story of this release.
What actually shipped
Opus 5 isn't a new tasting menu of its own. It's a way to order fewer courses from the same kitchen. The model ships with five effort levels: low, medium, high, xhigh, and max, and Anthropic recommends xhigh as the starting point for serious coding and agent work, though the default in the API and in Claude Code is one notch lower, at high. Order low and you get a quick, single course. Order max and you get most of the tasting menu back; per CursorBench, that task-level cost runs at roughly half of Fable 5's for the same work. Effort also replaces a chunk of the old model-picking ritual: instead of moving your whole application from Sonnet to Opus to Fable, you can change one setting inside a single model and choose how much you want on the plate.
The context window is 1 million tokens. That's both the default and the ceiling; Anthropic doesn't offer a smaller, cheaper context-window tier for Opus 5 the way some earlier models had one. Maximum output sits at 128,000 tokens in a normal synchronous request (the Batch API can push that to 300,000 with a beta header). Even Anthropic's own documentation warns about "context rot," though: accuracy and recall can degrade as the context fills up, so having a million tokens available doesn't mean using all of them is the right call. Adaptive thinking is on by default, and you can only turn it off at effort high or below. Turn it off anyway, and the docs warn the model may occasionally leak a tool call as plain text or expose its internal XML tags. That's a real integration wrinkle worth planning around.
It's available across every paid Claude plan: the default model in Claude Max, the strongest option in Claude Pro, and absent from the Free plan entirely. You'll also find it in Claude Code, Claude Cowork, the Claude API, and the major clouds: Amazon Bedrock, Google Cloud, and Microsoft Foundry.
The strongest evidence
CursorBench 3.2 is worth dwelling on precisely because the underlying results are Cursor's, not Anthropic's. The chart below is Anthropic's own visualization of those results, the one their announcement leans on to make this exact case: agentic coding performance, broken out by effort level.
Agentic coding by effort level (CursorBench)

| Effort | Opus 5 score | Opus 5 cost | GPT-5.6 Sol score | GPT-5.6 Sol cost |
|---|---|---|---|---|
| max | 70.0% | $8.23 | 67.2% | $5.69 |
| xhigh | 69.3% | $7.35 | 64.5% | $3.88 |
| high | 66.7% | $3.91 | 63.5% | $2.79 |
| medium | 64.3% | $3.29 | 60.0% | $1.95 |
| low | 62.8% | $2.55 | 52.6% | $1.01 |
For comparison, at their own best settings, Fable 5 scores 70.5% at $17.32 per task, Opus 4.8 scores 62.3% at $5.77, and Sonnet 5 (not shown on the chart above) scores 61.5% at $6.45.
GPT-5.6 Sol, a competing model, is cheaper than Opus 5 at every matching effort level, and it scores lower at every level too.
Two numbers tell you almost everything. Opus 5 at max effort trails Fable 5 by half a percentage point and costs 52% less. And Opus 5 at its cheapest, lowest-effort setting still edges out Opus 4.8 running at full effort, for 56% less money than that older model's best run. It's a little counterintuitive that spending less compute can beat a supposedly stronger model spending its maximum, but that's what the numbers say.
Cursor itself adds a caveat worth repeating: results carry variance, and small gaps like 70.0% versus 70.5% may not be statistically meaningful. Treat that half-point gap as noise, not a verdict.
The economics
The tasting-menu math gets specific fast, and the standard rate works out to almost exactly half.
| Model / mode | Input per 1M tokens | Output per 1M tokens |
|---|---|---|
| Opus 5 standard | $5 | $25 |
| Opus 5 Fast | $10 | $50 |
| Fable 5 standard | $10 | $50 |
| Sonnet 5 (promo, through Aug 31, 2026) | $2 | $10 |
Opus 5's standard rate is exactly half of Fable 5's, and the same rate as its own predecessor, Opus 4.8: the base API rate, nothing else folded in yet. Fast mode closes that gap again: it runs about 2.5 times quicker, at Fable 5's own $10 and $50 rate, though right now it only works through Anthropic's first-party API, not through Bedrock, Google Cloud, or Microsoft Foundry. Batch processing halves both rates again. Prompt caching adds three more numbers: a five-minute cache write costs $6.25 per million tokens, an hour-long write costs $10, and reading from an existing cache costs $0.50.
Run the numbers on a task that burns 1 million input tokens and 100,000 output tokens: Opus 5 standard comes out around $7.50, next to $15 for Fable 5 and $15 for Opus 5 running in Fast mode. Sonnet 5, at the promotional rate above, would run about $3 for the same shape of task, though that promotion ends September 1, 2026, when the rate rises to $3 and $15.
One more wrinkle before you compare any of this to older models: Opus 5 uses the same tokenizer introduced in Opus 4.7, which Anthropic says can produce roughly 30% more tokens for the same text than earlier models did. That doesn't distort the Opus 5 versus Opus 4.8 comparison here, since both use the newer tokenizer. It will make your old cost spreadsheets, built on pre-4.7 pricing, quietly wrong.
Where Anthropic claims the biggest gains
Beyond CursorBench, most of the strong numbers come from Anthropic's own system card, so read what follows as the company's claims, not independent findings. Anthropic reports 79.2% on SWE-bench Pro, 89.5% on SWE-bench Multilingual, and 90.8% on BrowseComp for agentic search. On OSWorld 2.0, a benchmark of long computer-operating tasks, Anthropic reports a score of 70.6% for Opus 5, which it says beats Fable 5's best result at a little over a third of the cost.
The company also reports a GDPval-AA v2 score of 1861 Elo, an evaluation from Artificial Analysis covering realistic work product across 44 professions and 9 industries, and says Opus 5 more than doubles Opus 4.8's score on Anthropic's own Frontier-Bench v0.1, at a lower cost per task. That test ran inside Anthropic's own agent harness, with five attempts allowed per task, and Opus 4.8 quietly handled any request that Opus 5 or Fable 5's safety filters rejected during that specific run, so the result describes a deployed system, not a 100% unassisted model. Keep that in mind: a related but different fallback shows up later, live in the product rather than inside a benchmark harness.
The independent counterpoint
Not every early test is flattering, and the most useful one comes from CodeRabbit, which ran roughly 100 known bug patterns pulled from verified real pull requests, each configuration tested three times.
At Opus 5's xhigh effort setting, precision among comments CodeRabbit judged concrete and actionable rose to 39.3%, up from 35.2% for its production baseline. That sounds like a clean win until you read the next line: detection of known issues dropped to 55.2% from 61.1%, and the count of low-value, nitpicky comments jumped from 23 to 92 (if you've ever watched a reviewer spend four minutes on a comma while missing the actual bug two lines down, you already know this feeling). Measured across the entire comment stream, not just the actionable subset, precision actually fell, to 28.6% from 32.8%, and the model burned roughly 60,500 input tokens and 9,500 output tokens per review, well above the baseline's 40,500 and 5,800.
CodeRabbit's conclusion is narrow but important: Opus 5 shouldn't automatically become anyone's sole code reviewer. It's better used as a precise, additional pass, or as the model actually writing the change, not the only one checking it afterward.
Safety and the fallback question
On safety, Anthropic calls Opus 5 its best-aligned model yet, reporting a score of 2.3 on its internal audit of undesirable behavior, its lowest, meaning best, result among recent releases. That's Anthropic's own audit; there's no independent replication of it yet. The company says Opus 5 sticks closer to its published Constitution than Opus 4.8, Sonnet 5, or even Fable 5, and behaves less deceptively.
Cybersecurity tells a similar story of producer-reported gains, not independently tested ones. Anthropic didn't specifically train Opus 5 on cybersecurity tasks, but general capability gains carried over anyway; on the OSS-Fuzz benchmark, Anthropic reports that it approaches Mythos 5 at finding vulnerabilities, while staying noticeably weaker at turning those findings into working exploits. Guardrails, which Anthropic says let it search source code for bugs but block vulnerability scanning of binary files, penetration testing, and exploit generation, also route flagged requests in Claude.ai, Claude Code, and Claude Cowork to Opus 4.8 by default.
That routing is a different mechanism from the Frontier-Bench fallback above, and the one that matters most day to day: live routing inside the product, not a benchmark-harness substitution. It only kicks in on flagged requests, and the API version is opt-in rather than automatic, but when it does trigger, whatever you thought was a benchmark or vibe check of "Opus 5" was partly measuring Opus 4.8 instead.
What it means for the market
Opus 5 puts pressure on both sides of Anthropic's own menu now. It closes to within half a percentage point of Fable 5 on real coding work, and CursorBench shows it costing roughly half as much per task to get there, which raises an honest question: why order the full tasting menu by default anymore? Anthropic's answer is Fable 5's edge in genuinely long, multi-day autonomous work, not any single short benchmark.
Sonnet 5 feels pressure too, even though it keeps its base price advantage: at its promotional $2 and $10 rate, it's still cheaper per token than Opus 5 at any effort level. What's changed is the size of the capability gap you give up to get that price. Opus 5 at low effort can already beat Opus 4.8 at full effort, which makes Sonnet's discount a bigger tradeoff than the price sheet alone suggests.
How you're supposed to order has changed too. It's no longer just which model; it's which model, at which effort level, with how much caching, batching, and fallback routing attached to the bill. The unit worth comparing has quietly become the whole meal, not the price of one ingredient.
So, is it half price?
On CursorBench specifically, the independent benchmark that put a dollar figure on real tasks, yes: the realized cost at max effort came out almost exactly half of Fable 5's. Anthropic itself recommends xhigh as the starting point for complex coding and agent work; for most new projects I'd actually start lower, at low or medium, and only reach for max once you've hit its ceiling, because the same CursorBench numbers say you'll get there later than you think.
I'd also wait before treating the safety and alignment numbers as settled. A 2.3 audit score is Anthropic grading Anthropic, and the CodeRabbit test is a useful reminder that a smarter model doesn't automatically fit every job in your pipeline. The real verdict on Opus 5 won't come from launch week. It'll come from whoever runs the next independent benchmark and checks which model was actually in the kitchen.
Sources
- Anthropic — Introducing Claude Opus 5 (including the "Agentic coding by effort level" charts on that page, source of the GPT-5.6 Sol comparison figures)
- Anthropic — Claude Opus 5 System Card
- Claude Platform Docs — Pricing
- Claude Platform Docs — Context windows
- Cursor — CursorBench 3.2
- CodeRabbit — Opus 5 for code review: Cleaner actionable comments, noisier overall