Claude Opus 5: Same Price per Token, a Different Response Shape, and Data Terms You Can Deploy Under
Opus 5 shipped on 24 July at the same $5 / $25 as Opus 4.8, with thinking on by default, five effort levels and a 1M-token window. The migration checklist a builder needs, the benchmark table read by who ran each row, the classifier flag rate that matters in production, and what EU residency actually costs.
This week Anthropic released Claude Opus 5. It arrived on Friday 24 July as claude-opus-5, at $5 and $25 per million tokens, the same as Opus 4.8, with a 1M-token context window, and as the new default model on the Max plan. The announcement's pitch was that "it comes close to the frontier intelligence of Claude Fable 5 at half the price". The part a builder needs first is in smaller print: thinking is now on unless you turn it off, and that changes both the bill and the shape of the response.
Every benchmark below is tagged with who ran it, because on this release that tag carries more information than the score.
What breaks when you change the model string
The rate card did not move. Opus 5 costs $5 per million input tokens and $25 per million output tokens, "unchanged from Claude Opus 4.8", with cache reads at $0.50 and the Batch API at half price. The 1M-token window "is both the default and the maximum", output tops out at 128K tokens, and the whole window is billed at the standard rate. Fast mode, a research preview on the Claude API only, runs at "up to 2.5x higher output tokens per second" for $10 and $50.
Anthropic's model IDs are pinned snapshots, the dateless ones included, so nothing changes in your system until you change the string. When you do, the migration guide lists what to expect. We would treat it as a checklist:
- Thinking runs when the field is omitted. On Opus 4.8 a request without a thinking field ran without thinking. On Opus 5 the same request runs with adaptive thinking. max_tokens "remains a hard limit on total output, thinking plus response text", so a budget tuned for 4.8 can now cut the answer short.
- The same price can mean a bigger bill. Thinking tokens "are billed as output tokens even when the thinking text is not returned to you". A workload that ran without thinking on 4.8 produces more output tokens per request at an identical rate.
- Position-based parsing breaks. A response can start with one or more thinking blocks, which arrive empty by default. The guide is blunt: code that reads content[0].text, or a stream handler that assumes the first block is text, "breaks on these responses". Select blocks by their type.
- You cannot disable thinking at the top two effort levels. thinking: {"type": "disabled"} with effort xhigh or max returns a 400. Below that it is allowed, but Anthropic warns the model can then write a tool call into plain text, and recommends keeping thinking on at a lower effort level instead.
- Tool loops must echo thinking blocks untouched. Pass them back "complete and unmodified" with the tool result. Edited, reordered or dropped blocks are rejected with a 400.
- Sampling parameters are rejected, and this one is inherited. A non-default temperature, top_p or top_k "returns a 400 error on Claude Opus 5, the same as on Claude Opus 4.7". It is no change from 4.8. It bites teams arriving from Opus 4.6 or earlier, together with the removed manual thinking budget, removed prefill and the newer tokenizer.
Then there is effort. Anthropic's announcement calls it "the model's effort setting"; "dial" and "toggle" are the press's words, and Fortune's write-up listed three levels. The docs list five: low, medium, high, xhigh and max, with high as the API default. The parameter is not new; Opus 4.5 to 4.8 had it. What Anthropic claims is new is that Opus 5 "converts additional effort into better results more reliably than any earlier Opus model". The guide adds that the token allocation behind each level has changed, tells you to "run a fresh effort sweep on your own evals", and concedes that max "may show diminishing returns". Its own benchmark run shows what that looks like.
"Same price" is a statement about the rate card. With thinking on by default, the number that moves is tokens per request.
Read the table by who ran it
The system card's summary table mixes four kinds of evidence in one grid. Pulled apart, it reads like this, with Opus 5 at maximum effort unless stated.
- Artificial AnalysisGDPval-AA v2, an Elo rating over 220 professional tasks: Opus 5 1861 at max and 1827 at xhigh, the latter with "25% fewer output tokens". Fable 5 1747, GPT-5.6 Sol 1736, Opus 4.8 1593. On AA-Briefcase, 1720 against 1574 for Fable 5.
- ARC Prize FoundationARC-AGI-3: a verified 30.16% at high effort, against 7.78% for GPT-5.6 Sol and 1.52% for Opus 4.8. The announcement says "three times" the next-best model and the system card says "roughly four times"; quote the scores. On ARC-AGI-2, Sol is ahead, 92.5 to 90.4.
- Harbor, the benchmark's authorsFrontierBench v0.1: Opus 5 43.3, GPT-5.6 Sol 34.4, Fable 5 33.8, Opus 4.8 21.1.
- AnthropicSWE-bench Pro 79.2, against 80 for Fable 5, 69.2 for Opus 4.8 and 64.6 for Sol. OSWorld 2.0 70.6 against 66.1 for Fable 5. On DeepSWE v1.1 Sol leads, 72.7 to 68.8.
FrontierBench exists in two sets of numbers and they get mixed up in secondary coverage, so to be explicit: the row above is Harbor's run, which is the one to use for comparing vendors. Fig. 01 uses Anthropic's own run on the mini-SWE-agent harness, which is the only one that reports effort levels; in that run Fable 5 scored 33.7 and Opus 4.8 18.7. Anthropic's claim that Opus 5 "more than doubles Opus 4.8's performance" holds on both.
The line from that section of the system card that we would take to a production review describes the classifier: "Opus 5 safety classifiers flagged and refused 5% of the API calls, in 4% of the total trials, falling back to Opus 4.8. Fable safety classifiers flagged 42% API calls on 26% of trials". This was on a benchmark of computational biology, physics simulation, CAD, formal proofs and GPU performance work. Ordinary science and engineering, in other words.
Anthropic's headline version is that Opus 5's cyber classifiers "intervene around 85% less often than they do for Fable 5". The benchmark detail says something sharper. A route that names Fable 5 for technical work was, on this evidence, often an Opus 4.8 route with a Fable 5 invoice line. Opus 5 mostly answers as itself.
The caveat on all of it came from Apidog's write-up the day after launch: "No third party has published a reproduction of any of these results at the time of writing." That is fair for the Anthropic rows. The Decoder, reading Artificial Analysis's data the same day, put the overall index at 61 for Opus 5, 60 for Fable 5 and 59 for GPT-5.6 Sol, which is a tie with a ranking on top, and gave AA-Briefcase cost per task as $17.79 at max and $10.41 at high.
Where it runs, and under what data terms
Opus 5 was on every platform on day one: the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry. For an EU buyer the list splits in two.
Anthropic's first-party API has no EU region. The data-residency docs are plain about it: for inference_geo, "Only "us" and "global" are available", and for workspace geo, "Only "us" is currently available". US-only inference carries a 1.1× multiplier. There is nothing to select for Europe.
EU residency therefore means a hyperscaler, and it costs ten percent. On Bedrock the EU inference profile spans Frankfurt, Zurich, Stockholm, Milan, Spain, Ireland, London and Paris, with Stockholm and Ireland also offered in-region, and "Regional endpoints carry a 10% pricing premium over global endpoints." On Google Cloud the eu multi-region endpoint carries the same premium, and models newer than Sonnet 4.6 cannot be pinned to a single region at all. On Microsoft Foundry the deployment types are Global Standard and US Data Zone Standard; no EU data zone is listed. Two details to verify before you promise anyone a region: the Bedrock docs list the models open to all customers and Opus 5 is not among them, with its row pointing to the access console instead, and the region table is per platform, not per model.
The premium is not the only cost of pinning. Off the first-party API you lose fast mode, server-side fallback and, on both Bedrock and Google Cloud, the Batch API with its 50% discount. That is a routing constraint, and it belongs in the policy file next to each route.
On data terms the news is good. The announcement says "Opus 5 does not have data retention requirements for general access", where Fable 5 requires 30-day retention and is closed to zero-data-retention customers unless Anthropic authorises it. On Bedrock, Anthropic describes the service as running with "zero operator access"; on Azure-hosted Foundry deployments, "prompts and completions remain within Azure" and only usage metadata and safety-flagged content leave. For a regulated EU workload that could not accept Fable's terms, Opus 5 was the first model of this generation at the top of the range that fits an existing data-processing agreement.
Our read
Opus 5 is the release from this summer that changes the most routes, and what does it is a combination: scores within a point of Fable 5 on the coding rows that matter, half the rate, no retention requirement, a classifier that leaves technical work alone, and EU-pinned endpoints on two hyperscalers. For most EU production systems that makes it the highest model you can actually deploy under the terms you already have.
It is not a drop-in, and we would not let anyone treat it as one. Change the model string behind an eval gate. Compare cost per task, not cost per token, because thinking now runs by default and the default effort is high. Sweep model and effort together: low and medium on Opus 5 compete with Sonnet 5, and xhigh competes with Fable, so the cheapest cell that passes your own gold set is the one to promote. Keep Opus 4.8 wired as the fallback; it is where flagged requests go anyway, and it is not going away before May 2027.
The honest limit: almost every score above is from the vendor or from a leaderboard that drifts, and none of it was measured on your data. A model card is a reason to run an eval and cannot stand in for the result of one.
If you want that eval gate built, or a routing policy that treats effort and region as first-class fields, that is what Sebrona does for teams inside the EU data boundary. Write to info@sebrona.com.
Reading
Where the prices, API behaviour and scores come from. Each benchmark row says who ran it.
- Claude Opus 5 announcementRelease, pricing, the "half the price" line, "the model's effort setting", Max default, the 85% classifier figure, the Opus 4.8 fallback, the data-retention sentence, and the claims on Frontier-Bench, CursorBench and ARC-AGI 3.
- Claude Opus 5 system cardTable 8.1.A and sections 8.5, 8.13 and 8.14: every score in § 02, the effort-level results in Fig. 01, who ran each benchmark, and the classifier flag rates in Fig. 02.
- Opus 5 docs: overview, what's new, migration guideModel ID, context and output limits, cache and batch prices, fast mode, pinned snapshots, thinking on by default, max_tokens, content[0].text, the 400 on disabled thinking at xhigh and max, thinking blocks in tool loops, sampling parameters, and the effort sweep advice.
- Effort, data residency, Bedrock, Google Cloud, Foundry, deprecationsThe five effort levels; inference_geo and workspace geo limits; EU regions, the 10% regional premium and the access note on Bedrock; the eu multi-region endpoint; Foundry deployment types and the Azure data statement; retirement floors for Opus 4.5 to 5.
- Press and reviewsFortune on the "toggle" and three levels (24 Jul); The Decoder on the Artificial Analysis index and AA-Briefcase cost per task (25 Jul); Apidog on the absence of reproductions (25 Jul).
- Not verifiedThe access criteria for Opus 5 on Bedrock, which effort levels the consumer apps expose, and Opus 5 availability in each individual EU region. Check the AWS console before you commit to one.