Fable 5 Returned Behind a Stricter Filter; Sonnet 5 Is Where the Traffic Belongs
On 30 June the export controls on Fable 5 were lifted and Claude Sonnet 5 launched at $2 / $10. The flagship came back with a classifier that Anthropic says flags benign coding more often, answered by Opus 4.8 when it fires. Why "over 99%" covers one technique, why Sonnet 5's real saving is about 13% per request, and why your logs should record the model that actually answered.
This week Anthropic got its flagship back. On Tuesday 30 June the US Commerce Department lifted the export controls it had placed on Claude Fable 5 eighteen days earlier, and on 1 July the model was serving users globally again, this time behind a new cyber classifier that Anthropic itself says flags ordinary coding requests more often. The same Tuesday, with far less noise, it released Claude Sonnet 5 at $2 and $10 per million tokens. The first event made the headlines. The second one is the model most production traffic should be tested on next.
This is the second follow-up to The 72-hour model, a week after Locked out by passport. The one calculation below that is ours is labelled as ours.
The flagship came back with a second model behind it
Commerce Secretary Howard Lutnick lifted the controls by letter to Anthropic's Tom Brown on 30 June. His public line: "Over the past two weeks, we have worked closely with Anthropic to analyze and approve Fable 5 to ensure alignment across the US Government and strengthen America's leadership in AI." Anthropic announced that "Fable 5 will be available starting tomorrow, Wednesday, July 1, to users globally on the Claude Platform", with AWS, Google Cloud and Microsoft Foundry to follow "as quickly as possible". A third-party test dated 1 July already found the model active on Amazon Bedrock, including in eu-west-1. Subscribers got it "for up to 50% of weekly usage limits through July 7".
The model itself did not change. It was still claude-fable-5 at $10 input and $50 output per million tokens, with a 1M-token context window. What changed was the filter in front of it, and the number attached to that filter needs reading slowly.
Anthropic's sentence is: "The new classifier means that the specific technique described in the Amazon report is blocked in over 99% of cases." One technique, the one Amazon's researchers reported. Infosecurity Magazine's account of the same post adds that in "a very small fraction of cases" the model may still respond, though not with anything "detailed enough to help a cyber attacker". Researchers at Commerce's Center for AI Standards and Innovation tested the old and the new safeguards and, per the same report, called them "extraordinarily strong".
"Over 99%" describes one reported technique. It is not a block rate for jailbreaks in general, and it says nothing about how often the filter stops ordinary work.
On that second point Anthropic was unusually direct: "The new classifier also comes at the cost of flagging benign requests more often during routine coding and debugging tasks." And when a request is flagged, "the request will instead be sent to Opus 4.8." On 2 July it published the policy behind the filter, in four tiers:
- Prohibited useRansomware, wipers, malware, command-and-control. Blocked.
- High-risk dual use"Hacking, penetration testing, red teaming, and bug bounties", exploit development. Blocked on Fable 5, which is a real constraint for a security team.
- Low-risk dual useOpen-source intelligence, vulnerability identification that existing tools can already do. Monitored, and sometimes blocked as a safety margin.
- Benign use"Secure coding, and fixing simple or already identified vulnerabilities in code", patching, incident response. Allowed, with monitoring. This is the tier where the extra false positives land.
So the model that came back on 1 July was two models in practice: Fable 5 for most requests, Opus 4.8 for whatever the classifier disliked, with the boundary drawn by a filter tuned in a hurry.
The model that shipped the day before
Claude Sonnet 5 went live on 30 June as claude-sonnet-5, "available everywhere today": the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry. It became the default model on the Free and Pro plans. The context window is 1M tokens, which is both the default and the maximum, with 128K tokens of output. Unlike Fable 5, it supports zero data retention for organisations that have that agreement.
The price was $2 per million input tokens and $10 per million output tokens, labelled introductory through 31 August, with a rise to $3 and $15 scheduled for 1 September. Cache reads cost $0.20 and the Batch API halves input and output to $1 and $5.
The benchmark figures are vendor-reported, from Anthropic's system card, with Sonnet 5 at maximum effort. On SWE-bench Pro it scored 63.2, against 58.1 for Sonnet 4.6, 58.6 for GPT-5.5 and 55.1 for Gemini 3.5 Flash. Opus 4.8 sits at 69.2, the same figure our June chart used. On OSWorld-Verified, Sonnet 5 reached 81.2 against 78.5 for its predecessor. On GDPval-AA v2, the Artificial Analysis leaderboard for professional knowledge work, its Elo was 1618 against 1395. Six points behind Opus 4.8 on the hardest coding row, at 40% of its per-token price.
Anthropic calls it a drop-in upgrade for Sonnet 4.6 with three behaviour changes: adaptive thinking is on by default, the old manual thinking budget returns a 400 error, and so does any non-default temperature, top_p or top_k. There is a fourth change that does not break anything and still moves the bill. The docs say: "The same input text produces approximately 30% more tokens than on Claude Sonnet 4.6." The announcement puts it at "roughly 1.0–1.35× depending on the content type".
Here is the arithmetic, and it is ours, done on Anthropic's figures. A price of $2 against $3 is a factor of 0.667. Thirty percent more tokens is a factor of 1.30. Multiply them and an equivalent request costs about 0.87 of what it cost on Sonnet 4.6: roughly 13% cheaper, not 33%. At the $3 and $15 scheduled for September, the same request will cost about 30% more than on the old model.
Thirteen percent cheaper for a better model is still a good trade. It is a different sentence from the one on the pricing slide, and the only way to know your own number is to recount your own prompts with the token-counting endpoint.
A silent fallback is a different model answering
The part of the Fable 5 redeployment that matters most in production is easy to miss. When the classifier fires, the user still gets an answer. It comes from Opus 4.8. Same prompt, different model, different output. For any client that never reads which model answered, the handoff is silent.
On the API the behaviour is explicit, and the documentation is precise about it. A declined request is an HTTP 200 with stop_reason: "refusal" and a stop_details.category of cyber, bio, frontier_llm, reasoning_extraction or general_harms. The docs add, for the cyber category: "Benign cybersecurity work can also trigger this category." You are "not billed for a refusal that arrives before any output", although the request still counts against your rate limits.
If you opt into the server-side fallbacks parameter, the API retries on another model inside the same call. The response names the model that answered in its top-level model field, marks the handoff with a fallback content block, and records each attempt in usage.iterations. Two limits apply. The parameter is a beta on the Claude API only; it is "not available on Amazon Bedrock, Google Cloud, or Microsoft Foundry", which are the platforms an EU team uses for regional pinning, so there you write the retry yourself or use the SDK middleware. And the docs warn that you must size the fallback model's rate limits for the refusal volume you expect, "or fallbacks degrade to refusals under load."
Three things follow for anyone running this in production:
- Log the model that served the call, not the model you asked for. The field exists. An audit trail that records only the requested model ID is wrong for every flagged request.
- Evaluate the route as a pair. An eval run that scored Fable 5 says nothing about the Opus 4.8 answers now mixed into the same route. Treat the pair as the unit under test.
- Check the data terms per route. Fable 5 is a Covered Model that requires 30-day data retention and is unavailable under zero data retention unless Anthropic expressly authorises it; without retention enabled the API returns a 400. Sonnet 5 runs under ZDR. For many regulated workloads that decides the route before any benchmark does.
Our read
The two events of this week make one argument. The flagship's return is a story about access, and access came back on the same terms it left: by letter, with nothing to stop the next one. An analyst quoted by CIO put it in one line: "Restored access is not restored certainty." A route that names claude-fable-5 needs a written fallback today for the same reason it needed one on 12 June.
Sonnet 5 is a story about where the volume should go. A model within six points of Opus 4.8 on SWE-bench Pro, at 40% of its per-token price, with a 1M-token window and no retention requirement, is the natural first candidate for most task classes in a routing policy. Candidate is the operative word. The figures are the vendor's, the tokenizer moves the cost, and thinking is now on by default, so the promotion step is the same as ever: run your own eval set on it, compare cost per request and not cost per token, and only then change the default. Fable 5 stays what it was before the ban, a per-route exception for the work that fails on everything cheaper.
If you would like a second pair of eyes on a routing policy, on what your logs record when a vendor swaps the model under you, or on an eval gate for a Sonnet 4.6 to Sonnet 5 move, that is work Sebrona does. The address is info@sebrona.com.
Reading
Where the figures, quotes and dates come from. Benchmarks are vendor-reported unless the row says otherwise.
- Anthropic, "Redeploying Claude Fable 5"The 1 July global return, cloud platforms "as quickly as possible", the 50% promo through 7 July, the "over 99%" sentence, the admitted cost on benign coding, and the Opus 4.8 fallback.
- Anthropic, cyber safeguards and jailbreak frameworkThe four usage tiers and what the classifier does in each.
- Claude Sonnet 5 announcement, docs and system cardRelease date, model ID, default-plan status, context and output limits, the three behaviour changes, the tokenizer figures, ZDR support, and the benchmark table (SWE-bench Pro, OSWorld-Verified, GDPval-AA v2). The Opus 4.8 figure of 69.2 is from the announcement.
- Anthropic pricing page and model docsAll list prices in Fig. 01, cache and batch rates, Fable 5's ID and limits, and Sonnet 5's introductory price schedule.
- Refusals and fallback; API and data retentionRefusal response shape and categories, billing of refusals, the fallbacks beta and where it is unavailable, the rate-limit warning, and the 30-day retention requirement for Covered Models.
- Press on the liftLutnick's statement and the letter to Tom Brown (NBC News, CIO); the "very small fraction of cases" and CAISI's "extraordinarily strong" (Infosecurity Magazine); Bedrock regions on 1 July (DevelopersIO); the analyst quote, Sanchit Vir Gogia of Greyhound Research (CIO).
- Left out on purposeTerminal-Bench 2.1. Third-party transcriptions of Anthropic's chart give two different Opus 4.8 figures and the system card does not list one, so we quote neither side of that comparison.