The Workhorse Tier Repriced Three Times in a Fortnight: Route on Price Ceilings, Not on a Vendor
Between 30 July and 13 August, OpenAI, Anthropic and Google each repriced their mid-tier model, while Meta and Alibaba put two Apache 2.0 models around 30B on the table. Some of those prices are permanent and some end on 31 December. Why a routing policy with a ceiling and an expiry date per route beats picking a vendor.
This week the tier of models that does most production work got cheaper from three directions at once. On Monday 10 August Anthropic added a note to its Claude Sonnet 5 launch post making the introductory price permanent, and Meta released Muse Glimmer under Apache 2.0. On Thursday Google shipped Gemini 3.7 Flash at half the price of a model it had launched three weeks earlier. Alibaba opened Qwen3.8 in between.
The lesson of the week is not which vendor won it. A mid tier that reprices three times in a fortnight cannot be handled by choosing a vendor. It needs a routing policy with a price ceiling on every route.
Three price moves in fourteen days
Start a little before the week. On 30 July OpenAI cut GPT-5.6 Terra by 20% to $2 in and $12 out per million tokens, and GPT-5.6 Luna by 80% to $0.20 / $1.20.
On 10 August Anthropic followed. Sonnet 5 had launched on 30 June at $2 / $10 as an introductory price through 31 August, with $3 / $15 due on 1 September. The editor's note reads: "Sonnet 5's introductory pricing of $2 per million input tokens and $10 per million output tokens is now permanent. The standard pricing of $3 input / $15 output previously set to take effect September 1 no longer applies." Anthropic gave no reason.
On 13 August Google shipped Gemini 3.7 Flash at $0.75 / $3.75, which its own recap described as "half the original 3.6 Flash cost per million tokens", arriving "just three weeks after our launch of 3.6 Flash". That older model had launched on 21 July at $1.50 / $7.50, and on the same Thursday it was repriced to match. The same day DeepSeek took V4-Pro to general availability under MIT, with off-peak rates at half the peak price from 16 August.
Two details keep these numbers honest. Google's price is introductory: it runs "through December 31, 2026", and the list price of $1.50 / $7.50 returns on 1 January 2027. And Anthropic's saving is smaller than it looks if you are coming from Sonnet 4.6, because its pricing page says models from Claude 4.7 on use a tokenizer that "produces approximately 30% more tokens for the same text". A third off the token price is not a third off the bill.
Two Apache 2.0 models around 30B in five days
The closed vendors cut prices. The open side moved the floor underneath them.
Muse Glimmer, released on 10 August, is a dense model of about 29.6B parameters including a vision encoder of roughly 1.8B, with 52 layers and a context of 131,072 tokens or more. The repository is meta-models/Muse-Glimmer-30B and its LICENSE file is the unmodified Apache 2.0 text: no user cap, no acceptable-use file, no EU carve-out. Meta built it for one GPU. The 17 GB quantisation targets 24 GB of VRAM at 1.0% average degradation across 15 benchmarks; the dynamic one targets 32 GB at 0.2%. On an RTX 5090 Meta reports 74.9 tokens per second, rising to 233.4 with its speculative drafter.
Meta's own table puts Glimmer ahead of Gemma4-31B and Qwen3.6-27B on MCP Atlas (75.5 against 54.2 and 62.5) and SWE-Bench Pro (51.2 against 36.9 and 50.2), and behind Qwen on TerminalBench 2.1 (51.7 against 60.7). Those are vendor-reported. One number on the card deserves a second look: an attack success rate of 28.4 on Siren AgentDojo. Roughly one indirect prompt injection in four lands, by the vendor's own count.
The release was also a reversal. Muse Spark had shipped closed in April. In an interview released on 13 May, Meta's chief AI officer Alexandr Wang called Spark "not suitable for open sourcing". On 10 August Mark Zuckerberg published an essay titled "The Future Is for Everyone", and he and Wang were reported as saying that open weights for Muse Spark 1.2 were coming "soon". We could not read those posts ourselves, so treat the wording as reported. No date was given and no licence was named.
Four days later Alibaba published Qwen3.8-27B: dense, vision-language, Apache 2.0, with 262,144 tokens of native context and a documented path to 1,000,000. Meta's comparison table had been drawn against Qwen3.6. It was out of date before the week ended.
Qwen3.8: one banner, two licences
Qwen's announcement on 14 August read: "We promised open weights for Qwen3.8. Now, time to meet them!" At least one trade outlet ran a headline saying the models shipped under Apache 2.0. The files say something more complicated.
- Qwen3.8-27BApache 2.0, with the standard licence text in the repository. This is the only member of the family that clears a permissive-licence floor.
- Qwen3.8-2.4T-A95BThe open form of Qwen3.8-Max, 2.4T parameters with 95B active. Its file, committed on 12 August, is the "Qwen3.8-Max License". A company running a "Model as a Service or AI Work Assistant business" with group revenue above US$50,000,000 in any twelve months "shall obtain a separate license from Qwen". Above 100,000,000 monthly active users or US$20,000,000 in monthly revenue, the model name must be shown in the interface. Internal use is exempt.
"AI Work Assistant" is Qwen's own term. The licence defines it as a standalone product built mainly for AI-assisted coding or office productivity, and names Qoder and QwenWork as examples. A coding copilot sold as a product falls inside it. An assistant that is one feature of a larger product does not.
There is a second gap between the banner and the file. The flagship's model card says the hosted Qwen3.8-Max is "the official version" with "more features, such as vision input & non-thinking support, 1M context length by default, official built-in tools". The weights you can download are a different, narrower thing than the API you can call at $2 / $6.
None of this makes Qwen3.8 unusable. It means "we run Qwen3.8" is not a sentence a licence review can approve. The repository name is.
Our read
Four vendors now sell a capable model at exactly $2 per million input tokens: Sonnet 5, GPT-5.6 Terra, Gemini 3.1 Pro in preview, and hosted Qwen3.8-Max. They separate on output, at $10, $12, $12 and $6. Agent loops are heavy on output and on cached input, so that is where a route should be priced.
A price with an end date is a discount. Budget on the list price.
What we would put in the routing policy after a fortnight like this one:
- A ceiling per route, in and out. Each route names the most it may cost per million tokens. When a vendor moves, the policy decides, and nobody has to call a meeting.
- An expiry column. The Flash price ends on 31 December. A route that depends on it needs a named fallback at list price before that date.
- Promotion by eval, never by price. A cheaper model earns a route only when your own eval set passes on it. A price cut is a reason to run the eval, and nothing more.
- Licence per repository. Apache 2.0 for Glimmer and Qwen3.8-27B. A revenue gate for the Qwen flagship.
- Nothing promised goes on the roadmap. Spark 1.2 has no date and no licence.
The honest limit on all of it: a price per token is not a price per task. Anthropic's newer tokenizer produces about 30% more tokens for the same text. Google says 3.6 Flash uses 17% fewer output tokens than 3.5 Flash. Two models at the same list price can differ by a third on the invoice. Set the ceiling in euros per completed task on your own traffic, and keep the per-token number as the alarm.
A routing policy with price ceilings, an expiry column and an eval gate is a short file and a standing habit. Writing it, and wiring it to logs an auditor can read, is work we do at Sebrona. If your 2027 budget rests on a price that ends on 31 December, write to info@sebrona.com.
Reading
Where each figure comes from. Prices are the vendors' published prices for this week.
- Introducing Claude Sonnet 5Launch date and introductory terms, and the 10 August editor's note quoted above.
- Claude pricingThe $2 / $10 line and the tokenizer note.
- Introducing Gemini 3.7 FlashRelease date, $0.75 / $3.75 through 31 December 2026, the list price from 1 January 2027.
- Gemini 3.6 Flash announcement and August recapThe 21 July launch at $1.50 / $7.50, the 17% output-token figure, and the "half the original 3.6 Flash cost" wording.
- Gemini API pricing and changelog; TokenCost price historyFlash and Pro prices, the 20 May price of 3.5 Flash and the 13 August repricing of 3.6 Flash.
- OpenAI model pagesPrices for Sol, Terra and Luna.
- meta-models/Muse-Glimmer-30BThe Apache 2.0 LICENSE file, architecture, quantisation and speed tables, and the vendor-reported benchmark and security figures.
- Qwen3.8 repositoriesThe two licence files quoted above, parameter counts, context lengths, and the note on how hosted Qwen3.8-Max differs from the open weights.
- Qwen3.8 coverageThe 14 August announcement wording. One outlet's headline says both models ship under Apache 2.0; the flagship's licence file says otherwise.
- Other price pagesHosted Qwen3.8-Max, Muse Glimmer on Together AI, and the DeepSeek change log for V4-Pro.