Skip to main content
ENSK
PracticeElektrikProJARVISBlog
← All posts

The Q3 2026 Model Floor: A Shortlist You Can Defend, and the Dates That Will Break It

A quarter-end routing review, dated 20 September 2026. Which models hold each role, from frontier reasoning to speech-to-text, with the price, the licence and the catch. Plus the two risks a shortlist does not show: a model ID that retires, and one that keeps answering while a different model does the work.

TopicField notes
Published20 Sept 2026
AuthorMiroslav Striško
Reading10 min

This week two vendor notices came due on the same Monday. Anthropic's weekly limits for Claude Code on subscription plans dropped by 17%. And DeepSeek, which had released V4.1 Flash under MIT four days earlier, was due to start sending every V4-Pro request to the new model at 04:00 UTC. By the time we read its pages, the change log said the opposite.

This one is a quarter-end routing review: what a defensible model shortlist looks like on 20 September 2026, by role, with the price, the licence and the catch for each. And the dates already on the calendar that will break it.


A new floor, a limit cut, and a notice taken back

DeepSeek V4.1 Flash shipped on 10 September. It is a multimodal mixture-of-experts model with a 552B backbone and an unusual shape: a 20-layer causal encoder feeding a 20-layer decoder, so that it activates 8B parameters per token while reading a prompt and 16B while writing. Context is 1M tokens. The LICENSE file is plain MIT. On DeepSeek's API the model is deepseek-flash, at $0.30 in and $1.20 out per million tokens at peak hours and half that off-peak. DeepSeek says the design cuts the KV cache to roughly a quarter of V4-Flash's, and states the reason plainly: "Cache-hit charges often account for a large share of agent costs."

The same announcement carried a second message. "We're phasing out V4-Pro. Starting at 04:00 UTC on Sept 14, 2026, all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates." That is four days' notice that a model ID would keep answering while a different model, with a third of the active parameters, did the work. DeepSeek's supporting claim was that "tests by multiple parties" put V4.1 Flash ahead of V4-Pro. It does not name the parties, and its own base-model table has V4-Pro ahead on SimpleQA-Verified (55.2 against 42.3), MATH (64.5 against 61.1) and LongBench-V2 (51.5 against 45.2).

Then it took the notice back. The change log entry for the same release now reads: "In response to user demand, we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged." When we read the pricing page this weekend, deepseek-v4-pro still pointed at DeepSeek-V4-Pro-0813 at its own price. We could not establish on which day the reversal was posted.

Anthropic's Monday was simpler and worse communicated. From 14 September, weekly limits in Claude Code on Pro, Max, Team and seat-based Enterprise plans are 25% above the original baseline. They replace a temporary 50% boost, so an index of 100 went to 150 and then to 125. Anthropic's first post on X said "we're permanently raising standard weekly limits in Claude Code by 25%". It was deleted. The clarification said: "Compared to today, this works out to a 17% reduction in weekly limits on Claude Code." This concerns subscription plans. Nothing we read says metered API limits changed, and we found no Anthropic-hosted page for the announcement, only BleepingComputer's report of 29 August.


The shortlist on 20 September, by role

Prices are US dollars per million tokens, input then output, standard tier, read from vendor pages this weekend.

  • Frontier reasoningClaude Fable 5.1 (claude-fable-5-1, released 1 September) and GPT-6 Astra (gpt-6-astra) both cost $10 / $50. Fable's cache reads fell to $0.25; Astra's cached input is $1.00. Anthropic's own guidance is to "start with Claude Opus 5 for most workloads", at $5 / $25, and move up only when "your evals on Claude Opus 5 at higher effort still fall short". Catch: five times the workhorse input price, and both are proprietary.proprietary · 1m context
  • WorkhorseClaude Sonnet 5 at $2 / $10, permanent since 10 August. GPT-5.6 Terra at $2 / $12. Gemini 3.8 Flash (gemini-3.8-flash, 2 September) at $0.75 / $3.75. Catch: the Flash price ends on 31 December 2026 and doubles on 1 January. Terra bills twice the input rate once a prompt passes 272K input tokens. Google's newest Pro is still gemini-3.1-pro-preview from February; its changelog has no Gemini 3.5 Pro entry between 1 July and 17 September.proprietary · check the end dates
  • Cheap routineGPT-5.6 Luna at $0.20 / $1.20. DeepSeek V4.1 Flash at $0.30 / $1.20 peak. Mistral Small 4 at $0.15 / $0.6. Claude Haiku 4.5 at $1 / $5 with a 200K context. Catch: Haiku 4.5 is now the dearest model in its own tier, and the only one here with a retirement floor this year.mixed licences
  • On-prem, permissive licenceMIT: DeepSeek V4.1 Flash and V4-Pro (1.6T total, 49B active). Apache 2.0: Mistral Small 4 (119B, 6.5B active, 256K context), Mistral Large 3 (675B), Qwen3.8-27B and Muse Glimmer 30B. We read each licence file. Mistral Small 4 is still the newest Small. Catch: the licence is clean and the hardware bill is not. Only Glimmer has a vendor-stated single-GPU footprint, 24 GB of VRAM at 1.0% degradation.apache 2.0 · mit
  • Speech-to-textGroq still lists whisper-large-v3-turbo as a production model at $0.04 per audio hour, with no deprecation entry. Mistral's Voxtral Mini Transcribe Realtime has Apache 2.0 weights and costs $0.006 per audio minute hosted, the same rate as OpenAI's gpt-4o-transcribe. Catch: a per-minute price of $0.006 is $0.36 per hour, nine times Groq's.groq · mistral · openai

The fourth row is short because most of what calls itself open does not belong on it.

Inkling, the subject of our July post, would join the left column once counsel has read its separate use policy. Three of the others deserve the exact words, because the marketing name hides them.

  • Mistral Medium 3.5A "Modified MIT License". You may not exercise any rights "if the global consolidated monthly revenue of your company (or that of your employer) exceeds $20 million" for the preceding month. Hosted, it is mistral-medium-3504 at $1.5 / $7.5. The weights are a dense 128B with a 256K context.monthly revenue gate
  • MiniMax M3The "MINIMAX COMMUNITY LICENSE" grants rights "for non-commercial purposes". Commercial use requires showing "Built with MiniMax M3" and sending a one-time notice, or "prior written authorization" above $20 million in yearly revenue. The prohibited-uses appendix includes "any military purpose". The hosted API costs $0.30 / $1.20 up to 512K input tokens.non-commercial grant
  • Qwen3.8 and Kimi K3Qwen3.8 is three licences: Apache 2.0 for the 27B, a $50M service-revenue gate for the 2.4T flagship, and no threshold at all for Flash-Next. Kimi K3's gate is $20M of group revenue over twelve months for anyone running it as a service. The Qwen clauses are quoted in the August post, Kimi K3's in the July one.name the repository

Retirement is an outage you can see coming

The quietest way for a production system to fail is for a model ID to stop existing. Anthropic's deprecations page is unambiguous about what retired means: "The model is no longer available for use. Requests to retired models will fail."

Three IDs crossed that line this year while still common in configuration files. claude-sonnet-4-20250514 and claude-opus-4-20250514 were deprecated on 14 April and retired on 15 June 2026. claude-opus-4-1-20250805 was deprecated on 5 June and retired on 5 August. Anthropic commits to "at least 60 days' notice before model retirement for publicly released models", sent by email and posted in the documentation. Both retirements kept to it, with 62 and 61 days.

Now the window ahead. The same table gives claude-sonnet-4-5-20250929 a retirement date "not sooner than September 29, 2026", claude-haiku-4-5-20251001 "not sooner than October 15, 2026" and claude-opus-4-5-20251101 not sooner than 24 November. Read that carefully, because the obvious reading is wrong. A floor is not a schedule. Today all three are listed as Active with no deprecation date. If Anthropic keeps its 60-day commitment, a notice posted today could not take effect before about 20 November. So Haiku 4.5 is not switching off on 15 October. It is the only model in Anthropic's current lineup whose floor falls this year, and the lineup table shows no successor for it.

A model ID is a dependency with an expiry date. Put it in an inventory and put the date in a calendar.

The control is dull and it works. Keep one inventory of every pinned model ID in every service, with the vendor's retirement date or floor beside it. Check it against the vendor's deprecations page on a calendar, monthly at least. Anthropic's Console will export usage as a CSV "broken down by API key and model", which finds the IDs nobody remembers pinning. Those dates cover Anthropic-operated platforms only; Amazon Bedrock and Google Cloud "set their own retirement schedules", so an ID reached through them needs its own row.

DeepSeek's notice was the other half of the same risk. There the ID survives and the model behind it changes: the legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are, in DeepSeek's words, "temporarily routed" to V4.1 Flash. An inventory of IDs will not catch that. A call log that records the model actually served will.


Our read

A shortlist is a snapshot and this one will be stale by the next review. That is the argument for writing it down. In one quarter Anthropic shipped Opus 5 and Fable 5.1, and Google shipped three Flash models in six weeks. A router pinned to a version string in June is behind on every tier, and nobody decided that.

What we would hold to at quarter end:

  1. Route by written policy, across closed and open models. Each role above has at least two vendors. A route with one candidate is a dependency you have not priced.
  2. Apache 2.0 or MIT for anything deployed on-prem. Six models clear that floor today. Name the repository, keep the licence file with the checkpoint, and do not let a family name stand in for either.
  3. Promote on your own eval set, never on the model card. DeepSeek's claim that V4.1 Flash beats V4-Pro sat next to a table of its own where it did not.
  4. Log every call with model, policy and retrieval set, and the model actually served. After this quarter that last field is what makes the log an audit trail.
  5. Treat residency as a design input. Anthropic's first-party API offers inference in us or global only, with workspace storage in us; an EU boundary means Bedrock or Google Cloud regional endpoints at a 10% premium. Mistral's default is that "your data is hosted in the European Union".
  6. Subscriptions are for people. Production runs on metered API. A seat limit can change by a sixth with about two weeks' notice and a deleted post.

The limit of this review is what we could not verify. We have left out EU region lists for Gemini, OpenAI, Groq and Alibaba, because we could not read a primary page for them, and we found no third-party EU host for most of the permissive-licence models. Both gaps matter to an EU buyer, and both need an answer from the vendor in writing before a route depends on them.

A quarter-end review like this is a half day of reading and a diff against the routing policy. Writing that policy, the eval gate behind it and the call log under it is what Sebrona does. If your shortlist still names a model from the spring, write to info@sebrona.com.


Reading

Every price, ID, licence name and date above comes from one of these pages, read this weekend unless another date is given.

  • DeepSeek V4.1 Flash announcement, change log and pricingRelease date, the V4-Pro phase-out notice, the reversal, model aliases, peak and off-peak prices.DeepSeek · 10 Sep 2026
  • deepseek-ai/DeepSeek-V4.1-Flash and DeepSeek-V4-Pro-0813MIT licence files, architecture, parameter counts and the vendor-reported base-model table.Hugging Face
  • Anthropic is cutting Claude Code's current weekly limits by 17 percentThe plans affected, the 100, 150, 125 arithmetic, the deleted post and the clarification.BleepingComputer · 29 Aug 2026
  • Claude models overview, pricing and data residencyModel IDs, prices, cache rates, context windows, the model-choice guidance, the inference_geo values and the regional premium.Anthropic documentation
  • Model deprecationsLifecycle definitions, the 60-day notice commitment, retired IDs with dates, retirement floors, and the usage export.Anthropic documentation
  • OpenAI pricing and model pages; staff post on SolIDs and prices for GPT-6 Astra and GPT-5.6, the long-context surcharge on Terra, the Sol promotion end date, transcription prices.OpenAI documentation · Developer Community 21 Aug 2026
  • Gemini API pricing, models and changelogGemini 3.8 Flash and 3.1 Pro IDs and prices, the introductory end date, and the absence of a Gemini 3.5 Pro entry.Google documentation
  • Mistral models overview, API pricing and licence filesSmall 4, Large 3 and Medium 3.5 IDs, prices and licences, the Modified MIT text, Voxtral pricing, and the data-hosting statement.Mistral · Hugging Face · help centre
  • Qwen3.8, Kimi K3, MiniMax M3 and Muse Glimmer licence filesThe licence names, thresholds and clauses quoted above, with parameter counts from the model cards.Hugging Face
  • MiniMax pay-as-you-go pricingThe M3 price line and its "Permanent 50% off" label.MiniMax documentation
  • Groq models and deprecationsProduction status and price of the Whisper models.Groq documentation