Skip to main content
ENSK
PracticeElektrikProJARVISBlog
← All posts

GPT-5.6 Luna, Terra and Sol: The Vendor Handed You a Routing Table

OpenAI made GPT-5.6 generally available on 9 July 2026 as three models priced from $1 / $6 to $5 / $30, thirteen days after a limited preview it started at the US government's request. The headline index win was measured by an evaluator that says it supported OpenAI before release, and OpenAI's own table has Sol behind Fable 5 on SWE-Bench Pro. The number that decides routing is cost per completed task on your own eval set.

TopicField notes
Published12 Jul 2026
AuthorMiroslav Striško
Reading7 min

This week OpenAI made GPT-5.6 generally available: three models, Luna, Terra and Sol, priced from $1 / $6 to $5 / $30 per million tokens. The release landed on Thursday 9 July, thirteen days after a limited preview that OpenAI said it had started at the US government's request. Most coverage led with a benchmark win. We think the price list was the more useful document.

What the vendor handed buyers is a three-row routing table, and the number that decides which row a task belongs in is one no vendor can publish for you.


Thirteen days behind a request nobody has documented

On Friday 26 June, OpenAI published a preview of GPT-5.6 and said plainly that it was not a normal launch. "At their request, we are starting with a limited preview for a small group of trusted partners whose participation has been shared with the government, before releasing more broadly." Access ran through the API and Codex only.

OpenAI objected in the same post: "We don't believe this kind of government access process should become the long-term default. It keeps the best tools from users, developers, enterprises, cyber defenders, and global partners who need them."

CNBC reported on 8 July that the public release would come "roughly two weeks" after the limited rollout. It added context: an AI executive order signed in June asks developers to voluntarily provide new models to the government ahead of a full release, and gave federal agencies 60 days to develop an evaluation process. Axios quoted Sam Altman the next day saying OpenAI made "many changes" after a "collaborative back and forth" with the administration.

That is the whole public record, and its edges matter. Neither OpenAI nor CNBC names the agency that made the request. Neither names the legal instrument behind it. OpenAI's wording describes a request, not an order. We found no government document that says what was reviewed or what approval consisted of, so we will not describe it.

For an EU buyer the practical reading does not depend on those details. For thirteen days a frontier model existed, had a price, and could be called only by organisations whose names had been shared with a foreign government. Day-one availability of a US model is an assumption your roadmap makes. Write it down as one.


Three models, one spec sheet, three prices

On paper the three models are the same machine in three sizes. Each carries a 1,050,000-token context window, a 922,000-token input ceiling, 128,000 output tokens, a knowledge cutoff of 16 February 2026, and the same reasoning.effort ladder: none, low, medium, high, xhigh, max. OpenAI's docs map them onto the old naming. Sol "roughly corresponds to the unsuffixed model tier", Terra to the mini tier, Luna to nano.

  • gpt-5.6-luna$1 input · $6 output per million tokens at launch.nano tier · 9 Jul 2026
  • gpt-5.6-terra$2.50 input · $15 output per million tokens at launch.mini tier · 9 Jul 2026
  • gpt-5.6-sol$5 input · $30 output per million tokens at launch. The gpt-5.6 alias routes here.flagship · 9 Jul 2026

Three pricing rules shipped with the family, and they move a cost model more than the headline rates do. Cache reads keep their 90% discount, but cache writes are now billed at 1.25× the uncached input rate, with a 30-minute minimum cache life. Artificial Analysis called these the "First OpenAI models with cache-write pricing". Prompts above 272K input tokens pay 2× on input and 1.5× on output for the whole request. Batch and Flex run at 50% of Standard.

One detail for anyone who pins versions: these IDs have no dated snapshot. gpt-5.6-sol is both the model ID and the snapshot name.

ChatGPT Work shipped the same day, described by OpenAI as "an agent in ChatGPT" that gathers information across apps and produces "finished materials like sheets, slides, docs, and web apps". It is the end-user product. For anyone building, the API family is the part that matters.


The index win, and the row beside it

OpenAI's headline claim rested on the Artificial Analysis Coding Agent Index. Sol at max effort scores 80, which is 2.8 points above Claude Fable 5 at 77.2, "while using less than half the output tokens, taking less than half the time, and costing about one-third less."

Two things belong next to that number. The first is a disclosure. Artificial Analysis is a third party, and its own write-up says: "We supported OpenAI with pre-release evaluation of GPT-5.6 Sol, Terra, and Luna." The index pairs each model with an agent harness, so Sol ran in Codex and the Claude models ran in Claude Code. That is a reasonable way to compare products. It is also a vendor-assisted measurement published on launch day, and the two parties do not quite agree on it: Artificial Analysis puts Sol's per-task cost at "~40%" below Fable 5, OpenAI at "about one-third".

The second is in OpenAI's own results table, in the same post. On SWE-Bench Pro, Sol scores 64.6%. Claude Fable 5 scores 80% and Opus 4.8 69.2%. On Toolathlon, Sol reaches 58% against Fable 5's 61.7%. All of these are vendor-reported, and OpenAI deserves credit for printing rows it lost.

A composite index tells you who won the average. Your workload is one row, and it may be the row the winner lost.

The more useful finding in the Artificial Analysis piece concerns the middle tier. "Luna and Sol are always on the Pareto frontier ahead of Terra. This means that for any Terra effort level, there is a Luna or Sol effort level that is more intelligent at no extra cost, or equally intelligent at lower cost." On its Intelligence Index a task cost $1.04 on Sol, $0.55 on Terra and $0.21 on Luna. By that measurement the routing table had two live rows at launch.

The rows are also less alike than the spec sheet suggests. All three list the same context window, yet on OpenAI's own long-context test (MRCR v2, 8-needle, 512K to 1M tokens) Sol scores 73.8%, Terra 72.5% and Luna 41.3%. Same window, different ability to use it.

So the number that matters is on none of these pages. It is cost per completed task on your own eval set: what a finished, correct task costs on Luna at high effort against Sol at low, counting retries and failures. Per-token price is one input to that figure. Six effort levels across three models make an eighteen-cell grid, and no published benchmark says which cell your contract-review route or your ticket-triage route belongs in.


Our read

A three-tier family on one API surface, with one context window, is a vendor conceding the routing argument. No single model is the right answer for every call, and the gap between rows is wide enough to pay for the work of choosing. Sol's input price is five times Luna's, and its output price five times too.

What we would check before putting this family into a production route:

  1. The tiers live in a written routing policy, one task class per row. A model earns a row when your own eval set passes on it. The model card does not count, and neither does a launch-day index the vendor helped prepare.
  2. Every call is logged with model, effort level and policy version. With undated IDs and prices the vendor can move at any time, that log is the only record of what ran and what it cost.
  3. The price in the policy is a ceiling per route. When the vendor reprices, a rule fires and someone reviews it. Nobody finds out from the invoice.
  4. The availability date is treated as a risk. A route that can fall back to a second provider, or to an open-weight model under an Apache-2.0 or MIT licence on hardware you control, turns a thirteen-day gate into a delay.

The honest limit of this post: we do not know which agency asked for the preview limit or under what instrument, and nobody relying on the public record does either. We also could not verify which EU data-residency options exist for these models, so we make no claim about them here.

Sebrona writes routing policies, eval gates and call logging for teams that run production AI inside the EU. If a three-row price list is about to become a line in your budget, write to info@sebrona.com.


Reading

Every figure above comes from one of these pages. Vendor-reported numbers are marked as such in the text.

  • GPT-5.6 launch postLaunch prices, cache rules, availability, the full results table including SWE-Bench Pro, Toolathlon and MRCR,. openai.com/index/gpt-5-6OpenAI · 9 Jul 2026 · vendor-reported benchmarks
  • Previewing GPT-5.6 SolThe limited preview, the "At their request" wording, and OpenAI's objection to the process. openai.com/index/previewing-gpt-5-6-solOpenAI · 26 Jun 2026
  • OpenAI API docsModel IDs, context window, input and output limits, knowledge cutoff, effort levels, tier mapping, long-context and Batch pricing rules, changelog dates. developers.openai.comOpenAI
  • Artificial AnalysisCoding Agent Index score of 80, the pre-release evaluation disclosure, cost per task on Sol, Terra and Luna, and the Pareto finding on Terra. artificialanalysis.aiArtificial Analysis · 9 Jul 2026
  • CNBCRelease timing, "roughly two weeks", and the description of the June executive order. cnbc.comCNBC · Ashley Capoot · 8 Jul 2026
  • AxiosAltman's "many changes" and "collaborative back and forth" remarks. axios.comAxios · 9 Jul 2026
  • ChatGPT Work announcementWhat the product is and how OpenAI describes it. openai.comOpenAI · 9 Jul 2026