Skip to main content
ENSK
PracticeElektrikProJARVISBlog
← All posts

A Promotional Price Is Not a Price: OpenAI's Training Pause and the GPT-5.6 Sol Cut

This week OpenAI disclosed a two-week pause in RL training, with monitoring that costs roughly 20% of the compute it watches, then cut GPT-5.6 Sol to $4 / $20 "at least through November 21, 2026". Budget on list price, treat the promotion as upside, and put a price ceiling per route in the routing policy so a reprice is a config change.

TopicField notes
Published23 Aug 2026
AuthorMiroslav Striško
Reading8 min

This week OpenAI did two things that belong in the same budget meeting. On Tuesday 18 August it published a post on pacing model development that disclosed a two-week pause in reinforcement-learning training and a monitoring system costing roughly 20% of the compute it watches. On Friday 21 August it cut GPT-5.6 Sol from $5 / $30 to $4 / $20 per million tokens, "at least through November 21, 2026".

Both announcements hand a buyer a number. One is a price with an end date on it. The other is the first public figure we know of for what agent oversight costs to run. Each deserves a line in a business case, and each is easy to misread.


A two-week pause, described in the past tense

OpenAI gave two reasons for slowing down: the Hugging Face incident, and "preliminary evidence that one of our upcoming models, Astra, may meet the Critical cybersecurity capability threshold under our Preparedness Framework". The response, in its words: "This included a two-week pause in reinforcement learning (RL) training on our latest models intended for deployment while we further hardened and red-teamed our research environments and expanded the coverage of our monitoring systems."

Two more sentences set the scope. "Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations". And for the unreleased model: "a significant number of workloads remain paused until they are fully migrated and enhanced to meet the new security bar."

Read that carefully and it is narrower than the headlines were. OpenAI did not pause all training, or all reinforcement learning. Sam Altman's post on X the same day, quoted by The New Stack, said "We have paused some frontier RL training", and a follow-up added: "We still expect to ship great new models soon; this impacts further-out releases."

The honest limit sits here. The start and end dates of the pause have not been published. OpenAI's post uses the past tense. GovInfoSecurity wrote a day later that OpenAI "will enter" a pause and quoted a spokesperson saying it "has already started and is ongoing". The same outlet noted that "OpenAI has not said why it chose a two-week period". Nobody outside the company can say which fourteen days these were.

That gap is what the critics went for. Max Tegmark of the Future of Life Institute called the move "a step in the right direction", then added: "a voluntary pause that the U.S. government can neither verify nor enforce isn't enough." Researcher Nathan Lambert wrote that "we should have independent organizations that can access the full details of these training runs for monitoring." A CISO will recognise the shape of the complaint. It is the question an auditor asks about your own agent sandbox: what evidence, held by whom, shows the control exists?


Twenty percent: a first public price for agent oversight

The most reusable sentence in the post is about cost. "Our current estimates put monitoring overhead at roughly 20% of the inference compute being monitored, though the cost varies substantially across training and evaluation workloads."

OpenAI describes what that buys. Activation classifiers, detectors that read the model's internal activity, "run at every sampled token" and escalate concerns to heavier automated investigators. Coverage is "required for all RL training and evaluations involving tools for models of Sol capability or higher", and from 7 August for all inference of Astra with tools. The response rule is specific: "We aim to issue an alert within 30 minutes", and if the paged teams "cannot conclusively determine within 30 minutes that the flag is a false positive, those teams are expected to pause the activity."

  • OverheadRoughly 20% of the inference compute being monitored. An estimate, and one that "varies substantially" by workload.self-reported · 18 Aug 2026
  • Alert targetWithin 30 minutes of concerning activity surfacing.detection
  • Default on doubtIf humans cannot clear the flag as a false positive in a further 30 minutes, the activity is paused.response

Treat the figure with care. It comes from a lab's training and evaluation clusters, which look nothing like a production agent answering tickets. It is self-reported, and OpenAI promises detail in a later post.

Still, anyone sizing oversight for their own agents has had no public reference point until now. This one sets an order of magnitude. Oversight that reads reasoning and actions at every step costs about a fifth on top; it does not vanish into the compute bill. Logging each call with its model, policy and retrieval set is the cheap part. Reading those logs fast enough to stop something is the part with a price, and the 30-minute rule is a pattern worth copying: when a person cannot clear a flag in time, the default is stop.


$5 / $30 became $4 / $20, with a date attached

Three days later the API changelog carried one line: "GPT-5.6 Sol now costs $4 per million input tokens and $20 per million output tokens, representing 20% lower input pricing and 33% lower output pricing. GPT-5.6 Sol's promotional pricing is available at least through November 21, 2026." Cached input moved from $0.50 to $0.40, and cache writes now cost $5.00.

The stated reason is one clause long. A staff post on the OpenAI developer forum reads: "As we continue to push the frontier of capabilities while improving efficiency, we're dropping API and credit pricing of GPT-5.6 Sol by over 20% for the next 3 months." The cut also covers Fast mode, long-context requests, Batch and Flex, plus ChatGPT Work and Codex credits, while "Pro, Plus, and Business subscription usage remains unchanged." That is everything OpenAI said. Several write-ups framed the move as an answer to a competitor's price list. We found no OpenAI statement to that effect, so we do not repeat it as fact.

What the market looked like that day, with each figure tied to its publisher:

  • GPT-5.6 Sol$4 input · $20 output. Promotional, at least through 21 Nov 2026. List price $5 / $30.OpenAI · gpt-5.6-sol
  • Claude Opus 5$5 input · $25 output. Standard price since launch on 24 Jul.Anthropic · claude-opus-5
  • Claude Sonnet 5$2 input · $10 output. Introductory until Anthropic made it the standard price on 10 Aug, cancelling a rise to $3 / $15 planned for 1 Sep.Anthropic · claude-sonnet-5
  • Claude Fable 5$10 input · $50 output. Standard price since launch on 9 Jun.Anthropic · claude-fable-5
  • Gemini 3.7 Flash$0.75 input · $3.75 output. Introductory through 31 Dec 2026, then $1.50 / $7.50 from 1 Jan 2027. Launched 13 Aug.Google · introductory

Prices as published and in force on 21 August. The rest of the GPT-5.6 family sat lower: Terra at $2 / $12 and Luna at $0.20 / $1.20, both permanent since 30 July.

By our arithmetic on those published prices, Sol went from level with Opus 5 on input and 20% dearer on output to 20% cheaper on both. A per-token table hides one asymmetry. OpenAI charges 2× on input and 1.5× on output above 272K input tokens, while Anthropic says its current models "include the full 1M token context window at standard pricing".

One more change shipped the same Friday and drew almost no attention. "API customers can now select regional processing for an individual request by using a prefixed domain with an API key from a project having Global geography." For an EU team that is the more durable of the two announcements.


Our read

A promotional price with an end date is a budgeting trap, and the trap is set by the spreadsheet rather than the vendor. Someone types $20 into the output column in August, the project is approved on that figure, and in late November the unit cost rises by half with no change in the system.

Write the list price into the business case. The promotion is upside, and it expires on a date you did not choose.

What we would put in place before leaning on a rate like this:

  1. Budget on list, report on actual. For Sol the planning number is $5 / $30 until OpenAI says otherwise. For Gemini 3.7 Flash it is $1.50 / $7.50 for anything running into 2027. Put 21 November in the calendar as a re-baseline date.
  2. A price ceiling per route, written into the routing policy. When a model's effective price crosses the ceiling, the route falls to the next model that passed its eval. A reprice then becomes a reviewed config change, in either direction, instead of a migration project.
  3. Log the policy version with every call. The ID gpt-5.6-sol has no dated snapshot and its price has now changed under the same string. The call log is the only place the two line up.
  4. Let the eval decide among near-equal prices. At $4 / $20 against $5 / $25, per-token price no longer separates the candidates. Cost per completed task on your own eval set does.
  5. Give oversight its own budget line. A lab that builds these models puts monitoring at roughly a fifth of monitored compute. Planning for zero is a choice, and an auditor will ask who made it.

The counter-argument deserves a hearing. Every repricing in this post went down, and the one planned rise was cancelled, so budgeting on list may overstate cost. We would still take that error over the other one. An overestimate returns money. An underestimate returns a conversation with finance in December.

Sebrona writes routing policies with price ceilings, builds the eval gates behind them, and sets up call logging that survives a reprice. If your 2027 budget has a promotional rate in it, write to info@sebrona.com.


Reading

Where each figure and quote comes from. Percentages we computed ourselves are marked as our arithmetic.

  • Pacing model developmentScope of the pause, the run on hold, the Astra workloads, monitoring design, the 30-minute rule, the 20% overhead estimate. openai.comOpenAI · 18 Aug 2026
  • OpenAI API changelog and pricingThe 21 August Sol entry, cached and cache-write rates, long-context rule, per-request regional processing. developers.openai.comOpenAI · 21 Aug 2026
  • OpenAI developer forumStaff announcement with the one-clause reason and the scope of the cut. community.openai.comOpenAI staff post · 21 Aug 2026
  • Anthropic pricing and release notesOpus 5, Sonnet 5 and Fable 5 prices; launch dates of 9 Jun and 24 Jul; the 10 August Sonnet 5 entry; the 1M-context pricing statement. platform.claude.comAnthropic
  • Gemini 3.7 Flash pricingIntroductory and 2027 rates from Google's pricing page; launch date and pricing confirmed by VentureBeat. ai.google.devGoogle · VentureBeat · 13 Aug 2026
  • GovInfoSecurityTegmark and Lambert quotes, the spokesperson's "already started and is ongoing", the unanswered two-week question. govinfosecurity.comISMG · 19 Aug 2026
  • The New StackText of Sam Altman's two X posts of 18 August. thenewstack.ioThe New Stack · Paul Sawers · 19 Aug 2026