Skip to main content
ENSK
PracticeElektrikProJARVISBlog
← All posts

GPT-6 Astra, One Week In: Your Agent Can Now Be Stopped by the Vendor

GPT-6 Astra costs $10 / $50, drops temperature, top_p and logprobs, needs the Responses API for tools, and can end an API task with a 403 and no resume. Its Critical cyber rating leaves default access completing 44.4% of patching tasks. The 99.9% ARC-AGI-3 score is 62.7% on the standard harness, and in week one Azure listed no EU Data Zone while Bedrock's Mantle endpoint was Oregon only.

TopicField notes
Published13 Sept 2026
AuthorMiroslav Striško
Reading10 min

This week teams had their first full working week with GPT-6 Astra. OpenAI announced the model on Thursday 3 September to "a limited set of organizations" and opened the API and paid ChatGPT plans the next day. It costs $10 / $50 per million tokens. It also ships with a failure mode that is new with this model: the vendor can stop your agent mid-run, and the API offers no way to resume.

The launch headlines were about AGI. For anyone who builds, the changes that matter sit in the API reference and the system card, so that is where this post starts. The benchmark claims and the AGI remark come later, at the length they have earned.


What changes in your code, and in your harness

Start with the reference page, because several of these changes break existing integrations.

  • Price$10 input · $50 output per million tokens. Above 272K input tokens the whole request is billed at $20 / $75. Cached input $1.00, cache writes $12.50, Batch and Flex at half.standard tier · 3 Sep 2026
  • Model IDgpt-6-astra, with no dated snapshot. Context window 1,050,000 tokens, 922,000 in, 128,000 out. Knowledge cutoff 30 April 2026.api docs
  • Reasoning effortlow, medium, high, xhigh, max. The none level that GPT-5.6 offered is gone.breaking
  • SamplingAstra "does not support custom temperature or top_p values or log probabilities (logprobs)".breaking
  • Tools"Tool calling requires the Responses API." Chat Completions still works, without tools.breaking

That list price matches Claude Fable 5 and Fable 5.1, and Artificial Analysis notes it is "2.5x GPT-5.6 Sol's current prices across the board".

The bigger change is operational. Astra ships with what OpenAI calls misalignment monitoring, an asynchronous system that reviews the model's reasoning and actions and "can stop a conversation when it identifies a potential issue". In ChatGPT and Codex a paused task can be reviewed and continued. In the API, per the launch post, "the task will stop." The developer guide spells out what that looks like:

  • The errorHTTP 403, type invalid_request_error, code misalignment_policy_violation.match the code, not the message
  • No retry"Do not automatically retry the blocked workflow."openai guidance
  • No resume"The API does not provide a general way to resume a conversation stopped by misalignment monitoring."openai guidance
  • No undo"A stopped request does not undo earlier actions." Monitoring is asynchronous, so a tool call may already have run.openai guidance

Coverage depends on the surface. Responses API requests that use persisted reasoning, WebSockets or OpenAI compaction are monitored and can be stopped automatically. Responses requests without those are monitored, but only produce webhook alerts. Chat Completions is "Not covered by this monitoring system". Alerts arrive on a safety.alert.created webhook, the reason field "can be null, including for Zero Data Retention (ZDR) requests", and OpenAI states that "Alert delivery and retrieval do not provide a complete audit history."

Your agent can now be stopped by someone else's classifier, halfway through a job, with no resume. Design for it the way you design for a crashed process.

OpenAI is open that this will hit legitimate work: the system "may occasionally flag legitimate activity", including "work that does not appear directly related to cybersecurity or tasks in which an agent is running for an extended period." How often is unknown. The public evidence from the first week is three GitHub issues on the Codex repository, and no rate has been published.

For a harness this means four things. Checkpoint task state outside the model session. Make each step idempotent, so a rerun after a stop cannot act twice. Keep a fallback route to a model that has passed the same eval. Add an operator queue, because the vendor's own guidance is to put a person in the loop before anything restarts.


"Critical" means most of it is switched off for you

On 1 September OpenAI designated Astra as meeting "the Critical cybersecurity capability threshold under our Preparedness Framework", adding: "It is the first model we are designating at this level". In plain terms, "with the right tools and access, it can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step."

The launch post states the consequence for ordinary users: "Astra will refuse to comply with more advanced cybersecurity tasks such as creating proof-of-concept exploits for vulnerabilities." Defenders, it says, can still use it for "secure code review and patching". The system card's own table shows how much of that survives the safeguards.

Read the second row twice. Without trusted access, Astra completes 44.4% of vulnerability-patching tasks and 66.7% of discovery and analysis. Proof-of-concept creation sits at 2.4%. These are OpenAI's numbers. The safeguards take a large share of defensive work down with the offensive.

The way up is identity. OpenAI's Trusted Access for Cyber, "also known as Daybreak access", takes applications from organisations and identity verification from individuals, who "must enable Advanced Account Security to use Daybreak Blue". OpenAI is also "taking additional steps to restrict access by high-risk entities and in high-risk jurisdictions." At launch, Astra's advanced cyber work went to "a small group of alpha testers, with access through Daybreak Blue expanding afterward". ChatGPT Enterprise had its own gate. OpenAI's enterprise docs say "Astra is off by default for ChatGPT Enterprise for the first two weeks after launch", and during the initial rollout an organisation needed Daybreak access before an admin could enable it.


99.9% on whose harness

OpenAI's launch post says Astra "saturates ARC-AGI-3 with a 99.9% score". A footnote adds that it was run with "our responses API harness". ARC Prize published both numbers the same day: "62.7% for $26K on ARC-AGI-3 Semi-Private with our Standard harness, and 99.9% for $19K with a Provider Adapter harness". The second harness "preserves opaque reasoning state between requests and uses compaction for longer conversations".

Both scores are state of the art, and neither is dishonest. The 37-point gap is the harness, and that harness is a feature of OpenAI's own API surface. So the higher number is a fact about the product you would be locked into as much as about the model.

The same hygiene applies to OpenAI's own results table, which is more candid than the prose above it. On the Artificial Analysis Intelligence Index v4.1.1 it lists Astra at 61.2, behind Claude Fable 5.1 at 65.7 and Opus 5 at 63.1. On Humanity's Last Exam with tools, Astra scores 57.2% against Fable 5.1's 65.0%. All vendor-reported, by the vendor that lost those rows.

Artificial Analysis then published its own write-up on 9 September. There Astra at max effort scores 53, tied with Fable 5.1, at "~40% of the cost per task ($3.26 vs $7.63)". The same piece records a drop of about 45 Elo against Sol on its GDPval-AA v2 benchmark, and puts Astra at max effort about 60% more expensive per task than Sol. OpenAI's table cites index v4.1.1; the 9 September write-up prints no version. Same evaluator, six days apart, two different pictures. Quote the version whenever you quote the score.

As for the headline most outlets ran: at the 3 September press briefing, OpenAI president Greg Brockman said, as reported by The Verge, "If we fast-forward a couple of years, and we look back and say, 'When was it, really, that AGI was created?' I think it's going to be about this time, and I think it might be about this model." We have nothing to add to that beyond the four paragraphs above it.


Week one from inside the EU

For a workload that has to stay in the EU, the launch-week picture was thin.

  • Microsoft FoundryAnnounced 3 Sep as "now generally available for all customers", "in both Global and US Data Zone geographies". No EU Data Zone is listed. US Data Zone pricing is $11 / $55.azure blog · 3 Sep 2026
  • Amazon BedrockGenerally available 8 Sep. OpenAI's Bedrock guide lists the Mantle endpoint as "Available in us-west-2 (Oregon)". We could not verify Bedrock Runtime regions. The same guide warns: "An AWS Region is not an OpenAI data residency jurisdiction."aws blog · openai docs
  • OpenAI directOpenAI's data-residency docs list gpt-6-astra for regional processing in "Europe (EEA + Switzerland)" via eu.api.openai.com. It needs eligibility through sales, approval for abuse-monitoring controls and a Modified Retention amendment, and carries a 10% uplift. "Fast mode is unavailable for GPT-6 Astra with EU data residency."openai docs

Through the two clouds named in the launch post, then, we found no documented in-region option for the newest frontier model in week one. The direct route exists on paper behind a contract gate, and we could not verify that it was live on 3 or 4 September. One point in Astra's favour for a DPO: it "supports Zero Data Retention for eligible API customers", while Anthropic's release notes say Fable 5.1 requires 30-day data retention unless Anthropic authorises otherwise.

The rollout itself was bumpy. The New Stack quoted Sam Altman opening a post with "First, sorry for the messy rollout", then on 4 September announcing API access. By 6:30 p.m. Eastern that day, Plus and Business users had the model too. On 10 September TechCrunch reported that OpenAI had paused new sign-ups to its $200-a-month Pro plan, citing demand for Astra. The status page for 3 to 13 September shows no incident that names Astra, though "Elevated errors for ChatGPT users in Europe" appears on both 11 and 13 September.

Ten days in, capacity is the first visible constraint, ahead of safety. The promised loosening, "less restrictive safeguards in the coming weeks" through Daybreak, has no date yet.


Our read

Astra makes the vendor's safety system part of your runtime. That is a reasonable response to a model its maker rates Critical. It also moves a class of failure from "our bug" to "their decision", and the API contract is explicit that the decision is final for that conversation.

What we would settle before promoting it into a production route:

  1. A stop is handled like a crash. Durable checkpoints, idempotent steps, a fallback model that passed the same eval set, and a human queue for anything that returns misalignment_policy_violation.
  2. Your log is the audit record. OpenAI says its alerts are not a complete history. Log every call with model, policy version, retrieval set and tool actions on your side.
  3. Promotion is per workload. Astra lists at 2.5× Sol and uses fewer tokens. Artificial Analysis recorded at least one regression against Sol. Only your own eval set says which side a given route falls on.
  4. Anything built on logprobs or temperature needs a plan. Confidence scoring from log probabilities stops working, and an undated snapshot means drift shows up only in your evals.
  5. Data residency is settled before design, contract in hand. Which route gives in-region processing, under which amendment, at what uplift, and without which features.

The honest limit: nobody outside OpenAI knows the false-positive rate of the new monitor. Three GitHub issues are an anecdote. If your agents run for hours, measure it on your own traffic before you depend on it.

Sebrona designs agent harnesses that survive a vendor-side stop, with checkpointing, fallback routing and call logging an auditor can read. If Astra is heading for your production path, write to info@sebrona.com.


Reading

Every figure and quote above traces to one of these. Vendor-reported results are labelled in the text.

  • GPT-6 Astra launch postAvailability wording, price, the refusal of advanced cyber tasks, "the task will stop", Zero Data Retention, and the full results table including the Artificial Analysis and Humanity's Last Exam rows. openai.com/index/gpt-6-astraOpenAI · 3 Sep 2026 · vendor-reported benchmarks
  • OpenAI API docsModel page, pricing tiers, misalignment-monitoring guide, data-residency table, Amazon Bedrock guide, changelog, and the Enterprise availability page on learn.chatgpt.com. developers.openai.comOpenAI
  • Path to AstraThe Critical designation, its plain-language meaning, the false-positive admission, the alpha-tester and Daybreak Blue sequence. openai.com/index/path-to-astraOpenAI · 1 Sep 2026
  • GPT-6 Astra system cardTable 21 completion rates and the Trusted Access for Cyber terms. deploymentsafety.openai.comOpenAI · 3 Sep 2026 · vendor-reported
  • ARC PrizeARC-AGI-3 results on both harnesses with costs, and the description of the Provider Adapter harness. arcprize.org/blog/astraARC Prize · 3 Sep 2026 · independent
  • Artificial AnalysisIndex scores, cost per task, the GDPval-AA regression, the "2.5x" price comparison. artificialanalysis.aiArtificial Analysis · 9 Sep 2026 · independent
  • Microsoft and AWSFoundry geographies and US Data Zone pricing; Bedrock general availability on 8 September. azure.microsoft.com · aws.amazon.comMicrosoft · 3 Sep 2026 · AWS · 8 Sep 2026
  • PressThe Verge for Brockman's remark; The New Stack for the 4 September rollout; TechCrunch for the Pro sign-up pause. theverge.com3–10 Sep 2026
  • OtherAnthropic release notes for Fable 5.1 pricing and retention; OpenAI status history for 3 to 13 September; Codex repository issues 43042, 43208 and 43781. github.com/openai/codexAnthropic · OpenAI · GitHub · Sep 2026