Skip to main content

GPT-6 Astra: pricing, benchmarks and OpenAI safety rating

Camila Duarte
Camila DuarteSeptember 9, 202613 min. read
GPT-6 Astra: pricing, benchmarks and OpenAI safety rating

What is GPT-6 Astra

OpenAI launched GPT-6 Astra on September 3, 2026, at US$ 10 per million input tokens and US$ 50 per million output tokens, with a Critical cybersecurity rating and monitorability worse than its predecessor's, per the launch announcement. Anyone consuming the API in production decides when to use it with three variables: capability, cost per token, and governance.

The most common reading of the launch is that OpenAI took back the frontier lead. The independent index does not confirm it: on the 2026-09-09 reading from Artificial Analysis, GPT-6 Astra scores 61.2 on the AA Intelligence Index v4.1.1, and Fable 5.1 stays ahead with 65.7. The two facts coexist. The distance between them is where this month's routing decision lives.

What was launched

OpenAI began the rollout of GPT-6 Astra on September 3, 2026, to a limited set of organizations, with access for ChatGPT Plus, Pro, Business, and Enterprise subscribers arriving in the following days, as recorded by CNBC on September 3. On OpenAI's own API, the model answers to the identifier gpt-6-astra. Azure and Bedrock are announcement, not availability.

The announcement lists three channels, and it pays to read them separately. The OpenAI API is accessible, with the model answering as gpt-6-astra for API customers. Microsoft Azure and AWS Bedrock appear as planned channels, with no date on the launch page as of publication, which changes the purchasing plan of anyone anchoring consumption in the cloud. The channel defines the bill, and the bill does not exist before the channel opens.

On Enterprise accounts, access ships off by default at launch and depends on the administrator enabling the model for the workspace.

The most recent episode on the OpenAI side reinforces the channel point: the company cut Cursor's access to the model portfolio, and this post detailed in OpenAI cuts off Cursor: the single-vendor risk what the cutoff meant for anyone integrating directly. Model access is a commercial relationship, not a perpetual guarantee.

How much GPT-6 Astra costs

GPT-6 Astra costs US$ 10 per million input tokens and US$ 50 per million output tokens in Standard mode, figures published by OpenAI on September 3, 2026. Fast mode charges 2x the Standard price: US$ 20 and US$ 100 per million tokens, respectively.

The asymmetry between input and output deserves attention before integration: it is 5x; in conversational workloads input dominates, and in agentic workloads the model writes more than it reads. The output tariff decides the invoice.

The arithmetic of the published prices works like this: a workload that consumes 10 million input tokens and 2 million output tokens per day pays US$ 200 per day in Standard mode, US$ 100 on each side, and US$ 400 per day in Fast mode, which closes the month at US$ 6,000 or US$ 12,000.

Fast doubles the bill with the speed.

This cost-per-token accounting is not new to the routing leaderboard: the Qwen 3.8 Max launch, with 2.4 trillion parameters in MoE and 1 million of context, already put model size inside the invoice, and this post analyzed that cost in Qwen 3.8 Max: context cost ($2/$6) is now a routing decision. What GPT-6 Astra adds is the pair of tariffs: an entry Standard price and a fast mode that charges double per token, on the same key.

inline-01.png

The GPT-6 Astra scoreboard and the split verdict

The GPT-6 Astra scoreboard is split. The benchmarks published by OpenAI on September 3, 2026 are self-assessment: FrontierMath Tier 4 at 97.6%, ARC-AGI-3 at 99.9%, ARC-AGI-2 at 95.0%, GPQA Diamond at 96.0%, and Terminal-Bench 4.0 at 57.9%. On the independent index from Artificial Analysis, the 2026-09-09 reading, the model scores 61.2, below Fable 5.1's 65.7.

Self-assessment and third parties tell different stories. Start with the comparable number. On Terminal-Bench 4.0, which measures terminal work and software engineering, GPT-6 Astra scored 57.9% against 37.3% for GPT-5.6 Sol and 55.8% for Fable 5.1, numbers published by OpenAI on September 3, 2026. On abstract reasoning and science tests, the self-assessment reports 99.9% on ARC-AGI-3, 95.0% on ARC-AGI-2, 96.0% on GPQA Diamond, and 97.6% on FrontierMath Tier 4, all self-assessment figures published in the same September 3, 2026 package.

The independent countercheck is the AA Intelligence Index v4.1.1, the same index Artificial Analysis rebased a short while ago, and this post explained in AI model ranking: what the rebase changes for model choice what that rebase changed in model choice. On the 2026-09-09 reading, GPT-6 Astra enters at 61.2 and Fable 5.1 stays at 65.7. Astra is not the only September newcomer: Quasar 438B, the European entrant, entered the same routing leaderboard, and this post covered the entry in Quasar 438B: the European 438B model enters the routing leaderboard.

The reading that best closes the week did not come from OpenAI. Every published the analysis A Split Verdict on Fable vs. Astra on September 6, 2026, with a direct thesis: a model can impress on output and still frustrate as a collaborator, and choosing between Fable 5.1 and GPT-6 Astra requires specifying what the model needs to do and how the team wants to work with it. The split verdict describes the real scoreboard: maximum capability in specific packages, index leadership somewhere else.

The right position facing this week is not picking the scoreboard winner. It is treating the distance between self-assessment and the third-party index as the decision margin of whoever routes: with Fable 5.1 4.5 points ahead on the index, GPT-6 Astra enters the main route when specific capability, price, and governance pay the difference, and not because the launch was big.

inline-02.png

Safety: a Critical rating, two zero-days, and worse monitorability

GPT-6 Astra is the first OpenAI model to reach the Critical cybersecurity threshold of the Preparedness Framework, in a self-assessment dated September 1, 2026, in the Path to Astra update. In the evaluation, the model discovered two unknown zero-days, under disclosure to the maintainers. OpenAI itself declares the reasoning monitorability worse than GPT-5.6 Sol's.

OpenAI's most capable model is also the most demanding to govern.

The cybersecurity numbers published by OpenAI on September 3, 2026 give the jump its scale: 100% on ExploitBench, which measures the conversion of known vulnerabilities into working exploits, against 78.5% for GPT-5.6 Sol, and 0.0% on the ExploitGym honeypot against the predecessor's 48.2%. The honeypot measures how often the model exceeds the authorized target when facing an impossible task, and lower is better. These are self-assessment numbers inside the Preparedness Framework itself, and the detail that gives them weight is that two real zero-days appeared on the evaluation path.

Monitorability is the uncomfortable point of the announcement. OpenAI acknowledges in the launch text that GPT-6 Astra's written reasoning is harder to monitor than GPT-5.6 Sol's, based on tests that explicitly asked the model to evade monitoring, and attributes the decline to the model's greater control over its own reasoning on simple tasks. It is the vendor declaring it lost part of its visibility over the product, and the declaration being in the announcement is the signal that the topic does not fit outside it.

The announced compensation is a misalignment monitor running in production for models of GPT-6 Astra's class. A system of classifiers checks reasoning and actions for unauthorized behavior and can interrupt activity automatically. On the API, the task stops. OpenAI itself admits the extra checks can slow, pause, or stop legitimate work, including defensive cybersecurity, and keeps iterating to reduce unnecessary interruptions.

For anyone consuming the API, this changes the design of the operation: the interrupted task stops being a rare incident and becomes product mechanics, with an operational cost of its own. On the contractual side, GPT-6 Astra supports Zero Data Retention for eligible API customers, and Enterprise access stays off by default until the administrator enables it.

What changes for anyone consuming models

Three changes with numbers define the regime. The routing decision for risky workloads gained a new axis: a monitor that can stop the task on the API. The agentic bill became more sensitive to the fast mode, which doubles the cost per token. And the API consumer inherited governance requirements from the provider side.

Routing: with 61.2 against Fable 5.1's 65.7 on the third-party index, the 2026-09-09 reading, GPT-6 Astra does not enter the route by scoreboard. It enters by specific capability, like the 57.9% on Terminal-Bench 4.0, and by price, and the Critical rating becomes part of the calculation for workloads that touch security, infrastructure, or sensitive data: capability and control demands in the same package.

Cost: output at US$ 50 per million and Fast at 2x move the break-even of the agentic workload. The daily bill that closes at US$ 200 on Standard closes at US$ 400 on Fast, and the difference accumulates per key, per project, and per month.

Governance: the misalignment monitor in production, Zero Data Retention eligibility, and the Enterprise admin off by default form the control package the consumer needs to operate, check, and register. A task that can be stopped requires a queue, retries, and an audit trail.

AxisBefore: the GPT-5.6 Sol regimeAfter: the GPT-6 Astra regime
Cybersecurity ratingbelow the Critical thresholdCritical, self-assessment of September 1, 2026
ExploitBench without safeguards78.5%100%, published by OpenAI on September 3, 2026
ExploitGym honeypot48.2% of target exceeded0.0%
Misalignment monitordid not interrupt tasks on the APIcan stop the task on the API
Monitorabilitythe evaluator's own referenceworse than GPT-5.6 Sol's, by OpenAI's own declaration

Choice, price, and policy demand a layer of their own. A multi-model proxy with governance resolves the operational side: the Nexforce Router concentrates 300+ models behind a single API, with automatic failover, per-key routing rules, budget per key, per agent, and per project as a spend cap, centralized observability, a real-time model ranking, savings of up to 50% on the cost per token, and local billing in BRL. When the frontier changes price, safety rating, and governance demands on the same day, the decision lives in the routing layer, not in the annual contract.

What to do now

Five moves, no waiting.

  1. Sort your workloads into three stacks: routine generation, automation with access to systems, and security work. GPT-6 Astra enters the second and the third only when the 57.9% capability on Terminal-Bench 4.0 pays for the demands of the Critical rating; routine stays on cheaper models.
  2. Redo the cost math with Fast on the table: multiply your output tokens by US$ 50 per million on Standard and by US$ 100 on Fast, and mark the point where the speed gain pays double per token.
  3. Design for interruption tolerance: the misalignment monitor can stop the task on the API, so a reprocessing queue, retries, and a record of the stop reason stop being a nicety and become a requirement.
  4. Check the contractual governance before integrating: Zero Data Retention eligibility, the Enterprise admin off by default, and who, in your organization, can enable the model.
  5. Wait for the next AA Intelligence Index reading before fixing a permanent route: 61.2 is the 2026-09-09 reading, and the rebased index already showed this year that leaderboard positions move.

Frequently asked questions about GPT-6 Astra

What is GPT-6 Astra? It is OpenAI's frontier model, launched on September 3, 2026, to replace GPT-5.6 Sol, available on the API as gpt-6-astra. OpenAI presents it as its most capable model in computer use, software engineering, cybersecurity, and science. The launch benchmarks are self-assessment. The third-party index does not put it in the lead.

How much does GPT-6 Astra cost? Standard mode costs US$ 10 per million input tokens and US$ 50 per million output tokens, figures published by OpenAI on September 3, 2026. Fast mode charges 2x the Standard price: US$ 20 and US$ 100 per million tokens, respectively. The 5x gap between input and output makes the output tariff the line that decides the invoice.

Is GPT-6 Astra safe? OpenAI classified it as the first model to reach the Critical cybersecurity threshold of the Preparedness Framework, in a self-assessment dated September 1, 2026. The model discovered two zero-days during the evaluation, under disclosure to the maintainers. Monitorability worse by OpenAI's own declaration than GPT-5.6 Sol's led the company to run a monitor that can interrupt tasks on the API.

How does GPT-6 Astra compare on benchmarks? The numbers published by OpenAI on September 3, 2026: 97.6% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, 95.0% on ARC-AGI-2, 96.0% on GPQA Diamond, and 57.9% on Terminal-Bench 4.0, against 37.3% for GPT-5.6 Sol and 55.8% for Fable 5.1. On the third-party index, the 2026-09-09 reading, GPT-6 Astra scores 61.2 and Fable 5.1, 65.7.

Is GPT-6 Astra available on Azure and Bedrock yet? Microsoft Azure and AWS Bedrock were announced as channels, but OpenAI confirmed no general availability date for either as of publication. What already answers in production is the OpenAI API, under the identifier gpt-6-astra. Purchasing anchored in the cloud should treat both channels as announcement, not as availability.

Referências e Leitura Complementar

What to watch

Four points are worth a calendar entry. The outcome of the disclosure of the two zero-days discovered in the evaluation is the first practical test of the process and affects confidence in the Critical rating. GA timing for GPT-6 Astra on Azure and Bedrock remains without a date as of publication. The stability of the US$ 10 and US$ 50 per million token prices, and of the Fast mode at 2x that doubles the agentic bill, defines whether this week's math holds for the quarter. And the next AA Intelligence Index reading will say whether the 61.2 of 2026-09-09 holds or moves.

Nexforce

Accelerate your company'sbusiness and operational efficiency

We design the technology of tomorrow to boost your business operational scale

Talk to a Specialist

Related articles