OpenAI announced on September 3, 2026, the launch of GPT-6 Astra — a new flagship model that the company calls "the smartest and most aligned in the world." The model's key feature is a 100% result on ExploitBench, after which Astra became the first to cross the critical cybersecurity threshold of OpenAI's Preparedness Framework. Access to the model is being rolled out in phases: first to enterprise contractors of the Daybreak program, and in the coming days the model will become available to ChatGPT subscribers, developers via the API, and AWS customers.

image

What Happened

OpenAI provided the following reported results for GPT-6 Astra: FrontierMath Tier 4 v2 — 97.6%, DeepSWE v1.1 — 74.1%, BenchCAD — 95.9%, GPQA Diamond — 96%, ExploitBench — 100%, ARC-AGI-3 — 98.6%. The company clarified that ARC-AGI-3 was measured in a Responses API configuration. According to Aidan Clark, this is OpenAI's largest launch to date: post-training was conducted for the first time on more than 100,000 GPUs at the Stargate facility in Texas, and Astra became one of the first models of its class in which previous models substantially participated in supervision during training. API pricing is set at $10 per million input tokens and $50 per million output tokens. Access is being rolled out in phases: enterprise contractors of the Daybreak program receive it starting Thursday, and in the coming days the model will appear for ChatGPT Plus, Pro, Business, and Enterprise subscribers, as well as in the OpenAI API and on AWS.

Context

OpenAI's Preparedness Framework sets capability thresholds for frontier models: the cybersecurity threshold is considered crossed when a model finds unknown vulnerabilities and exploit chains without continuous human guidance. To work with such models, the company has the Daybreak program and its extended version, Daybreak Blue, with monitoring and ZDR access for contractors. GPT-6 Astra appears in the API under the name gpt-6-astra. The methodology deserves separate attention: according to OpenAI, this is one of the first models of its class in which previous models substantially participated in supervision during post-training, and the company considers this a new vector of training methodology, designated as model-supervised training, rather than simply a scale increase.

Why This Matters for the Industry

The 100% result on ExploitBench changes the industry standard for risk assessment: Astra became the first model to cross the critical cybersecurity threshold of the Preparedness Framework, meaning it autonomously finds unknown vulnerabilities and exploit chains, and the access and operation framework for such models is defined by the extended Daybreak Blue program with monitoring and ZDR access. At the same time, the claimed leadership in agentic coding is not confirmed: Astra's DeepSWE v1.1 is 74.1% versus 75.4% for Meta Muse Spark 1.3, and the difference of about 1.3 percentage points falls within benchmark noise. The price of $10 per million input tokens and $50 per million output tokens puts Astra on par with Anthropic Fable 5.1 and makes it 2.5 times more expensive than the promotional OpenAI Sol, fixing the upper price anchor of the frontier class and turning model selection into a question of unit economics for specific scenarios. For security teams, a model capable of autonomously finding vulnerabilities becomes a new factor in threat models, requiring reassessment of exposure.

Why This Matters for Users

In the coming days, gpt-6-astra will appear in the OpenAI API and in ChatGPT, with a price benchmark of $10 per million input tokens and $50 per million output tokens. Enterprise contractors of the Daybreak program received access starting Thursday, and ChatGPT Plus, Pro, Business, and Enterprise subscribers can expect the model to appear in the coming days. Since Astra is 2.5 times more expensive than the promotional OpenAI Sol, it is not advisable to move production workloads to it without measuring latency and throughput on your own traffic; it is more reasonable to first prepare evals and test specific scenarios. Astra's confirmed strong suit is mathematics and cybersecurity, which have been confirmed by independent analyses, while in coding the model does not yet outperform its closest competitors. The statements about the "era of AGI" are words from OpenAI President Greg Brockman at a press briefing, not a formally verified statement: OpenAI itself calls AGI a gray, ambiguous concept.

What Is Still Unknown / Limitations

All reported figures are vendor-reported data from OpenAI without a full description of the methodology. Independent analyses to date have only confirmed results in mathematics and cybersecurity, while for other benchmarks, including DeepSWE v1.1 and 95.9% on BenchCAD, independent reproducibility is not yet complete. The 98.6% result on ARC-AGI-3 was measured in a specific Responses API configuration, so direct comparison with other models without accounting for the harness and configuration is incorrect. The model's latency and throughput under real production load have not been published.

Sources

Author

Look at AI, editorial team