GPT-6 Astra is now live
OpenAI has released GPT-6 Astra, its most capable model to date. The company is now willing to state that this system qualifies as Artificial General Intelligence, or at least sits within reach of that definition. By their own metrics, an AI system that outperforms humans at most economically valuable work has arrived.
Greg Brockman, the company’s president, confirmed the launch. Access is rolling out first to select organisations via the Daybreak program. Broader availability for ChatGPT Plus, Pro, Business, and Enterprise customers is expected in the coming days. The model is also accessible through the API and cloud platforms including AWS Bedrock and Microsoft Azure. Subscribers to Pro, Business, and Enterprise tiers receive access to GPT-6 Astra Pro, a higher-performance variant. Enterprise workspace administrators must activate this model manually.
The training run
Astra was pretrained on more than 100,000 GPUs at the Stargate facility in Texas. OpenAI researcher Aidan Clark described this as the company’s largest training run ever. He noted that the jump from the previous model, GPT-5.6 Sol, to Astra represents a bigger capability gain than the jump to Sol from earlier models. Clark said this improvement is partly because previous AI models played a role in monitoring the training process.
Performance benchmarks
In benchmarks OpenAI published, Astra scores well above its predecessor Sol and competitors including Anthropic‘s Fable models. The model hits top marks across a range of disciplines. Logical reasoning scores 99.9 percent on ARC-AGI-3, though under its own test conditions. Math reaches 97.6 percent on FrontierMath Tier 4 v2. Software engineering stands at 74.1 percent on DeepSWE v1.1. Expert knowledge is 96 percent on GPQA Diamond. Engineering scores 95.9 percent on BenchCAD. Cybersecurity reaches 100 percent on ExploitBench.
Token prices are 2.5 times higher than the previous generation and are on par with Anthropic’s Fable 5.1. OpenAI argues that the cost per completed task is actually lower depending on the use case.
Computer use
| Benchmark | Astra | Sol | Fable 5.1 | Fable 5 |
|---|---|---|---|---|
| Agents’ Final Exam | 59.3% | 53.6% | 48.7% | 55.5% |
| OSWorld 2.0 (offline, partial) | 72.6% | 65.7% | 70.2% | — |
| ScreenSpot-Pro (no tools) | 92.7% | 76.9% | 87.3% | — |
Professional tasks
| Benchmark | Astra | Sol | Fable 5.1 | Fable 5 |
|---|---|---|---|---|
| AutomationBench | 41.4% | 18.1% | 31.4% | 17.4% |
| BenchCAD | 95.9% | 83.3% | 84.3% | 67.5% |
| BrowseComp | 91.5% | 90.4% | 87.4% | 90.8% |
| OpenScore String Quartets | 0.84 | 0.19 | — | — |
| Internal Design Tasks | 50.0% | 47.4% | 35.8% | — |
| Internal Data Science Tasks | 40.9% | 30.5% | 34.7% | — |
| AA Intelligence Index v4.1.1 | 61.2 | 60.9 | 65.7 | 62.1 |
BenchCAD cost is approximately 43 percent below Sol and 86 percent below Fable 5.1.
Coding
| Benchmark | Astra | Sol | Fable 5.1 | Fable 5 |
|---|---|---|---|---|
| Terminal Bench 4.0 | 57.7% | 37.3% | 55.8% | 42.0% |
| DeepSWE v1.1 | 74.1% | 72.7% | 67.4% | 69.9% |
| FrontierCode 1.1 Extended | 64.5% | 60.6% | 63.6% | 64.9% |
| FrontierCode 1.1 Main | 53.3% | 47.5% | 50.9% | 53.5% |
| Internal Database Migration | 63.9% | 42.7% | 57.8% | 50.3% |
| AA Coding Agent Index v1.4 | 67.0 | 65.1 | 67.2 | 68.1 |
Terminal-Bench 4.0 cost is approximately 9 percent below Sol and 63 percent below Fable 5.1.
Academic work
| Benchmark | Astra | Sol | Fable 5.1 | Fable 5 |
|---|---|---|---|---|
| Terminal-Bench Science 0.1 | 64.6% | 22.4% | 52.6% | 21.4% |
| FrontierMath Tier 4 (v2) | 97.6% | 83.0% | 87.8% | 87.8% |
| GPQA Diamond | 96.0% | 94.6% | 93.7% | 92.6% |
| Humanity’s Last Exam (tools) | 57.2% | 65.0% | 63.8% | 63.6% |
Lower-cost settings show Terminal-Bench Science at 61.1 percent with about 27 percent lower cost. GPQA Diamond reaches 94.9 percent at approximately 37 percent lower cost. Prime gaps improved from 240 to 186, and a large-gap bound term improved for the first time in over 80 years.
Science and health
| Benchmark | Astra | Sol | Fable 5.1 | Fable 5 |
|---|---|---|---|---|
| GeneBench Pro | 37.8% | 28.7% | — | — |
| MedChemBench (internal) | 49.3% | 47.4% | — | — |
| LifeSciBench | 60.3% | 59.9% | — | — |
| HealthBench Professional | 63.4% | 60.5% | 56.6% | 60.9% |
Fable 5 and 5.1 are not included in LifeSciBench, GeneBench Pro, and MedChemBench because they reject most questions.
Cybersecurity
| Benchmark | Astra | Sol | Fable 5.1 | Fable 5 |
|---|---|---|---|---|
| ExploitBench | 100.0% | 78.5% | 70% | — |
| ExploitGym | 42.4% | 30.3% | 30.4% | 28.4% |
| ExploitBench (Jun–Aug 2026) | 39.0% | 5.5% | — | — |
| SRE-Bench | 88.0% | 55.9% | 12.5% | — |
| SEC-Bench Pro | 85.4% | 79.1% | — | — |
SRE-Bench within four attempts reached 99.2 percent versus 68.7 percent for Sol. Astra found




