GPT-6 Astra is the first model making OpenAI willing to declare the “AGI era”

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane September 3, 2026 2 min read
GPT-6 Astra is the first model making OpenAI willing to declare the “AGI era”


GPT-6 Astra is now live

OpenAI has released GPT-6 Astra, its most capable model to date. The company is now willing to state that this system qualifies as Artificial General Intelligence, or at least sits within reach of that definition. By their own metrics, an AI system that outperforms humans at most economically valuable work has arrived.

Greg Brockman, the company’s president, confirmed the launch. Access is rolling out first to select organisations via the Daybreak program. Broader availability for ChatGPT Plus, Pro, Business, and Enterprise customers is expected in the coming days. The model is also accessible through the API and cloud platforms including AWS Bedrock and Microsoft Azure. Subscribers to Pro, Business, and Enterprise tiers receive access to GPT-6 Astra Pro, a higher-performance variant. Enterprise workspace administrators must activate this model manually.

The training run

Astra was pretrained on more than 100,000 GPUs at the Stargate facility in Texas. OpenAI researcher Aidan Clark described this as the company’s largest training run ever. He noted that the jump from the previous model, GPT-5.6 Sol, to Astra represents a bigger capability gain than the jump to Sol from earlier models. Clark said this improvement is partly because previous AI models played a role in monitoring the training process.

Performance benchmarks

In benchmarks OpenAI published, Astra scores well above its predecessor Sol and competitors including Anthropic‘s Fable models. The model hits top marks across a range of disciplines. Logical reasoning scores 99.9 percent on ARC-AGI-3, though under its own test conditions. Math reaches 97.6 percent on FrontierMath Tier 4 v2. Software engineering stands at 74.1 percent on DeepSWE v1.1. Expert knowledge is 96 percent on GPQA Diamond. Engineering scores 95.9 percent on BenchCAD. Cybersecurity reaches 100 percent on ExploitBench.

Token prices are 2.5 times higher than the previous generation and are on par with Anthropic’s Fable 5.1. OpenAI argues that the cost per completed task is actually lower depending on the use case.

Computer use

BenchmarkAstraSolFable 5.1Fable 5
Agents’ Final Exam59.3%53.6%48.7%55.5%
OSWorld 2.0 (offline, partial)72.6%65.7%70.2%
ScreenSpot-Pro (no tools)92.7%76.9%87.3%

Professional tasks

BenchmarkAstraSolFable 5.1Fable 5
AutomationBench41.4%18.1%31.4%17.4%
BenchCAD95.9%83.3%84.3%67.5%
BrowseComp91.5%90.4%87.4%90.8%
OpenScore String Quartets0.840.19
Internal Design Tasks50.0%47.4%35.8%
Internal Data Science Tasks40.9%30.5%34.7%
AA Intelligence Index v4.1.161.260.965.762.1

BenchCAD cost is approximately 43 percent below Sol and 86 percent below Fable 5.1.

Coding

BenchmarkAstraSolFable 5.1Fable 5
Terminal Bench 4.057.7%37.3%55.8%42.0%
DeepSWE v1.174.1%72.7%67.4%69.9%
FrontierCode 1.1 Extended64.5%60.6%63.6%64.9%
FrontierCode 1.1 Main53.3%47.5%50.9%53.5%
Internal Database Migration63.9%42.7%57.8%50.3%
AA Coding Agent Index v1.467.065.167.268.1

Terminal-Bench 4.0 cost is approximately 9 percent below Sol and 63 percent below Fable 5.1.

Academic work

BenchmarkAstraSolFable 5.1Fable 5
Terminal-Bench Science 0.164.6%22.4%52.6%21.4%
FrontierMath Tier 4 (v2)97.6%83.0%87.8%87.8%
GPQA Diamond96.0%94.6%93.7%92.6%
Humanity’s Last Exam (tools)57.2%65.0%63.8%63.6%

Lower-cost settings show Terminal-Bench Science at 61.1 percent with about 27 percent lower cost. GPQA Diamond reaches 94.9 percent at approximately 37 percent lower cost. Prime gaps improved from 240 to 186, and a large-gap bound term improved for the first time in over 80 years.

Science and health

BenchmarkAstraSolFable 5.1Fable 5
GeneBench Pro37.8%28.7%
MedChemBench (internal)49.3%47.4%
LifeSciBench60.3%59.9%
HealthBench Professional63.4%60.5%56.6%60.9%

Fable 5 and 5.1 are not included in LifeSciBench, GeneBench Pro, and MedChemBench because they reject most questions.

Cybersecurity

BenchmarkAstraSolFable 5.1Fable 5
ExploitBench100.0%78.5%70%
ExploitGym42.4%30.3%30.4%28.4%
ExploitBench (Jun–Aug 2026)39.0%5.5%
SRE-Bench88.0%55.9%12.5%
SEC-Bench Pro85.4%79.1%

SRE-Bench within four attempts reached 99.2 percent versus 68.7 percent for Sol. Astra found

Scroll to Top