Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane September 4, 2026 1 min read
Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward

OpenAI‘s GPT-6 Astra receives mixed scores from independent evaluators. Epoch AI ranks the model first with 169 points, yet Artificial Analysis places it behind Claude Fable 5.1 and no higher than its own predecessor. The most notable result comes from ARC-AGI-3, where Astra demonstrates human-beating efficiency for the first time. François Chollet, chief of the ARC Prize, does not label this achievement as proof of artificial general intelligence. He does note that the progress is occurring twice as fast as his previous predictions. This shift causes him to move his AGI forecast forward.

The contradiction between broad capability tests and specific efficiency metrics highlights the difficulty in defining intelligence. High scores on standard datasets do not guarantee performance on complex reasoning tasks. Efficiency gains on specialised benchmarks suggest the model handles novel problems differently than earlier versions. This divergence forces researchers to reconsider how they measure advancement.

* Epoch AI scores Astra at 169 points
* Artificial Analysis ranks it below Claude Fable 5.1
* Chollet doubles his expected speed of progress

Scroll to Top