Artificial Analysis overhauls its Intelligence Index after GPT-6 Astra scoring drew skepticism

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane September 5, 2026 1 min read
Artificial Analysis overhauls its Intelligence Index after GPT-6 Astra scoring drew skepticism

Artificial Analysis has released version 4.2 of its Intelligence Index following criticism that earlier benchmarks failed to capture GPT-6 Astra’s actual progress. Previous evaluations, including OpenAI’s own, placed Astra well ahead of the field, yet the initial index scored it just on par with its predecessor. The updated ranking now shows a four-point gain for Astra over its previous version. Anthropic‘s Claude Fable 5.1 still leads the chart, followed by Astra in second and Meta in third. The report notes that Astra uses fewer tokens per task than every other frontier model, while Anthropic, OpenAI, Meta, and Zhipu AI share the lead on cost-to-performance ratio.

The organisation added two new benchmarks, AA-Briefcase for real-world knowledge work and GDP.pdf from Surge AI for PDF document analysis, while dropping GPQA-Diamond because models have solved it. Private test data now makes up 40 percent of the weighting to make gaming harder. Artificial Analysis also fixed scoring errors across several benchmarks and tweaked its grading systems for more stable results. The company stated it held off on updates to keep scores stable during major model launches but an interim update was necessary because the top of the leaderboard moved so fast.

  • Claude Fable 5.1 retains the number one position
  • Private test data weighting increased to forty percent
  • Version 5 is scheduled to roll out in stages
Scroll to Top