Meta closes in on the top with Muse Spark 1.3, and undercuts rivals on price

In this articleThe release and the tiersPrice and performanceWhere the model gains groundPerformance outside agentic tasks Meta launches Muse Spark 1.3 with…

By Vane September 3, 2026 2 min read
Meta closes in on the top with Muse Spark 1.3, and undercuts rivals on price


Meta launches Muse Spark 1.3 with a price undercut and a narrower gap to the leaders

The release and the tiers

Meta has released Muse Spark 1.3, its fourth model in five months. The xhigh tier is available now, while the more powerful max version runs as a limited preview.

The series launched in April, with version 1.1 following in July and 1.2 in August. Meta distributes the model through Muse Code and the Meta Model API. The max tier arrives only after further safety testing and currently runs as a limited partner preview, according to Artificial Analysis.

Price and performance

At an unchanged $1.25 and $4.25 per million input and output tokens, one index task costs $0.55. No model scoring 59 points or higher is cheaper, and rivals at the same index level run between $0.94 and $1.23. Muse Spark 1.3 does cost more than version 1.2, which ran $0.40.

Where the model gains ground

On the Intelligence Index, max scores 62 points and xhigh 61, up from 57 in August and 53 in July. The jump comes down to how the index weights its tests. GDPval-AA v2 counts for 20 percent, Terminal-Bench 2.1 for 16 percent, and τ³-Bench Banking for 14 percent, and Meta’s biggest gains land in exactly those three tests.

On τ³-Banking, where agents operate tools in a simulated banking scenario, max hits 52 percent. That is number one right now, according to Artificial Analysis, and it is the only outright lead the model holds. The available xhigh tier reaches 47 percent, tying Claude Fable 5.1 (max) and GLM-5.3-Flash rather than leading. The predecessor 1.2 sat at 35 percent.

Terminal-Bench 2.1, which tests coding in the terminal, climbs from 80 to 85 percent on xhigh and 86 on max, but Claude Fable 5.1 still holds the top spot at 91.4 percent in its max tier, 91.0 at xhigh, and 89.9 at high. On the index’s highest-weighted test, GDPval-AA v2, Meta improves from 1,615 to 1,709 and 1,754 on a scale calibrated to human expert performance at 1,000 across 220 real-world professional tasks. Claude Fable 5.1 (max) sits at 1,853. Meta buys the max variant’s edge with compute, burning 62 percent more reasoning tokens than xhigh.

Performance outside agentic tasks

On GPQA Diamond, which poses expert-level science questions, Muse Spark rises from 90 to 94 percent. That is the top group, but still below Gemini 3.8 Flash (high) at 95.3 and Grok 4.6 (high) at 94.9. CritPt, covering research physics, jumps from 18 to 26 percent, well behind GPT-5.6 Sol (max) at 32.3 and Claude Fable 5.1 (xhigh) at 31.1 percent.

Two scores actually drop against 1.2. AA-LCR falls from 83 to 79 percent. Factual accuracy in AA-Omniscience slips by up to three points, because the model more often declines to answer when it is unsure.

Neither Meta nor Artificial Analysis has named a price for the max variant yet. Larger models and an open-weights release are on the way.


Scroll to Top