Alibaba’s Qwen3.8 Max achieved a score of 56 on the Artificial Analysis Intelligence Index, matching Claude Opus 4.8 performance while trailing Kimi K3 at 57. The new model improved by ten points from its predecessor, yet it requires sixty-four inference steps per task compared to fourteen previously. This approach forces the system to resend full conversation history repeatedly, causing input tokens to increase fifteenfold. Consequently, the price per task rose to $1.14, more than double the $0.53 cost for Qwen3.7 Max, despite lower individual token rates.
The efficiency loss means the model works more thoroughly but runs slower and costs more than competitors like GLM-5.2. Performance also regressed in specific areas where the previous version excelled. Long context retrieval accuracy dropped two points, and the ability to admit ignorance fell ten points. The hallucination rate jumped from 23 to 40 percent, indicating the model guesses far more often than before.
* Input token growth reached fifteen times the original volume
* Task cost doubled to $1.14 despite cheaper per-token rates
* Hallucination rate increased from 23 to 40 percent



