Moonshot’s Kimi K3 outperforms Fable 5 in frontend code but lags far behind in complex math

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane July 19, 2026 1 min read
Moonshot’s Kimi K3 outperforms Fable 5 in frontend code but lags far behind in complex math

Moonshot’s Kimi K3 model has taken the top spot in the Code Arena Frontend benchmark with a score of 1,679. This performance beats Claude Fable 5, GPT-5.6 Sol, and every other tested model by a wide margin. It marks the first time a Chinese AI has claimed the leading position on this specific human preference rating test.

The results look different when testing complex mathematical reasoning. Data from Epoch AI shows Kimi K3 achieves only about 39 percent accuracy on FrontierMath Tier 4 tasks. This benchmark covers the hardest expert-level math problems where OpenAI and Anthropic models score close to 90 percent. The gap highlights a significant divergence in capabilities between the new Chinese system and established Western leaders in advanced calculation.

  • Chinese AI models now lead in frontend code generation preferences
  • Western models maintain a large lead on expert math benchmarks
  • The performance gap on Tier 4 math is roughly 50 percentage points
Scroll to Top