Frontier Radar #4: China has caught up, so what’s left of the Western AI lead?

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane August 20, 2026 8 min read
Frontier Radar #4: China has caught up, so what’s left of the Western AI lead?

Chinese models Kimi K3 and GLM-5.3 have closed the gap on top-tier US systems to a level that threatens investor confidence. Western labs point to distillation as a primary cause, though the conclusion remains unchanged: a model performance lead cannot be defended.

THE DECODER publishes Frontier Radar six times a year to examine core AI topics. Issue #4 is our most detailed yet, analysing how Chinese models caught up, where Western labs still hold ground, why the industry’s moat has shifted, and why Europe is losing two races simultaneously. Previous issues covered AI agents, productivity metrics, and the emerging token economy.

A year and a half ago, DeepSeek R1 caused a market shock. A Chinese lab suddenly competed with OpenAI’s o1, the first commercial reasoning model, reportedly for far less capital. Billions in market value evaporated within days. Markets questioned whether the planned infrastructure buildout was overblown.

The situation was murkier than the headlines suggested. DeepSeek’s own report showed R1 beating o1 on individual tests like AIME 2024 but trailing clearly on others, such as factual knowledge. Later benchmarks exposed more gaps. Chinese models only reached the top in individual disciplines, not across the board.

As recently as late June, Z.ai’s GLM-5.2 showed the same pattern. The landscape changed with the latest Chinese open-weights models: Moonshot’s Kimi K3, Alibaba’s Qwen3.8-Max, and GLM-5.3.

These models now sit near the top of almost every broad, demanding evaluation. They handle long knowledge tasks, code across many steps, and coordinate tools far more reliably than their predecessors.

Measured by common benchmarks, the often-cited gap of a few months has shrunk enough to become an investor problem. The Wall Street Journal reports that Anthropic is fielding uncomfortable questions ahead of its upcoming IPO and points to its remaining lead at the top as its primary defence. Below that tier, the field belongs largely to open, far cheaper models from China.

Investors worry that raw model performance can barely carry a business anymore. Whatever a model can do exclusively today, a freely downloadable one can do a few months later.

Two accusations are in play. Chinese labs allegedly tapped Western models as teachers, a practice known as distillation. They also allegedly tune their models for strong benchmark scores without the broad capabilities to match, a phenomenon called benchmaxxing.

The American lead has not disappeared. It has retreated to a few, ever-narrower areas of the so-called frontier, the leading edge of what is technically possible. Where that lead still sits, and what it is worth economically, is the first question of this issue.

The second question follows from it. If the model alone can make a clear difference anymore, what does the lead rest on? Our thesis is that it rests less and less on the model, and more and more on the overall system around it. This means the system where ongoing work produces the next models.

What’s left of the lead

At launch, Artificial Analysis had K3 in third place on its Intelligence Index with 57 points, right behind then-leaders GPT-5.5 and Opus 4.8. K3 improved most on agentic tasks, so exactly the hands-on work that matters in enterprise use. On AutomationBench-AA, K3 even debuted in first place until Anthropic answered with Opus 5.

On CEO-Bench, where an agent runs a fictional software company for 500 simulated days, K3 posted the best published single run at $22.15 million. Chinese models, including its predecessor K2.7, had regularly failed these long hauls. Qwen3.8-Max reaches a similarly high overall level. One caveat remains: newer Chinese models sometimes burn far more tokens than Western ones, which eats into part of their price advantage. The cost-per-task math still applies.

The edge of current model capabilities has been growing at uneven speeds for a while, what researchers call the “jagged frontier.” Individual capabilities advance at different rates, which is why the much-quoted months-long gap always depended on who measured what. What is new is where a Western lead is still measurable at all. Three areas remain.

The first is abstract specialty tests. Shortly after K3’s launch, Opus 5 retook the top of the index with 61 points, a small gap over K3’s 57. On ARC-AGI-1, a test of abstract pattern recognition using small puzzle grids, K3 and Fable 5 – Anthropic’s flagship line for coding and agent work – sit practically even at 94.5 and 98.5 percent. Only on ARC-AGI-2 does the gap widen: 60.4 versus 89.2 percent. But ARC-AGI deliberately measures abstract pattern recognition far removed from everyday tasks. There is no guarantee this gap predicts practical differences. It could simply stay economically irrelevant in most cases.

The second is reliability. The AA-AnalystAgent benchmark, launched August 12, tests agentic data analysis on real tables and documents. It only counts a task as solved if a model gets it right in five out of five independent runs, a measure the operators call pass^5. Opus 5 leads with 54 percent, ahead of GPT-5.5 at 50. K3 is the best open model at 39 percent.

The gap comes almost entirely from poor repeatability. K3 solves 73 percent of tasks at least once in five attempts, practically even with Opus 5 at 74. Reliability does not follow the usual intelligence-index rankings either. Even GPT-5.6 Sol falls behind its own predecessor here.

Commercially, reliability weighs heavily, as the operators explain in their launch article. An analyst agent only saves work when its answers hold up without review. Multiple runs and checkers can compensate, but they drive up the cost per accepted result.

The third is cybersecurity. Here the gap is best documented. A joint assessment by the UK’s AISI and the US CAISI found that K3 lags far behind leading US models in offensive cyber capabilities.

On ExploitBench, a test for developing exploits, K3 scored 32 percent. The top US models averaged about 76. K3 failed all 41 tasks that required executing code on a target system; leading US models solved 20 on average. And in a simulated attack, K3 made it to step 17 of 32, the US leaders to 28.5 on average.

But this gap is shrinking too. GLM-5.3, unveiled August 14, scores 54.4 percent on ExploitBench by Z.ai’s own measurement, landing on more than double of its predecessor GLM-5.2. That cuts the distance to the US leaders roughly in half within a month. On finding and validating vulnerabilities in source code (CyberGym), GLM-5.3 even edges past the leading US models.

At the same time, cybersecurity is the one area where providers do not openly sell their strongest capabilities. Anthropic’s best cyber model, Mythos 5, which hits 78 percent on ExploitBench, is only available under controlled conditions through Project Glasswing. Its public sibling Fable 5 effectively stays at the level of the older Opus 4.8 at 40 percent, because upstream safeguards intercepted 407 of 410 test episodes.

OpenAI takes a similar approach with Daybreak and specialized cyber variants. GPT-5.6 Sol reaches 73.5 percent on ExploitBench in OpenAI’s own evaluation but stays below Critical, the highest level on OpenAI’s risk scale. With the upcoming Astra, OpenAI cannot rule out for the first time that one of its own models crosses that threshold. If it does, the Preparedness Framework kicks in much harder, up to a full development halt. OpenAI has already paused some internal Astra work.

And with GLM-5.3, a Chinese lab is adopting this pattern for the first time. Z.ai is delaying the weights release by about two weeks for extra safety work and plans to limit the most sensitive cyber functions to verified users.

The remaining lead adds up like this. The gap in abstract specialty tests is measurable, but its economic value is unclear. The reliability gap matters most commercially but is the least settled – when a newer model falls behind its own predecessor within the same lab (GPT-5.6 Sol vs. GPT-5.5), that ranking is clearly in flux.

The cybersecurity gap involves capabilities almost no customer can buy through regular channels, and even that gap has narrowed sharply of late. Only the few dangerous capabilities can be walled off. The commercially valuable rest, like coding, research, agent work, has to stay accessible through APIs, or there is no business. And that is where the lead shrinks within months.

The obvious suspicion is that the two are connected. Chinese labs could be using Western models as teachers through their APIs (the distillation mentioned earlier) and the lead would then hold exactly where that access is missing. There are good reasons for this suspicion, and we will get into how Chinese labs might profit from the knowledge inside Western models below.

But first, for the strategic assessment, the question of guilt is almost secondary. If the suspicion is true, every capability sold eventually migrates to the competition, because there is still no reliable way to prevent the skimming. If it is not true, Chinese labs caught up on their own, and the lead was worth even less.

Both readings lead to the same place. A model lead only holds where the product is not broadly offered in the first place, and you cannot build a business on that. A close look at distillation is still worthwhile, because its mechanics decide whether any defensible model capabilities exist at all.

Distillation explains the pace, not the breadth

OpenAI and Anthropic accuse Chinese companies of using their models at scale to build their own. According to Anthropic, industrial campaigns by DeepSeek, Moonshot, and MiniMax ran more than 16 million interactions through around 24,000 fraudulent accounts. The campaign attributed to Moonshot targeted agent reasoning, tool use, coding, and reasoning traces with more than 3.4 million interactions. That fits with Together AI finding a 0.72 correlation between the task-level success rates of K3 and Fable 5 on real software problems, and with a style analysis by Typebulb placing K3 closest to Fable 5.

Critics counter that the window between Fable 5’s release and K3’s was too short to influence training. But Opus 4.8, a closely related model, was available much longer, and the same studies show clear overlaps there too. OpenAI describes a similar pattern with DeepSeek in its memorandum to the US Congress. None of the labs has published verifiable evidence, though, and our inquiries with OpenAI and Anthropic turned up nothing new.

Still, the known accusations and several research papers make it possible to reconstruct how Moonshot, DeepSeek, and others could benefit from distillation during

Scroll to Top