China’s Largest AI Model Is Being Developed at Bytedance

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane August 7, 2026 1 min read
China’s Largest AI Model Is Being Developed at Bytedance

Bytedance is training an artificial intelligence model with up to ten trillion parameters, a figure reported by the Financial Times. This capacity is three times larger than Moonshot’s Kimi K3, which currently holds the title of the biggest Chinese model. The project places the TikTok parent company in the same performance bracket as Anthropic‘s Mythos 5, estimated at around eight trillion parameters. Anthropic has not confirmed its own numbers. Three insiders told the newspaper that the system is in pretraining, a phase typically lasting three to six months. Parameters determine storage capacity, but performance also depends on data quality and training methods. One source says Bytedance has avoided distillation, meaning training on outputs from other companies’ models, for over a year. Founder Zhang Yiming told the 2,000-person Seed team internally to aim for world-leading model capabilities over the long term. xAI is also training Grok variants with six and ten trillion parameters on its Colossus 2 cluster, according to Elon Musk.

This development matters because it signals a shift in how Chinese technology firms approach large language models. Moving away from distillation suggests a willingness to invest heavily in proprietary data and training infrastructure rather than relying on existing outputs. The scale indicates a commitment to matching or exceeding current global leaders in raw computational capacity. Such moves could influence the trajectory of research and development within the region.

* Pretraining phase typically lasts three to six months
* Distillation has been avoided for over a year
* Seed team consists of 2,000 staff members

Scroll to Top