Deepseek has released open-source programming tools designed for Huawei’s Ascend chips, aiming to offer a simpler alternative to Nvidia’s CUDA platform. The core of this release is TileLang, a language developed by Peking University researchers that Deepseek has used for roughly a year and now distributes freely.
In this article
A post on the developer’s official WeChat channel confirmed the collaboration, with Huawei stating it “fully supported” the work. The software suite includes libraries for computation and data movement across chips. Deepseek and Huawei also optimised a supernode cluster containing 128 Ascend 950 chips.
Deepseek argues that building an independent software ecosystem requires a universal language that balances ease of use with hardware performance. TileLang is now the company’s primary tool for artificial general intelligence work. The team previously tested the language on older Nvidia hardware.
Software will decide how far China’s chip push goes
The partnership addresses a critical bottleneck for China’s AI sector: domestic chips require software capable of extracting maximum performance. Nvidia’s market dominance relies not just on chip design but on an estimated four million developers worldwide using CUDA. This ecosystem acts as a barrier that rivals like AMD have struggled to breach, even when their hardware matched Nvidia on paper.
Chinese model makers such as Z.ai and Moonshot AI have historically moved faster than domestic chip manufacturers. Huawei aims to close this gap. Two weeks prior to this announcement, the company unveiled new AI processors and supernode systems intended for widespread model training next year.
Huawei also acknowledged it cannot meet domestic demand. The company plans to sell fewer chips abroad, citing US export controls. Huawei’s current rotating chairman, Eric Xu, noted the company cannot accept a future dependent on whether other nations are willing to sell chips to China.
CUDA’s moat is weakening, but Nvidia still leads on agent workloads
Research firm SemiAnalysis examined the remaining strength of Nvidia’s software advantage after testing Jalapeño, OpenAI‘s inference chip. The analysts described the CUDA moat as “potentially dead” because OpenAI deploys new models on its own hardware with significant speed. Jalapeño outperformed Nvidia’s Blackwell in most tested scenarios regarding performance per watt.
OpenAI models assisted in the chip’s design, and those same models run on Nvidia GPUs. However, the firm added caveats. Tests focused on relatively simple optimisation scenarios involving 8,000 input tokens and 1,000 output tokens. The team has not yet run AgentX, a benchmark measuring AI agent performance on multistep tasks.
In August, SemiAnalysis found Nvidia well ahead in that specific area. With AMD’s current software stack, Nvidia would remain cheaper per token even if AMD gave away its hardware. The authors do not view Nvidia’s lasting advantage as residing in the silicon itself but in the software that links many chips into a single system.
Huawei’s chips were not included in the AgentX comparison. An earlier analysis of DeepSeek V4, however, noted that Huawei’s CANN software stack was the only option besides CUDA to support the model on day one.
What it means
For developers and model builders in China, this release removes a major friction point. Previously, teams using Huawei hardware were forced to rely on limited or proprietary tools that did not match the flexibility of the global standard. Now, the availability of TileLang allows engineers to write code for Ascend chips using a more intuitive framework. This shift enables faster iteration and potentially lowers the barrier for smaller teams to deploy models on domestic hardware without needing deep expertise in restricted software stacks.




