An AI chip on its own does nothing. To run a model, developers need software that turns their code into instructions the hardware understands. NVIDIA‘s version of that software is called CUDA. It has spent almost twenty years becoming the industry default, which is why a faster-looking rival chip can still sit unused. If the world’s existing code will not run on it, the board is a paperweight.
That lock-in is what people mean when they call CUDA NVIDIA’s moat: a ditch around the castle that is very expensive to cross. DeepSeek‘s TileLang, released today for Huawei’s Ascend chips, is the first serious attempt at a bridge.
What Is Ascend?
When Washington blocked NVIDIA from selling its best AI chips to China, Beijing had to find a substitute.
Ascend is Huawei’s line of AI accelerators, which the company prefers to call NPUs, neural processing units. The 910B and 910C are the cards already in Chinese data centres. The 950PR began shipping in 2026, with a 950DT due later this year. The 960 is on the 2027 roadmap, pulled forward at Huawei Connect in September.
On raw silicon the line has become competitive enough to run ten-thousand-card clusters. The problem was never only the chip. It was everything sitting on top of it.
The Missing Piece: DeepSeek’s TileLang
Huawei’s answer to CUDA is CANN, short for Compute Architecture for Neural Networks. Although it does work, developers have long said it is awkward to write for, which is why many steered clear of Ascend even when they could get the boards.
On September 30, DeepSeek released a free toolkit for that stack. The centrepiece is an Ascend backend for TileLang, a tiled kernel language that already existed as open source. DeepSeek says it is simpler to write than CUDA. TileLang sits above CANN: it lowers into Ascend C and the CANN toolchain rather than replacing them.
The lab also shipped matching libraries (i.e. DeepGEMM, DeepEP, TileKernels, FlashMLA and DeepSelect) covering matrix maths, inter-card communication, vector ops, long-context attention and data selection, with a second backend that targets the Ascend 950.
For the first time there is a serious Chinese package from chip to high-level kernels that does not owe NVIDIA a licence. That is what people mean by a CUDA-free stack. It is not the first software on Ascend; CANN, MindSpore and torch_npu already existed.
What The Stack Does Not Settle
A working toolkit lowers the cost of leaving CUDA. It gives Chinese labs a way to grow on chips they can still buy, and it shows every other region, Europe included, that an independent stack is a software problem as much as a silicon one. It does not mean Ascend has caught NVIDIA. Independent estimates still put Huawei behind on efficiency, and volume is a fraction of NVIDIA’s. Switching remains painful for anyone whose models, staff and vendors already speak CUDA.
Author: Grace Sharp

