What happened
DeepSeek published a full software stack for Huawei’s Ascend AI accelerators on September 30 via its WeChat channel. The suite comprises TileLang, a high‑level language positioned as a domestic alternative to Nvidia’s CUDA, plus a set of and communication libraries—DeepGEMM, FlashMLA, TileKernel, DeepSelect and DeepEP. The tools are optimized for the Ascend 950 chip and support a “supernode” mode that can link 128 accelerators. All components are offered as free downloads, and benchmarking utilities are included to help developers measure and tune performance. Huawei is reported to have provided extensive technical support during development, and DeepSeek plans a data‑center in Inner Mongolia that will host up to 160,000 Ascend accelerators.
On September 30, DeepSeek used its official WeChat account to release an open‑source programming stack aimed at Huawei’s Ascend AI accelerators. The package includes TileLang, a high‑level language designed to mirror the functionality of Nvidia’s CUDA, and a collection of libraries—DeepGEMM for matrix multiplication, FlashMLA for memory management, TileKernel for kernel execution, DeepSelect for data selection, and DeepEP for inter‑chip communication. The tools are specifically tuned for the Ascend 950 chip and support a supernode configuration that can interconnect 128 chips, enabling large‑scale parallel processing.
All components are freely downloadable, reflecting DeepSeek’s strategy to lower entry barriers for developers who might otherwise default to Nvidia’s ecosystem. The release also bundles benchmarking utilities that allow developers to assess kernel performance and optimize code for the Ascend architecture. Huawei is said to have supplied extensive technical assistance throughout the development process.
DeepSeek announced plans to build a data‑center in Inner Mongolia that will house up to 160,000 Ascend accelerators, indicating a long‑term commitment to scaling the hardware‑software stack. This infrastructure aim aligns with the broader goal of creating a self‑sufficient AI ecosystem within China.
Source details: cryptobriefing.com ↗
Why it matters
The release tackles a strategic bottleneck for China’s AI ecosystem. Nvidia’s CUDA framework has become the de‑facto standard for AI model training and , and U.S. export controls have limited Chinese access to Nvidia’s latest chips. By delivering a complete, free programming environment for Huawei’s Ascend hardware, DeepSeek lowers the cost and technical barrier for developers to adopt domestic silicon, potentially reducing dependence on foreign tooling. The toolkit’s high‑level language and performance‑critical libraries enable existing CUDA‑based codebases to be ported with less re‑engineering effort, accelerating the migration to Huawei chips. If widely adopted, the stack could spur a broader AI software ecosystem around Ascend, influencing hardware procurement decisions and reshaping competitive dynamics in the global AI chip market.
Nvidia’s CUDA has dominated AI software development for years, making it a strategic vulnerability for countries restricted from accessing Nvidia’s latest chips. By providing a domestic alternative, DeepSeek’s toolkit directly addresses this dependency, offering Chinese developers a path to continue AI research and product development on locally produced hardware.
The free, open‑source nature of the stack reduces financial and licensing hurdles, encouraging broader community contributions and faster iteration. This could accelerate the maturation of a Chinese AI software ecosystem, potentially leading to new models, tools, and services built on Ascend hardware.
If the toolkit delivers performance comparable to CUDA, it may influence procurement decisions for enterprises and cloud providers within China, shifting market share away from Nvidia‑centric solutions and reshaping the global AI chip landscape.
Interactive Mechanism: How It Actually Works
Explore the underlying technology behind this development interactively.
In AI, what are a model's "parameters"?
What to watch next
Key indicators to monitor include adoption rates among Chinese AI startups, performance benchmarks comparing TileLang‑based workloads to CUDA equivalents, and any policy shifts that might further restrict or enable foreign chip software. The planned Inner Mongolia data‑center’s rollout timeline and the scale of its accelerator deployment will also signal the practical impact of the toolkit. Additionally, follow‑up announcements from Huawei or DeepSeek about updates to TileLang or new library versions could reveal how quickly the ecosystem matures.
Adoption metrics: number of projects, repositories, or companies publicly using TileLang and the associated libraries.
Performance benchmarks: independent tests comparing TileLang‑based workloads on Ascend chips to CUDA workloads on Nvidia GPUs, especially for large‑scale training and tasks.
Policy environment: any new export controls, subsidies, or regulatory incentives that could affect the attractiveness of domestic versus foreign AI tooling.
Infrastructure rollout: progress on the Inner Mongolia data‑center, including the timeline for deploying the announced 160,000 Ascend accelerators.
Future releases: updates to TileLang or additional libraries from DeepSeek or Huawei that expand functionality or improve performance.