Back to News
ProductAI Understanding briefing

DeepSeek releases open‑source software stack for Huawei Ascend AI chips

DeepSeek announced a free, open‑source toolkit—including TileLang, DeepGEMM, FlashMLA and other libraries—tailored for Huawei’s Ascend 950 accelerators, aiming to cut reliance on Nvidia’s CUDA ecosystem.

4 min readRead the linked source
Source-provided image accompanying DeepSeek releases open‑source software stack for Huawei Ascend AI chips
Source referenceSource recorded
Publisher
cryptobriefing.com
Source link
cryptobriefing.comhttps://cryptobriefing.com/deepseek-huawei-ai-chip-software/
Source type
Linked source — primary-source status has not been established.
ContextUnderstand this in 60 seconds

Start here

Key terms

Memory (Agent Memory)
Stored context an AI agent uses across steps or sessions to improve continuity.
Inference
The runtime phase where a trained model generates predictions or outputs.
Compute
The processing resources required to train and run models, often measured in FLOPS or GPU hours.
Test yourselfAI Models Explained Quiz

What happened

DeepSeek published a full software stack for Huawei’s Ascend AI accelerators on September 30 via its WeChat channel. The suite comprises TileLang, a high‑level language positioned as a domestic alternative to Nvidia’s CUDA, plus a set of and communication libraries—DeepGEMM, FlashMLA, TileKernel, DeepSelect and DeepEP. The tools are optimized for the Ascend 950 chip and support a “supernode” mode that can link 128 accelerators. All components are offered as free downloads, and benchmarking utilities are included to help developers measure and tune performance. Huawei is reported to have provided extensive technical support during development, and DeepSeek plans a data‑center in Inner Mongolia that will host up to 160,000 Ascend accelerators.

On September 30, DeepSeek used its official WeChat account to release an open‑source programming stack aimed at Huawei’s Ascend AI accelerators. The package includes TileLang, a high‑level language designed to mirror the functionality of Nvidia’s CUDA, and a collection of libraries—DeepGEMM for matrix multiplication, FlashMLA for memory management, TileKernel for kernel execution, DeepSelect for data selection, and DeepEP for inter‑chip communication. The tools are specifically tuned for the Ascend 950 chip and support a supernode configuration that can interconnect 128 chips, enabling large‑scale parallel processing.

All components are freely downloadable, reflecting DeepSeek’s strategy to lower entry barriers for developers who might otherwise default to Nvidia’s ecosystem. The release also bundles benchmarking utilities that allow developers to assess kernel performance and optimize code for the Ascend architecture. Huawei is said to have supplied extensive technical assistance throughout the development process.

DeepSeek announced plans to build a data‑center in Inner Mongolia that will house up to 160,000 Ascend accelerators, indicating a long‑term commitment to scaling the hardware‑software stack. This infrastructure aim aligns with the broader goal of creating a self‑sufficient AI ecosystem within China.

Source details: cryptobriefing.com ↗

Why it matters

The release tackles a strategic bottleneck for China’s AI ecosystem. Nvidia’s CUDA framework has become the de‑facto standard for AI model training and , and U.S. export controls have limited Chinese access to Nvidia’s latest chips. By delivering a complete, free programming environment for Huawei’s Ascend hardware, DeepSeek lowers the cost and technical barrier for developers to adopt domestic silicon, potentially reducing dependence on foreign tooling. The toolkit’s high‑level language and performance‑critical libraries enable existing CUDA‑based codebases to be ported with less re‑engineering effort, accelerating the migration to Huawei chips. If widely adopted, the stack could spur a broader AI software ecosystem around Ascend, influencing hardware procurement decisions and reshaping competitive dynamics in the global AI chip market.

Nvidia’s CUDA has dominated AI software development for years, making it a strategic vulnerability for countries restricted from accessing Nvidia’s latest chips. By providing a domestic alternative, DeepSeek’s toolkit directly addresses this dependency, offering Chinese developers a path to continue AI research and product development on locally produced hardware.

The free, open‑source nature of the stack reduces financial and licensing hurdles, encouraging broader community contributions and faster iteration. This could accelerate the maturation of a Chinese AI software ecosystem, potentially leading to new models, tools, and services built on Ascend hardware.

If the toolkit delivers performance comparable to CUDA, it may influence procurement decisions for enterprises and cloud providers within China, shifting market share away from Nvidia‑centric solutions and reshaping the global AI chip landscape.

Interactive Mechanism

Interactive Mechanism: How It Actually Works

Explore the underlying technology behind this development interactively.

Document Size:128K tokens
Needle Placement Depth (Location in document):50% into text
Attention Context Buffer Map:
Target Fact (50%)
Equivalent Pages~320Standard book pages
Retrieval Accuracy99.9%Needle recall score
RAM / KV Cache5.1 GBMemory overhead
Prompt CachingActive~80% discount on reuse
Core takeaway: Million-token context windows allow querying whole codebases or legal archives in one prompt. However, KV cache memory scales with context length, making prompt caching crucial for real-time production.
Interactive Concept Check+10 Points
AI Models Explained Quiz

In AI, what are a model's "parameters"?

What to watch next

Key indicators to monitor include adoption rates among Chinese AI startups, performance benchmarks comparing TileLang‑based workloads to CUDA equivalents, and any policy shifts that might further restrict or enable foreign chip software. The planned Inner Mongolia data‑center’s rollout timeline and the scale of its accelerator deployment will also signal the practical impact of the toolkit. Additionally, follow‑up announcements from Huawei or DeepSeek about updates to TileLang or new library versions could reveal how quickly the ecosystem matures.

Adoption metrics: number of projects, repositories, or companies publicly using TileLang and the associated libraries.

Performance benchmarks: independent tests comparing TileLang‑based workloads on Ascend chips to CUDA workloads on Nvidia GPUs, especially for large‑scale training and tasks.

Policy environment: any new export controls, subsidies, or regulatory incentives that could affect the attractiveness of domestic versus foreign AI tooling.

Infrastructure rollout: progress on the Inner Mongolia data‑center, including the timeline for deploying the announced 160,000 Ascend accelerators.

Future releases: updates to TileLang or additional libraries from DeepSeek or Huawei that expand functionality or improve performance.

Related guides & quizzes

AI Models ExplainedAI TrainingFuture of AITest what you know — try a free AI quizLook up an AI term in our glossaryFollow the AI model release tracker
Found this useful?