Back to News
ProductAI Understanding briefing

DeepSeek and Huawei open‑source tools for Ascend AI chips aim to cut Nvidia reliance

DeepSeek and Huawei have released an open‑source software stack—including compute and communication libraries and TileLang support—for Huawei’s Ascend AI processors, targeting developers who want to avoid Nvidia’s CUDA ecosystem.

4 min readRead the original reporting
Source-provided image accompanying DeepSeek and Huawei open‑source tools for Ascend AI chips aim to cut Nvidia reliance
Attributed reportingSource recorded
Publisher
tomshardware.com
Source link
tomshardware.comhttps://www.tomshardware.com/tech-industry/artificial-intelligence/deepseek-and-huawei-release-open-source-ascend-ai-programming-tools-to-reduce-reliance-on-nvidia-ecosystem-tools-include-compute-and-communication-libraries-as-well-as-ascend-support-for-tilelang
Source type
Reporting by a news outlet — not a first-party document.

What we could not confirm independently: This claim is attributed to the named outlet. We did not verify it against a first-party document. (tomshardware.com)

ContextUnderstand this in 60 seconds

Start here

Key terms

Large Language Model (LLM)
A language model trained on massive text corpora to generate and analyze text.
Benchmark
A standardized test or dataset used to measure and compare model performance.
Inference
The runtime phase where a trained model generates predictions or outputs.
Test yourselfAI Models Explained Quiz

What happened

DeepSeek announced the release of a suite of open‑source programming tools for Huawei’s Ascend AI chips. The bundle includes libraries that handle AI‑specific computation and chip‑to‑chip communication, as well as support for the high‑level language TileLang on Ascend hardware. According to a Reuters report cited by Tom’s Hardware, the tools were co‑developed with Huawei, which provided full engineering support. The companies also optimized the stack for a super‑node configuration built around 128 Ascend 950 accelerators, addressing both efficient per‑chip calculation and fast inter‑chip data movement required for large‑scale AI workloads.

DeepSeek, a Chinese AI startup, released an open‑source software stack for Huawei’s Ascend AI processors. The stack comprises two primary libraries: one for AI‑specific computation kernels and another for high‑throughput chip‑to‑chip communication. Both libraries are intended to run efficiently on Ascend hardware without requiring Nvidia’s CUDA drivers.

In addition to the low‑level libraries, the release adds Ascend support for TileLang, a high‑level programming language designed to simplify AI model development. TileLang abstracts hardware details, allowing developers to write code that can be compiled for multiple accelerator architectures.

The collaboration between DeepSeek and Huawei also involved performance tuning for a super‑node system that links 128 Ascend 950 chips. The optimization focuses on two critical challenges for large AI models: maximizing per‑chip compute utilization and minimizing data‑transfer latency across the node.

The tools are hosted publicly, with source code and build instructions available for developers. DeepSeek states that Huawei provided full engineering support throughout development, but no pricing or commercial licensing details were disclosed. The release is positioned as a community‑driven effort to reduce reliance on Nvidia’s proprietary software stack.

Source details: tomshardware.com ↗

Why it matters

The release directly challenges the dominance of Nvidia’s CUDA ecosystem by giving developers a viable, open‑source alternative for high‑performance AI workloads on Ascend silicon. By lowering the software barrier, the stack could broaden the adoption of Huawei’s AI hardware, especially in regions or organizations that are seeking to diversify away from Nvidia for cost, geopolitical, or supply‑chain reasons. If the tools deliver the promised performance, they may shift market dynamics, encourage more competition in AI‑accelerator software, and spur further open‑source contributions to non‑CUDA ecosystems. However, the actual impact will depend on community uptake, documentation quality, and real‑world results, none of which have been independently verified yet.

Reducing dependence on Nvidia’s CUDA ecosystem can lower costs for organizations that currently pay licensing fees or face supply constraints for Nvidia GPUs. An open‑source alternative also mitigates geopolitical risks for companies operating in regions where Nvidia hardware may be restricted.

By providing a ready‑to‑use programming model (TileLang) and performance‑critical libraries, the stack lowers the technical barrier for developers to experiment with Ascend chips. This could accelerate the growth of a software ecosystem around Huawei’s hardware, which has historically lagged behind Nvidia’s extensive tooling.

If the stack delivers comparable or superior performance on large models, it may encourage cloud providers and enterprises to consider Ascend‑based offerings, diversifying the AI‑accelerator market. Such diversification can foster innovation, as hardware vendors compete not only on raw performance but also on the quality and openness of their software ecosystems.

The open‑source nature invites community contributions, potentially leading to rapid bug fixes, feature additions, and broader compatibility with AI frameworks like PyTorch or TensorFlow. However, the lack of independent data means the actual performance gains remain unverified.

Interactive Mechanism

Interactive Mechanism: How It Actually Works

Explore the underlying technology behind this development interactively.

Model Parameter Size:8B Parameters
VRAM Required5.5 GBGPU memory footprint
Target HardwareMacBook / Single GPUDeployment tier
Privacy100% Air-GappedLocal device capability
Core takeaway: Small, quantized models (3B–8B) now run directly inside smartphones and laptops with complete data privacy, while mammoth 400B+ models remain the domain of datacenter clusters.
Interactive Concept Check+10 Points
AI Models Explained Quiz

In AI, what are a model's "parameters"?

What to watch next

Key indicators to monitor include developer adoption rates, performance benchmarks comparing Ascend + TileLang against Nvidia CUDA on comparable models, and any subsequent updates from DeepSeek or Huawei expanding the toolset. Industry reactions—particularly from cloud providers and AI‑chip manufacturers—will reveal whether the stack can erode Nvidia’s market share. Watch for announcements of additional hardware support, integration with popular AI frameworks, and any licensing or support policies that could affect enterprise deployment.

Developer uptake: number of GitHub stars, forks, and contributions to the repository over the next few months.

releases: independent performance comparisons of Ascend + TileLang versus Nvidia CUDA on standard AI workloads (e.g., large language model ).

Enterprise announcements: cloud providers or AI service firms stating support for Ascend chips using the new stack.

Further tooling: any follow‑up releases from DeepSeek or Huawei that add support for additional frameworks, model formats, or hardware generations.

Policy and supply‑chain shifts: reactions from governments or industry groups that may influence hardware procurement decisions away from Nvidia.

Related guides & quizzes

AI Models ExplainedAI TrainingFuture of AITest what you know — try a free AI quizLook up an AI term in our glossaryFollow the AI model release tracker
Found this useful?