DeepSeek has open-sourced DeepGEMM-Ascend — a port of the DeepGEMM matrix kernel library to Huawei Ascend NPUs. The release is distributed under the MIT license, is fully API-compatible with the original, and supports the Ascend 950 series. On the Ascend 950DT chip, developers claim up to 99.8% hardware utilization of the hardware limit for dense GEMM. For the industry, this is a practical step toward training and inference of DeepSeek models outside the NVIDIA CUDA stack.
What happened
DeepSeek published the DeepGEMM-Ascend repository — a version of the DeepGEMM matrix kernel library rewritten for Huawei Ascend NPUs. The first release is dated September 30, 2026, and supports the Ascend 950 series. The API is fully compatible with the original: it uses the same deep_gemm package, so calling code does not need to be rewritten. The kernels have been reworked for Ascend matrix instructions with JIT compilation and coroutine-based pipelining. It supports GEMM in FP4, FP8, and BF16 formats, grouped GEMM for MoE, MQA logits for the DeepSeek Lightning Indexer, MegaMoE, and HC Prenorm GEMM. On Ascend 950DT with CANN 9.20, dense GEMM, according to DeepSeek, achieves 98.3–99.8% of the hardware limit: 1701 TFLOPS in FP4, 861 TFLOPS in FP8, and 431 TFLOPS in BF16. The team separately thanked Huawei for engineering support.
Context
Matrix multiplication (GEMM) is a fundamental operation in transformers, largely determining the speed of training and inference. Until recently, the ecosystem of low-level kernels for it existed virtually only in the NVIDIA CUDA stack: there were no equivalents of CUTLASS-class libraries for other platforms, and performance outside NVIDIA was controlled by the vendor itself. Against this backdrop, the claimed peaks of DeepGEMM-Ascend relate as 4:2:1 (1701 to 861 to 431 TFLOPS), which is typical for tensor architectures and indirectly supports the consistency of the numbers themselves. Notably, the coverage is not an abstract GEMM but primitives of the actual DeepSeek architecture, meaning it is designed for end-to-end training and inference, not synthetic tests. The publication of the optimization techniques for the new architecture — from sparse data loading to the UE8M0 scaling factor ABI — is of particular value to engineers: usually such details remain hidden inside closed vendor stacks.
Why this matters for the industry
For the industry, this is the first verifiable signal of the erosion of the CUDA moat at the level of low-level matrix kernels: a high-performance matrix layer on Huawei NPUs is turning from a unique hardware feature into a replaceable component of a working stack. The combination of the MIT license, API compatibility, and the same deep_gemm package allows teams already using DeepGEMM to port research code with minimal changes and compare CUDA and NPU implementations on identical calls. The result also confirms that Ascend 950 can handle FP4 and FP8 at near-peak loads, opening a practical path to training and inference of DeepSeek models outside NVIDIA, and in the future — to inference services with dual-vendor deployment around MoE models. For startups, compute is no longer a de facto exclusive of a single vendor, so assumptions about compute costs in financial models should be revisited. If support in CANN is solidified and external contributions appear, the port could grow into a full-fledged CUTLASS-level kernel ecosystem for Ascend.
Why this matters for users
The code is open under the MIT license, so any engineer can examine how extreme optimizations for a foreign architecture are built. Owners of Ascend 950 hardware only need CANN 9.20, torch_npu, and Python 3.10+ to clone the repository, build benchmarks from the tests/ directory, and verify the claimed 1701, 861, and 431 TFLOPS on their own hardware. Thanks to API compatibility, this is a cheap and honest first step: at the current stage, benchmarking and PoC integration are appropriate, not production deployment. For other readers, the value for now lies in reading the source code as research material; the community around the project is just forming — at the time of publication, the repository has one commit.
What is still unknown / limitations
The key figures (98.3–99.8% of the hardware peak and the claimed TFLOPS) are self-declared by DeepSeek: there are no independent measurements yet. A near-peak result on dense regular GEMM does not prove the maturity of the entire CANN compiler stack, because this is structurally the most optimization-friendly case; maturity should be demonstrated on irregular matrix shapes, grouped MoE cases, and end-to-end loads, and such data is not available in the sources. The library covers only the GEMM layer: attention, all-reduce, and KV-cache remain outside its scope, so it is not yet sufficient for production serving. The repository is new, with a single commit, and the community is just forming; whether the port will take root depends on Huawei's roadmap, CANN stability, and community activity, while the economic effect depends on the price and availability of Ascend 950.
Sources
Author
Look at AI, editorial team
