💻 DeepSeek ports DeepGEMM to Huawei Ascend and squeezes out 99.8% of hardware peak

DeepSeek has released DeepGEMM-Ascend — a port of the DeepGEMM matrix kernel library to Huawei Ascend NPU. The first release on September 30, 2026 supports the Ascend 950 series: GEMM in FP4, FP8, and BF16, grouped GEMM for MoE and MegaMoE. The API is fully compatible with the original. On Ascend 950DT (CANN 9.20), dense GEMM reaches 98.3–99.8% of the hardware limit: FP4 — 1701 TFLOPS, FP8 — 861, BF16 — 431.

🌍 The Ascend stack has gained a low-level library at the CUTLASS level, and Ascend 950 has demonstrated near-peak FP4/FP8 performance. A real step toward training and inference of DeepSeek models outside the CUDA stack.

👤 The code is open under the MIT license and API-compatible with the original: it reveals the device of extreme optimizations — coroutines, sparse data loading, ABI UE8M0. Ascend 950 owners only need CANN 9.20, torch_npu, and Python 3.10+ — benchmarks are in tests/.

Source 1: https://github.com/deepseek-ai/DeepGEMM-Ascend