According to Bloomberg, DeepSeek is preparing an order for at least 160,000 Huawei Ascend 950DT accelerators for a new data center with a capacity of about 1 GW in Ulanqab (Inner Mongolia). If the plans are realized, this will become the largest known cluster built on Huawei chips.

What happened

The order is aimed at inference — running DeepSeek models, while the company continues to train on Nvidia accelerators. The stated specifications of the Ascend 950DT are 144 GB of HBM memory (HiZQ 2.0) and 4 TB/s of memory bandwidth: in terms of memory volume, the chip is close to the Nvidia H200 (141 GB, 4.8 TB/s), but in terms of memory bandwidth, it is about 20% lower. DeepSeek wants to launch part of the capacity by the end of 2027 — the beginning of 2028, depending on Huawei's production capacity. Bloomberg estimates the level of the 950DT to be approximately at the level of the Hopper generation (H100/H200).

Context

The mechanism for the emergence of such an order is a combination of US export restrictions and the Chinese state program for chip localization. Until now, Huawei has mainly covered pilot implementations, and here for the first time there is talk of mass production loads. An unusual detail: Huawei positioned the 950DT as a chip for training, but DeepSeek is reserving it specifically for inference. For DeepSeek's MoE models, this is logical: inference is heavily dependent on the volume and bandwidth of HBM, since the active part of the model must fit into the accelerator's memory. At the same time, the main limiting factor is not demand, but memory supply: CXMT has only just begun producing small batches of HBM3E and is lagging behind Samsung, SK Hynix, and Micron, which are already producing HBM4, by 3–5 years.

Why this matters for the industry

For the industry, this is the first time a Chinese inference stack has moved beyond pilots into mass production use and a signal of the formation of a second largest source of inference capacity. The market is getting specific guidelines for long-term planning: 160,000 Ascend 950DT accelerators, about 1 GW, launching part of the cluster by the end of 2027 — the beginning of 2028. If part of the capacity comes online, by 2028 the first measurable production metrics of the Chinese stack will appear — tokens per second per chip, token cost, reliability — which will translate the marketing claim of Hopper level into numbers and set a new benchmark for large-scale MoE inference without Nvidia. The event does not affect the current capabilities of models and public benchmarks.

Why this matters for users

There is no direct impact on readers now: no new APIs, pricing, latency, or anything that can be tried today has appeared. The practical meaning is long-term: launching part of the cluster by the end of 2027 — the beginning of 2028 will provide the first real production measurements of inference on the Ascend 950DT at a scale of 160,000 chips, and for the first time it will be possible to measure the Chinese stack in real inference loads of such scale, not in benchmarks. A possible reduction in the cost of inference for DeepSeek models may make them economically more accessible in the long term, but this is a scenario, not a guarantee.

What is still unknown / limitations

The order has not yet been officially confirmed, and the figure of 160,000 is a Bloomberg estimate in the wording "at least". There are no measurable parameters of the future cluster: no latency, no cost, no real uptime. The estimate of "approximately Hopper level" is media expertise, not an independent benchmark: there may be a gap between the stated specifications and real performance. Launch deadlines depend on Huawei's production capacity and HBM supply from CXMT, which may shift the launch beyond the end of 2027 — the beginning of 2028. The question remains open as to how the 950DT will perform in mass inference of MoE models and how it really differs from the Nvidia H200 in workloads.

Sources

Author

Look at AI, editorial team