Industry

DeepSeek's Latest Model Arrives With Native Support for China's Chip Ecosystem

DeepSeek has unveiled DeepSeek-V3.2-Exp with optimizations built in from day one for Huawei, Cambricon, and Hygon accelerators, signaling a strategic pivot away from Nvidia's CUDA dominance.

2 min read
DeepSeek’s new AI model debuts with support for China-native chips and CANN, a replacement for Nvidia's CUDA — Chinese chipmakers Huawei, Cambricon, and Hygon get first-class support

The response from Chinese technology companies has been swift and coordinated. DeepSeek, a prominent artificial intelligence firm based in China, has introduced DeepSeek-V3.2-Exp, its newest large language model, featuring integrated support for Huawei's Ascend processors and the CANN software framework. This release represents a deliberate reorientation toward enabling cutting-edge models to function effectively on homegrown accelerators, moving away from dependence on Nvidia's CUDA platform.

On September 29, DeepSeek made the announcement public, sharing model code and weights via Hugging Face along with accompanying technical documentation. According to the company, V3.2-Exp functions as an "intermediate step toward our next-generation architecture," with the goal of reducing expenses associated with processing extended-context sequences. The architecture incorporates a sparse attention mechanism designed to lower both memory consumption and computational load without sacrificing result quality.

Huawei's Ascend division and the broader vLLM-Ascend community responded rapidly to integrate the new model. Within the vLLM-Ascend repository, developers documented procedures for installing custom operators and packaging kernels tailored for Ascend neural processing units to enable V3.2-Exp compatibility. The CANN team subsequently released an inference guide, preparing the model for rapid rollout across Huawei's hardware lineup.

Additional domestic chipmakers have contributed to this effort. Cambricon released a refreshed version of its vLLM-MLU variant that supports V3.2-Exp, asserting that pairing its inference platform with the model's sparse attention design substantially reduces the expense of handling long-sequence workloads. Hygon announced that its DCU processors had undergone optimization for "zero-wait" deployment via its DTK software framework.

https://twitter.com/cantworkitout/status/1972701949095997940

The velocity of this adoption demonstrates that China's artificial intelligence sector is actively preparing for a scenario where Nvidia hardware availability cannot be guaranteed. While Nvidia's CUDA framework continues to dominate both model training and inference operations, DeepSeek's current release stands out as among the first from a significant Chinese organization to ship with optimization for alternative, non-CUDA software stacks from its initial release.

The synchronized mobilization across Ascend, Cambricon, and Hygon represents the most tangible evidence yet that Chinese technology companies are treating Beijing's push for AI self-sufficiency as a genuine priority. Rather than retrofitting existing hardware for compatibility after development, these firms are now treating domestic platforms as primary targets from the outset.

Source: Tom's Hardware

Source: Tom's Hardware · Reporting supplemented by The Silicon Ledger staff.