China’s leading model developers are still relying on Nvidia chips to train some of their most advanced models, despite Beijing’s push to replace foreign semiconductors with domestic alternatives. The problem is not simply whether Chinese-made chips can handle the workload. Switching hardware also means changing the software systems built around Nvidia’s CUDA platform, creating a costly engineering challenge for model developers. Huawei has developed its own alternative, Compute Architecture for Neural Networks (CANN), for its Ascend chips. But developers accustomed to CUDA may need to rewrite and optimize substantial portions of their existing code before moving their training workloads to Ascend. One researcher involved in model development told the South China Morning Post that his team’s existing training pipelines depend on CUDA, making a hardware switch far from straightforward. “CUDA code cannot run directly on Ascend and requires extensive rewriting.” Software keeps Nvidia ahead James Wang, who develops models at a research institute affiliated with a Shanghai-based university, estimated that moving his team’s workflows to Huawei’s Ascend chips could increase the time and cost of the process by at least 50%. The burden varies depending on how open the model is and how mature the supporting software ecosystem has become. A Beijing-based engineer involved in moving model workloads to Ascend told SCMP that open-source models can be migrated more easily because developers can use existing ecosystem support. For models such as DeepSeek and its distilled versions, the engineer estimated that two or three developers could complete the transition with roughly an additional month of work. The situation becomes considerably more complicated for models whose underlying source code is not publicly available. A model such as Moonshot AI’s Kimi K3, which has released its weights but not its source code, could require about 10 engineers and more than six months of additional work to make the switch, according to the engineer. That gap highlights why China’s semiconductor self-sufficiency drive involves more than producing domestic chips. Developers also need software tools, libraries and optimized workflows that can compete with the ecosystem Nvidia has built around CUDA. Training remains hardest hurdle The distinction between training and inference is also important. Training a model requires enormous computing resources and involves running complex workloads repeatedly across large clusters. Inference, by comparison, uses an already trained model to generate outputs from new data and is generally easier to adapt to different hardware. Chinese developers have already demonstrated that some advanced models can run on domestic hardware. DeepSeek-V4 and Kimi K3 have reportedly been adapted for Chinese platforms including Huawei and Alibaba’s systems, although neither company has disclosed the chips used to train those models. Kimi K3 was reportedly trained using Nvidia hardware, including advanced Blackwell processors, according to a July report by The Information. The report also indicated that Chinese developers continued to use Nvidia chips for training advanced models despite efforts to expand domestic alternatives. There are signs that domestic hardware is beginning to take a larger role. In June, Meituan released LongCat-2.0, a 1.6-trillion-parameter model that the company said was fully trained and operated on a 50,000-card domestic computing cluster, without identifying the chip supplier.The challenge for China is therefore moving beyond simply building alternatives to Nvidia. Domestic chips must also support the software ecosystem and development workflows needed to train increasingly large models at competitive speed and cost.Recommended ArticlesGet the latest in engineering, tech, space & science - delivered daily to your inbox.With over a decade-long career in journalism, Neetika Walter has worked with The Economic Times, ANI, and Hindustan Times, covering politics, business, technology, and the clean energy sector. Passionate about contemporary culture, books, poetry, and storytelling, she brings depth and insight to her writing. When she isn’t chasing stories, she’s likely lost in a book or enjoying the company of her dogs.
Why is Nvidia still powering China’s model race? Huawei chips face costly software wall
Full Article
Original Source
Read the full article at Interestingengineering →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.