Chinese AI vendor Z.ai said it used 100,000 China-made chips to handle online queries to its latest AI model, GLM-5.3 Flash.The move shows how Chinese vendors are reducing their dependence on U.S. chipmaker Nvidia, while also highlighting the trend toward greater optimization of AI models.The new open-weight, multimodal model is low-cost, with 320 billion parameters. Z.ai originally previewed the model under the code name Ox Alpha on August 20 and officially released it under the MIT open source license on August 26. The model is specialized for long-context processing, vision-driven agentic tasks and code synthesis.GLM-5.3 Flash costs $0.075 per million input tokens and $0.25 per million output tokens from now until Sept. 9. After that, the price will be $0.15 per million input tokens and $0.50 per million output tokens. Comparatively, GPT-5.6 Luna from OpenAI costs $0.20 per million input tokens and $2 per million output tokens. Anthropic Claude Opus 5 is $5 per million input tokens and $5 per million output tokens.Related:Qwen 3.8 Flash-Next is Cheap, But There Are Complicating FactorsThe Chinese vendor’s decision to use purely Chinese chip providers comes as China aims to rely less on Nvidia. China’s pivot toward self-reliance comes after years in which the U.S. tightened export controls to prevent Chinese tech vendors from obtaining the most powerful Nvidia chips. Although some of these controls have eased, in September 2025, China’s government ordered tech giants such as Alibaba and ByteDance to stop buying Nvidia AI chips, and in November, Beijing banned all foreign AI chips from state-funded data centers.The Advancement of Chinese HardwareDespite these moves, many observers still saw China as behind in the infrastructure contest. However, Z.ai’s strategy suggests that the gap between China and the U.S. in AI chips may be narrowing, especially with the increased focus on inference over the past year amid the sharp rise of agentic AI.“For training the largest frontier models, Nvidia still has a significant advantage in chip performance, networking and the broader software ecosystem,” said Kashyap Kompella, CEO and founder of RPA2AI Research. “But inference is a different story.”He added that Z.ai’s ability to serve GLM-5.3-Flash at scale using Chinese chips indicates that Chinese hardware is already good enough for many high-volume inference workloads.“This distinction is important because inference is likely to become the larger AI chip market over time,” Kompella said. “Chinese companies may not need exact chip-for-chip parity with Nvidia if they can compensate through efficient model architectures, software optimization and large clusters of domestically made AI chips.”Related:Cost Challenges With Perplexity Portable ComputerHowever, the first signals of Chinese vendors’ move toward independence began with DeepSeek rewriting Nvidia CUDA drivers to make inference faster, said Bradley Shimmin, an analyst at Futurum Group.He said that while Chinese vendors are responding to constraints first imposed by U.S. government and later by the Chinese government, another factor is that companies “need to do more with the watt hours they have available to them.”Vendors such as Google, with its TPU series and Amazon, with its Trainium chips, are focusing on vertical integration, in which the AI stack is optimized to use AI models, infrastructure and software together for the most efficient AI processing.“What we’re witnessing here is this real realization that AI value is driven by economics of the stack, as much as it is about the capabilities of the models,” Shimmin said. “What these guys are doing is simply showcasing the value of a little bit of investment in optimizing that hardware.”The ModelAs for GLM-5.3-Flash itself, the fact that it received strong developer mindshare in a blind test is significant, Kompella said. Many developers flocked to the model early this month and used it without knowing it came from a Chinese vendor; it quickly became the most downloaded model on OpenRouter.Related:Mistral and Saudi Vendor to Advance Sovereign AI in Middle East“That is more interesting than another benchmark claiming that a Chinese model is close to a U.S. frontier model,” Kompella said, adding that it shows that Chinese open-weight models “continue to function as a counterweight to higher industry pricing.”For enterprises, this means they need to design their AI system “so workloads can be routed across models based on capability, quality, cost and strategic considerations rather than becoming dependent on a single model provider,” he continued.Although the model attracted a large number of developers right away is a good sign for Z.ai, the vendor still faces challenges ahead. For one, it must prove that the chips it is using can continue to deliver reliability, economy and scale, Kompella said.“At the absolute frontier of model training, access to compute also remains a meaningful constraint because Nvidia’s advantage is much harder to replicate there,” he said.The vendor will also need to convert U.S. and European enterprises that remain cautious about using Chinese-hosted AI services due to data security, regulatory and procurement concerns, Kompella continued.About the AuthorNews Writer, AI BusinessEsther Shittu has covered AI technologies and industry trends since 2021. As co-host of the Targeting AI podcast, she talks with experts, thought leaders and practitioners exploring critical AI developments. Before AI Business, she wrote for SearchEnterpriseAI, the New York Daily News, Bklyner and the Brooklyn Daily Eagle. When she's not diving deep into the world of AI, she spends her time on passion projects and raising her three daughters.
Z.AI's Use of Chinese Chips for New Model is About Optimization
Full Article
Original Source
Read the full article at Aibusiness →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.