Why Zhipu Ai Big Bet On Homegrown Silicon Should Concern Nvidia

Why Zhipu Ai Big Bet On Homegrown Silicon Should Concern Nvidia

When Bloomberg reported that Beijing-based AI startup Z.AI, formerly Zhipu AI, built and began operating a 1-gigawatt data center using zero Nvidia chips, market watchers took notice. Zhipu's stock surged over 36% on the Hong Kong Stock Exchange shortly after the news broke.

For the past three years, the tech industry assumed that training top-tier AI models required silicon from Santa Clara. American export restrictions were designed to freeze Chinese AI labs in place by cutting off access to Nvidia's H100 and A100 processors. Instead, Zhipu built a massive compute site powered entirely by domestic chips, running 10,000-chip clusters to train its next-generation GLM models.

If you're tracking hardware, cloud infrastructure, or tech stocks, this move isn't just another headline about China's tech ecosystem. It marks a fundamental shift in how frontier AI models get built when access to top-shelf hardware disappears.

The 1 Gigawatt Reality Check

To put a 1-gigawatt data center into context, that's enough power to run roughly 750,000 homes simultaneously. It's a massive footprint for any computing facility, let alone one dedicated strictly to training large language models.

Most people assumed Chinese AI labs would hit a hard ceiling without Nvidia's high-bandwidth memory and unified CUDA architecture. Operating clusters at this scale requires immense inter-chip communication speed. When you string tens of thousands of GPUs together, bottlenecks kill performance long before raw compute power becomes the limiting factor.

Zhipu proved that raw engineering effort can bridge part of that hardware gap. By linking clusters of more than 10,000 domestic chips together, the company created enough collective throughput to handle heavy training workloads for its GLM-4.6 and GLM-5.2 models. They aren't just using local chips for light inference tasks. They're actively using domestic stacks to train foundation models from scratch.

Why Domestic Silicon Is Suddenly Good Enough

American trade controls forced a brutal choice on Chinese software labs. They could either buy smuggled or downgraded Nvidia chips at inflated prices, or they could make domestic hardware work. Zhipu chose the latter, partnering heavily with local chipmakers like Huawei and Cambricon.

Early Chinese AI accelerators suffered from buggy drivers, poor software optimization, and high failure rates during long training runs. Running a thousand-chip cluster for weeks without a crash was nearly impossible two years ago.

That dynamic changed fast. Driven by necessity, Chinese chipmakers drastically improved their software development kits and runtime libraries. Huawei's Ascend ecosystem and Cambricon's hardware stacks matured because developers had no choice but to file bug reports and write custom kernels.

👉 See also: this story

Zhipu didn't wait for domestic chips to match Nvidia's flagship performance on a 1-to-1 basis. They simply scaled out. If a single domestic accelerator offers 60% of an H100's performance, you build a bigger cluster, optimize your parallel processing software, and accept slightly higher power bills. When electricity is cheap and state backing is strong, that math works out remarkably well.

Software Bridge That Makes Mixed Hardware Work

One of the biggest hurdles in dropping Nvidia hardware is CUDA, Nvidia's proprietary software framework. Almost every major AI model over the last decade was built natively for CUDA. Switching to alternative hardware usually means rewriting code from scratch and running into severe performance drops.

Zhipu addressed this software problem directly by acquiring Zhongke Jiahe, also known as XCore Sigma.

XCore Sigma specializes in heterogeneous computing software. Their tech creates compiler layers and runtime systems that allow AI models to run across different chip architectures without requiring developers to rewrite core algorithms.

This acquisition gives Zhipu three distinct advantages:

  • Hardware Agnosticism: They can deploy workloads across chips from Huawei, Cambricon, and Moore Threads inside the same broader network.
  • Lower Inference Costs: Custom compilation tools optimize how models run on cheaper, localized hardware, driving down API costs for users.
  • Fast Deployment: New model checkpoints can be deployed to enterprise customers without waiting for single-vendor hardware supply chains.

By solving the software translation layer, Zhipu removed the primary operational barrier preventing companies from ditching Nvidia.

Enterprise Demand Is Driving Real Revenue

Wall Street and venture capital firms often criticize AI startups for high cash burn and weak revenue. Zhipu is turning out to be an exception in the Chinese market.

The company is on track to hit $1 billion in annual recurring revenue, having met its full-year sales targets ahead of schedule. A massive portion of that revenue comes from on-premises deployments for Chinese state-owned enterprises, financial institutions, and government bodies that are legally required to use domestic technology.

At the same to time, global usage of Zhipu's open-platform APIs is spiking. Their GLM-5.2 model saw daily token usage jump 27 times during its first week on major model aggregation platforms, driven by developers looking for cost-effective alternatives to Western models.

When you combine mandatory local adoption in China with competitive global pricing for coding and reasoning tasks, Zhipu's revenue model starts looking unusually durable compared to hype-driven AI peers.

What This Shift Means for Global AI Investments

If you're evaluating tech portfolios or watching semiconductor supply chains, Zhipu's 1-gigawatt milestone offers several critical takeaways.

First, Nvidia's addressable market in China is shrinking permanently. Every time a major lab optimizes its software stack for Huawei or Cambricon, the cost to switch back to Nvidia increases. Even if US export controls were lifted tomorrow, Chinese firms have built working alternatives they control completely.

Second, the assumption that AI development will be dominated by a single hardware standard is dead. We are moving toward a multi-architecture world where smart compiler software abstracts away the underlying chip.

Finally, state-level computing initiatives in China, backed by over 2 trillion yuan in infrastructure funding, are producing functional technical outcomes. The gap between Western frontier models and Chinese open-weight models is closing faster than consensus estimates suggested.

Steps to Take Now

If you manage tech infrastructure, invest in enterprise software, or build AI applications, adapt your strategy to account for domestic hardware scaling:

  1. Audit your hardware dependencies: Relying exclusively on CUDA-specific libraries creates supply chain risks. Evaluate abstraction tools like Triton or PyTorch 2.0 to keep your software stack hardware-agnostic.
  2. Benchmark open-weight models: Test Zhipu's GLM series alongside open models like DeepSeek or Llama. For specialized tasks like coding and structured data extraction, lower token prices can cut cloud bills significantly.
  3. Re-evaluate semiconductor exposure: Factor structural losses in Chinese market share into long-term valuations for Western chipmakers.
IL

Isabella Liu

Isabella Liu is a meticulous researcher and eloquent writer, recognized for delivering accurate, insightful content that keeps readers coming back.