RSS Amplifier

Weijin Research · Aug 7, 2026

AI Competition: 3 Trillion Parameters at Home, GW-Scale Compute Abroad

0
Sign in to vote or save

Weijin Research · Weijin Research

This article originally appeared on Weijin Research on Sina on July 22, 2026. Original Chinese title: 「AI竞争,内卷参数3万亿,外拼算力GW级」. It has been translated and adapted for an English-speaking audience.

China's AI competition is shifting from chip-model synergy toward the coordination of AGI capabilities with GW-scale computing infrastructure.

The large-scale buildout of China's AI infrastructure is quietly changing gears. For the past few years, the core of China's AI industry competition was "chip-model synergy" — tightly coupled optimization of chips and models. But as open-source models gradually approach the global frontier, the key competitive factor is turning toward whether a company can possess the infrastructure to support continuous model evolution and large-scale deployment. This is no longer a contest of single-point efficiency but a broader system-level coordination.

Now, the players driving this infrastructure race are no longer just traditional tech giants. Emerging AI companies like DeepSeek and Zhipu AI have also stepped to the front. They are exploring the construction of gigawatt-level domestic computing clusters, which could become the key infrastructure for the next phase of China's push toward higher intelligence levels and expanding the supply of intelligence.

Zhipu AI, which announced its "Reaching Higher" initiative, has been reported to be building a GW-scale domestic computing cluster, parts of which are already operational. At the same time, Zhipu AI has built or is operating multiple computing clusters exceeding 10,000 cards in scale. According to Shanghai Securities News, citing sources close to Zhipu AI, this GW-level cluster uses entirely domestic AI chips. Meanwhile, Zhipu AI completed the acquisition of XCore Sigma, a domestic AI heterogeneous computing software company, further bolstering its capabilities across the heterogeneous computing software stack — a move that also fits with the relatively fragmented reality of China's domestic AI chip ecosystem.

Zhipu AI is far from alone in aiming higher. DeepSeek targeted AGI from the very start and was the earliest in Asia to build an Nvidia A100 cluster, laying the computing foundation that would create its "DeepSeek moment." This year, as DeepSeek began fundraising, it started openly recruiting IDC design and planning engineers who "will have the opportunity to participate in the planning and construction of infrastructure ranging from MW to GW scale." It has also been reported to be exploring the development of its own AI inference chips.

Compared with the market's intense focus on overseas AI infrastructure, if measured purely by power capacity, the global AI infrastructure race has, in truth, only just crossed the GW threshold. According to Epoch AI data, there is currently not a single fully operational GW-scale AI computing cluster anywhere in the world.

Among the largest AI infrastructure projects today, SpaceXAI's Colossus 2 has been built to roughly 946 MW, though its computing power is shared among Anthropic, SpaceXAI, Cursor, and other parties. Amazon's New Carlisle project, built for Anthropic, is planned at roughly 1.9 GW, with about 910 MW actually in operation now — making it one of the largest AI infrastructure deployments occupied by a single user. In addition, Microsoft and OpenAI's shared Fairwater Atlanta, as well as Meta's Prometheus (self-used for now but possibly opened for leasing in the future), also sit above the 500 MW scale.

Horizontal bar chart comparing IT power consumption (in MW) across global AI data centers by country and facility. Data centers are ranked by power usage, with facilities color-coded by location: United States (teal), China (red), and Malaysia (pink). Source: Q1 2023 data.

This means that even the world’s top AI companies are still moving from the hundreds-of-megawatts scale to the gigawatt scale in their AI infrastructure. The gigawatt-scale domestic compute clusters that Chinese AI companies are exploring are not about catching up to a fully mature industry model; they are actively participating in defining the next stage of AI infrastructure. Of course, gigawatts are only one measure of AI infrastructure competition. Simply comparing power capacity does not fully capture the level of AI computing power.

This infrastructure competition is not just about capital expansion. It is being driven jointly by leaps in open-source model capabilities and the expansion of AI demand. Zhipu aims to complete pretraining of a frontier model close to the Mythos level before year-end, which will require even larger compute clusters. Meituan’s LongCat-2.0 has already successfully trained a trillion-parameter model on a domestic compute cluster.

Moonshot AI’s Kimi K3 shows another possibility. When near-frontier intelligence is offered at a lower cost, more companies and developers may begin to experiment with AI, shifting the demand curve rightward and creating opportunities to discover application scenarios that did not exist before. Only through continuous interaction and feedback in the real world can the model then be continuously refined.

In an internal discussion at the AI analysis firm Semianalysis, Kimi K3 was hailed as possibly the world’s second-best model. Fable often forces users back to Opus because of safety policies or access restrictions, a problem Kimi K3 does not have. Except for one thing: compared to other leading U.S. models, Kimi K3 is simply too slow.

As intelligence improves, the practical challenge for Chinese open-source models closing in on the frontier may be how to supply that intelligence at scale. In fact, since the release of Kimi K3, Moonshot AI’s compute consumption has been pushing its capacity limits, forcing a halt to new consumer subscriptions. The company promises to expand its compute capacity at full speed.

The bigger challenge goes beyond economic discussions like the Jevons paradox. Chinese model builders are continuously improving GPU and HBM utilization efficiency through new attention mechanisms, sparser architectures, and other methods. But as model sizes keep ballooning, optimizing per-card efficiency is far from enough. AI competition is irreversibly moving into a phase of system-level co-optimization.

Semianalysis’s analysis hits the nail on the head. Kimi K3 has 2.8 trillion parameters. Even with MXFP4 low-precision computation, each forward pass still requires around 1.5 TB of HBM bandwidth. The model includes 896 experts, and actual serving relies on WideEP, which distributes experts in parallel across multiple GPUs and uses high-speed networking for token distribution and result aggregation. To benefit inference efficiency from larger high-bandwidth communication domains, the team officially recommends deploying on supernodes with 64 or more accelerators.

As one of the largest open-weight models in China today, Kimi K3 once again proves that scaling laws still hold force at this stage. Larger parameter scales, larger training scales, and more complex inference systems remain key paths to driving leaps in model capabilities. As models enter the trillion-parameter era, AI competition is shifting from mere model rivalry to a systemic contest that integrates compute, networking, and power.

Timeline chart showing total parameters of flagship AI models from Jul 2025 to Jul 2026, tracking releases by company including Kimi, Moonshot AI, DeepSeek, Llama, and others. Y-axis shows parameter count; x-axis shows time progression with release dates and latest model sizes labeled.

The race to scale model parameters is already underway. Alibaba recently previewed its latest flagship model, Qwen-3.8, with 2.4 trillion parameters. The company says its overall capabilities are second only to Fable 5. MiniMax is also reportedly training its next-generation flagship model, with parameter counts expected to reach 2.5 trillion to 3 trillion.

In a sense, this concentrated push by Chinese model companies toward larger parameter scales coincides with a critical inflection point for domestic chips and domestic super-nodes: the shift from "usable" to "scalable and usable." The growth in model scale drives infrastructure upgrades, and the maturation of infrastructure in turn unlocks space for model innovation. This two-way feedback loop is replacing the old model of chips merely adapting to models. Venture capital firm Qiming Venture Partners projects that AI infrastructure will face persistent structural shortages over the next two years. Compute reserves will become a core strategic asset for AI companies, and new-architecture compute chips and large super-node clusters will emerge during this period.

The per-chip performance of domestic chips is still breaking new ground. SMIC's N+3 process packs roughly 113 million transistors per square millimeter, with logic density now comparable to TSMC's N6 node at 108 million per square millimeter. Huawei has proposed "Tao's Law," replacing geometric scaling with "time compaction." By using logic folding to continuously compress signal propagation latency, it further boosts effective transistor density and overall system performance.

Comparison of chip layouts for Kirin 9030 (N+3) and Helio G99 (N6) processors shown in X-ray or cross-section images. The four panels display different views or magnifications of the two chipsets side by side.

At the 2026 World Artificial Intelligence Conference (WAIC), nearly all domestic computing power vendors showcased super-nodes as a central theme. Huawei, ZTE, Muxi, Biren Technology, and other companies unveiled system-level solutions for large-scale AI training and inference.

Huawei’s Ascend Atlas 950 super-node further expanded interconnect scale, introducing 800G optical interconnects and high-density optical fiber technology to increase the interconnect scale of a single super-node from the previous generation’s 384 cards to 1,024 cards. Alibaba’s T-Head launched the Zhenwu M890 × Panjiu AL128 super-node and open-sourced its self-developed AI software stack, T-Head SAIL, aiming to lower the barrier for third-party models to access the domestic computing ecosystem.

Meanwhile, Sugon unveiled the Sugon 8000 (Dengfeng) all-domestic 100,000-card AI super-cluster, demonstrating the capability of domestic computing power to move toward large-scale cluster construction. Biren Technology released the BLink 2.0 super-node interconnect protocol, which uses memory-semantic interconnects, efficient communication scheduling, and other methods to enable up to 1,024 GPUs to form tighter computational coordination.

The watershed in future AI competition may not just be about having the strongest model, but about the ability to continuously produce intelligence, supply intelligence, and diffuse it into the economic system. This capability is reflected not only in model parameters but also in the collaborative efficiency of the computing power, networks, and power systems that support intelligent production. Today, frontier model companies are competing around this complete intelligent production system, trying to control the critical closed loop from computing infrastructure to model iteration.

Read the original on weijinresearch.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.