Open-source large models have never been this close to the world's top tier—but they have also never proven so dramatically that the demand for computing power is far from peaking.
On July 17, Moonshot AI released its new flagship model, Kimi K3. In the authoritative Text Arena comprehensive ranking, the model scored 1486 points, placing 9th, making it the first open-source model to enter the global top ten.
Even more striking, in the Fronted Code Arena rankings, Kimi K3 scored 1679 points, surpassing Anthropic's Claude Fable 5 to claim the top spot, ranking first in 6 of the 7 sub-directions of frontend code.
However, within 48 hours of K3's release, user request volumes far exceeded estimates and approached the cluster's capacity limits, forcing Moonshot AI to announce on July 19 a suspension of new user subscriptions on the C-end—the launch of an open-source model ended with "insufficient computing power."
This seemingly contradictory phenomenon precisely reveals the core debate in the current AI industry: do more efficient models reduce or amplify hardware demand?
From 38th to 9th: Open Source Breaks into the Top Ten for the First Time
Kimi K3's ranking leap is extremely significant. The previous generation, Kimi K2.6, ranked only 38th in Text Arena, while K3 jumped to 9th.
China Merchants Securities Strategy Research pointed out in its latest weekly report that the previous highest-ranking open-source model, Qwen 3.7-max-preview, ranked 18th. K3's emergence marks the first time an open-source model has entered the same tier as the world's top closed-source models in comprehensive text evaluation.
By category, K3 entered the top ten in three core areas: creative writing, code generation, and instruction following. It ranked first in professional fields such as physical and social sciences, law and government, and healthcare, and third in business management and financial operations.
On the technical side, K3 has a total of 2.8 trillion parameters, employing Kimi Delta Attention and Attention Residuals architecture, combined with Stable LatentMoE to achieve high sparsity with only 16 experts activated per token out of 896. According to official estimates, the extension efficiency of the entire architecture is about 2.5 times higher than Kimi K2.
Topping the Charts Then "Fuse Tripping": Reverse Evidence of Computing Power Bottleneck
The most dramatic scene after K3's release was not the ranking, but the subscription fuse. Kimi disclosed in an announcement that user requests over the past 48 hours had approached the cluster's capacity limit, and the company was immediately suspending new user subscriptions on the C-end to push forward with computing power expansion.
This directly conflicts with Goldman Sachs' earlier prediction. Goldman Sachs partner Rich Privorotsky had noted that a Chinese lab unable to match the largest Western pre-training computing power had rapidly closed the gap with top US models through architectural innovation, and warned that "scaling is no longer the only path to victory."
But Kimi's subscription fuse provides a counter-factual: when model performance truly gains market recognition, the computing power gap appears even more violently.
South Korea's Meritz Securities also clearly responded to this debate in a report on July 19. The report distinguished K3 from the DeepSeek impact: DeepSeek's core narrative was ultra-low-cost training at $6 million, while K3 did not disclose training costs, and the official recommendation is at least a supernode environment consisting of 64 high-performance chips for deployment.
From a usage cost perspective, K3's cost per task is $0.95, in the same order of magnitude as GPT-5.6 Sol's $1.04 and Claude Fable 5's $2.75.
Closed-Source Moat Narrows, But the Jevons Paradox Is Taking Effect
China Merchants Securities Research further pointed out that the most direct impact of open-source models on closed-source is performance catch-up.
Stanford AI Index 2026 shows that as of March, top closed-source models lead top open-source models by about 3.3%, a gap insufficient to sustain high premiums for closed-source in all scenarios. Epoch AI research shows that the strongest open-weight model lags behind the strongest closed-source model by an average of about 4 months.
From token consumption trends, the Silicon Data LLM Token Expenditure Index has been declining since June, reflecting that AI services are experiencing a development path similar to cloud computing—expanding usage scale through cost reduction.
According to OpenRouter data, as of July 6, global weekly call volume reached 52.6 trillion tokens. Open-source models represented by DeepSeek and Qwen have repeatedly led globally in weekly call volume due to their cost-effectiveness.
Citi analysts view K3 as a potential case of the Jevons Paradox in the AI industry chain: after technological efficiency reduces unit costs, usage volume may increase significantly, ultimately causing total resource consumption to rise rather than fall.
Bank of America Securities also pointed out that if Chinese open-source models continue to approach, US frontier AI labs are more likely to increase rather than decrease computing power investment to maintain differentiation.
Rise of Supernodes: Entering System-Level Competition
When computing power demand continues to expand and single-chip performance can no longer independently determine system capability, supernodes become the core direction of AI infrastructure evolution.
China Merchants Securities introduced in its weekly report that supernodes are new high-density computing systems that connect hundreds or even thousands of AI accelerator chips into a unified logical computing unit through high-speed, low-latency Scale-up interconnects.
At WAIC 2026, Huawei Ascend 950 supernode was publicly displayed for the first time, supporting up to 1024-card deployment, providing 1 EFLOPS-level FP8 computing power and 256TB of unified memory address space.
According to BCC Research, the global data center network technology market will grow from $45.8 billion in 2025 to $103 billion in 2030, with a compound annual growth rate of about 17.6%.
Research and Markets expects the AI high-speed interconnect market to grow from $10.31 billion in 2025 to $30.17 billion in 2030, with a compound annual growth rate of about 23.9%, significantly faster than the traditional data center network.
For investors, the release of Kimi K3 sends a clear signal: breakthroughs in open-source model capability will not weaken demand for underlying hardware. The new workloads released by efficiency improvements are shifting the industry's competitive focus from the models themselves to the computing infrastructure that supports model operation.



