A developer working at monitors in an open-plan office at dusk

China’s Open-Weight Lead Exposes America’s Real AI Gap: Substrate, Not Capability

Kenny Le Avatar


AcadeResearch Economic Report

Executive Summary

This report extends the analysis published here on July 24, 2026 of the Microsoft- and Nvidia-hosted coalition letter “Open Weights and American AI Leadership,” and the July 28 follow-up on Dario Amodei’s response. Those reports documented the positions. This one examines what the underlying data says about the gap both sides are arguing over (AcadeResearch, 2026a, 2026b).

The ATOM Report, a peer-reviewable measurement study of the open model ecosystem by Nathan Lambert and Florian Brand, provides the most rigorous public accounting available. Its findings are stark. Alibaba’s Qwen overtook Meta’s Llama in cumulative Hugging Face downloads in September 2025 and by March 2026 had nearly doubled it, 942.1 million to 476.0 million. Chinese models’ share of new derivatives rose from 10 percent in November 2023 to 70 percent by February 2026. Chinese open models’ share of inference tokens on OpenRouter went from 2.8 percent to over 70 percent in fourteen months (Lambert & Brand, 2026).

Key finding. The American gap is not in capability. It is in substrate position — whose model the rest of the world builds on top of. The single most telling number in the ATOM data is not Qwen’s 942.1 million downloads but this: all U.S. new open-model entrants combined account for roughly 56 million. The gap is roughly seventeen to one. And it is a supply gap, not a demand gap — when America ships open weights, they are adopted. The policy debate documented in the two prior reports is a debate about whether to restrict. The data says restriction was never the binding constraint.

A coalition of 77 companies has asked Washington not to restrict open-weight models. Anthropic has replied that the real answers are chip controls, distillation enforcement, and safety testing. Both positions treat openness as something to be permitted or constrained. Neither treats it as something America has largely stopped supplying — which is what the measurement data actually shows.

Where the Prior Reports Left Off

The July 24 report examined the coalition letter co-hosted by Microsoft and Nvidia and signed initially by 25 organizations, which argued that open-weight models are essential to American AI leadership and should be encouraged rather than restricted. It noted that Chinese-origin models accounted for 46.4 percent of tokens routed through OpenRouter as of mid-2026, up from 4.5 percent in the first half of 2025 (AcadeResearch, 2026a).

The July 28 follow-up tracked the letter’s expansion to 77 signatories, OpenAI’s and Google’s accession, and Amodei’s July 27 response — which conceded that a blanket ban is wrong policy while proposing chip export controls, distillation enforcement, and mandatory safety testing for sufficiently capable models regardless of openness (AcadeResearch, 2026b).

Both reports were faithful to the debate as conducted. The purpose of this one is to point out that the debate as conducted has a hole in it. Read the two positions side by side and a shared premise emerges: that the American open-weight question is fundamentally about permission. One side argues against restriction; the other argues for calibrated restriction. Neither is primarily arguing about production.

The Measurement Record

The ATOM Report, released in April 2026 and revised in May, tracks roughly 1,500 mainline open models using Hugging Face download data obtained directly from Hugging Face, an independent daily scraper, derivative tracking via base-model tags, and OpenRouter inference data. It covers November 2023 through March 2026 and more than three billion downloads (Lambert & Brand, 2026).

Measure Then Now (to Mar 2026)
Qwen cumulative downloads 325.4M (Sep 2025) 942.1M
Llama cumulative downloads 323.7M (Sep 2025) 476.0M
Qwen share of new fine-tunes 1% (Jan 2024) 69% (Feb 2026)
Meta share of new fine-tunes 44% peak (Aug 2024) 11%
China share of new derivatives 10% (Nov 2023) 70% (Feb 2026)
EU share of new derivatives 58% peak (Jan 2024) 4% (Feb 2026)
Chinese open-model token share (OpenRouter) 2.8% >70% (14 months)
All U.S. new entrants combined ~56M vs Qwen’s 942.1M

The regional download gap widened from 23 million in August 2025, when Chinese models first overtook American ones, to 428 million by March 2026 — an order-of-magnitude divergence in seven months.

Why Derivatives Are the Number That Matters

Download counts are a weak metric and the ATOM authors say so plainly: automated pipelines, bots, and repeated pulls inflate them, and small models dominate the totals. The authors exclude re-packaged GGUF and MLX uploads from derivative counts precisely to avoid that distortion.

Derivatives are different in kind. A fine-tune or adapter is not a curiosity download; it is a record that some team spent compute, data, and engineering hours building a product on top of a specific foundation. It carries switching costs. It commits tooling, evaluation harnesses, deployment pipelines, and institutional familiarity to one base model family.

The compounding mechanism. When 69 percent of new fine-tunes worldwide are built on Qwen, the consequence is not merely that Alibaba has users. It is that a generation of applied AI engineers is learning Qwen’s architecture, that the tutorials and Stack Overflow answers and quantisation recipes are written for Qwen, that vendor tooling defaults to Qwen, and that the next model most likely to be adopted is the one that is drop-in compatible with Qwen. Substrate position is self-reinforcing in a way that benchmark leadership is not.

Qwen became the primary base-model choice as early as June 2024 — well before the token-share shift became visible and roughly a year before it became a Washington talking point. The adoption data led the policy conversation by about eighteen months.

The Gap Is Supply, Not Demand

The strongest evidence that this is a production problem rather than a preference problem is what happened when America did ship. The ATOM authors note that OpenAI’s gpt-oss models accumulated more adoption than long-established open model organisations such as Mistral AI — a new entrant outperforming an incumbent within months of release (Lambert & Brand, 2026).

Developers were not refusing American open weights. There were very few American open weights to use. Ai2’s OLMo and IBM’s Granite, at roughly 8.6 million downloads for Granite, represent genuine and growing efforts at a scale two orders of magnitude below the leader. The combined U.S. new-entrant total of approximately 56 million against Qwen’s 942.1 million is the clearest statement of the problem available in public data.

A second structural point compounds it. The ATOM data show that sub-10-billion-parameter models account for roughly 75 percent of all downloads, with the smallest bucket alone representing about a third. The volume of the open ecosystem is in small models that run on a laptop, an edge device, or a single GPU. American frontier strategy has concentrated on very large systems delivered through APIs. The United States has been competing at the top of the parameter distribution while the ecosystem forms at the bottom of it.

Why Restriction Cannot Close a Diffusion Gap

Recent American policy has been overwhelmingly restrictive in orientation. Within a single week in July 2026: Treasury threatened sanctions over industrial-scale distillation; the White House accused Moonshot AI of distilling a U.S. frontier model to build Kimi K3; the FCC added foreign-produced advanced robotic devices and power inverters to its Covered List under a domestic-content test; and a bipartisan AI Kill Switch Act was introduced. Chip export controls remain the central instrument.

These measures address a coherent concern — the transfer of American capability to strategic competitors. But they are the wrong instrument for the gap the data describes, for a structural reason: restriction operates on outflows, and the diffusion deficit is an inflow problem. Nothing in an export control, a Covered List entry, or a distillation enforcement action causes a developer in São Paulo, Jakarta, or Lagos to build on an American base model. The choice of substrate is made on availability, licence terms, size options, and documentation. Restriction changes none of those variables.

There is also a plausible reflexive effect worth naming carefully, because it is an argument rather than a measured finding. To the extent restrictive measures make American technology feel conditional or revocable to foreign developers, they raise the perceived risk of building on it — which is the precise variable open weights are valuable for neutralising. Weights already downloaded cannot be withdrawn. That property is a large part of why sovereignty-conscious buyers prefer them, and it favours whichever supplier is actually shipping.

What the Coalition Letter Got Right, and What It Left Out

On the evidence, the letter’s central empirical premise holds: open-weight models are strategically consequential, and the United States is not leading in them. The signatories are correct that a restrictive posture toward open weights would compound a position already deteriorating.

What the letter is, however, is a defensive instrument. It asks government to refrain. It does not commit its signatories to release frontier weights, and with the notable exceptions of Meta, Mistral, and OpenAI’s gpt-oss line, most of the largest signatories do not publish competitive open frontier models. A coalition of 77 organisations asking Washington not to restrict a category in which most of them do not compete is a real policy position, but it is not a supply response.

Amodei’s rebuttal has the mirror-image property. Chip controls, distillation enforcement, and capability-triggered safety testing are all defensible on their own terms, and the safety-testing proposal is genuinely openness-neutral. But none of the three increases the quantity of American open weights available to the world’s developers. Both documents are arguments about the terms of restriction.

What the Data Does Not Say

Downloads and derivatives are not revenue, and not capability. Substrate position is a strategic asset, but it is not the same as frontier capability, and the ATOM authors do not claim otherwise. On the most demanding reasoning tasks, leading closed American systems remain competitive or ahead. A country can lose the diffusion contest and still hold the capability lead — for a while.

The metrics have known distortions. The authors are explicit that download counts are inflated by automation, that large enterprise deployments may register as a single download, and that small classifier models can dominate counts while contributing little technologically. This is why the derivative and inference-token series matter more than the raw totals, and why all three pointing the same direction is the meaningful signal rather than any one of them.

The data ends in March 2026. Four months have passed. Kimi K3’s weight release in late July and subsequent launches are not captured. Directionally these would widen rather than narrow the gap, but that is inference, not measurement.

This report does not resolve the safety question. Whether widely released frontier weights create unacceptable misuse risk is a separate argument on which reasonable analysts disagree, and nothing in adoption data settles it. The claim here is narrower: that the diffusion gap is real, that it is a supply gap, and that the current policy debate is not addressing it.

What to Watch

Whether any signatory converts advocacy into release. The letter’s credibility test is whether the 77 signatories ship competitive open weights, not whether they win the policy argument. Watch for a frontier-class American open release with a permissive licence and a full size range.

The small-model tier. Since roughly 75 percent of downloads sit below 10 billion parameters, an American strategy that ships only flagship-scale open models will not move the derivative share. Watch whether U.S. labs release full families rather than single large checkpoints.

Derivative share, monthly. It is the leading indicator. It moved eighteen months before the policy conversation did, and it will register any genuine American recovery before download totals or token share do.

Whether public funding follows. Sustained open-model production at competitive scale is expensive and has no direct revenue model. If the United States treats this as infrastructure, it will show up as procurement or research funding rather than as coalition letters.

Conclusion. The July 24 letter and the July 27 rebuttal are a serious debate about the right level of restriction on open-weight AI. The measurement record suggests that debate is being held over the wrong variable. China did not take 70 percent of new derivatives because America restricted anything. It took them because it shipped usable models in every size, under permissive licences, continuously, for two years, while the American frontier moved behind APIs. Substrate position is won by supply. On present evidence, only one country is treating it that way.

References

AcadeResearch. (2026a, July 24). Open weights and the open-closed choice: What the July 24 Microsoft and Nvidia coalition letter argues, and what the data shows. https://acaderesearch.com/open-weights-american-ai-leadership-huang-nvidia-analysis-2026/

AcadeResearch. (2026b, July 28). The open-weights debate just got its second side. https://acaderesearch.com/open-weights-debate-amodei-rebuttal-july-27/

AcadeResearch. (2026c, July 26). What is AI distillation, and why did a 2015 machine learning technique become a national security fight in one week? https://acaderesearch.com/ai-distillation-explained-kimi-k3-policy-fight-2026/

AcadeResearch. (2026d, July 29). The FCC did not ban robots. It made U.S. market access conditional on American parts. https://acaderesearch.com/fcc-covered-list-robots-power-inverters-market-analysis/

Lambert, N., & Brand, F. (2026). The ATOM Report: Measuring the open language model ecosystem (arXiv:2604.07190v2). arXiv. https://arxiv.org/abs/2604.07190

The White House. (2025, July). Winning the race: America’s AI action plan. https://www.whitehouse.gov/wp-content/uploads/2025/07/Americas-AI-Action-Plan.pdf


How to cite this paper

Le, K. (2026, July 30). China’s Open-Weight Lead Exposes America’s Real AI Gap: Substrate, Not Capability. AcadeResearch. http://acaderesearch.com/china-open-weight-lead-america-ai-diffusion-gap/