Home/Blog/Evaluation & evidence/1 to 2% of the chips, about 30% of the tokens

1 to 2% of the chips, about 30% of the tokens

Export controls were designed to constrain compute and have. The measure that moved instead was usage, and the two have separated sharply.

TL;DR. On the compute measure the controls have done what they were designed to do. Estimates put Chinese advanced chip production at roughly 1 to 4% of US capacity in 2025 and 1 to 2% in 2026, with a projected US advantage in 2026-produced AI compute of 21 to 49 times depending on the performance metric used. Shipments in 2025 ran approximately 800,000 Huawei Ascend units against about five million Nvidia Blackwell units. And on a different measure the picture inverts. One count puts Chinese models' share of global AI token usage at roughly 1% in 2025 and about 30% in 2026. Both figures come from parties arguing for stronger controls. The compute gap and the usage share are measuring different things, and a policy evaluated in one currency produced its most visible effect in another.

---

Status: contested, and this article does not adjudicate. Figures are attributed to their sources, several of which are advocacy organisations or named analysts with stated positions on the policy. This corpus does not take positions on contested political questions. What follows sets out what each side measures and why the measures diverge, so a reader can navigate rather than be told.

---

The case that the controls are working

Chinese advanced chip production is estimated at roughly 1 to 4% of US capacity in 2025, declining to 1 to 2% in 2026 as US and allied manufacturers scale.

Shipment volumes tell the same story. Approximately 800,000 Huawei Ascend 910C units in 2025 against roughly five million Nvidia Blackwell units.

Projected 2026 compute advantage: 21 to 49 times, depending on whether FP4 or FP8 performance is used for Blackwell-generation chips.

Manufacturing constraints compound it. SMIC's 5nm release has been repeatedly delayed on independent analysis, its 7nm process reportedly runs poor yields with reliability problems, and one account reports more than 22,000 Chinese semiconductor companies shutting down over five years. Huawei is not expected to match an H200-class chip before Q4 2027 at the earliest.

The conclusion drawn from this is that China can build capable systems at a cost premium, estimated at roughly 50% for training and one to five times for inference, rather than closing the gap.

The case that they are being routed around

SMIC has reached high-volume production of a 5nm-class node without EUV, by stretching deep ultraviolet multi-patterning to its limits. Yields are estimated at 30 to 40% against 80% or better at TSMC, and the difference is being absorbed as state subsidy on national-security grounds rather than judged commercially.

That process supports a current flagship smartphone chipset.

And Huawei's Ascend 950 series, released in Q1 2026, is reported as the first Chinese accelerator line with integrated in-house high-bandwidth memory, which addresses one of the three components usually named as gating: HBM, advanced packaging, and logic fabrication.

The conclusion drawn from this is that targeted restriction acted as an industrial policy catalyst, producing domestic capability that would otherwise have been bought.

The number that does not fit either story

One count puts Chinese models' share of global AI token usage at approximately 1% in 2025 and approximately 30% in 2026.

That figure appears in material from an organisation arguing for stronger export controls, which cites it as evidence of urgency rather than of failure. It is presented here with that provenance attached.

If it is approximately right, it does not sit comfortably with either case. It is not the compute-gap story, because a thirtyfold rise in usage share did not require a thirtyfold rise in chip supply. And it is not straightforwardly the neutralisation story either, because the compute gap is simultaneously reported as widening.

The reconciliation is that compute production and token service are different quantities. Training a frontier model is compute-intensive and concentrated. Serving tokens is cheaper, more distributable, and can run on older or less efficient hardware at a cost penalty the state may absorb. Efficient models trained once can serve very large volumes, and open-weight release decouples usage from the infrastructure of whoever trained the model.

So a policy that successfully constrained the first measure has coincided with a large move in the second, and whether it caused, failed to prevent, or is unrelated to that move is not established by any source here.

Why the measures diverge

This is a scope boundary problem with policy consequences.

Compute production measures manufacturing capacity. It answers: how many advanced chips can be made, and by whom.

Token usage share measures adoption. It answers: whose models are people actually running.

These are linked and they are not the same, and the link is weaker than an intuitive account suggests. A model trained once on constrained hardware, released with open weights, and served from anywhere generates usage that appears nowhere in a chip production statistic.

Which raises the evaluative question this article can pose and not answer: what were the controls for? If the objective was to slow the training of frontier systems, the compute figures are the relevant measure and they look favourable to the policy. If the objective was to limit the diffusion and influence of models from a particular origin, the usage figure is the relevant measure and it does not.

Both objectives have been stated at different times by different officials, which is why the same evidence supports opposite conclusions depending on which is taken as the target.

Three things this establishes

A policy's success measure and its stated objective can drift apart. Compute share is measurable, reported quarterly and moving in the intended direction. Adoption is harder to measure and moved sharply the other way. When those diverge, whichever is easier to count tends to become the way the policy is discussed.

Cost penalties are not barriers when a state absorbs them. Yields of 30 to 40% against 80% are commercially disqualifying and strategically tolerable, and an analysis assuming commercial logic will mispredict what gets built.

And nearly every figure here comes from an interested party. The compute-gap numbers come from analysts and organisations arguing the controls work or should be strengthened. The capability claims appear in syndicated financial content with limited attribution. A reader has to weigh provenance on both sides, and this article has tried to make that possible rather than resolve it.

What it does not establish

That the controls succeeded or failed. Both conclusions are available from the evidence assembled here, which is why the article does not draw one.

That the token-share figure is reliable. It comes from a single advocacy source, its methodology is not stated, and no independent replication is cited. It may be wrong, and the article's argument depends on it.

That capability claims about Chinese manufacturing are verified. The 5nm-class production and in-house HBM reports appear in syndicated commercial content without primary attribution, and independent analysts dispute the timeline.

And nothing about what policy should be. This corpus does not take positions on contested political questions, and export control policy is squarely one.

What is unresolved

Whether the token-share figure holds up. It is the load-bearing number in this article and it rests on one interested source.

What the objective actually is. Slowing frontier training and limiting model diffusion imply different policies and different success measures, and both have been articulated.

Whether subsidised low-yield production is sustainable. Absorbing a fifty-point yield gap is a fiscal choice that can persist for a long time or stop abruptly.

And how much moves through enforcement gaps. Estimates of diverted hardware exist, vary enormously, and none is independently verifiable, so this article does not use any.

The counter-argument

Token share is a poor proxy for anything strategic. Usage counts include consumer chat, cheap inference and workloads with no security or economic significance, while the compute gap bears directly on who can train the next frontier system. Elevating an adoption metric over a capability metric may simply be measuring the wrong thing more precisely.

The 30% figure may be an artefact. A shift from 1% to 30% in a year is extraordinary and could reflect a change in measurement, in what counts as a Chinese model, or in which providers report. A number that moves thirtyfold in twelve months deserves more scepticism than this article gives it.

The controls have not had time to work. Manufacturing capability responds over years, and assessing a policy designed to compound over a decade at the three-year mark repeats the error the labour article warned about with technology waves.

And treating both sides as symmetric may be false balance. The compute figures are numerous, come from several analysts, and agree. The neutralisation account rests substantially on syndicated content with weak attribution. Presenting them as two cases of comparable standing overstates the second, and this article does exactly that in the interest of neutrality.

The short version

On compute the controls have done what they were built to do. Chinese advanced chip production at roughly 1 to 4% of US capacity in 2025 and 1 to 2% in 2026, shipments of about 800,000 Huawei Ascend units against roughly five million Nvidia Blackwell, and a projected 21 to 49 times US advantage in 2026-produced AI compute.

On capability the picture is less clean. SMIC reaching a 5nm-class node without EUV at 30 to 40% yields against 80% or better elsewhere, subsidised as national security rather than judged commercially, and an accelerator line reported as the first Chinese one with integrated in-house high-bandwidth memory.

And one count puts Chinese models' share of global AI token usage at about 1% in 2025 and about 30% in 2026. That figure comes from an organisation arguing for stronger controls, cited as evidence of urgency.

If it is roughly right, it fits neither story. A thirtyfold rise in usage share did not require a thirtyfold rise in chip supply, because training compute and token service are different quantities: a model trained once on constrained hardware and released openly generates usage that appears in no chip production statistic.

Which leaves the question the evidence cannot settle: what were the controls for? Slowing frontier training makes the compute figures the measure, and they look favourable. Limiting the diffusion of models from a given origin makes the usage figure the measure, and it does not. Both objectives have been stated, which is why the same evidence supports opposite conclusions.

Almost every number here comes from an interested party on one side or the other, and this article has tried to make that visible rather than pick a winner.

Common questions

Have the export controls limited Chinese AI compute? On the available estimates, yes, substantially. Chinese advanced chip production is put at roughly 1 to 4% of US capacity in 2025 and 1 to 2% in 2026, with 2025 shipments of approximately 800,000 Huawei Ascend units against about five million Nvidia Blackwell units, and a projected US advantage in 2026-produced AI compute of 21 to 49 times depending on the performance metric used.

Has China developed capability anyway? Reports indicate SMIC reaching high-volume production of a 5nm-class node without EUV lithography, by stretching deep ultraviolet multi-patterning, at estimated yields of 30 to 40% against 80% or better at leading foundries, with the difference absorbed as state subsidy. Huawei's Ascend 950 series is reported as the first Chinese accelerator line with integrated in-house high-bandwidth memory. These claims appear largely in syndicated commercial content with limited primary attribution, and independent analysts dispute the timelines.

What is the token usage figure? One count, appearing in material from an organisation arguing for stronger export controls, puts Chinese models' share of global AI token usage at approximately 1% in 2025 and approximately 30% in 2026. Its methodology is not stated and no independent replication is cited. It is the load-bearing number in this article and it rests on a single interested source, which is stated rather than glossed.

How can compute share fall while usage share rises? Because they measure different things. Compute production measures manufacturing capacity: how many advanced chips can be made and by whom. Token usage measures adoption: whose models people actually run. Training a frontier model is compute-intensive and concentrated, while serving tokens is cheaper, distributable, and can run on older hardware at a cost penalty a state may absorb. A model trained once and released with open weights generates usage that appears in no chip production statistic.

So did the controls work? That depends on what they were for, and both answers are available. If the objective was slowing the training of frontier systems, the compute figures are the relevant measure and they favour the policy. If it was limiting the diffusion and influence of models from a particular origin, the usage figure is relevant and it does not. Both objectives have been articulated by officials at different times, which is why the same evidence supports opposite conclusions.

Is the 30% figure trustworthy? It should be treated cautiously. A shift from 1% to 30% in twelve months is extraordinary and could reflect a change in measurement, in what counts as a model of a given origin, or in which providers are counted. The article's own strongest objection to itself is that a number moving thirtyfold in a year deserves more scepticism than it receives here.

Why does yield matter so much? Because it determines whether production is commercially viable. Yields of 30 to 40% against 80% or better mean far more wafers are discarded per working chip, which is disqualifying for a company optimising profit and tolerable for a state treating capability as security. An analysis that assumes commercial logic will mispredict what gets built and at what scale.

Does this article take a position on export control policy? No. This corpus does not take positions on contested political questions, and export controls are squarely one. What it does is set out what each side measures, why those measures diverge, and where each figure comes from, so a reader can weigh the evidence rather than be handed a conclusion.

Sources

Primary documents only. Where a claim rests on a single report, the entry says so.

  1. America's chip export controls are working Noah Smith, January 2026 The compute-gap case: production at 1 to 4% of US capacity falling to 1 to 2%, a projected 21 to 49 times advantage, SMIC's delayed 5nm and yield problems. A named analyst arguing a position, which is stated.
  2. Strengthening Export Controls: A Critical National Security Priority for Congress FDD Action, April 2026 The shipment comparison of roughly 800,000 Ascend against about five million Blackwell units, and the token usage share figures of about 1% in 2025 and 30% in 2026. An advocacy organisation arguing for stronger controls, citing the usage figure as evidence of urgency.
  3. Silicon Sovereignty: How Huawei and SMIC are Neutralizing US Export Controls in 2026 Syndicated commercial content, January 2026 The capability claims: 5nm-class production without EUV at 30 to 40% yields, and an accelerator line with integrated in-house HBM. Appears across multiple syndication outlets with limited primary attribution, which is why it is named rather than linked and weighted accordingly.

Further reading

The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.

  • Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — operationalisation determining what a measurement can support. :: https://arxiv.org/abs/2111.15366 Scope Boundary
  • Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — the difference between overall performance and performance on the population that matters. :: https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307 Scope Boundary

Learn the concepts

← All posts