For most of this year, the top of OpenRouter’s token leaderboard has been Chinese. Not one model sneaking in at number 8, the top of it, week after week, at something like two thirds of everything flowing through the platform. A year before that the same providers were under 2%.
I have been circling this for a month now. I started with lighters and solar panels and the idea that overcapacity is a strategy rather than an accident. Then Geely, and the discovery that you do not climb the value-chain, you buy the top of it and industrialise underneath. Both pieces actually come together in this one. Yes, this is part three! And it is what happens when that machine gets pointed at AI models.
I wanted to write about what is happening and found that I needed quite a long windup to get here.
Nobody is selling
Geely needed a willing seller for their capture of the up part of the value-chain. Ford wanted out of Volvo in 2010, so there was a deal to be done. In AI the equivalent purchase is the middle of the stack, the leading edge logic and the memory sitting next to it, and that is not for sale at any price. ASML, 40 minutes down the road from some of my old clients in Veldhoven, is not for sale. Neither is TSMC. The tooling underneath both is under export control and has been tightened repeatedly.
You can see part of it it reflected in the fabrication. SMIC is still stuck at 7nm. But the actual constraint, well that is why I think, is not the logic or CPU at all, it is the memory. Huawei’s Ascend parts still depend on HBM stacks (a 3D package made from multiple layers of DRAM) from SK Hynix and Samsung, and domestic HBM production is what sets this ceiling. You can fabricate a perfectly good compute die and still have nothing to ship… which is a fairly humbling place to be after 20 years of being told you cannot be stopped.
So the country that owned the upstream in polysilicon, in lithium refining, in cathode and anode chemistry, does not own the upstream here. In every other example I have written about this month they held the bottom of the value chain and climbed outward. In this one they are locked out of the middle.
Except that is only half of it,
If you look at what actually goes into an AI datacentre, you may see it. Boards, power distribution, cooling loops and cold plates, racks and sleds, connectors, interconnect, PDUs, the whole thermal stack. China’s datacentre infrastructure manufacturing supply chain runs to something like CNY 500 billion. That is not a niche, that is a moat with the water still going in.
And that is the lighter skill! Taking a hundredth of a cent out of a valve seat is the same muscle as taking cost out of a cold plate. It does not transfer to the die at all. It transfers beautifully to everything the die is bolted to, and in a datacentre buildout that is a very large share of what you are actually paying for. Let alone their electricity surplus.
Then look at the other end, where Qwen has passed a billion downloads on Hugging Face and overtaken Llama as the most downloaded open model in the world.
Strong at the bottom. Strong at the top. Blocked in the middle. Take Stan Shih’s smiling curve from last week and knock the trough out of it, and you get what I have started calling the hollow smile: a firm that owns both ends of a value chain and cannot buy, build or borrow the part in between.
Because when you cannot climb out of the middle, there is only one other move available. You make the middle worthless.
It is not dumping
The obvious thing to read into this, is that this is solar again. Subsidised product going out below cost until everyone else is dead. That is the path I started on, and if I am honest, it does not survive contact with the evidence.
DeepSeek published its own inference numbers and claimed a theoretical cost profit margin of 545%. The caveats matter, because that excludes training and R&D entirely, most access is free or discounted off-peak, and actual revenue is far below the theoretical figure. But it is not the balance sheet of a business burning money to kill rivals. The cheapness looks architectural, in the sparse routing and the attention work and the serving layer, and that is a fairly different animal.
That distinction took me an embarrassingly long time to see it. Totally opposite of the consolidation now grinding through Chinese solar, efficiency does not end.
Qwen 3.8-Max, the most capable model in the Qwen family to date.
(This also marks the first time they reveal the open-source the weights of a Qwen-Max-class model. Built upon the architectural foundation of Qwen 3.5, Qwen 3.8-Max scales to 2.4 trillion parameters, delivering comprehensive improvements across coding, work, research, and long-horizon tasks)
Twelve years ago I wrote a small piece asking what the endless IaaS price cuts were signalling and decided they were a loss leader. Cheap compute to get you in the store, real money on the layer above. I have now watched that move twice more, and the open weight release is the third time. So giving the model away is not charity, and I do not think it is a price war either. It is what a hollow smile does. You commoditise the layer you cannot own so the money moves to the layer you do.
The trap on your desk
It gets awkward for anyone running European IT at this point, and I do mean awkward, because I cannot find a clean answer in it anywhere.
Running an open weight model on your own infrastructure is genuinely the strongest sovereignty position available to you right now. The weights sit on your iron. Nobody can reprice them at renewal, restrict them by region, deprecate the version you built against, or change the terms because a government asked, and every one of those has happened to somebody with a proprietary API in the last 18 months. On pure supply chain resilience, self-hosting beats renting, and it is the kind of fragility you can already see if you go looking rather than trying to forecast which shock arrives first. And although I am still very interested in Mistral it will probably be a Chinese Qwen model).
And it is simultaneously your worst possible audit position. You cannot bench test weights the way you can bench test a solar panel, so you cannot verify what is in the training data, what refusals are baked in, or what the model quietly will not say. EU AI Act Article 10 wants data governance documentation for high risk systems, enforcement is live from August, and the fines run to EUR 35 million or 6% of global turnover. I argued a few weeks ago that regulation is a design input rather than a PDF you file somewhere
Which raises the question: Who in your organisation has actually read the licence on the model you are about to standardise on, and are they the same person who signed the DPA (Data Processing Addendum)? In most places I have worked those are 2 different people who have never met, and one of them thinks the other one handled it.
The only concrete thing I have is this. Stop treating model choice as a procurement decision and start treating substitution cost as a design property. Write down, per integration, what it would cost you in weeks to swap the model underneath it. If you cannot answer that, you have already made the commitment you thought you were still deciding on. The prompt scaffolding, the eval suite, the tool schemas, the fine tunes, the retrieval assumptions (especially the retrieval assumptions) are actual the lock-in.
The panels were cheap too. We all remember how that went for whoever was still trying to make them here.
So here is what I keep coming back to, and I do not have a good answer. If the middle of the stack is being deliberately hollowed out, and neither the chips nor the compliance story are yours, what exactly is the durable thing you are building on top of it?
Leave a Reply