The American AI industry built its business on one bet above all others: that the most capable models, kept behind an API and priced to match, would always stay a step ahead of anything a rival was willing to give away for free. This week, on an American company’s own platform, developers called that bet, and the model they reached for instead was Chinese.
Vercel, an American web development company whose AI Gateway routes traffic to dozens of models, published figures showing that open-weight systems had overtaken proprietary ones on its platform.
By August 18 they accounted for 54% of all token volume against 46% for closed models, and by August 22 the split had widened to a record 62% to 38%, according to numbers the company’s chief executive, Guillermo Rauch, posted publicly.
At the top of that leaderboard sat DeepSeek-V4-Flash.
What Is DeepSeek-V4-Flash, and Why Is It Winning?
DeepSeek-V4-Flash is the cheaper half of the V4 family that the Chinese lab previewed in April and shipped in an updated build at the end of July.
Under the bonnet it is a mixture-of-experts model with 284 billion parameters, of which only 13 billion fire on any given token, which is the trick that lets it run fast and cost almost nothing. It carries a one-million-token context window, and more importantly its weights are published on Hugging Face under an MIT licence, which means anyone can download the thing and run it on their own hardware without asking DeepSeek’s permission for anything.
Then there is the price. Until the middle of August DeepSeek charged a flat $0.14 per million input tokens and $0.28 per million output, rates that make the frontier American models look like a luxury import. For the work that now dominates real production systems, autonomous agents that chew through millions of tokens as they reason, write code and call tools, that gap stops being insignificant and starts being the difference between a product that ships and one that does not.
DeepSeek-V4-Flash ranked first by token volume, and the next places were dominated by other Chinese systems including StepFun’s Step 3.7 Flash and Zhipu’s GLM-5.2, with a second DeepSeek build also high on the list. OpenAI’s GPT-5.6 Luna was the only American model in the top five.
Is DeepSeek-V4-Flash About to Get More Expensive?
Having won its audience on price, DeepSeek is now testing how much of that loyalty was ever really about the model.
On 16 August the flat rate disappeared, replaced by peak and off-peak pricing that pushes V4-Flash to $0.44 per million input and $1.32 output during busy hours, before settling back to $0.22 and $0.66 the rest of the time.
It is still remarkably cheap. It is also no longer the offer that pulled developers across in the first place, and the timing, landing just as the usage numbers went vertical, is not what anyone would call subtle.
Is One Platform’s Data Enough to Call It?
Vercel’s AI Gateway is a single platform, tilted heavily towards developers building code and agents. Token volume is a metric that flatters cheap models almost by design. The entire appeal of a cheap model is that you feel free to spend tokens as though they were free.
Measured by revenue, or by paying users, or by enterprise contracts with a compliance team attached, the picture would look very different, and the American labs would look a good deal healthier. A snapshot assembled from a chief executive’s social posts is a signal rather than a census.
That being said, a swing from 28% to 62% in two months is too large to wave away as noise and too structural to write off as a launch promotion, because the cost advantage is built into the architecture rather than bolted on for a moment. When the cheapest option on the board is also the one you are legally free to host yourself, the reasons to reach for it compound, and they do not evaporate the moment the marketing does.
Can America’s Closed Models Hold the Line?
The American labs have spent three years betting that capability is the moat. Keep the best model behind an API, charge what the frontier is worth, and let everyone else fight over the scraps. That bet holds only while the frontier is what developers are buying. The Vercel numbers suggest most of them are buying something else: enough capability to finish the job at a price that works at scale.
That should worry Washington more than any benchmark. It has spent two years trying to wall China off from the chips that train frontier models. In the process it pushed Beijing to build its own homegrown answers to ASML. The reward is a Chinese lab winning the developer market with a model it simply gave away. Today’s defaults become tomorrow’s habits. The ground floor, where most of the world’s software actually gets written, is turning Chinese and free.
The closed bet is not lost. The frontier still charges a premium and regulated buyers still pay for a model with a compliance desk behind it. What took a public knock this week, on an American platform, is the assumption that free and open would always mean second best.
See Also:
Was Hugging Face Breached by AI Agents?
What is DeepSeek? China’s $45 Billion Bet Threatening OpenAI
