AI: The ‘Race to Zero’ that really Isn’t. AI-RTZ #1169

AI: The ‘Race to Zero’ that really Isn’t. AI-RTZ #1169

Today’s discussion is about what AI models cost in this AI Tech Wave, and about two opposite things being true at the same time. The pricing spectrum for AI globally is widening at both ends. A non-intuitive phenomenon.

Axios frames it dramatically in one direction with ‘DeepSeek’s new bargain model accelerates AI’s race to zero’.

“Some of the smartest software on Earth is rapidly becoming a commodity.”

They cover the landscape, led by Chinese open source model in particular, with the prices going in one direction.

They cover through the flurry of releases.

DeepSeek shipped V4 Flash on Friday. On complex coding and autonomous software tasks it performs close to Anthropic’s Claude Opus 4.8, which is about as high as the bar goes right now. Then lean in on the difference for effect.

“The price gap is staggering: DeepSeek charges about 28 cents for the same amount of output that costs $25 on Opus 4.8.”

That is a 99% discount, on work a frontier lab was charging frontier money for three weeks ago.

Meanwhile, Anthropic’s lead US competitor, ahead of mutual mega-AI IPOs for both, is leaning on lower prices for market share metrics. Something I’ve discussed of late.

As a reminder, OpenAI cut the price of GPT-5.6 Luna by 80% recently. Luna had launched only three weeks earlier, so it’s a current ‘latest and greatest’ in its class and category.

Google shipped three new Gemini ‘flash’ models, all of them built around efficiency. As I’ve discussed, Google is playing a different game vs the frontier models, going for price and speed efficiency, since they cater to billions rather than the tens and hundreds of billions for Anthropic and OpenAI at the higher, business end.

One week. Three of the largest AI companies in the world, all moving the same direction on price. It makes a great headline.

So the commodity read is understandable. Axios puts it plainly.

“When a product becomes a commodity, buyers care less about who made it and more about what it costs.”

And Zack Kass, OpenAI’s former head of go to market, gives it the sharpest version.

“At some point, the next model doesn’t matter to you.”

Here is where my Take is a bit different, and it is not a small difference.

‘Race to zero’ describes a line. One price, heading down, everybody chasing it.

What is actually happening is a spectrum, and it is widening at both ends at the same time. Widening price gap and entirely different capabilities on both ends.

I’ve long said in these pages that two opposite things can be true at the same time. This is one of the cleanest examples I’ve seen in this wave.

AI token prices are falling hard. AI token bills are rising hard. Both. Same market, same quarter, same customers in many cases.

I wrote the second half of that over the weekend in ‘The Meter’s Running’ on Long AI Agents. Anthropic’s Astra ran up roughly $2,000 in API charges working decade old math problems.

Expensive and a bargain at once. That was the whole point of the part of the discussion.

Unit prices go down. Units consumed go up faster. The bill goes up. For really different types of work. None of those drivers cancel the others.

It’s something I’ve been underlining for a while now.

  • March 2024. In Open Source AI momentum continues I wrote that ‘open and closed AI will both make all the difference. Just all together, not separate.’ That was two and a half years ago.

  • November 2025. Smaller AIs gain momentum had GPT-5 Nano at around 10 cents per million tokens against GPT-5 at around $3.44. My line then was that growth is at both ends of the AI model spectrum.

  • March 2026. Add AI Tokens to US vs China AI metrics race put it thus: ‘It’s not either/or for users globally. But BOTH.’

What has changed since is the width of the gap, not the direction of the argument.

The menu customers are actually shopping from today, top to bottom.

  • Frontier, closed. Claude Opus 4.8 and its peers. $25 for the output DeepSeek will do for 28 cents. People are still paying it, and the reason they get value at that end. Far from a commodity.

  • Cheap, closed. GPT-5.6 Luna after an 80% cut. Three Gemini flash models. This tier did not exist in a serious way two years ago and it is where the volume is going.

  • China open. DeepSeek V4 Flash, MiniMax, Moonshot’s Kimi K3. Two to three dollars per million output tokens against roughly $15 for Claude Sonnet 4.5. A near sixfold gap.

  • US open. Nvidia’s Nemotron, Reflection, Thinking Machines. Early, well funded, and explicitly aimed at the gap.

  • Small and local. GPT-5 Nano at roughly 10 cents per million. On device, in the browser, inside the appliance.

Five tiers. Every one of them is growing. Providing different uses across a spectrum of model classes, categories and capabilities. That is not what a race to zero looks like. Certainly not ‘commodities’.

A barrel of gas is a barrel of gas. A bale of hay is a bale of hay.

A race to zero has one survivor at the bottom. This has five live businesses at five prices, and customers buying from more than one of them in the same week. We are far from a saturated market for AI. Still early days.

Now the part that Axios gets right and that deserves more weight. China’s edge here is structural, not promotional.

Cheaper energy and more efficient models, and the pricing follows from both.

I laid the numbers out in Add AI Tokens to US vs China AI metrics race in March. A Hong Kong developer spending $50 a day running Kimi, against the roughly $900 a day the same work would cost on Claude alone.

MiniMax’s M2.5 was up 476% month over month on token usage at the time.

US companies have been quietly building on Chinese open models for a while now. I covered that in December, and it has not slowed down.

China has led the US in open source models for over a year. I wrote that in June 2025 and nothing since has reversed it.

Which brings me to Nvidia, and to why Jensen Huang is spending his own attention on open models rather than leaving it to the labs.

Nvidia put it in its own words when it assembled the Open Secure AI Alliance recently, 37 founding partners deep. The world needs both closed and open models.

Needs repeating to sink in indeed.

Most of the world is not going to buy frontier compute at frontier prices. Somebody’s open models are going to fill that gap.

Nvidia does not need to know whose. It needs the tokens to get generated somewhere, on its silicon.

And here is the part people keep getting backwards. Cheaper models do not mean fewer tokens. They mean more.

A 99% price cut does not shrink a market that is still this early. It opens it to everyone who was priced out yesterday.

So Nemotron, and the backing for Reflection and Thinking Machines, are not charity and they are not a hedge. They are demand generation.

I made the longer case for this in May, in How Nvidia & Apple can be the Global, US Open Source AI Champions vs China. It reads better now than it did then.

There is a second reason the top tier holds, and it is the oldest story in technology pricing.

A price umbrella. The premium vendor holds a high price, and that price is what makes the space underneath it worth entering.

Jeff Bezos said the quiet part decades ago. Your margin is my opportunity.

I walked through this on ARD #114. Anthropic pressing an a la carte price advantage while the clouds and the open labs quietly test how much of that margin they can take.

The last time this ran to completion, CompuServe had the better product and AOL had the good enough one at the price that moved. Good enough won, and it was not close.

But notice what actually happened there. The premium tier did not vanish. It got smaller, richer, and more specialized.

That is the likely shape here too.

Meanwhile the top of the market keeps finding new ways to charge.

I called it Velvet Rope AI in April. The very best models, a la carte, for customers who barely ask how much.

And on the other side, the pushback is real. Customers started asking how much about tokenmaxxing in May. Meta stepped back from it in June.

High prices and low prices, both under pressure, in opposite directions, at the same time.

My overall Take.

‘Race to zero’ is the right observation attached to the wrong picture.

The observation is that the floor is collapsing, and it is. Twenty eight cents against twenty five dollars is not a discount, it is a different business.

Jevons Paradox accelerates the uses and applications.

The wrong picture is that everything gets dragged to the floor with it.

Markets commoditize from the bottom when the top has stopped improving. The top here has not stopped improving. It is improving fast enough that a frontier model from three weeks ago is already the cheap tier.

So both things run at once. The floor falls and the ceiling rises, and the distance between them is where five separate businesses are now sitting.

For customers, the useful question this year is not which model wins. It is which tier each job belongs in, and most companies have not done that sorting yet.

The one I’d watch is the middle. The frontier is safe for now and the floor is safe for ever new applications. Many not yet invented. And the two frontier leaders will contest here in particular.

It is the tier in between, the cheap closed models, that has the least defensible position and the most competitors.

Axios ends on the right question, worth stressing.

“The U.S. and China are both racing to make intelligence abundant. Now someone has to prove abundance can still be profitable.”

Both of those races can be won. By different companies, at different prices, at the same time.

An AI Tech Wave pricing trends worth watching at both ends. Stay tuned.


SOURCES

Primary, this post:

  • Axios, August 1, 2026 — DeepSeek’s new bargain model accelerates AI’s race to zero

  • Zack Kass, former OpenAI head of go to market, quoted in the Axios piece above

  • HuggingFace, DeepSeek V4 Flash model card, for the performance comparison cited by Axios

  • Arena.ai crowdsourced leaderboard, for the front end coding comparison cited by Axios

  • Financial Times — ‘The rise of China’s hottest new commodity: AI tokens’, source of the $2 to $3 vs $15 per million output token gap, discussed in RTZ #1041

  • The Wall Street Journal — source of the GPT-5 Nano vs GPT-5 per token figures, discussed in RTZ #894

For more curious AI readers, MP’s own ‘prior discussed’ threads on this topic:

(NOTE: The discussions here are for information purposes only, and not meant as investment advice at any time. Thanks for joining us here)





Want the latest?

Sign up for Michael Parekh's Newsletter below:


Subscribe Here