AI: The Global State of Open Source AI Models (part 2). AI-RTZ #1185

AI: The Global State of Open Source AI Models (part 2). AI-RTZ #1185

Yesterday’s part 1 was the ‘zoom-in’ close-up: China’s week in downloads and valuations. Today, the zoom-out for this AI Tech Wave globally: the full picture of open source AI worldwide, through Hugging Face’s biannual State of Open Models report, published this week. It is a noteworthy overview of the open AI ecosystem the world over with reasonable proxies for AI ‘census’ data. And its charts and graphs reward a slow walk. Here are the five findings that matter most, with my takes on each.

First, the shape of the ecosystem itself. The Hugging Face hub grew from 2.43 to 2.96 million public model repositories between January and August, with datasets crossing a million. But the distribution underneath is extreme: 85.6% of models have fewer than 200 lifetime downloads, and 1.5% of repositories account for 99.2% of all downloads.

My take: open source AI is a hits business, more concentrated than music or venture returns. That is precisely why ‘default workflow’ status, the thing Alibaba Qwen just won in part 1, is the whole game: in a 99.2/1.5 world, distribution compounds and the long tail is a rounding error.

Second, the parameter ceilings. In almost every month of 2026, the largest open model from a Chinese lab was bigger than anything American labs released. China’s monthly ceiling ran between 754 billion and 2.78 trillion parameters. America’s stayed under 130 billion in five of seven months, the exceptions being Nvidia’s Nemotron 3 Ultra at 561 billion and Thinking Machines’ Inkling. I wrote about them and Nvidia recently.

My take: the report’s own framing is the right one: size is now a statement of intent, not capability. Two camps have formed. China’s Moonshot, MiniMax, Xiaomi and Z.ai publish almost nothing below 70 billion parameters, staking everything on benchmark position. Alibaba and Tencent cover the whole range, from under 1 billion up, bidding to be the family developers standardize on. Different prizes, both rational. And the community’s quantization layer, which makes giant models runnable within days, quietly referees the whole contest.

Third, the most surprising finding: America’s top open source publishers are now its chipmakers. AMD and Nvidia each released more than 200 new model repositories this year, far ahead of everyone else, with Google and Meta ranking well below. As the report puts it:

“Hardware vendors have realized that open models are a way to sell chips: a model optimized for your hardware and freely available is the clearest proof that the hardware works.”

My take: this is the Nvidia kingmaker thesis confirmed from primary data, and it marks a structural handoff: open source AI has moved from the model labs to the hardware and infrastructure companies. Meta, which defined open model publishing for years, has drifted toward closed flagships, a shift we covered in ‘Open Source AI: Rockets Are Hard’. The US is still fully present in open source. It just moved down the stack, closer to the silicon.

Fourth, the interleaving nobody talks about. Most US model releases above 100 billion parameters this year are not new models: they are built on top of Chinese open models. The short list of original American entries at that scale: Thinking Machines’ Inkling at 952 billion, Nvidia’s Nemotron 3 Ultra and Super, and Arcee’s Trinity-Large. Meanwhile AMD’s conversion work makes trillion-parameter Chinese models run efficiently on American hardware, and Chinese models are increasingly optimized for domestic chips, the same competition in reverse.

My take: this is the strongest anti-‘US/China race’ datapoint in the report. The stack is interleaving, not decoupling: American silicon running Chinese weights, Chinese fabs like the CXMT story from part 1 racing to serve domestic models, and developers everywhere fine-tuning whatever runs best. As I argued on America’s 250th July 4th, a race needs separable lanes. There are none here.

Fifth, attention is not adoption. Hugging Face compared the top 25 model repositories by downloads with the top 25 by likes. Exactly one repository appears on both lists.

My take: ‘likes’ are marketing, downloads are production. The models that trend on launch day are almost never the ones quietly doing the world’s work a quarter later. It is the same lesson as the valuation twist in part 1: the loud signal and the load-bearing signal are different signals. Measure what runs.

The zoom-out, in one paragraph: open source AI is now global infrastructure, with China supplying the gravitational mass of models, America supplying the silicon and the optimization layers, and a concentrated handful of model families carrying nearly all the traffic.

The watch list from here: whether Meta’s closed turn holds, the AMD and Nvidia publishing cadence, Qwen’s derivative ecosystem compounding, and the domestic-chip optimization loop in China. None of that is a lap time either. The pie expands, and this report is the census of it doing so.

An impact of open source/weight AI the world over. This AI Tech Wave is just beginning. Stay tuned.

(NOTE: The discussions here are for information purposes only, and not meant as investment advice at any time. Thanks for joining us here.)





Want the latest?

Sign up for Michael Parekh's Newsletter below:


Subscribe Here