Pangram verdict · v3.3
We believe this text is mainly AI, with some human-written content.
AI likelihood · overall
AIArticle text · 1,282 words · 4 segments analyzed
In the AI world, time feels compressed. A few months after our spring report in our biannual analysis worked through the ecosystem, there are quite a few findings that we have observed until this summer. This report lays out these observations from January to August 2026 and presents the data behind each one. Models and datasets on HF hub are growing on a daily basis. Public model repositories grew from 2.43 to 2.96 million over the period, datasets from 711,000 to 1 million, Spaces from 1.00 to 1.44 million. The distribution underneath stays extreme, roughly 85.6% of models have fewer than 200 lifetime downloads, and 1.5% of repositories account for 99.2% of all downloads. Everything below happens inside that shape. 1. The frontier is moving fast There used to be a clear progression path: labs would start by releasing smaller models and gradually work their way toward the top end of the scale. In 2026, several Chinese labs skipped this progression entirely. In almost every month of 2026, the largest and most performant open model from a Chinese lab was larger than any model an American lab released. China's monthly ceiling ran between 754B and 2.78 trillion parameters; U.S. models stayed under 130B in five of seven months, the exception being NVIDIA's Nemotron 3 Ultra at 561B in May and June, and Inkling from Thinking Machines Lab. The chart splits the labs into two camps. Moonshot, MiniMax, Xiaomi and Z.ai publish almost nothing below 70B, so a developer's first encounter with them is a model too large to run on anything they own. Tencent and Alibaba Qwen cover the whole range instead, from under 1B upward. Two things made the first camp possible. Building large stopped being a differentiator. Xiaomi, Ant Group and Meituan all cleared a trillion parameters this year, and neither was a household name in open weights twelve months ago. And a lab no longer has to ship a small model to be reachable, because the community's quantization layer will make a large one runnable within days, a dependency we return to below. That leaves the size profile as a statement of intent rather than of capability. A frontier only portfolio stakes everything on benchmark position and API demand. A full spectrum portfolio is a bid to be the family developers standardise on. Both are rational, they are playing for different prizes. The United States is not absent from open source. The two organizations publishing the most new open models this year are also the companies making the hardware: AMD and NVIDIA. Each released more than 200 new model repositories, far ahead of the rest of the field, with LiquidAI ranking third at around 100. Hardware vendors have realized that open models are a way to sell chips: a model optimized for your hardware and freely available is the clearest proof that the hardware works. When smaller models and embedding models are included, where Google, Microsoft, IBM Granite, and OpenAI’s older vision and speech models generate hundreds of millions of downloads annually, U.S. open source AI is growing. More hardware and infrastructure organizations such as NVIDIA are training and open-weighting competitive models. NVIDIA's Nemotron model family boasts high performance. Long-time leaders such as Meta reignite open roots with Meta's Muse Glimmer. At the frontier scale, some U.S. model releases above 100B parameters this year are built on top of Chinese models or leverage artifacts from Chinese labs, such as Thinking Machines’ Inkling (952B). Major original American models include NVIDIA’s Nemotron 3 Ultra (561B), Nemotron 3 Super (124B), and Arcee AI’s Trinity-Large (399B). AMD contributed many conversions. This work is important: it enables trillion-parameter models to run efficiently on U.S. hardware. This represents a distribution and optimization layer. Meanwhile, Chinese open models are increasingly optimized for domestic chips in China, the same competition in reverse, where models are designed around specific hardware ecosystems. 2. Attention ≠ Adoption We took the top 25 model repositories by downloads accumulated this year and the top 25 by likes. Exactly one repository appears in both lists. We counted downloads inside the window rather than lifetime, so nothing is credited for merely having existed longer, and controlling for age makes the split sharper. Not one model published in 2026 reaches the download top 25, while thirteen of the twenty-five date from 2022. all-MiniLM-L6-v2 was pulled 1.55 billion times in seven months against 5,156 likes; Kimi-K3 was pulled about 60 times per like it received. The two numbers record different acts. A like says a release matters, and goes to frontier models in the weeks after they ship. A download says something is wired into a pipeline that runs on a schedule, and accrues to small, stable models over years. Likes are the right instrument for reading what the field is excited about, downloads for reading what it currently depends on. Treating either as a proxy for the other is the most common mistake we see in coverage of the Hub, including our own earlier work. The same split appears at the level of the publisher. China's frontier labs are the only accounts on the Hub where the heavy band carries the volume. Effectively all of MiniMax's 2026 downloads are of models above 70B, along with 88% of Moonshot's, 55% of DeepSeek's and 39% of Z.ai's. No large American account looks like this: Google, Microsoft and IBM Granite record essentially none of their 2026 downloads above 70B, and NVIDIA and Meta only 14% and 9%. The difference becomes clearer in total downloads. Moonshot’s frontier-only portfolio recorded 37M downloads over the year, while Qwen’s broader release strategy across model sizes reached 2,045M (across repositories with declared parameter counts, 2,061M including all repositories) , about 55 times more. The continued expansion of the family, from the 2.4T-parameter Qwen 3.8 Max to smaller variants such as 27B, shows the same focus on coverage across different use cases. Time also plays an important role. Most models experience a sharp decline in usage after release, followed by a long tail of steady activity. A model’s adoption is largely determined within its first few months. This helps explain why today’s download volume is often driven not by the newest releases, but by a smaller group of models that have become established infrastructure over time. 3. Open weights shift where value accumulates If frontier models were a licensing business, you would expect the biggest releases to carry the tightest terms. However, the data below shows a different story. Of 178 Chinese releases above 20B parameters this year, 59% carry Apache 2.0 and 22% carry MIT, and almost none carry non-commercial restrictions.
However, in the last few weeks, we started to see a change on this trend for the really large models, with Kimi K3 and Qwen 3.8 2.4T starting to include some non-commercial restrictions and revenue share requirements to their licenses DeepSeek and Z.ai ship models between 700 billion and 1.65 trillion parameters under plain MIT.
Chinese labs license their largest models about as permissively as their smallest, and more permissively than American labs license theirs: on the American side of the same size band, 29% is Apache or MIT, 41% sits under custom terms and 30% declares nothing at all. Whatever these releases are for, it is not licence revenue. The weights are given away on the most permissive terms available. The return has to come from somewhere else: API and cloud business, hardware and platform positioning, or the ecosystem position itself.
For instance, the valuations of Z.ai and Kimi point to an effective open source strategy, getting traction and growth opportunities in the community. Going forward, however, the industry is likely to shift toward clearer monetization paths from open-source adoption. 4.