Open-Weight Token Report: September 2026
Tokens served on open-weight models through a large public LLM router in August 2026, by model family and provider. Compiled after the fact, on October 10, 2026.
What moved in August
Open-weight volume rose 28% a day. 242T open-weight tokens were served through the router in August 2026, 7.8T a day against 6.1T in July 13-31. The router is one channel among many: direct and enterprise traffic is not included.
DeepSeek was the largest family, with 35%. It also gained the most share: +11.9 points.
DeepInfra added the most volume among providers serving other labs' models. Its open-weight tokens rose from 140B a day in July 13-31 to 499B in August. Next: Baidu Qianfan (+209B a day) and CoreWeave (+190B a day).
Daily volume by model family
- DeepSeek
- GLM
- Hunyuan
- MiMo
- Nemotron
- MiniMax
- Other families
Model families
| Family | Lab | Tokens, August | Share | Share chg, pts | Per day chg | Top provider | Lab serves |
|---|---|---|---|---|---|---|---|
| DeepSeek | DeepSeek | 84.7T | 35.0% | +11.9 | +93% | NovitaAI 17% | 14% |
| Hunyuan | Tencent | 38.9T | 16.1% | -1.9 | +14% | Tencent Cloud 99% | 99% |
| MiMo | Xiaomi | 31.7T | 13.1% | -9.6 | -26% | Xiaomi 99% | 99% |
| GLM | Z.ai (Zhipu) | 25.6T | 10.6% | +1.9 | +55% | Z.ai 33% | 33% |
| Nemotron | NVIDIA | 20.7T | 8.6% | +1.3 | +51% | NVIDIA 98% | 98% |
| MiniMax | MiniMax | 11.0T | 4.5% | -2.1 | -13% | MiniMax 52% | 52% |
| Other open-weight | Various | 10.6T | 4.4% | +1.5 | +96% | Poolside 76% | n/a |
| Kimi | Moonshot AI | 7.2T | 3.0% | -0.4 | +13% | Moonshot AI 36% | 36% |
| Step | StepFun | 3.6T | 1.5% | -1.9 | -44% | StepFun 100% | 100% |
| Gemma | 3.1T | 1.3% | -0.3 | +7% | DeepInfra 25% | n/a | |
| gpt-oss | OpenAI | 2.7T | 1.1% | -0.1 | +16% | CoreWeave 33% | n/a |
| Qwen | Alibaba (Qwen team) | 1.3T | 0.5% | 0.0 | +28% | AkashML 22% | 12% |
| Mistral | Mistral AI | 760B | 0.3% | -0.4 | -43% | DeepInfra 74% | 17% |
| Llama | Meta | 341B | 0.1% | 0.0 | +36% | Groq 64% | n/a |
Lab serves: share of the family's tokens served by the lab that made it; n/a where the lab sells no API through the router.
Provider leaderboard
| # | Provider | Tokens, August | Share | Share, July 13-31 | Per day chg | Main family |
|---|---|---|---|---|---|---|
| 1 | Tencent Cloudlab | 38.4T | 15.8% | 5.4% | +273% | Hunyuan |
| 2 | Xiaomilab | 31.4T | 13.0% | 22.4% | -26% | MiMo |
| 3 | NVIDIAlab | 20.4T | 8.4% | 7.1% | +51% | Nemotron |
| 4 | NovitaAI | 18.0T | 7.4% | 20.9% | -55% | DeepSeek |
| 5 | DeepInfra | 15.5T | 6.4% | 2.3% | +257% | DeepSeek |
| 6 | DeepSeeklab | 11.9T | 4.9% | 6.0% | +4% | DeepSeek |
| 7 | Baidu Qianfan | 11.0T | 4.5% | 2.4% | +143% | DeepSeek |
| 8 | GMICloud | 10.9T | 4.5% | 3.1% | +82% | DeepSeek |
| 9 | StreamLake | 9.9T | 4.1% | 4.5% | +17% | DeepSeek |
| 10 | Z.ailab | 8.5T | 3.5% | 0.7% | +566% | GLM |
| 11 | Poolsidelab | 8.0T | 3.3% | 1.0% | +307% | Other open-weight |
| 12 | CoreWeave | 6.6T | 2.7% | 0.4% | +778% | DeepSeek |
| 13 | MiniMaxlab | 5.7T | 2.4% | 5.3% | -42% | MiniMax |
| 14 | Relace | 5.0T | 2.1% | 0.0% | new | DeepSeek |
| 15 | Alibaba Cloud Int. | 4.3T | 1.8% | 2.6% | -14% | DeepSeek |
Share of all open-weight tokens. Lab: a lab serving its own models.
Fastest-growing providers
| Provider | Per day, July 13-31 | Per day, August | Added per day | Change | Last 7 days of August |
|---|---|---|---|---|---|
| Tencent Cloudlab | 332B | 1.2T | +906B | +273% | 1.5T |
| DeepInfra | 140B | 499B | +359B | +257% | 448B |
| Z.ailab | 41B | 275B | +234B | +566% | 1.0T |
| NVIDIAlab | 437B | 659B | +222B | +51% | 901B |
| Baidu Qianfan | 147B | 356B | +209B | +143% | 301B |
| Poolsidelab | 64B | 259B | +195B | +307% | 215B |
| CoreWeave | 24B | 214B | +190B | +778% | 178B |
| Relace | 0 | 162B | +162B | - | 509B |
| GMICloud | 193B | 351B | +158B | +82% | 695B |
| DigitalOcean | 21B | 83B | +62B | +297% | 114B |
Ranked by open-weight tokens added a day, among providers serving at least 20B a day in August.
Other editions
- October 2026 edition (September 2026 data, 368T open-weight tokens)
Discuss a transaction
Serving open-weight models on your own GPUs, or lending to a provider that does? Tell us about the cluster and the financing you need.
Prefer email? hello@amcompute.com