Inference Providers
Who serves open-weight model tokens, and how many.
What it shows
Tokens each provider served over Oct 3 to Oct 9, 2026, through a large public LLM router. That is one channel: direct, cloud-marketplace and enterprise traffic is not included.
Who serves the most open-weight tokens
Tokens served for other labs' models; a lab serving its own model is left out of its own total.
| # | Provider | HQ | Tokens, 7d | Share | Largest family | Open models |
|---|---|---|---|---|---|---|
| 1 | Together | United States | 11.5T | 18.4% | DeepSeek | 14 |
| 2 | Relace | - | 9.54T | 15.3% | DeepSeek | 6 |
| 3 | DeepInfra | United States | 6.43T | 10.3% | DeepSeek | 69 |
| 4 | inference.net | United States | 4.13T | 6.6% | DeepSeek | 7 |
| 5 | NovitaAI | United States | 3.26T | 5.2% | DeepSeek | 51 |
| 6 | Parasail | United States | 2.67T | 4.3% | DeepSeek | 38 |
| 7 | Wafer | United States | 2.57T | 4.1% | DeepSeek | 10 |
| 8 | StreamLake | China | 2.53T | 4.1% | DeepSeek | 19 |
| 9 | Fireworks | United States | 2.20T | 3.5% | DeepSeek | 4 |
| 10 | Sail Research | United States | 1.71T | 2.7% | DeepSeek | 6 |
| 11 | CoreWeave | United States | 1.55T | 2.5% | DeepSeek | 19 |
| 12 | GMICloud | United States | 1.45T | 2.3% | DeepSeek | 26 |
| 13 | DigitalOcean | United States | 1.37T | 2.2% | DeepSeek | 16 |
| 14 | AtlasCloud | United States | 1.17T | 1.9% | DeepSeek | 23 |
| 15 | Morph | United States | 1.02T | 1.6% | DeepSeek | 5 |
| 16 | Decart | United States | 981B | 1.6% | GLM | 5 |
| 17 | SiliconFlow | Singapore | 893B | 1.4% | GLM | 40 |
| 18 | Open Inference | United States | 759B | 1.2% | GLM | 4 |
| 19 | Baseten | United States | 598B | 1.0% | GLM | 8 |
| 20 | Mistral | France | 576B | 0.9% | GLM | 11 |
| 21 | Baidu Qianfan | China | 500B | 0.8% | DeepSeek | 11 |
| 22 | Alibaba Cloud Int. | Singapore | 466B | 0.7% | Qwen | 18 |
| 23 | Venice | United States | 465B | 0.7% | DeepSeek | 38 |
| 24 | Modal | United States | 460B | 0.7% | GLM | 5 |
| 25 | Makora | - | 414B | 0.7% | DeepSeek | 5 |
| 26 | Friendli | United States | 350B | 0.6% | GLM | 7 |
| 27 | Phala | United States | 311B | 0.5% | DeepSeek | 20 |
| 28 | ModelRun [by Modular] | United States | 283B | 0.5% | Qwen | 3 |
| 29 | Cohere | United States | 283B | 0.5% | DeepSeek | 2 |
| 30 | Reka AI | - | 280B | 0.4% | GLM | 8 |
| 31 | Crusoe | United States | 204B | 0.3% | GLM | 6 |
| 32 | DekaLLM | Indonesia | 164B | 0.3% | Qwen | 10 |
| 33 | Groq | United States | 161B | 0.3% | gpt-oss | 6 |
| 34 | Inceptron | Sweden | 160B | 0.3% | DeepSeek | 6 |
| 35 | NextBit | Spain | 154B | 0.2% | Gemma | 7 |
| 36 | Cloudflare | United States | 114B | 0.2% | GLM | 18 |
| 37 | Darkbloom | - | 108B | 0.2% | Gemma | 9 |
| 38 | Ionstream | United States | 103B | 0.2% | DeepSeek | 3 |
| 39 | io.net | United States | 99B | 0.2% | MiMo | 6 |
| 40 | AkashML | - | 89B | 0.1% | gpt-oss | 7 |
Tokens by model family
| Family | Tokens, 7d | Share | Week on week | Providers | Top provider | Lab serves |
|---|---|---|---|---|---|---|
| DeepSeek | 48.3T | 47.0% | +34% | 39 | Together 18% | 12% |
| GLM | 16.4T | 16.0% | +20% | 42 | Together 16% | 8% |
| MiMo | 13.1T | 12.7% | +11% | 11 | Xiaomi 91% | 91% |
| Hunyuan | 10.1T | 9.8% | +14% | 7 | Tencent Cloud 97% | 97% |
| Nemotron | 7.85T | 7.6% | +12% | 11 | NVIDIA 100% | 100% |
| Kimi | 2.22T | 2.2% | +20% | 33 | inference.net 28% | 15% |
| Qwen | 1.47T | 1.4% | +5% | 31 | Alibaba Cloud Int. 47% | 47% |
| MiniMax | 1.32T | 1.3% | +0% | 17 | MiniMax 86% | 86% |
| Gemma | 882B | 0.9% | +4% | 19 | DeepInfra 22% | n/a |
| gpt-oss | 733B | 0.7% | -5% | 22 | Groq 20% | n/a |
| Mistral | 323B | 0.3% | -2% | 7 | DeepInfra 78% | 8% |
| Llama | 117B | 0.1% | -6% | 12 | Groq 66% | n/a |
Lab serves: the share the model's own lab served through its API; n/a where it has none.
Price alone does not win share. On DeepSeek V4.1 Flash, the most-used open model, Together served 23% of tokens at $1.20 per million output tokens. The cheapest listing, Decart at $0.18, served 0.3%.
Prices and uptime. Each family page lists every provider's price and speed per model. Daily history is in the price index and the Inference Uptime Tracker; monthly totals in the Open-Weight Token Report; gateways compared in LLM routers.
Finance your GPUs
Serving tokens on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.
Prefer email? hello@amcompute.com
Questions
- What is an inference provider?
- A company that runs AI models on its own or rented GPUs and sells the output by the token through an API. Open-weight models such as DeepSeek, GLM, Kimi and Qwen can be served by anyone, so dozens of providers compete on price, speed and uptime for the same model.