Inference Providers

Who serves open-weight model tokens, and how many.

What it shows

Tokens each provider served over Oct 3 to Oct 9, 2026, through a large public LLM router. That is one channel: direct, cloud-marketplace and enterprise traffic is not included.

Who serves the most open-weight tokens

Tokens served for other labs' models; a lab serving its own model is left out of its own total.

#ProviderHQTokens, 7dShareLargest familyOpen models
1TogetherUnited States11.5T18.4%DeepSeek14
2Relace-9.54T15.3%DeepSeek6
3DeepInfraUnited States6.43T10.3%DeepSeek69
4inference.netUnited States4.13T6.6%DeepSeek7
5NovitaAIUnited States3.26T5.2%DeepSeek51
6ParasailUnited States2.67T4.3%DeepSeek38
7WaferUnited States2.57T4.1%DeepSeek10
8StreamLakeChina2.53T4.1%DeepSeek19
9FireworksUnited States2.20T3.5%DeepSeek4
10Sail ResearchUnited States1.71T2.7%DeepSeek6
11CoreWeaveUnited States1.55T2.5%DeepSeek19
12GMICloudUnited States1.45T2.3%DeepSeek26
13DigitalOceanUnited States1.37T2.2%DeepSeek16
14AtlasCloudUnited States1.17T1.9%DeepSeek23
15MorphUnited States1.02T1.6%DeepSeek5
16DecartUnited States981B1.6%GLM5
17SiliconFlowSingapore893B1.4%GLM40
18Open InferenceUnited States759B1.2%GLM4
19BasetenUnited States598B1.0%GLM8
20MistralFrance576B0.9%GLM11
21Baidu QianfanChina500B0.8%DeepSeek11
22Alibaba Cloud Int.Singapore466B0.7%Qwen18
23VeniceUnited States465B0.7%DeepSeek38
24ModalUnited States460B0.7%GLM5
25Makora-414B0.7%DeepSeek5
26FriendliUnited States350B0.6%GLM7
27PhalaUnited States311B0.5%DeepSeek20
28ModelRun [by Modular]United States283B0.5%Qwen3
29CohereUnited States283B0.5%DeepSeek2
30Reka AI-280B0.4%GLM8
31CrusoeUnited States204B0.3%GLM6
32DekaLLMIndonesia164B0.3%Qwen10
33GroqUnited States161B0.3%gpt-oss6
34InceptronSweden160B0.3%DeepSeek6
35NextBitSpain154B0.2%Gemma7
36CloudflareUnited States114B0.2%GLM18
37Darkbloom-108B0.2%Gemma9
38IonstreamUnited States103B0.2%DeepSeek3
39io.netUnited States99B0.2%MiMo6
40AkashML-89B0.1%gpt-oss7

Tokens by model family

FamilyTokens, 7dShareWeek on weekProvidersTop providerLab serves
DeepSeek48.3T47.0%+34%39Together 18%12%
GLM16.4T16.0%+20%42Together 16%8%
MiMo13.1T12.7%+11%11Xiaomi 91%91%
Hunyuan10.1T9.8%+14%7Tencent Cloud 97%97%
Nemotron7.85T7.6%+12%11NVIDIA 100%100%
Kimi2.22T2.2%+20%33inference.net 28%15%
Qwen1.47T1.4%+5%31Alibaba Cloud Int. 47%47%
MiniMax1.32T1.3%+0%17MiniMax 86%86%
Gemma882B0.9%+4%19DeepInfra 22%n/a
gpt-oss733B0.7%-5%22Groq 20%n/a
Mistral323B0.3%-2%7DeepInfra 78%8%
Llama117B0.1%-6%12Groq 66%n/a

Lab serves: the share the model's own lab served through its API; n/a where it has none.

Price alone does not win share. On DeepSeek V4.1 Flash, the most-used open model, Together served 23% of tokens at $1.20 per million output tokens. The cheapest listing, Decart at $0.18, served 0.3%.

Prices and uptime. Each family page lists every provider's price and speed per model. Daily history is in the price index and the Inference Uptime Tracker; monthly totals in the Open-Weight Token Report; gateways compared in LLM routers.

Finance your GPUs

Serving tokens on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.

Prefer email? hello@amcompute.com

Questions

What is an inference provider?
A company that runs AI models on its own or rented GPUs and sells the output by the token through an API. Open-weight models such as DeepSeek, GLM, Kimi and Qwen can be served by anyone, so dozens of providers compete on price, speed and uptime for the same model.