Qwen API Providers

Who serves Qwen, at what price and speed.

Qwen is Alibaba's model family. The open-weight releases (dense and mixture-of-experts, from small to Qwen3.8 27B and larger) are served by many providers; the Plus and Max tiers are Alibaba-only.

Alibaba Cloud Int. served 47% of the 1.47T Qwen tokens that went through a large public LLM router over Oct 3 to Oct 9, 2026. Alibaba's own API served 47%; 31 providers list the family.

Who serves Qwen tokens

ProviderTokens, 7dShare
Alibaba Cloud Int. (lab)541B46.9%
ModelRun [by Modular]218B18.9%
Wafer80B6.9%
DekaLLM49B4.3%
Phala33B2.9%
Parasail33B2.9%
SiliconFlow31B2.6%
Reka AI25B2.2%
Nebius Token Factory20B1.7%
Chutes19B1.7%
AkashML17B1.5%
Venice16B1.4%
Darkbloom14B1.2%
NovitaAI14B1.2%
StreamLake9B0.8%
14 others34B2.9%

All Qwen models combined, Oct 3 to Oct 9, 2026.

Qwen3.8 27B: prices and speed by provider

Output prices run from $1.35 to $4.70 per million tokens across 19 providers, median $2.30.

ProviderShareInput $/MOutput $/MCache read $/MTokens/sLatencyUptime 3dContext
ModelRun [by Modular]37.9%$0.70$4.70$0.1352310.09s100.0%262K
Alibaba Cloud Int.21.5%$0.425$2.55$0.085471.81s98.8%1M
Wafer13.9%$0.04$2.30$0.021240.57s99.7%262K
DekaLLM8.1%$0.049$3.00$0.021401.08s99.8%262K
Reka AI4.4%$0.14$1.40$0.035620.43s98.7%262K
Phala3.2%$0.15$1.88$0.037669.81s95.9%1M
Parasail3.1%$0.24$2.20$0.05610.62s100.0%262K
Chutes2.8%$0.24$2.20$0.024352.26s99.8%262K
NovitaAI2.4%$0.42$3.00$0.085304.57s99.2%1M
Ionstream0.8%$0.089$2.35$0.085481.02s99.2%262K
AkashML0.5%$0.225$1.98$0.05520.59s99.7%262K
Darkbloom0.4%$0.05$2.20$0.025242.42s99.1%262K
NEAR AI0.3%$0.04$1.35$0.018290.67s100.0%262K
Venice0.3%$0.45$3.20–1030.68s99.7%262K
Cerebras0.3%$0.99$1.49$0.995600.14s100.0%66K
Mancer0.2%$0.20$2.50–690.60s98.8%262K
DeepInfra-$0.15$1.88$0.037580.69s94.7%262K
CoreWeave-$0.40$3.00$0.151100.48s98.8%262K
Cloudflare-$0.45$3.20$0.05410.57s93.9%262K

List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.

Qwen3.8 Flash: prices and speed by provider

Output prices run from $0.47 to $0.47 per million tokens across 1 providers, median $0.47.

ProviderShareInput $/MOutput $/MCache read $/MTokens/sLatencyUptime 3dContext
Alibaba Cloud Int.100.0%$0.15$0.47$0.016603.58s100.0%1M

List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.

All Qwen models

ModelTokens, 7dProvidersOutput $/MContextReleased
Qwen3.8 27B508B19$1.35 to $4.701M2026-08-14
Qwen3.8 Flash418B1$0.471M2026-08-26
Qwen3 235B A22B Instruct 250776B6$0.35 to $1.00262K2025-07-21
Qwen3.8 2.4T A95B76B7$6.001.05M2026-08-12
Qwen3.6 35B A3B70B10$0.70 to $1.80262K2026-04-27
Qwen3 30B A3B Instruct 250742B3$0.30262K2025-07-29
Qwen3.5-9B39B6$0.13 to $0.25262K2026-03-10
Qwen3.5 397B A17B39B10$2.34 to $4.50262K2026-02-16
Qwen3 Coder Next31B1$0.80262K2026-02-04
Qwen3 Next 80B A3B Instruct26B3$1.10 to $1.20262K2025-09-11
Qwen3 VL 235B A22B Instruct23B3$0.88 to $1.90262K2025-09-23
Qwen3.6 27B15B7$2.00 to $3.25262K2026-04-27
Qwen3 32B15B3$0.28 to $0.57131K2025-04-28
Qwen3 Coder 480B A35B14B3$1.00 to $1.80262K2025-07-23
Qwen3.5-35B-A3B13B7$0.75 to $1.80262K2026-02-25
Qwen3.5-27B12B6$1.56 to $2.60262K2026-02-25
Qwen3.5-122B-A10B10B4$2.08 to $3.20262K2026-02-25
Qwen3 VL 30B A3B Instruct9B3$0.60 to $1.00262K2025-10-06
Qwen3 VL 8B Instruct8B1$0.75262K2025-10-14
Qwen3 Coder 30B A3B Instruct6B2$0.28 to $0.60262K2025-07-31
Qwen3 14B6B2$0.22 to $0.2441K2025-04-28
Qwen2.5 72B Instruct3B1$0.4033K2024-09-19
Qwen2.5 7B Instruct3B1$0.2033K2024-10-16
Qwen3 30B A3B2B1$0.5041K2025-04-28
Qwen2.5 VL 72B Instruct2B1$1.00128K2025-02-01
Qwen3 235B A22B Thinking 25072B1$3.50128K2025-07-25
Qwen3 Next 80B A3B Thinking222M1$1.20262K2025-09-11
Qwen2.5 Coder 32B Instruct134M1$1.0033K2024-11-11
Qwen3 VL 30B A3B Thinking113M1$1.00262K2025-10-06

Finance your GPUs

Serving Qwen on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.

Prefer email? hello@amcompute.com