Qwen API Providers
Who serves Qwen, at what price and speed.
Qwen is Alibaba's model family. The open-weight releases (dense and mixture-of-experts, from small to Qwen3.8 27B and larger) are served by many providers; the Plus and Max tiers are Alibaba-only.
Alibaba Cloud Int. served 47% of the 1.47T Qwen tokens that went through a large public LLM router over Oct 3 to Oct 9, 2026. Alibaba's own API served 47%; 31 providers list the family.
Who serves Qwen tokens
| Provider | Tokens, 7d | Share |
|---|---|---|
| Alibaba Cloud Int. (lab) | 541B | 46.9% |
| ModelRun [by Modular] | 218B | 18.9% |
| Wafer | 80B | 6.9% |
| DekaLLM | 49B | 4.3% |
| Phala | 33B | 2.9% |
| Parasail | 33B | 2.9% |
| SiliconFlow | 31B | 2.6% |
| Reka AI | 25B | 2.2% |
| Nebius Token Factory | 20B | 1.7% |
| Chutes | 19B | 1.7% |
| AkashML | 17B | 1.5% |
| Venice | 16B | 1.4% |
| Darkbloom | 14B | 1.2% |
| NovitaAI | 14B | 1.2% |
| StreamLake | 9B | 0.8% |
| 14 others | 34B | 2.9% |
All Qwen models combined, Oct 3 to Oct 9, 2026.
Qwen3.8 27B: prices and speed by provider
Output prices run from $1.35 to $4.70 per million tokens across 19 providers, median $2.30.
| Provider | Share | Input $/M | Output $/M | Cache read $/M | Tokens/s | Latency | Uptime 3d | Context |
|---|---|---|---|---|---|---|---|---|
| ModelRun [by Modular] | 37.9% | $0.70 | $4.70 | $0.135 | 231 | 0.09s | 100.0% | 262K |
| Alibaba Cloud Int. | 21.5% | $0.425 | $2.55 | $0.085 | 47 | 1.81s | 98.8% | 1M |
| Wafer | 13.9% | $0.04 | $2.30 | $0.02 | 124 | 0.57s | 99.7% | 262K |
| DekaLLM | 8.1% | $0.049 | $3.00 | $0.02 | 140 | 1.08s | 99.8% | 262K |
| Reka AI | 4.4% | $0.14 | $1.40 | $0.035 | 62 | 0.43s | 98.7% | 262K |
| Phala | 3.2% | $0.15 | $1.88 | $0.037 | 66 | 9.81s | 95.9% | 1M |
| Parasail | 3.1% | $0.24 | $2.20 | $0.05 | 61 | 0.62s | 100.0% | 262K |
| Chutes | 2.8% | $0.24 | $2.20 | $0.024 | 35 | 2.26s | 99.8% | 262K |
| NovitaAI | 2.4% | $0.42 | $3.00 | $0.085 | 30 | 4.57s | 99.2% | 1M |
| Ionstream | 0.8% | $0.089 | $2.35 | $0.085 | 48 | 1.02s | 99.2% | 262K |
| AkashML | 0.5% | $0.225 | $1.98 | $0.05 | 52 | 0.59s | 99.7% | 262K |
| Darkbloom | 0.4% | $0.05 | $2.20 | $0.025 | 24 | 2.42s | 99.1% | 262K |
| NEAR AI | 0.3% | $0.04 | $1.35 | $0.018 | 29 | 0.67s | 100.0% | 262K |
| Venice | 0.3% | $0.45 | $3.20 | – | 103 | 0.68s | 99.7% | 262K |
| Cerebras | 0.3% | $0.99 | $1.49 | $0.99 | 560 | 0.14s | 100.0% | 66K |
| Mancer | 0.2% | $0.20 | $2.50 | – | 69 | 0.60s | 98.8% | 262K |
| DeepInfra | - | $0.15 | $1.88 | $0.037 | 58 | 0.69s | 94.7% | 262K |
| CoreWeave | - | $0.40 | $3.00 | $0.15 | 110 | 0.48s | 98.8% | 262K |
| Cloudflare | - | $0.45 | $3.20 | $0.05 | 41 | 0.57s | 93.9% | 262K |
List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.
Qwen3.8 Flash: prices and speed by provider
Output prices run from $0.47 to $0.47 per million tokens across 1 providers, median $0.47.
| Provider | Share | Input $/M | Output $/M | Cache read $/M | Tokens/s | Latency | Uptime 3d | Context |
|---|---|---|---|---|---|---|---|---|
| Alibaba Cloud Int. | 100.0% | $0.15 | $0.47 | $0.016 | 60 | 3.58s | 100.0% | 1M |
List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.
All Qwen models
| Model | Tokens, 7d | Providers | Output $/M | Context | Released |
|---|---|---|---|---|---|
| Qwen3.8 27B | 508B | 19 | $1.35 to $4.70 | 1M | 2026-08-14 |
| Qwen3.8 Flash | 418B | 1 | $0.47 | 1M | 2026-08-26 |
| Qwen3 235B A22B Instruct 2507 | 76B | 6 | $0.35 to $1.00 | 262K | 2025-07-21 |
| Qwen3.8 2.4T A95B | 76B | 7 | $6.00 | 1.05M | 2026-08-12 |
| Qwen3.6 35B A3B | 70B | 10 | $0.70 to $1.80 | 262K | 2026-04-27 |
| Qwen3 30B A3B Instruct 2507 | 42B | 3 | $0.30 | 262K | 2025-07-29 |
| Qwen3.5-9B | 39B | 6 | $0.13 to $0.25 | 262K | 2026-03-10 |
| Qwen3.5 397B A17B | 39B | 10 | $2.34 to $4.50 | 262K | 2026-02-16 |
| Qwen3 Coder Next | 31B | 1 | $0.80 | 262K | 2026-02-04 |
| Qwen3 Next 80B A3B Instruct | 26B | 3 | $1.10 to $1.20 | 262K | 2025-09-11 |
| Qwen3 VL 235B A22B Instruct | 23B | 3 | $0.88 to $1.90 | 262K | 2025-09-23 |
| Qwen3.6 27B | 15B | 7 | $2.00 to $3.25 | 262K | 2026-04-27 |
| Qwen3 32B | 15B | 3 | $0.28 to $0.57 | 131K | 2025-04-28 |
| Qwen3 Coder 480B A35B | 14B | 3 | $1.00 to $1.80 | 262K | 2025-07-23 |
| Qwen3.5-35B-A3B | 13B | 7 | $0.75 to $1.80 | 262K | 2026-02-25 |
| Qwen3.5-27B | 12B | 6 | $1.56 to $2.60 | 262K | 2026-02-25 |
| Qwen3.5-122B-A10B | 10B | 4 | $2.08 to $3.20 | 262K | 2026-02-25 |
| Qwen3 VL 30B A3B Instruct | 9B | 3 | $0.60 to $1.00 | 262K | 2025-10-06 |
| Qwen3 VL 8B Instruct | 8B | 1 | $0.75 | 262K | 2025-10-14 |
| Qwen3 Coder 30B A3B Instruct | 6B | 2 | $0.28 to $0.60 | 262K | 2025-07-31 |
| Qwen3 14B | 6B | 2 | $0.22 to $0.24 | 41K | 2025-04-28 |
| Qwen2.5 72B Instruct | 3B | 1 | $0.40 | 33K | 2024-09-19 |
| Qwen2.5 7B Instruct | 3B | 1 | $0.20 | 33K | 2024-10-16 |
| Qwen3 30B A3B | 2B | 1 | $0.50 | 41K | 2025-04-28 |
| Qwen2.5 VL 72B Instruct | 2B | 1 | $1.00 | 128K | 2025-02-01 |
| Qwen3 235B A22B Thinking 2507 | 2B | 1 | $3.50 | 128K | 2025-07-25 |
| Qwen3 Next 80B A3B Thinking | 222M | 1 | $1.20 | 262K | 2025-09-11 |
| Qwen2.5 Coder 32B Instruct | 134M | 1 | $1.00 | 33K | 2024-11-11 |
| Qwen3 VL 30B A3B Thinking | 113M | 1 | $1.00 | 262K | 2025-10-06 |
Finance your GPUs
Serving Qwen on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.
Prefer email? hello@amcompute.com