DeepSeek API Providers
Who serves DeepSeek, at what price and speed.
DeepSeek releases its models with open weights, from V3 and R1 to the V4 Flash and Pro generation and V4.1 Flash.
Together served 18% of the 48.3T DeepSeek tokens that went through a large public LLM router over Oct 3 to Oct 9, 2026. DeepSeek's own API served 12%; 39 providers list the family.
Who serves DeepSeek tokens
| Provider | Tokens, 7d | Share |
|---|---|---|
| Together | 8.67T | 18.2% |
| Relace | 7.40T | 15.5% |
| DeepSeek (lab) | 5.96T | 12.5% |
| DeepInfra | 5.75T | 12.0% |
| inference.net | 2.56T | 5.4% |
| StreamLake | 2.17T | 4.5% |
| NovitaAI | 1.68T | 3.5% |
| Fireworks | 1.49T | 3.1% |
| Sail Research | 1.45T | 3.0% |
| Parasail | 1.27T | 2.7% |
| Wafer | 1.26T | 2.6% |
| DigitalOcean | 1.25T | 2.6% |
| AtlasCloud | 1.11T | 2.3% |
| CoreWeave | 940B | 2.0% |
| Morph | 813B | 1.7% |
| 27 others | 3.98T | 8.3% |
All DeepSeek models combined, Oct 3 to Oct 9, 2026.
DeepSeek V4.1 Flash: prices and speed by provider
Output prices run from $0.18 to $2.40 per million tokens across 29 providers, median $1.13.
| Provider | Share | Input $/M | Output $/M | Cache read $/M | Tokens/s | Latency | Uptime 3d | Context |
|---|---|---|---|---|---|---|---|---|
| Together | 22.9% | $0.30 | $1.20 | $0.0060 | 217 | 0.30s | 99.8% | 1.05M |
| DeepSeek | 14.6% | $0.15 | $0.60 | $0.0030 | 136 | 1.05s | 100.0% | 1.05M |
| DeepInfra | 14.2% | $0.14 | $0.42 | $0.0042 | 49 | 1.35s | 100.0% | 1.05M |
| Relace | 11.9% | $0.0021 | $0.60 | $0.0050 | 115 | 0.82s | 100.0% | 1.05M |
| inference.net | 6.8% | $0.105 | $0.60 | $0.015 | 97 | 0.35s | 100.0% | 1.04M |
| Fireworks | 4.0% | $0.30 | $1.20 | $0.0060 | 33 | 4.88s | 82.8% | 1.05M |
| NovitaAI | 3.8% | $0.195 | $0.78 | $0.0039 | 68 | 2.03s | 100.0% | 1.05M |
| DigitalOcean | 2.9% | $0.165 | $0.66 | $0.0060 | 162 | 0.80s | 99.8% | 1.05M |
| Parasail | 2.9% | $0.30 | $1.20 | $0.0060 | 188 | 0.74s | 100.0% | 1.05M |
| AtlasCloud | 2.7% | $0.30 | $1.20 | $0.03 | 148 | 1.50s | 99.7% | 1.05M |
| Wafer | 2.3% | $0.05 | $0.65 | $0.049 | 87 | 0.96s | 100.0% | 1.05M |
| Morph | 2.1% | $0.029 | $1.00 | $0.0090 | 72 | 1.16s | 97.9% | 1.05M |
| Sail Research | 1.8% | $0.08 | $0.40 | $0.01 | 50 | 1.35s | 99.2% | 1.05M |
| StreamLake | 1.5% | $0.15 | $0.60 | $0.0030 | 113 | 1.50s | 100.0% | 1.02M |
| CoreWeave | 1.2% | $0.20 | $0.65 | $0.03 | 179 | 0.79s | 99.5% | 1.05M |
| Makora | 0.7% | $0.27 | $1.15 | $0.0090 | 134 | 1.07s | 99.9% | 1.05M |
| Venice | 0.6% | $0.30 | $1.20 | $0.0075 | 170 | 0.78s | 100.0% | 1M |
| Open Inference | 0.5% | $0.0028 | $0.36 | $0.0028 | 26 | 1.31s | 90.0% | 1.05M |
| Baseten | 0.4% | $0.30 | $1.20 | $0.0070 | 207.5 | 0.38s | 100.0% | 1.05M |
| Phala | 0.4% | $0.21 | $0.84 | $0.0042 | 70 | 1.61s | 97.7% | 1.05M |
| Decart | 0.3% | $0.09 | $0.18 | $0.018 | 44 | 1.33s | 99.8% | 1.05M |
| GMICloud | 0.3% | $0.18 | $0.72 | $0.0036 | 112 | 2.14s | 99.9% | 1.05M |
| Alibaba Cloud Int. | 0.3% | $0.30 | $1.20 | $0.03 | 68 | 1.58s | 97.9% | 1M |
| Modal | 0.2% | $0.30 | $1.20 | $0.03 | 167 | 0.49s | 100.0% | 1.05M |
| Ionstream | 0.2% | $0.10 | $1.10 | $0.01 | 52 | 0.86s | 98.5% | 1.05M |
| SiliconFlow | 0.2% | $0.30 | $1.20 | $0.0060 | 93 | 2.21s | 99.8% | 1.05M |
| Crusoe | 0.1% | $0.29 | $1.20 | $0.0070 | 117 | 0.74s | 99.9% | 1.05M |
| Baidu Qianfan | <0.1% | $0.30 | $1.20 | $0.0060 | 114 | 1.51s | 97.1% | 1.05M |
| DekaLLM | <0.1% | $0.12 | $1.20 | $0.0050 | 78 | 1.81s | 99.4% | 1.05M |
List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.
DeepSeek V4 Flash 0731: prices and speed by provider
Output prices run from $0.18 to $1.32 per million tokens across 24 providers, median $0.42.
| Provider | Share | Input $/M | Output $/M | Cache read $/M | Tokens/s | Latency | Uptime 3d | Context |
|---|---|---|---|---|---|---|---|---|
| Relace | 31.7% | $0.0063 | $1.28 | $0.0063 | 57 | 0.54s | 100.0% | 1.05M |
| StreamLake | 15.2% | $0.088 | $0.264 | $0.0028 | 37 | 1.50s | 100.0% | 1.02M |
| Sail Research | 13.8% | $0.10 | $0.30 | $0.03 | 19 | 2.08s | 96.7% | 1.05M |
| CoreWeave | 8.3% | $0.13 | $0.28 | $0.07 | 108 | 0.41s | 100.0% | 262K |
| DeepInfra | 5.8% | $0.06 | $0.18 | $0.015 | 43 | 0.71s | 100.0% | 1.05M |
| Cohere | 5.5% | $0.14 | $0.28 | $0.07 | 121 | 0.41s | 99.5% | 1.05M |
| Wafer | 4.1% | $0.13 | $0.23 | $0.12 | 84 | 0.91s | 100.0% | 1.05M |
| GMICloud | 2.8% | $0.286 | $0.858 | $0.0091 | 54 | 2.12s | 100.0% | 1.05M |
| Reka AI | 2.3% | $0.021 | $0.528 | $0.0056 | 102 | 0.68s | 98.6% | 262K |
| Parasail | 1.5% | $0.14 | $0.28 | $0.05 | 94 | 0.71s | 100.0% | 1.05M |
| Baidu Qianfan | 1.4% | $0.44 | $1.32 | $0.014 | 48 | 1.22s | 99.9% | 1.05M |
| DigitalOcean | 1.3% | $0.119 | $0.238 | $0.024 | 56 | 0.63s | 100.0% | 1.05M |
| Together | 1.3% | $0.14 | $0.28 | $0.03 | 59 | 0.71s | 98.9% | 1.05M |
| Inceptron | 1.2% | $0.061 | $0.60 | $0.027 | 47 | 0.96s | 99.6% | 1.05M |
| Open Inference | 0.9% | $0.0069 | $0.27 | $0.0050 | 19 | 1.36s | 99.2% | 1.05M |
| Baseten | 0.9% | $0.13 | $0.26 | $0.028 | 177 | 0.40s | 100.0% | 1.05M |
| NovitaAI | 0.4% | $0.409 | $1.23 | $0.026 | 46 | 1.46s | 100.0% | 1.05M |
| Venice | 0.4% | $0.13 | $0.26 | $0.028 | 40 | 1.24s | 99.8% | 1M |
| Phala | 0.3% | $0.44 | $1.32 | $0.028 | 72 | 2.64s | 99.0% | 1.05M |
| AtlasCloud | 0.3% | $0.44 | $1.32 | $0.028 | 47 | 1.71s | 99.8% | 1.05M |
| SiliconFlow | 0.2% | $0.22 | $0.66 | $0.028 | 27 | 1.43s | 99.8% | 1.05M |
| Cloudflare | <0.1% | $0.44 | $1.32 | $0.014 | 50 | 2.42s | 99.6% | 1.05M |
| Mancer | <0.1% | $0.20 | $0.60 | – | 45 | 1.30s | 99.9% | 1.05M |
| Alibaba Cloud Int. | - | $0.352 | $1.06 | $0.035 | 56 | 1.43s | 100.0% | 1M |
List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.
All DeepSeek models
| Model | Tokens, 7d | Providers | Output $/M | Context | Released |
|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | 37.6T | 29 | $0.18 to $2.40 | 1.05M | 2026-09-10 |
| DeepSeek V4 Flash 0731 | 5.19T | 24 | $0.18 to $1.32 | 1.05M | 2026-07-31 |
| DeepSeek V4 Flash 0423 | 3.08T | 16 | $0.165 to $1.32 | 1.05M | 2026-04-24 |
| DeepSeek V4 Pro 0813 | 1.12T | 19 | $1.98 to $5.00 | 1.05M | 2026-08-12 |
| DeepSeek V4 Pro 0423 | 743B | 15 | $1.90 to $10.50 | 1.05M | 2026-04-24 |
| DeepSeek V3.2 | 336B | 12 | $0.31 to $4.50 | 164K | 2025-12-01 |
| DeepSeek V4 Flash Vision Exp | 88B | 3 | $0.647 to $1.32 | 1.05M | 2026-08-21 |
| DeepSeek V3.1 | 51B | 5 | $0.95 to $1.70 | 164K | 2025-08-21 |
| DeepSeek V3 0324 | 27B | 2 | $1.00 to $1.14 | 164K | 2025-03-24 |
| DeepSeek V3 | 22B | 2 | $0.89 to $1.03 | 164K | 2024-12-26 |
| DeepSeek V3.2 Exp | 13B | 1 | $0.41 | 164K | 2025-09-29 |
| R1 0528 | 11B | 4 | $2.15 to $2.50 | 164K | 2025-05-28 |
| DeepSeek V3.1 Terminus | 6B | 2 | $1.00 to $1.03 | 164K | 2025-09-22 |
| R1 | 2B | 1 | $2.50 | 64K | 2025-01-20 |
Finance your GPUs
Serving DeepSeek on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.
Prefer email? hello@amcompute.com