DeepSeek API Providers

Who serves DeepSeek, at what price and speed.

DeepSeek releases its models with open weights, from V3 and R1 to the V4 Flash and Pro generation and V4.1 Flash.

Together served 18% of the 48.3T DeepSeek tokens that went through a large public LLM router over Oct 3 to Oct 9, 2026. DeepSeek's own API served 12%; 39 providers list the family.

Who serves DeepSeek tokens

ProviderTokens, 7dShare
Together8.67T18.2%
Relace7.40T15.5%
DeepSeek (lab)5.96T12.5%
DeepInfra5.75T12.0%
inference.net2.56T5.4%
StreamLake2.17T4.5%
NovitaAI1.68T3.5%
Fireworks1.49T3.1%
Sail Research1.45T3.0%
Parasail1.27T2.7%
Wafer1.26T2.6%
DigitalOcean1.25T2.6%
AtlasCloud1.11T2.3%
CoreWeave940B2.0%
Morph813B1.7%
27 others3.98T8.3%

All DeepSeek models combined, Oct 3 to Oct 9, 2026.

DeepSeek V4.1 Flash: prices and speed by provider

Output prices run from $0.18 to $2.40 per million tokens across 29 providers, median $1.13.

ProviderShareInput $/MOutput $/MCache read $/MTokens/sLatencyUptime 3dContext
Together22.9%$0.30$1.20$0.00602170.30s99.8%1.05M
DeepSeek14.6%$0.15$0.60$0.00301361.05s100.0%1.05M
DeepInfra14.2%$0.14$0.42$0.0042491.35s100.0%1.05M
Relace11.9%$0.0021$0.60$0.00501150.82s100.0%1.05M
inference.net6.8%$0.105$0.60$0.015970.35s100.0%1.04M
Fireworks4.0%$0.30$1.20$0.0060334.88s82.8%1.05M
NovitaAI3.8%$0.195$0.78$0.0039682.03s100.0%1.05M
DigitalOcean2.9%$0.165$0.66$0.00601620.80s99.8%1.05M
Parasail2.9%$0.30$1.20$0.00601880.74s100.0%1.05M
AtlasCloud2.7%$0.30$1.20$0.031481.50s99.7%1.05M
Wafer2.3%$0.05$0.65$0.049870.96s100.0%1.05M
Morph2.1%$0.029$1.00$0.0090721.16s97.9%1.05M
Sail Research1.8%$0.08$0.40$0.01501.35s99.2%1.05M
StreamLake1.5%$0.15$0.60$0.00301131.50s100.0%1.02M
CoreWeave1.2%$0.20$0.65$0.031790.79s99.5%1.05M
Makora0.7%$0.27$1.15$0.00901341.07s99.9%1.05M
Venice0.6%$0.30$1.20$0.00751700.78s100.0%1M
Open Inference0.5%$0.0028$0.36$0.0028261.31s90.0%1.05M
Baseten0.4%$0.30$1.20$0.0070207.50.38s100.0%1.05M
Phala0.4%$0.21$0.84$0.0042701.61s97.7%1.05M
Decart0.3%$0.09$0.18$0.018441.33s99.8%1.05M
GMICloud0.3%$0.18$0.72$0.00361122.14s99.9%1.05M
Alibaba Cloud Int.0.3%$0.30$1.20$0.03681.58s97.9%1M
Modal0.2%$0.30$1.20$0.031670.49s100.0%1.05M
Ionstream0.2%$0.10$1.10$0.01520.86s98.5%1.05M
SiliconFlow0.2%$0.30$1.20$0.0060932.21s99.8%1.05M
Crusoe0.1%$0.29$1.20$0.00701170.74s99.9%1.05M
Baidu Qianfan<0.1%$0.30$1.20$0.00601141.51s97.1%1.05M
DekaLLM<0.1%$0.12$1.20$0.0050781.81s99.4%1.05M

List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.

DeepSeek V4 Flash 0731: prices and speed by provider

Output prices run from $0.18 to $1.32 per million tokens across 24 providers, median $0.42.

ProviderShareInput $/MOutput $/MCache read $/MTokens/sLatencyUptime 3dContext
Relace31.7%$0.0063$1.28$0.0063570.54s100.0%1.05M
StreamLake15.2%$0.088$0.264$0.0028371.50s100.0%1.02M
Sail Research13.8%$0.10$0.30$0.03192.08s96.7%1.05M
CoreWeave8.3%$0.13$0.28$0.071080.41s100.0%262K
DeepInfra5.8%$0.06$0.18$0.015430.71s100.0%1.05M
Cohere5.5%$0.14$0.28$0.071210.41s99.5%1.05M
Wafer4.1%$0.13$0.23$0.12840.91s100.0%1.05M
GMICloud2.8%$0.286$0.858$0.0091542.12s100.0%1.05M
Reka AI2.3%$0.021$0.528$0.00561020.68s98.6%262K
Parasail1.5%$0.14$0.28$0.05940.71s100.0%1.05M
Baidu Qianfan1.4%$0.44$1.32$0.014481.22s99.9%1.05M
DigitalOcean1.3%$0.119$0.238$0.024560.63s100.0%1.05M
Together1.3%$0.14$0.28$0.03590.71s98.9%1.05M
Inceptron1.2%$0.061$0.60$0.027470.96s99.6%1.05M
Open Inference0.9%$0.0069$0.27$0.0050191.36s99.2%1.05M
Baseten0.9%$0.13$0.26$0.0281770.40s100.0%1.05M
NovitaAI0.4%$0.409$1.23$0.026461.46s100.0%1.05M
Venice0.4%$0.13$0.26$0.028401.24s99.8%1M
Phala0.3%$0.44$1.32$0.028722.64s99.0%1.05M
AtlasCloud0.3%$0.44$1.32$0.028471.71s99.8%1.05M
SiliconFlow0.2%$0.22$0.66$0.028271.43s99.8%1.05M
Cloudflare<0.1%$0.44$1.32$0.014502.42s99.6%1.05M
Mancer<0.1%$0.20$0.60–451.30s99.9%1.05M
Alibaba Cloud Int.-$0.352$1.06$0.035561.43s100.0%1M

List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.

All DeepSeek models

ModelTokens, 7dProvidersOutput $/MContextReleased
DeepSeek V4.1 Flash37.6T29$0.18 to $2.401.05M2026-09-10
DeepSeek V4 Flash 07315.19T24$0.18 to $1.321.05M2026-07-31
DeepSeek V4 Flash 04233.08T16$0.165 to $1.321.05M2026-04-24
DeepSeek V4 Pro 08131.12T19$1.98 to $5.001.05M2026-08-12
DeepSeek V4 Pro 0423743B15$1.90 to $10.501.05M2026-04-24
DeepSeek V3.2336B12$0.31 to $4.50164K2025-12-01
DeepSeek V4 Flash Vision Exp88B3$0.647 to $1.321.05M2026-08-21
DeepSeek V3.151B5$0.95 to $1.70164K2025-08-21
DeepSeek V3 032427B2$1.00 to $1.14164K2025-03-24
DeepSeek V322B2$0.89 to $1.03164K2024-12-26
DeepSeek V3.2 Exp13B1$0.41164K2025-09-29
R1 052811B4$2.15 to $2.50164K2025-05-28
DeepSeek V3.1 Terminus6B2$1.00 to $1.03164K2025-09-22
R12B1$2.5064K2025-01-20

Finance your GPUs

Serving DeepSeek on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.

Prefer email? hello@amcompute.com