MiniMax API Providers

Who serves MiniMax, at what price and speed.

MiniMax publishes open weights for its M-series models (M2 through M3), which it also serves through its own API.

MiniMax served 86% of the 1.32T MiniMax tokens that went through a large public LLM router over Oct 3 to Oct 9, 2026. MiniMax's own API served 86%; 17 providers list the family.

Who serves MiniMax tokens

ProviderTokens, 7dShare
MiniMax (lab)1.07T85.7%
GMICloud104B8.3%
CoreWeave55B4.4%
Together9B0.8%
ModelRun [by Modular]3B0.3%
AtlasCloud3B0.3%
MARA1B0.1%
SambaNova1B<0.1%
Friendli603M<0.1%
Groq532M<0.1%
Prime Intellect69M<0.1%
Venice00.0%
StreamLake00.0%

All MiniMax models combined, Oct 3 to Oct 9, 2026.

MiniMax M3: prices and speed by provider

Output prices run from $0.96 to $3.00 per million tokens across 13 providers, median $1.20.

ProviderShareInput $/MOutput $/MCache read $/MTokens/sLatencyUptime 3dContext
MiniMax87.0%$0.30$1.20$0.06711.26s99.9%524K
GMICloud6.8%$0.24$0.96$0.048592.98s99.8%1.05M
CoreWeave4.6%$0.23$0.96$0.05680.78s99.9%262K
Together0.8%$0.30$1.20$0.06430.85s99.8%524K
ModelRun [by Modular]0.3%$0.75$3.00$0.15115.51.77s99.9%1.05M
AtlasCloud0.3%$0.30$1.20$0.06811.01s98.7%524K
MARA0.1%$0.60$2.40–1411.71s96.7%1.05M
SambaNova<0.1%$0.60$2.40$0.06321.47s98.3%1.05M
DeepInfra-$0.28$1.10$0.056401.15s99.5%524K
StreamLake-$0.30$1.20$0.06591.77s99.4%1M
Venice-$0.30$1.20$0.06851.00s96.5%524K
Parasail-$0.30$1.20$0.06741.60s99.9%1.05M
NovitaAI-$0.30$1.20$0.06602.22s99.9%1M

List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.

MiniMax M2.7: prices and speed by provider

Output prices run from $0.84 to $2.40 per million tokens across 5 providers, median $1.20.

ProviderShareInput $/MOutput $/MCache read $/MTokens/sLatencyUptime 3dContext
GMICloud51.9%$0.21$0.84$0.042411.68s99.1%197K
MiniMax46.8%$0.30$1.20$0.06441.34s100.0%205K
Groq1.2%$0.60$1.80–4190.12s96.2%197K
NovitaAI-$0.27$1.08$0.054381.45s99.9%205K
AtlasCloud-$0.30$1.20$0.0638.51.91s98.5%197K

List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.

All MiniMax models

ModelTokens, 7dProvidersOutput $/MContextReleased
MiniMax M31.23T13$0.96 to $3.001.05M2026-05-31
MiniMax M2.765B5$0.84 to $2.40205K2026-03-18
MiniMax M2.514B8$0.95 to $2.40205K2026-02-12
MiniMax M2.11B1$1.20 to $2.40205K2025-12-23
MiniMax M2876M2$1.02 to $1.20205K2025-10-23
MiniMax-01171M1$1.101M2025-01-15

Finance your GPUs

Serving MiniMax on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.

Prefer email? hello@amcompute.com