MiniMax API Providers
Who serves MiniMax, at what price and speed.
MiniMax publishes open weights for its M-series models (M2 through M3), which it also serves through its own API.
MiniMax served 86% of the 1.32T MiniMax tokens that went through a large public LLM router over Oct 3 to Oct 9, 2026. MiniMax's own API served 86%; 17 providers list the family.
Who serves MiniMax tokens
| Provider | Tokens, 7d | Share |
|---|---|---|
| MiniMax (lab) | 1.07T | 85.7% |
| GMICloud | 104B | 8.3% |
| CoreWeave | 55B | 4.4% |
| Together | 9B | 0.8% |
| ModelRun [by Modular] | 3B | 0.3% |
| AtlasCloud | 3B | 0.3% |
| MARA | 1B | 0.1% |
| SambaNova | 1B | <0.1% |
| Friendli | 603M | <0.1% |
| Groq | 532M | <0.1% |
| Prime Intellect | 69M | <0.1% |
| Venice | 0 | 0.0% |
| StreamLake | 0 | 0.0% |
All MiniMax models combined, Oct 3 to Oct 9, 2026.
MiniMax M3: prices and speed by provider
Output prices run from $0.96 to $3.00 per million tokens across 13 providers, median $1.20.
| Provider | Share | Input $/M | Output $/M | Cache read $/M | Tokens/s | Latency | Uptime 3d | Context |
|---|---|---|---|---|---|---|---|---|
| MiniMax | 87.0% | $0.30 | $1.20 | $0.06 | 71 | 1.26s | 99.9% | 524K |
| GMICloud | 6.8% | $0.24 | $0.96 | $0.048 | 59 | 2.98s | 99.8% | 1.05M |
| CoreWeave | 4.6% | $0.23 | $0.96 | $0.05 | 68 | 0.78s | 99.9% | 262K |
| Together | 0.8% | $0.30 | $1.20 | $0.06 | 43 | 0.85s | 99.8% | 524K |
| ModelRun [by Modular] | 0.3% | $0.75 | $3.00 | $0.15 | 115.5 | 1.77s | 99.9% | 1.05M |
| AtlasCloud | 0.3% | $0.30 | $1.20 | $0.06 | 81 | 1.01s | 98.7% | 524K |
| MARA | 0.1% | $0.60 | $2.40 | – | 141 | 1.71s | 96.7% | 1.05M |
| SambaNova | <0.1% | $0.60 | $2.40 | $0.06 | 32 | 1.47s | 98.3% | 1.05M |
| DeepInfra | - | $0.28 | $1.10 | $0.056 | 40 | 1.15s | 99.5% | 524K |
| StreamLake | - | $0.30 | $1.20 | $0.06 | 59 | 1.77s | 99.4% | 1M |
| Venice | - | $0.30 | $1.20 | $0.06 | 85 | 1.00s | 96.5% | 524K |
| Parasail | - | $0.30 | $1.20 | $0.06 | 74 | 1.60s | 99.9% | 1.05M |
| NovitaAI | - | $0.30 | $1.20 | $0.06 | 60 | 2.22s | 99.9% | 1M |
List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.
MiniMax M2.7: prices and speed by provider
Output prices run from $0.84 to $2.40 per million tokens across 5 providers, median $1.20.
| Provider | Share | Input $/M | Output $/M | Cache read $/M | Tokens/s | Latency | Uptime 3d | Context |
|---|---|---|---|---|---|---|---|---|
| GMICloud | 51.9% | $0.21 | $0.84 | $0.042 | 41 | 1.68s | 99.1% | 197K |
| MiniMax | 46.8% | $0.30 | $1.20 | $0.06 | 44 | 1.34s | 100.0% | 205K |
| Groq | 1.2% | $0.60 | $1.80 | – | 419 | 0.12s | 96.2% | 197K |
| NovitaAI | - | $0.27 | $1.08 | $0.054 | 38 | 1.45s | 99.9% | 205K |
| AtlasCloud | - | $0.30 | $1.20 | $0.06 | 38.5 | 1.91s | 98.5% | 197K |
List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.
All MiniMax models
| Model | Tokens, 7d | Providers | Output $/M | Context | Released |
|---|---|---|---|---|---|
| MiniMax M3 | 1.23T | 13 | $0.96 to $3.00 | 1.05M | 2026-05-31 |
| MiniMax M2.7 | 65B | 5 | $0.84 to $2.40 | 205K | 2026-03-18 |
| MiniMax M2.5 | 14B | 8 | $0.95 to $2.40 | 205K | 2026-02-12 |
| MiniMax M2.1 | 1B | 1 | $1.20 to $2.40 | 205K | 2025-12-23 |
| MiniMax M2 | 876M | 2 | $1.02 to $1.20 | 205K | 2025-10-23 |
| MiniMax-01 | 171M | 1 | $1.10 | 1M | 2025-01-15 |
Finance your GPUs
Serving MiniMax on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.
Prefer email? hello@amcompute.com