gpt-oss API Providers
Who serves gpt-oss, at what price and speed.
gpt-oss is OpenAI's open-weight model line (gpt-oss-120b and gpt-oss-20b). OpenAI does not serve it through its own API, so every token comes from third-party providers.
Groq served 20% of the 733B gpt-oss tokens that went through a large public LLM router over Oct 3 to Oct 9, 2026. 22 providers list the family.
Who serves gpt-oss tokens
All gpt-oss models combined, Oct 3 to Oct 9, 2026.
gpt-oss-120b: prices and speed by provider
Output prices run from $0.15 to $0.95 per million tokens across 21 providers, median $0.55.
| Provider | Share | Input $/M | Output $/M | Cache read $/M | Tokens/s | Latency | Uptime 3d | Context |
|---|---|---|---|---|---|---|---|---|
| DeepInfra | 22.0% | $0.037 | $0.17 | – | 73 | 0.27s | 98.5% | 131K |
| Groq | 20.1% | $0.15 | $0.60 | $0.075 | 342 | 0.19s | 100.0% | 131K |
| CoreWeave | 15.0% | $0.03 | $0.17 | $0.03 | 42 | 0.36s | 99.8% | 131K |
| AkashML | 11.9% | $0.03 | $0.187 | $0.03 | 91 | 0.62s | 99.9% | 131K |
| Cerebras | 8.4% | $0.35 | $0.75 | $0.35 | 452 | 0.80s | 100.0% | 131K |
| DekaLLM | 6.4% | $0.03 | $0.18 | $0.03 | 40 | 0.72s | 99.9% | 131K |
| Crusoe | 6.0% | $0.05 | $0.25 | $0.05 | 220 | 0.27s | 96.1% | 131K |
| Baseten | 3.2% | $0.10 | $0.50 | $0.10 | 184 | 0.32s | 100.0% | 128K |
| Mancer | 2.2% | $0.05 | $0.25 | – | 46 | 0.71s | 98.8% | 131K |
| Together | 1.4% | $0.15 | $0.60 | – | 189 | 0.27s | 91.7% | 131K |
| Amazon Bedrock | 1.4% | $0.15 | $0.60 | – | 145 | 0.42s | 100.0% | 131K |
| SambaNova | 1.0% | $0.14 | $0.95 | – | 352 | 0.89s | 99.8% | 131K |
| Nebius Token Factory | 0.8% | $0.15 | $0.60 | – | 273 | 0.31s | 98.1% | 131K |
| MARA | 0.2% | $0.15 | $0.75 | – | 210 | 0.93s | 97.4% | 131K |
| Venice | - | $0.03 | $0.15 | $0.03 | 13 | 2.11s | 100.0% | 128K |
| NovitaAI | - | $0.05 | $0.25 | – | 158 | 0.75s | 99.6% | 131K |
| Google Vertex | - | $0.09 | $0.36 | – | 196 | 0.27s | 13.3% | 131K |
| DigitalOcean | - | $0.06 | $0.42 | $0.012 | 58 | 0.91s | 100.0% | 128K |
| SiliconFlow | - | $0.15 | $0.60 | $0.075 | 14 | 1.56s | 77.1% | 131K |
| Phala | - | $0.15 | $0.60 | – | 128 | 1.06s | 99.4% | 131K |
| Parasail | - | $0.10 | $0.75 | $0.055 | 129 | 0.45s | 100.0% | 131K |
List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.
gpt-oss-20b: prices and speed by provider
Output prices run from $0.09 to $0.30 per million tokens across 10 providers, median $0.15.
| Provider | Share | Input $/M | Output $/M | Cache read $/M | Tokens/s | Latency | Uptime 3d | Context |
|---|---|---|---|---|---|---|---|---|
| CoreWeave | 29.3% | $0.03 | $0.13 | $0.03 | 128 | 0.15s | 100.0% | 131K |
| Groq | 19.6% | $0.075 | $0.30 | $0.037 | 362 | 0.34s | 83.1% | 131K |
| Darkbloom | 16.4% | $0.018 | $0.09 | $0.0090 | 23 | 3.96s | 100.0% | 131K |
| Parasail | 14.7% | $0.03 | $0.15 | $0.02 | 65 | 1.51s | 99.9% | 131K |
| Amazon Bedrock | 12.0% | $0.07 | $0.15 | – | - | 0.40s | 100.0% | 131K |
| DekaLLM | 4.0% | $0.029 | $0.14 | $0.029 | 20 | 1.42s | 99.5% | 131K |
| AkashML | 3.8% | $0.02 | $0.10 | – | 32 | 1.00s | 99.9% | 131K |
| DeepInfra | - | $0.03 | $0.14 | – | 122 | 0.26s | 99.7% | 131K |
| SiliconFlow | - | $0.04 | $0.18 | – | 48 | 1.17s | 86.9% | 131K |
| Google Vertex | - | $0.07 | $0.25 | – | 258 | 0.44s | 99.0% | 131K |
List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.
All gpt-oss models
| Model | Tokens, 7d | Providers | Output $/M | Context | Released |
|---|---|---|---|---|---|
| gpt-oss-120b | 484B | 21 | $0.15 to $0.95 | 131K | 2025-08-05 |
| gpt-oss-20b | 247B | 10 | $0.09 to $0.30 | 131K | 2025-08-05 |
| gpt-oss-safeguard-20b | 3B | 1 | $0.30 | 131K | 2025-10-29 |
Finance your GPUs
Serving gpt-oss on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.
Prefer email? hello@amcompute.com