Gemma API Providers
Who serves Gemma, at what price and speed.
Gemma is Google's open-weight model family. Google Vertex serves Gemma 4, but third-party providers serve nearly all Gemma tokens.
DeepInfra served 22% of the 882B Gemma tokens that went through a large public LLM router over Oct 3 to Oct 9, 2026. 19 providers list the family.
Who serves Gemma tokens
All Gemma models combined, Oct 3 to Oct 9, 2026.
Gemma 4 31B: prices and speed by provider
Output prices run from $0.34 to $1.15 per million tokens across 12 providers, median $0.40.
| Provider | Share | Input $/M | Output $/M | Cache read $/M | Tokens/s | Latency | Uptime 3d | Context |
|---|---|---|---|---|---|---|---|---|
| DeepInfra | 37.3% | $0.09 | $0.34 | $0.05 | 35 | 0.69s | 100.0% | 262K |
| CoreWeave | 22.4% | $0.10 | $0.34 | $0.10 | 68 | 0.38s | 100.0% | 262K |
| ModelRun [by Modular] | 14.7% | $0.75 | $1.00 | $0.20 | 53 | 0.07s | 99.5% | 262K |
| Friendli | 11.3% | $0.14 | $0.40 | – | 105 | 0.51s | 98.4% | 262K |
| Parasail | 6.3% | $0.15 | $0.40 | $0.06 | 24 | 1.99s | 99.9% | 262K |
| Venice | 5.1% | $0.12 | $0.36 | $0.09 | 40 | 0.48s | 97.5% | 256K |
| Chutes | 1.7% | $0.12 | $0.37 | $0.012 | 10 | 4.24s | 78.0% | 131K |
| Crusoe | 0.7% | $0.14 | $0.40 | $0.14 | 26 | 0.59s | 98.7% | 262K |
| SambaNova | 0.4% | $0.38 | $1.15 | – | 74 | 1.92s | 87.8% | 131K |
| io.net | 0.2% | $0.361 | $1.09 | $0.18 | 38 | 0.50s | 95.5% | 262K |
| NovitaAI | - | $0.14 | $0.40 | – | 24 | 0.83s | 60.7% | 262K |
| SiliconFlow | - | $0.75 | $1.00 | $0.25 | 24 | 1.50s | 95.3% | 262K |
List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.
Gemma 4 26B A4B : prices and speed by provider
Output prices run from $0.21 to $0.60 per million tokens across 13 providers, median $0.33.
| Provider | Share | Input $/M | Output $/M | Cache read $/M | Tokens/s | Latency | Uptime 3d | Context |
|---|---|---|---|---|---|---|---|---|
| NextBit | 27.9% | $0.068 | $0.225 | $0.037 | 84 | 0.29s | 100.0% | 262K |
| Darkbloom | 22.3% | $0.042 | $0.22 | $0.021 | 35 | 1.29s | 100.0% | 131K |
| Cloudflare | 15.3% | $0.10 | $0.30 | $0.05 | 45 | 0.50s | 99.7% | 256K |
| Makora | 12.9% | $0.08 | $0.32 | $0.032 | 212 | 0.30s | 99.9% | 256K |
| DekaLLM | 10.7% | $0.06 | $0.33 | $0.04 | 62 | 0.77s | 100.0% | 262K |
| io.net | 8.1% | $0.048 | $0.21 | $0.02 | 56 | 0.56s | 100.0% | 262K |
| Venice | 2.2% | $0.13 | $0.40 | $0.05 | 54 | 0.90s | 99.6% | 256K |
| CoreWeave | 0.6% | $0.10 | $0.30 | $0.05 | 63 | 0.30s | 99.8% | 262K |
| DeepInfra | - | $0.07 | $0.34 | – | 69 | 0.53s | 99.2% | 262K |
| Parasail | - | $0.13 | $0.40 | $0.05 | 59 | 0.43s | 99.6% | 262K |
| NovitaAI | - | $0.13 | $0.40 | – | 92 | 0.72s | 99.8% | 262K |
| SiliconFlow | - | $0.14 | $0.40 | $0.05 | 41 | 1.25s | 94.0% | 262K |
| Google Vertex | - | $0.15 | $0.60 | – | 32 | 1.35s | 96.0% | 262K |
List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.
All Gemma models
| Model | Tokens, 7d | Providers | Output $/M | Context | Released |
|---|---|---|---|---|---|
| Gemma 4 31B | 445B | 12 | $0.34 to $1.15 | 262K | 2026-04-02 |
| Gemma 4 26B A4B | 363B | 13 | $0.21 to $0.60 | 262K | 2026-04-03 |
| Gemma 3 27B | 60B | 5 | $0.16 to $0.45 | 131K | 2025-03-12 |
| Gemma 3 12B | 13B | 2 | $0.15 | 131K | 2025-03-13 |
| Gemma 3 4B | 664M | 1 | $0.10 | 131K | 2025-03-13 |
| Gemma 2 27B | 28M | 1 | $0.65 | 8K | 2024-07-13 |
Finance your GPUs
Serving Gemma on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.
Prefer email? hello@amcompute.com