Gemma API Providers

Who serves Gemma, at what price and speed.

Gemma is Google's open-weight model family. Google Vertex serves Gemma 4, but third-party providers serve nearly all Gemma tokens.

DeepInfra served 22% of the 882B Gemma tokens that went through a large public LLM router over Oct 3 to Oct 9, 2026. 19 providers list the family.

Who serves Gemma tokens

ProviderTokens, 7dShare
DeepInfra158B21.5%
CoreWeave97B13.1%
NextBit77B10.5%
Parasail63B8.5%
ModelRun [by Modular]62B8.5%
Darkbloom59B8.0%
Friendli48B6.5%
Cloudflare40B5.5%
Makora34B4.7%
DekaLLM29B3.9%
Venice27B3.7%
io.net22B3.0%
Chutes7B0.9%
Nebius Token Factory7B0.9%
Crusoe3B0.4%
2 others2B0.2%

All Gemma models combined, Oct 3 to Oct 9, 2026.

Gemma 4 31B: prices and speed by provider

Output prices run from $0.34 to $1.15 per million tokens across 12 providers, median $0.40.

ProviderShareInput $/MOutput $/MCache read $/MTokens/sLatencyUptime 3dContext
DeepInfra37.3%$0.09$0.34$0.05350.69s100.0%262K
CoreWeave22.4%$0.10$0.34$0.10680.38s100.0%262K
ModelRun [by Modular]14.7%$0.75$1.00$0.20530.07s99.5%262K
Friendli11.3%$0.14$0.40–1050.51s98.4%262K
Parasail6.3%$0.15$0.40$0.06241.99s99.9%262K
Venice5.1%$0.12$0.36$0.09400.48s97.5%256K
Chutes1.7%$0.12$0.37$0.012104.24s78.0%131K
Crusoe0.7%$0.14$0.40$0.14260.59s98.7%262K
SambaNova0.4%$0.38$1.15–741.92s87.8%131K
io.net0.2%$0.361$1.09$0.18380.50s95.5%262K
NovitaAI-$0.14$0.40–240.83s60.7%262K
SiliconFlow-$0.75$1.00$0.25241.50s95.3%262K

List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.

Gemma 4 26B A4B : prices and speed by provider

Output prices run from $0.21 to $0.60 per million tokens across 13 providers, median $0.33.

ProviderShareInput $/MOutput $/MCache read $/MTokens/sLatencyUptime 3dContext
NextBit27.9%$0.068$0.225$0.037840.29s100.0%262K
Darkbloom22.3%$0.042$0.22$0.021351.29s100.0%131K
Cloudflare15.3%$0.10$0.30$0.05450.50s99.7%256K
Makora12.9%$0.08$0.32$0.0322120.30s99.9%256K
DekaLLM10.7%$0.06$0.33$0.04620.77s100.0%262K
io.net8.1%$0.048$0.21$0.02560.56s100.0%262K
Venice2.2%$0.13$0.40$0.05540.90s99.6%256K
CoreWeave0.6%$0.10$0.30$0.05630.30s99.8%262K
DeepInfra-$0.07$0.34–690.53s99.2%262K
Parasail-$0.13$0.40$0.05590.43s99.6%262K
NovitaAI-$0.13$0.40–920.72s99.8%262K
SiliconFlow-$0.14$0.40$0.05411.25s94.0%262K
Google Vertex-$0.15$0.60–321.35s96.0%262K

List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.

All Gemma models

ModelTokens, 7dProvidersOutput $/MContextReleased
Gemma 4 31B445B12$0.34 to $1.15262K2026-04-02
Gemma 4 26B A4B 363B13$0.21 to $0.60262K2026-04-03
Gemma 3 27B60B5$0.16 to $0.45131K2025-03-12
Gemma 3 12B13B2$0.15131K2025-03-13
Gemma 3 4B664M1$0.10131K2025-03-13
Gemma 2 27B28M1$0.658K2024-07-13

Finance your GPUs

Serving Gemma on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.

Prefer email? hello@amcompute.com