Llama API Providers

Who serves Llama, at what price and speed.

Llama is Meta's open-weight family. Third-party providers still serve Llama 3.x and Llama 4.

Groq served 66% of the 117B Llama tokens that went through a large public LLM router over Oct 3 to Oct 9, 2026. 12 providers list the family.

Who serves Llama tokens

ProviderTokens, 7dShare
Groq33B65.6%
DigitalOcean10B20.1%
AkashML6B12.5%
SambaNova722M1.4%
Cloudflare189M0.4%
Prime Intellect937790.0%
CoreWeave00.0%
Crusoe00.0%
Nebius Token Factory00.0%

All Llama models combined, Oct 3 to Oct 9, 2026.

Llama 3.1 8B Instruct: prices and speed by provider

Output prices run from $0.04 to $0.287 per million tokens across 5 providers, median $0.08.

ProviderShareInput $/MOutput $/MCache read $/MTokens/sLatencyUptime 3dContext
Groq99.2%$0.05$0.08$0.0254410.14s100.0%131K
Cloudflare0.8%$0.152$0.287–210.45s99.8%32K
DeepInfra-$0.02$0.04–390.80s99.7%131K
NovitaAI-$0.02$0.05–1000.52s97.8%16K
CoreWeave-$0.22$0.22$0.221430.27s99.9%131K

List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.

Llama 3.3 70B Instruct: prices and speed by provider

Output prices run from $0.32 to $2.25 per million tokens across 10 providers, median $0.72.

ProviderShareInput $/MOutput $/MCache read $/MTokens/sLatencyUptime 3dContext
Groq54.9%$0.59$0.79$0.2952350.34s100.0%131K
AkashML40.5%$0.20$0.52$0.10340.46s99.6%131K
DeepInfra-$0.10$0.32–100.52s95.3%131K
NovitaAI-$0.135$0.40–351.21s96.5%12K
Parasail-$0.22$0.50$0.11500.62s100.0%131K
CoreWeave-$0.71$0.71$0.71630.88s99.7%128K
Google Vertex-$0.72$0.72–1021.59s-128K
sambanova-turbo-$0.45$0.90–1290.54s99.8%131K
Together-$1.04$1.04–57.50.41s97.7%131K
Cloudflare-$0.293$2.25–510.37s99.8%24K

List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.

All Llama models

ModelTokens, 7dProvidersOutput $/MContextReleased
Llama 3.1 8B Instruct47B5$0.04 to $0.287131K2024-07-23
Llama 3.3 70B Instruct40B10$0.32 to $2.25131K2024-12-06
Llama 4 Maverick16B4$0.652 to $1.151.05M2025-04-05
Llama 4 Scout11B3$0.30 to $0.701.31M2025-04-05
Llama 3.1 70B Instruct1B2$0.40 to $0.72131K2024-07-23
Llama Guard 4 12B1B1$0.18164K2025-04-30
Llama 3.2 3B Instruct866M2$0.33 to $0.335131K2024-09-25
Llama 3.2 1B Instruct258M1$0.20160K2024-09-25

Finance your GPUs

Serving Llama on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.

Prefer email? hello@amcompute.com