Llama API Providers
Who serves Llama, at what price and speed.
Llama is Meta's open-weight family. Third-party providers still serve Llama 3.x and Llama 4.
Groq served 66% of the 117B Llama tokens that went through a large public LLM router over Oct 3 to Oct 9, 2026. 12 providers list the family.
Who serves Llama tokens
| Provider | Tokens, 7d | Share |
|---|---|---|
| Groq | 33B | 65.6% |
| DigitalOcean | 10B | 20.1% |
| AkashML | 6B | 12.5% |
| SambaNova | 722M | 1.4% |
| Cloudflare | 189M | 0.4% |
| Prime Intellect | 93779 | 0.0% |
| CoreWeave | 0 | 0.0% |
| Crusoe | 0 | 0.0% |
| Nebius Token Factory | 0 | 0.0% |
All Llama models combined, Oct 3 to Oct 9, 2026.
Llama 3.1 8B Instruct: prices and speed by provider
Output prices run from $0.04 to $0.287 per million tokens across 5 providers, median $0.08.
| Provider | Share | Input $/M | Output $/M | Cache read $/M | Tokens/s | Latency | Uptime 3d | Context |
|---|---|---|---|---|---|---|---|---|
| Groq | 99.2% | $0.05 | $0.08 | $0.025 | 441 | 0.14s | 100.0% | 131K |
| Cloudflare | 0.8% | $0.152 | $0.287 | – | 21 | 0.45s | 99.8% | 32K |
| DeepInfra | - | $0.02 | $0.04 | – | 39 | 0.80s | 99.7% | 131K |
| NovitaAI | - | $0.02 | $0.05 | – | 100 | 0.52s | 97.8% | 16K |
| CoreWeave | - | $0.22 | $0.22 | $0.22 | 143 | 0.27s | 99.9% | 131K |
List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.
Llama 3.3 70B Instruct: prices and speed by provider
Output prices run from $0.32 to $2.25 per million tokens across 10 providers, median $0.72.
| Provider | Share | Input $/M | Output $/M | Cache read $/M | Tokens/s | Latency | Uptime 3d | Context |
|---|---|---|---|---|---|---|---|---|
| Groq | 54.9% | $0.59 | $0.79 | $0.295 | 235 | 0.34s | 100.0% | 131K |
| AkashML | 40.5% | $0.20 | $0.52 | $0.10 | 34 | 0.46s | 99.6% | 131K |
| DeepInfra | - | $0.10 | $0.32 | – | 10 | 0.52s | 95.3% | 131K |
| NovitaAI | - | $0.135 | $0.40 | – | 35 | 1.21s | 96.5% | 12K |
| Parasail | - | $0.22 | $0.50 | $0.11 | 50 | 0.62s | 100.0% | 131K |
| CoreWeave | - | $0.71 | $0.71 | $0.71 | 63 | 0.88s | 99.7% | 128K |
| Google Vertex | - | $0.72 | $0.72 | – | 102 | 1.59s | - | 128K |
| sambanova-turbo | - | $0.45 | $0.90 | – | 129 | 0.54s | 99.8% | 131K |
| Together | - | $1.04 | $1.04 | – | 57.5 | 0.41s | 97.7% | 131K |
| Cloudflare | - | $0.293 | $2.25 | – | 51 | 0.37s | 99.8% | 24K |
List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.
All Llama models
| Model | Tokens, 7d | Providers | Output $/M | Context | Released |
|---|---|---|---|---|---|
| Llama 3.1 8B Instruct | 47B | 5 | $0.04 to $0.287 | 131K | 2024-07-23 |
| Llama 3.3 70B Instruct | 40B | 10 | $0.32 to $2.25 | 131K | 2024-12-06 |
| Llama 4 Maverick | 16B | 4 | $0.652 to $1.15 | 1.05M | 2025-04-05 |
| Llama 4 Scout | 11B | 3 | $0.30 to $0.70 | 1.31M | 2025-04-05 |
| Llama 3.1 70B Instruct | 1B | 2 | $0.40 to $0.72 | 131K | 2024-07-23 |
| Llama Guard 4 12B | 1B | 1 | $0.18 | 164K | 2025-04-30 |
| Llama 3.2 3B Instruct | 866M | 2 | $0.33 to $0.335 | 131K | 2024-09-25 |
| Llama 3.2 1B Instruct | 258M | 1 | $0.201 | 60K | 2024-09-25 |
Finance your GPUs
Serving Llama on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.
Prefer email? hello@amcompute.com