Groq: Models and Prices
Inference cloud on its own LPU chips and NVIDIA systems.
Groq served 161B tokens of other labs' open-weight models over Oct 3 to Oct 9, 2026 through a large public LLM router, #33 among providers.
Open-weight models
| Model | Family | Share of model | Tokens, 7d | Input $/M | Output $/M | Tokens/s | Uptime 3d | Context |
|---|---|---|---|---|---|---|---|---|
| gpt-oss-120b | gpt-oss | 20.1% | 84B | $0.15 | $0.60 | 342 | 100.0% | 131K |
| gpt-oss-20b | gpt-oss | 19.6% | 40B | $0.075 | $0.30 | 362 | 83.1% | 131K |
| Llama 3.1 8B Instruct | Llama | 99.2% | 25B | $0.05 | $0.08 | 441 | 100.0% | 131K |
| Llama 3.3 70B Instruct | Llama | 54.9% | 9B | $0.59 | $0.79 | 235 | 100.0% | 131K |
| gpt-oss-safeguard-20b | gpt-oss | 100.0% | 3B | $0.075 | $0.30 | 710 | 99.5% | 131K |
| MiniMax M2.7 | MiniMax | 1.2% | 532M | $0.60 | $1.80 | 419 | 96.2% | 197K |
List prices per million tokens as of October 9, 2026. Share: of all tokens served on that model; a dash where the volume is not reported.
Company
- Website
- groq.com
- Headquarters
- United States
- Compute
- Says it runs 13 data centers with 54 MW and plans to pass 200 MW in 2027. NVIDIA licensed its LPU technology in December 2025.
- Latest financing
- $350M at a $3.5B valuation (August 2026), after a $650M raise in June 2026.
- Status page
- status.groq.com
Sources: Channel Insider.
Finance your GPUs
Serving tokens on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.
Prefer email? hello@amcompute.com