Together: Models and Prices
Serverless token API for open-weight models, dedicated endpoints, fine-tuning and GPU clusters.
Together served 11.5T tokens of other labs' open-weight models over Oct 3 to Oct 9, 2026 through a large public LLM router, #1 among providers.
Open-weight models
| Model | Family | Share of model | Tokens, 7d | Input $/M | Output $/M | Tokens/s | Uptime 3d | Context |
|---|---|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | DeepSeek | 22.9% | 8.59T | $0.30 | $1.20 | 217 | 99.8% | 1.05M |
| GLM 5.3 Flash | GLM | 20.8% | 2.32T | $0.15 | $0.50 | 97 | 99.8% | 1.05M |
| Kimi K3 | Kimi | 9.7% | 186B | $2.70 | $13.50 | 49 | 92.5% | 1.05M |
| GLM 5.3 | GLM | 3.5% | 116B | $1.40 | $4.40 | 134 | 96.0% | 1.05M |
| GLM 5.2 | GLM | 7.5% | 114B | $1.40 | $4.40 | 180 | 99.8% | 1.05M |
| DeepSeek V4 Flash 0731 | DeepSeek | 1.3% | 64B | $0.14 | $0.28 | 59 | 98.9% | 1.05M |
| DeepSeek V4 Pro 0813 | DeepSeek | 1.9% | 20B | $1.32 | $3.96 | 109 | 98.6% | 1.05M |
| Muse Glimmer 30B | Other open-weight | 75.0% | 13B | $0.35 | $1.50 | 99 | 99.4% | 131K |
| MiniMax M3 | MiniMax | 0.8% | 9B | $0.30 | $1.20 | 43 | 99.8% | 524K |
| Inkling | Other open-weight | 3.3% | 9B | $1.00 | $4.05 | 118 | 88.9% | 524K |
| Qwen3.8 2.4T A95B | Qwen | 19.1% | 8B | $2.00 | $6.00 | 173.5 | 99.6% | 1.01M |
| gpt-oss-120b | gpt-oss | 1.4% | 6B | $0.15 | $0.60 | 189 | 91.7% | 131K |
| Llama 3.3 70B Instruct | Llama | - | - | $1.04 | $1.04 | 57.5 | 97.7% | 131K |
| Qwen3.5-9B | Qwen | - | - | $0.17 | $0.25 | 80 | 99.1% | 262K |
List prices per million tokens as of October 9, 2026. Share: of all tokens served on that model; a dash where the volume is not reported.
Company
- Website
- together.ai
- Headquarters
- San Francisco, CA
- Compute
- Historically rented much of its capacity from GPU clouds; reported to be moving into its own facilities, with a Maryland site live since July 2025.
- Latest financing
- $800M Series C led by Aramco Ventures at an $8.3B valuation (July 2026).
- Status page
- status.together.ai
Finance your GPUs
Serving tokens on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.
Prefer email? hello@amcompute.com