Together: Models and Prices

Serverless token API for open-weight models, dedicated endpoints, fine-tuning and GPU clusters.

Together served 11.5T tokens of other labs' open-weight models over Oct 3 to Oct 9, 2026 through a large public LLM router, #1 among providers.

Open-weight models

ModelFamilyShare of modelTokens, 7dInput $/MOutput $/MTokens/sUptime 3dContext
DeepSeek V4.1 FlashDeepSeek22.9%8.59T$0.30$1.2021799.8%1.05M
GLM 5.3 FlashGLM20.8%2.32T$0.15$0.509799.8%1.05M
Kimi K3Kimi9.7%186B$2.70$13.504992.5%1.05M
GLM 5.3GLM3.5%116B$1.40$4.4013496.0%1.05M
GLM 5.2GLM7.5%114B$1.40$4.4018099.8%1.05M
DeepSeek V4 Flash 0731DeepSeek1.3%64B$0.14$0.285998.9%1.05M
DeepSeek V4 Pro 0813DeepSeek1.9%20B$1.32$3.9610998.6%1.05M
Muse Glimmer 30BOther open-weight75.0%13B$0.35$1.509999.4%131K
MiniMax M3MiniMax0.8%9B$0.30$1.204399.8%524K
InklingOther open-weight3.3%9B$1.00$4.0511888.9%524K
Qwen3.8 2.4T A95BQwen19.1%8B$2.00$6.00173.599.6%1.01M
gpt-oss-120bgpt-oss1.4%6B$0.15$0.6018991.7%131K
Llama 3.3 70B InstructLlama--$1.04$1.0457.597.7%131K
Qwen3.5-9BQwen--$0.17$0.258099.1%262K

List prices per million tokens as of October 9, 2026. Share: of all tokens served on that model; a dash where the volume is not reported.

Company

Headquarters
San Francisco, CA
Compute
Historically rented much of its capacity from GPU clouds; reported to be moving into its own facilities, with a Maryland site live since July 2025.
Latest financing
$800M Series C led by Aramco Ventures at an $8.3B valuation (July 2026).

Sources: Sacra, DCD.

Finance your GPUs

Serving tokens on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.

Prefer email? hello@amcompute.com