Baseten: Models and Prices

Inference platform for dedicated and shared model deployments.

Baseten served 598B tokens of other labs' open-weight models over Oct 3 to Oct 9, 2026 through a large public LLM router, #19 among providers.

Open-weight models

ModelFamilyShare of modelTokens, 7dInput $/MOutput $/MTokens/sUptime 3dContext
GLM 5.3 FlashGLM1.8%202B$0.15$0.5012299.1%1.05M
DeepSeek V4.1 FlashDeepSeek0.4%148B$0.30$1.20207.5100.0%1.05M
GLM 5.3GLM3.7%123B$1.40$4.406699.8%1.05M
DeepSeek V4 Flash 0731DeepSeek0.9%45B$0.13$0.26177100.0%1.05M
GLM 5.2GLM2.3%35B$1.40$4.4084100.0%1.05M
Kimi K3Kimi1.0%20B$3.00$15.0027100.0%1.05M
gpt-oss-120bgpt-oss3.2%13B$0.10$0.50184100.0%128K
Nemotron 3 UltraNemotron0.2%12B$0.60$2.40103100.0%203K

List prices per million tokens as of October 9, 2026. Share: of all tokens served on that model; a dash where the volume is not reported.

Company

Website
baseten.co
Headquarters
San Francisco, CA
Latest financing
$300M Series E at a $5B valuation co-led by IVP and CapitalG (early 2026); a larger round was reported in June 2026.

Sources: Lets Data Science.

Finance your GPUs

Serving tokens on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.

Prefer email? hello@amcompute.com