Modal: Models and Prices

Serverless GPU compute and sandboxes, with per-second billing.

Modal served 460B tokens of other labs' open-weight models over Oct 3 to Oct 9, 2026 through a large public LLM router, #24 among providers.

Open-weight models

ModelFamilyShare of modelTokens, 7dInput $/MOutput $/MTokens/sUptime 3dContext
GLM 5.3 FlashGLM1.8%196B$0.15$0.5091100.0%1.05M
Kimi K3Kimi5.8%111B$3.00$15.0056100.0%1.05M
GLM 5.3GLM2.3%76B$1.40$4.4011099.9%1.05M
DeepSeek V4.1 FlashDeepSeek0.2%73B$0.30$1.20167100.0%1.05M
Qwen3.8 2.4T A95BQwen11.6%5B$2.00$6.00153.599.2%1M

List prices per million tokens as of October 9, 2026. Share: of all tokens served on that model; a dash where the volume is not reported.

Company

Website
modal.com
Headquarters
New York, NY
Latest financing
$355M Series C at a $4.65B valuation led by General Catalyst and Redpoint (May 2026).
Status page
status.modal.com

Sources: DCD.

Finance your GPUs

Serving tokens on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.

Prefer email? hello@amcompute.com