Modal: Models and Prices
Serverless GPU compute and sandboxes, with per-second billing.
Modal served 460B tokens of other labs' open-weight models over Oct 3 to Oct 9, 2026 through a large public LLM router, #24 among providers.
Open-weight models
| Model | Family | Share of model | Tokens, 7d | Input $/M | Output $/M | Tokens/s | Uptime 3d | Context |
|---|---|---|---|---|---|---|---|---|
| GLM 5.3 Flash | GLM | 1.8% | 196B | $0.15 | $0.50 | 91 | 100.0% | 1.05M |
| Kimi K3 | Kimi | 5.8% | 111B | $3.00 | $15.00 | 56 | 100.0% | 1.05M |
| GLM 5.3 | GLM | 2.3% | 76B | $1.40 | $4.40 | 110 | 99.9% | 1.05M |
| DeepSeek V4.1 Flash | DeepSeek | 0.2% | 73B | $0.30 | $1.20 | 167 | 100.0% | 1.05M |
| Qwen3.8 2.4T A95B | Qwen | 11.6% | 5B | $2.00 | $6.00 | 153.5 | 99.2% | 1M |
List prices per million tokens as of October 9, 2026. Share: of all tokens served on that model; a dash where the volume is not reported.
Company
- Website
- modal.com
- Headquarters
- New York, NY
- Latest financing
- $355M Series C at a $4.65B valuation led by General Catalyst and Redpoint (May 2026).
- Status page
- status.modal.com
Sources: DCD.
Finance your GPUs
Serving tokens on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.
Prefer email? hello@amcompute.com