NovitaAI: Models and Prices
Model API for 200+ open models plus on-demand, spot and bare-metal GPU instances.
NovitaAI served 3.26T tokens of other labs' open-weight models over Oct 3 to Oct 9, 2026 through a large public LLM router, #5 among providers.
Open-weight models
| Model | Family | Share of model | Tokens, 7d | Input $/M | Output $/M | Tokens/s | Uptime 3d | Context |
|---|---|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | DeepSeek | 3.8% | 1.42T | $0.195 | $0.78 | 68 | 100.0% | 1.05M |
| GLM 5.3 Flash | GLM | 6.5% | 724B | $0.084 | $0.28 | 30 | 95.6% | 1.05M |
| MiMo-V2.6-Flash | MiMo | 3.0% | 324B | $0.14 | $0.28 | 30 | 96.7% | 1.05M |
| MiMo-V2.6-Pro | MiMo | 25.7% | 313B | $0.435 | $0.87 | 34 | 94.9% | 1.05M |
| Hy4 preview | Hunyuan | 2.1% | 168B | $0.834 | $2.50 | 29 | 100.0% | 1M |
| Ling 3.0 Flash | Other open-weight | 100.0% | 28B | $0.021 | $0.063 | 152 | 100.0% | 262K |
| DeepSeek V4 Flash 0731 | DeepSeek | 0.4% | 23B | $0.409 | $1.23 | 46 | 100.0% | 1.05M |
| GLM 5.3 | GLM | 0.4% | 14B | $0.70 | $2.20 | 51 | 98.4% | 1.05M |
| Qwen3.8 27B | Qwen | 2.4% | 14B | $0.42 | $3.00 | 30 | 99.2% | 1M |
| DeepSeek V4 Pro 0423 | DeepSeek | 1.8% | 11B | $1.60 | $3.20 | 42 | 100.0% | 1.05M |
| Hy3 | Hunyuan | - | - | $0.14 | $0.58 | 41 | 99.0% | 262K |
| GLM 5.2 | GLM | - | - | $0.65 | $2.04 | 68 | 100.0% | 1.05M |
| MiniMax M3 | MiniMax | - | - | $0.30 | $1.20 | 60 | 99.9% | 1M |
| DeepSeek V4 Pro 0813 | DeepSeek | - | - | $0.99 | $2.97 | 44 | 100.0% | 1.05M |
| MiMo-V2.5 | MiMo | - | - | $0.168 | $0.336 | 23 | 91.8% | 1.05M |
| gpt-oss-120b | gpt-oss | - | - | $0.05 | $0.25 | 158 | 99.6% | 131K |
| Gemma 4 31B | Gemma | - | - | $0.14 | $0.40 | 24 | 60.7% | 262K |
| Gemma 4 26B A4B | Gemma | - | - | $0.13 | $0.40 | 92 | 99.8% | 262K |
| Step 3.7 Flash | Step | - | - | $0.20 | $1.15 | 65 | 99.9% | 262K |
| MiMo-V2.5-Pro | MiMo | - | - | $0.48 | $0.961 | 27 | 98.5% | 1.05M |
| Kimi K2.6 | Kimi | - | - | $0.80 | $3.40 | 34 | 99.8% | 262K |
| Qwen3.8 2.4T A95B | Qwen | - | - | $2.00 | $6.00 | 27 | 100.0% | 1M |
| MiniMax M2.7 | MiniMax | - | - | $0.27 | $1.08 | 38 | 99.9% | 205K |
| Gemma 3 27B | Gemma | - | - | $0.119 | $0.20 | 35 | 95.5% | 98K |
| Ling 3.0 Flash VL | Other open-weight | - | - | $0.021 | $0.062 | 151 | 100.0% | 262K |
| GLM 4.7 | GLM | - | - | $0.54 | $1.98 | 23 | 100.0% | 205K |
| GLM 5 | GLM | - | - | $1.00 | $3.20 | 38.5 | 100.0% | 203K |
| Kimi K2.7 Code | Kimi | - | - | $0.912 | $3.84 | 37 | 100.0% | 262K |
| Llama 3.1 8B Instruct | Llama | - | - | $0.02 | $0.05 | 100 | 97.8% | 16K |
| Kimi K2.5 | Kimi | - | - | $0.57 | $2.85 | 43 | 98.1% | 262K |
| Llama 3.3 70B Instruct | Llama | - | - | $0.135 | $0.40 | 35 | 96.5% | 12K |
| Qwen3.5 397B A17B | Qwen | - | - | $0.60 | $3.60 | 50 | 97.1% | 262K |
| GLM 5.1 | GLM | - | - | $1.38 | $4.40 | 39.5 | 100.0% | 205K |
| GLM 4.6 | GLM | - | - | $0.55 | $2.20 | 24 | 100.0% | 205K |
| Llama 4 Maverick | Llama | - | - | $0.27 | $0.85 | 39 | 97.6% | 1.05M |
| GLM 4.7 Flash | GLM | - | - | $0.07 | $0.40 | 67 | 7.5% | 200K |
| MiniMax M2.5 | MiniMax | - | - | $0.30 | $1.20 | 50 | 95.8% | 205K |
| GLM 4.5 Air | GLM | - | - | $0.13 | $0.85 | 63 | 100.0% | 131K |
| Qwen3.5-27B | Qwen | - | - | $0.30 | $2.40 | 36 | 100.0% | 262K |
| R1 0528 | DeepSeek | - | - | $0.70 | $2.50 | 28 | 100.0% | 164K |
| Llama 4 Scout | Llama | - | - | $0.18 | $0.59 | 30.5 | 93.0% | 131K |
| Qwen3.5-122B-A10B | Qwen | - | - | $0.40 | $3.20 | 44 | 99.8% | 262K |
| Qwen3 VL 30B A3B Instruct | Qwen | - | - | $0.20 | $0.70 | 7 | 97.2% | 131K |
| Kimi K2 0905 | Kimi | - | - | $0.60 | $2.50 | 47 | 100.0% | 262K |
| Nemotron 3 Nano 30B A3B | Nemotron | - | - | $0.05 | $0.20 | 197 | 98.8% | 262K |
| Kimi K2 Thinking | Kimi | - | - | $0.60 | $2.50 | 40 | 100.0% | 262K |
| Kimi K2 0711 | Kimi | - | - | $0.57 | $2.30 | 45.5 | 100.0% | 131K |
| R1 | DeepSeek | - | - | $0.70 | $2.50 | 22 | 100.0% | 64K |
| WizardLM-2 8x22B | Other open-weight | - | - | $0.62 | $0.62 | 17 | 100.0% | 66K |
| GLM 4.6V | GLM | - | - | $0.30 | $0.90 | 26 | 99.6% | 131K |
| GLM 4.5V | GLM | - | - | $0.60 | $1.80 | 19 | 100.0% | 66K |
List prices per million tokens as of October 9, 2026. Share: of all tokens served on that model; a dash where the volume is not reported.
Company
- Website
- novita.ai
- Headquarters
- San Francisco, CA
- Status page
- status.novita.ai
Sources: Novita AI.
Finance your GPUs
Serving tokens on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.
Prefer email? hello@amcompute.com