NovitaAI: Models and Prices

Model API for 200+ open models plus on-demand, spot and bare-metal GPU instances.

NovitaAI served 3.26T tokens of other labs' open-weight models over Oct 3 to Oct 9, 2026 through a large public LLM router, #5 among providers.

Open-weight models

ModelFamilyShare of modelTokens, 7dInput $/MOutput $/MTokens/sUptime 3dContext
DeepSeek V4.1 FlashDeepSeek3.8%1.42T$0.195$0.7868100.0%1.05M
GLM 5.3 FlashGLM6.5%724B$0.084$0.283095.6%1.05M
MiMo-V2.6-FlashMiMo3.0%324B$0.14$0.283096.7%1.05M
MiMo-V2.6-ProMiMo25.7%313B$0.435$0.873494.9%1.05M
Hy4 previewHunyuan2.1%168B$0.834$2.5029100.0%1M
Ling 3.0 FlashOther open-weight100.0%28B$0.021$0.063152100.0%262K
DeepSeek V4 Flash 0731DeepSeek0.4%23B$0.409$1.2346100.0%1.05M
GLM 5.3GLM0.4%14B$0.70$2.205198.4%1.05M
Qwen3.8 27BQwen2.4%14B$0.42$3.003099.2%1M
DeepSeek V4 Pro 0423DeepSeek1.8%11B$1.60$3.2042100.0%1.05M
Hy3Hunyuan--$0.14$0.584199.0%262K
GLM 5.2GLM--$0.65$2.0468100.0%1.05M
MiniMax M3MiniMax--$0.30$1.206099.9%1M
DeepSeek V4 Pro 0813DeepSeek--$0.99$2.9744100.0%1.05M
MiMo-V2.5MiMo--$0.168$0.3362391.8%1.05M
gpt-oss-120bgpt-oss--$0.05$0.2515899.6%131K
Gemma 4 31BGemma--$0.14$0.402460.7%262K
Gemma 4 26B A4B Gemma--$0.13$0.409299.8%262K
Step 3.7 FlashStep--$0.20$1.156599.9%262K
MiMo-V2.5-ProMiMo--$0.48$0.9612798.5%1.05M
Kimi K2.6Kimi--$0.80$3.403499.8%262K
Qwen3.8 2.4T A95BQwen--$2.00$6.0027100.0%1M
MiniMax M2.7MiniMax--$0.27$1.083899.9%205K
Gemma 3 27BGemma--$0.119$0.203595.5%98K
Ling 3.0 Flash VLOther open-weight--$0.021$0.062151100.0%262K
GLM 4.7GLM--$0.54$1.9823100.0%205K
GLM 5GLM--$1.00$3.2038.5100.0%203K
Kimi K2.7 CodeKimi--$0.912$3.8437100.0%262K
Llama 3.1 8B InstructLlama--$0.02$0.0510097.8%16K
Kimi K2.5Kimi--$0.57$2.854398.1%262K
Llama 3.3 70B InstructLlama--$0.135$0.403596.5%12K
Qwen3.5 397B A17BQwen--$0.60$3.605097.1%262K
GLM 5.1GLM--$1.38$4.4039.5100.0%205K
GLM 4.6GLM--$0.55$2.2024100.0%205K
Llama 4 MaverickLlama--$0.27$0.853997.6%1.05M
GLM 4.7 FlashGLM--$0.07$0.40677.5%200K
MiniMax M2.5MiniMax--$0.30$1.205095.8%205K
GLM 4.5 AirGLM--$0.13$0.8563100.0%131K
Qwen3.5-27BQwen--$0.30$2.4036100.0%262K
R1 0528DeepSeek--$0.70$2.5028100.0%164K
Llama 4 ScoutLlama--$0.18$0.5930.593.0%131K
Qwen3.5-122B-A10BQwen--$0.40$3.204499.8%262K
Qwen3 VL 30B A3B InstructQwen--$0.20$0.70797.2%131K
Kimi K2 0905Kimi--$0.60$2.5047100.0%262K
Nemotron 3 Nano 30B A3BNemotron--$0.05$0.2019798.8%262K
Kimi K2 ThinkingKimi--$0.60$2.5040100.0%262K
Kimi K2 0711Kimi--$0.57$2.3045.5100.0%131K
R1DeepSeek--$0.70$2.5022100.0%64K
WizardLM-2 8x22BOther open-weight--$0.62$0.6217100.0%66K
GLM 4.6VGLM--$0.30$0.902699.6%131K
GLM 4.5VGLM--$0.60$1.8019100.0%66K

List prices per million tokens as of October 9, 2026. Share: of all tokens served on that model; a dash where the volume is not reported.

Company

Website
novita.ai
Headquarters
San Francisco, CA
Status page
status.novita.ai

Sources: Novita AI.

Finance your GPUs

Serving tokens on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.

Prefer email? hello@amcompute.com