Venice: Models and Prices

Private, uncensored AI app and API on self-hosted open-weight models.

Venice served 465B tokens of other labs' open-weight models over Oct 3 to Oct 9, 2026 through a large public LLM router, #23 among providers.

Open-weight models

ModelFamilyShare of modelTokens, 7dInput $/MOutput $/MTokens/sUptime 3dContext
DeepSeek V4.1 FlashDeepSeek0.6%239B$0.30$1.20170100.0%1M
GLM 5.3 FlashGLM0.5%55B$0.15$0.508999.0%1.05M
DeepSeek V4 Flash 0423DeepSeek1.7%53B$0.097$0.1932596.3%1M
Gemma 4 31BGemma5.1%22B$0.12$0.364097.5%256K
DeepSeek V4 Flash 0731DeepSeek0.4%19B$0.13$0.264099.8%1M
Mistral Small 3.2 24BMistral100.0%19B$0.094$0.253099.8%256K
MiMo-V2.5MiMo2.4%18B$0.40$2.001799.4%1M
Qwen3.5-9BQwen90.4%14B$0.10$0.155299.8%256K
MiMo-V2.6-FlashMiMo0.1%11B$0.175$0.351698.5%1M
Gemma 4 26B A4B Gemma2.2%6B$0.13$0.405499.6%256K
GLM 4.6GLM100.0%3B$0.43$1.756999.8%198K
GLM 5.3GLM<0.1%2B$1.40$4.407396.9%1M
Nemotron 3 UltraNemotron<0.1%2B$0.625$3.1310496.7%256K
Qwen3.8 27BQwen0.3%2B$0.45$3.2010399.7%262K
GLM 5.2GLM<0.1%2B$1.40$4.408599.7%1M
DeepSeek V3.2DeepSeek0.5%1B$0.268$0.391999.7%160K
MiniMax M3MiniMax--$0.30$1.208596.5%524K
DeepSeek V4 Pro 0813DeepSeek--$1.65$4.9576.593.8%1M
DeepSeek V4 Pro 0423DeepSeek--$1.65$3.305296.5%1M
gpt-oss-120bgpt-oss--$0.03$0.1513100.0%128K
Kimi K2.6Kimi--$0.75$3.501290.4%256K
Qwen3 235B A22B Instruct 2507Qwen--$0.15$0.752787.4%128K
Qwen3.8 2.4T A95BQwen--$2.00$6.006999.4%262K
Qwen3.6 35B A3BQwen--$0.10$1.009799.8%256K
GLM 4.7GLM--$0.40$1.931385.5%198K
GLM 5GLM--$1.00$3.209099.4%198K
Kimi K2.7 CodeKimi--$0.75$3.5044100.0%256K
Kimi K2.5Kimi--$0.532$3.335299.9%256K
Qwen3.5 397B A17BQwen--$0.75$4.505493.4%128K
GLM 5.1GLM--$1.40$4.408495.4%200K
Qwen3 VL 235B A22B InstructQwen--$0.21$1.903084.4%128K
Qwen3.6 27BQwen--$0.325$3.255398.5%256K
GLM 4.7 FlashGLM--$0.06$0.404597.6%128K
Qwen3 Coder 480B A35BQwen--$0.35$1.507474.8%256K
MiniMax M2.5MiniMax--$0.27$0.952694.9%198K
Qwen3.5-35B-A3BQwen--$0.15$1.0014499.9%256K
Qwen3 235B A22B Thinking 2507Qwen--$0.45$3.501491.9%128K
UncensoredOther open-weight--$0.20$0.9060100.0%128K

List prices per million tokens as of October 9, 2026. Share: of all tokens served on that model; a dash where the volume is not reported.

Company

Website
venice.ai
Headquarters
United States
Compute
Said it will use new capital to own its GPUs rather than rent them.
Latest financing
$65M Series A at a $1B valuation led by Dragonfly (July 2026).

Sources: GeekWire.

Finance your GPUs

Serving tokens on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.

Prefer email? hello@amcompute.com