Venice: Models and Prices
Private, uncensored AI app and API on self-hosted open-weight models.
Venice served 465B tokens of other labs' open-weight models over Oct 3 to Oct 9, 2026 through a large public LLM router, #23 among providers.
Open-weight models
| Model | Family | Share of model | Tokens, 7d | Input $/M | Output $/M | Tokens/s | Uptime 3d | Context |
|---|---|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | DeepSeek | 0.6% | 239B | $0.30 | $1.20 | 170 | 100.0% | 1M |
| GLM 5.3 Flash | GLM | 0.5% | 55B | $0.15 | $0.50 | 89 | 99.0% | 1.05M |
| DeepSeek V4 Flash 0423 | DeepSeek | 1.7% | 53B | $0.097 | $0.193 | 25 | 96.3% | 1M |
| Gemma 4 31B | Gemma | 5.1% | 22B | $0.12 | $0.36 | 40 | 97.5% | 256K |
| DeepSeek V4 Flash 0731 | DeepSeek | 0.4% | 19B | $0.13 | $0.26 | 40 | 99.8% | 1M |
| Mistral Small 3.2 24B | Mistral | 100.0% | 19B | $0.094 | $0.25 | 30 | 99.8% | 256K |
| MiMo-V2.5 | MiMo | 2.4% | 18B | $0.40 | $2.00 | 17 | 99.4% | 1M |
| Qwen3.5-9B | Qwen | 90.4% | 14B | $0.10 | $0.15 | 52 | 99.8% | 256K |
| MiMo-V2.6-Flash | MiMo | 0.1% | 11B | $0.175 | $0.35 | 16 | 98.5% | 1M |
| Gemma 4 26B A4B | Gemma | 2.2% | 6B | $0.13 | $0.40 | 54 | 99.6% | 256K |
| GLM 4.6 | GLM | 100.0% | 3B | $0.43 | $1.75 | 69 | 99.8% | 198K |
| GLM 5.3 | GLM | <0.1% | 2B | $1.40 | $4.40 | 73 | 96.9% | 1M |
| Nemotron 3 Ultra | Nemotron | <0.1% | 2B | $0.625 | $3.13 | 104 | 96.7% | 256K |
| Qwen3.8 27B | Qwen | 0.3% | 2B | $0.45 | $3.20 | 103 | 99.7% | 262K |
| GLM 5.2 | GLM | <0.1% | 2B | $1.40 | $4.40 | 85 | 99.7% | 1M |
| DeepSeek V3.2 | DeepSeek | 0.5% | 1B | $0.268 | $0.39 | 19 | 99.7% | 160K |
| MiniMax M3 | MiniMax | - | - | $0.30 | $1.20 | 85 | 96.5% | 524K |
| DeepSeek V4 Pro 0813 | DeepSeek | - | - | $1.65 | $4.95 | 76.5 | 93.8% | 1M |
| DeepSeek V4 Pro 0423 | DeepSeek | - | - | $1.65 | $3.30 | 52 | 96.5% | 1M |
| gpt-oss-120b | gpt-oss | - | - | $0.03 | $0.15 | 13 | 100.0% | 128K |
| Kimi K2.6 | Kimi | - | - | $0.75 | $3.50 | 12 | 90.4% | 256K |
| Qwen3 235B A22B Instruct 2507 | Qwen | - | - | $0.15 | $0.75 | 27 | 87.4% | 128K |
| Qwen3.8 2.4T A95B | Qwen | - | - | $2.00 | $6.00 | 69 | 99.4% | 262K |
| Qwen3.6 35B A3B | Qwen | - | - | $0.10 | $1.00 | 97 | 99.8% | 256K |
| GLM 4.7 | GLM | - | - | $0.40 | $1.93 | 13 | 85.5% | 198K |
| GLM 5 | GLM | - | - | $1.00 | $3.20 | 90 | 99.4% | 198K |
| Kimi K2.7 Code | Kimi | - | - | $0.75 | $3.50 | 44 | 100.0% | 256K |
| Kimi K2.5 | Kimi | - | - | $0.532 | $3.33 | 52 | 99.9% | 256K |
| Qwen3.5 397B A17B | Qwen | - | - | $0.75 | $4.50 | 54 | 93.4% | 128K |
| GLM 5.1 | GLM | - | - | $1.40 | $4.40 | 84 | 95.4% | 200K |
| Qwen3 VL 235B A22B Instruct | Qwen | - | - | $0.21 | $1.90 | 30 | 84.4% | 128K |
| Qwen3.6 27B | Qwen | - | - | $0.325 | $3.25 | 53 | 98.5% | 256K |
| GLM 4.7 Flash | GLM | - | - | $0.06 | $0.40 | 45 | 97.6% | 128K |
| Qwen3 Coder 480B A35B | Qwen | - | - | $0.35 | $1.50 | 74 | 74.8% | 256K |
| MiniMax M2.5 | MiniMax | - | - | $0.27 | $0.95 | 26 | 94.9% | 198K |
| Qwen3.5-35B-A3B | Qwen | - | - | $0.15 | $1.00 | 144 | 99.9% | 256K |
| Qwen3 235B A22B Thinking 2507 | Qwen | - | - | $0.45 | $3.50 | 14 | 91.9% | 128K |
| Uncensored | Other open-weight | - | - | $0.20 | $0.90 | 60 | 100.0% | 128K |
List prices per million tokens as of October 9, 2026. Share: of all tokens served on that model; a dash where the volume is not reported.
Company
Finance your GPUs
Serving tokens on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.
Prefer email? hello@amcompute.com