Nebius Token Factory: Models and Prices
Inference provider serving open-weight models.
Nebius Token Factory served 37B tokens of other labs' open-weight models over Oct 3 to Oct 9, 2026 through a large public LLM router, #45 among providers.
Open-weight models
| Model | Family | Share of model | Tokens, 7d | Input $/M | Output $/M | Tokens/s | Uptime 3d | Context |
|---|---|---|---|---|---|---|---|---|
| Qwen3 30B A3B Instruct 2507 | Qwen | 61.2% | 15B | $0.10 | $0.30 | 45 | 92.4% | 262K |
| Gemma 3 27B | Gemma | 15.3% | 7B | $0.10 | $0.30 | 49 | 93.8% | 110K |
| Qwen3 235B A22B Instruct 2507 | Qwen | 97.2% | 5B | $0.20 | $0.60 | 57.5 | 86.9% | 262K |
| gpt-oss-120b | gpt-oss | 0.8% | 4B | $0.15 | $0.60 | 273 | 98.1% | 131K |
| Nemotron 3 Nano 30B A3B | Nemotron | 61.4% | 3B | $0.06 | $0.24 | 145 | 98.1% | 262K |
| GLM 5.3 | GLM | <0.1% | 999M | $1.40 | $4.40 | 168 | 89.4% | 1.02M |
| GLM 5.2 | GLM | <0.1% | 527M | $1.40 | $4.40 | 160 | 89.1% | 1.05M |
| GLM 5.1 | GLM | 5.6% | 494M | $1.40 | $4.40 | 40 | 97.5% | 203K |
| Hermes 4 405B | Other open-weight | 100.0% | 438M | $1.00 | $3.00 | 42 | 100.0% | 131K |
| Kimi K2.7 Code | Kimi | <0.1% | 28M | $0.95 | $4.00 | 17 | 96.2% | 262K |
List prices per million tokens as of October 9, 2026. Share: of all tokens served on that model; a dash where the volume is not reported.
Company
- Website
- docs.nebius.com
- Headquarters
- Netherlands
Finance your GPUs
Serving tokens on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.
Prefer email? hello@amcompute.com