inference.net: Models and Prices

OpenAI-compatible API for open-weight models plus custom model training and deployment for enterprises. Formerly Kuzco.

inference.net served 4.13T tokens of other labs' open-weight models over Oct 3 to Oct 9, 2026 through a large public LLM router, #4 among providers.

Open-weight models

ModelFamilyShare of modelTokens, 7dInput $/MOutput $/MTokens/sUptime 3dContext
DeepSeek V4.1 FlashDeepSeek6.8%2.56T$0.105$0.6097100.0%1.04M
Kimi K3Kimi29.7%571B$0.64$13.5067100.0%1.05M
GLM 5.3 FlashGLM4.5%505B$0.07$0.508399.8%1.05M
GLM 5.3GLM8.8%291B$0.08$4.40108100.0%1.05M
GLM 5.2GLM10.4%159B$0.18$4.4068100.0%1.05M
Schematron V2 TurboOther open-weight100.0%194M$0.03$0.157699.0%128K
Schematron V2 SmallOther open-weight100.0%67M$0.05$0.2332100.0%128K

List prices per million tokens as of October 9, 2026. Share: of all tokens served on that model; a dash where the volume is not reported.

Company

Headquarters
United States
Latest financing
$11.8M seed led by Multicoin Capital and a16z CSX (October 2025).

Sources: Techleap.

Finance your GPUs

Serving tokens on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.

Prefer email? hello@amcompute.com