Wafer: Models and Prices

Inference provider whose agents rewrite kernels, batching, scheduling and memory layout for open-weight models.

Wafer served 2.57T tokens of other labs' open-weight models over Oct 3 to Oct 9, 2026 through a large public LLM router, #7 among providers.

Open-weight models

ModelFamilyShare of modelTokens, 7dInput $/MOutput $/MTokens/sUptime 3dContext
DeepSeek V4.1 FlashDeepSeek2.3%881B$0.05$0.6587100.0%1.05M
GLM 5.3GLM21.3%703B$0.039$4.80100100.0%1.05M
GLM 5.2GLM18.6%284B$0.06$7.0015199.7%1.05M
DeepSeek V4 Flash 0731DeepSeek4.1%209B$0.13$0.2384100.0%1.05M
GLM 5.3 FlashGLM1.4%154B$0.07$0.5043100.0%1.05M
DeepSeek V4 Pro 0813DeepSeek13.0%139B$0.30$5.00115.5100.0%1.05M
Kimi K3Kimi4.8%91B$2.80$14.005398.3%1.05M
Qwen3.8 27BQwen13.9%80B$0.04$2.3012499.7%262K
DeepSeek V4 Flash 0423DeepSeek0.9%28B$0.035$0.174898.8%1.05M
Nemotron 3.5 LightningNemotron--$0.035$0.1321299.9%262K

List prices per million tokens as of October 9, 2026. Share: of all tokens served on that model; a dash where the volume is not reported.

Company

Website
wafer.ai
Headquarters
United States
Latest financing
About $40M Series A (2026).

Sources: Runtime Wire.

Finance your GPUs

Serving tokens on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.

Prefer email? hello@amcompute.com