Phala: Models and Prices

Confidential inference on GPUs running in trusted execution environments.

Phala served 311B tokens of other labs' open-weight models over Oct 3 to Oct 9, 2026 through a large public LLM router, #27 among providers.

Open-weight models

ModelFamilyShare of modelTokens, 7dInput $/MOutput $/MTokens/sUptime 3dContext
DeepSeek V4.1 FlashDeepSeek0.4%132B$0.21$0.847097.7%1.05M
GLM 5.3 FlashGLM0.6%72B$0.113$0.37521100.0%1.05M
Kimi K3Kimi1.7%32B$2.55$12.754799.8%1.05M
Qwen3.8 27BQwen3.2%18B$0.15$1.886695.9%1M
GLM 5.3GLM0.4%14B$0.84$2.646199.9%1.05M
DeepSeek V4 Flash 0731DeepSeek0.3%14B$0.44$1.327299.0%1.05M
Qwen3.5 397B A17BQwen100.0%12B$0.55$3.502599.1%262K
DeepSeek V3.2DeepSeek1.7%5B$1.00$1.002599.6%164K
Muse Glimmer 30BOther open-weight25.0%4B$0.30$1.1062100.0%131K
GLM 5.2GLM0.2%3B$1.26$3.0081100.0%1.05M
Qwen3.6 35B A3BQwen6.2%2B$0.20$1.272381.0%262K
DeepSeek V4 Pro 0813DeepSeek0.1%1B$1.45$4.363599.4%1.05M
Hy3Hunyuan<0.1%650M$0.15$0.6463100.0%262K
Qwen3.6 27BQwen10.0%385M$0.32$3.256097.7%262K
Qwen2.5 7B InstructQwen100.0%331M$0.10$0.20132100.0%33K
Nemotron 3.5 LightningNemotron--$0.07$0.20240.599.8%262K
gpt-oss-120bgpt-oss--$0.15$0.6012899.4%131K
Kimi K2.6Kimi--$1.09$4.603399.8%262K
GLM 5.1GLM--$1.21$4.202283.7%203K
Qwen3.5-27BQwen--$0.30$2.401892.9%262K

List prices per million tokens as of October 9, 2026. Share: of all tokens served on that model; a dash where the volume is not reported.

Company

Headquarters
United States

Sources: Phala.

Finance your GPUs

Serving tokens on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.

Prefer email? hello@amcompute.com