GLM API Providers

Who serves GLM, at what price and speed.

GLM is Z.ai's (formerly Zhipu AI) family of open-weight models for coding and agents, from GLM-4.5 to GLM-5.3 and its Flash variants.

Together served 16% of the 16.4T GLM tokens that went through a large public LLM router over Oct 3 to Oct 9, 2026. Z.ai's own API served 8%; 42 providers list the family.

Who serves GLM tokens

ProviderTokens, 7dShare
Together2.55T15.9%
Relace2.01T12.5%
Z.ai (lab)1.30T8.1%
Parasail1.15T7.1%
Wafer1.14T7.1%
inference.net955B5.9%
Decart842B5.2%
NovitaAI738B4.6%
Mistral576B3.6%
Fireworks571B3.5%
SiliconFlow467B2.9%
Open Inference467B2.9%
Baseten360B2.2%
StreamLake350B2.2%
CoreWeave329B2.0%
26 others2.28T14.2%

All GLM models combined, Oct 3 to Oct 9, 2026.

GLM 5.3 Flash: prices and speed by provider

Output prices run from $0.25 to $1.60 per million tokens across 29 providers, median $0.50.

ProviderShareInput $/MOutput $/MCache read $/MTokens/sLatencyUptime 3dContext
Together20.8%$0.15$0.50$0.03970.68s99.8%1.05M
Relace13.8%$0.04$0.50$0.013461.42s100.0%1.05M
Parasail10.3%$0.15$0.50$0.03951.11s99.9%1.05M
Z.ai8.0%$0.15$0.50$0.03343.50s100.0%1.05M
NovitaAI6.5%$0.084$0.28$0.017301.69s95.6%1.05M
inference.net4.5%$0.07$0.50$0.04830.66s99.8%1.05M
Decart4.3%$0.093$0.31$0.019481.79s99.9%1.05M
Fireworks4.2%$0.15$0.50$0.03751.26s98.2%1.05M
Open Inference4.2%$0.044$0.45$0.01234.16s99.6%1.05M
GMICloud2.6%$0.09$0.30$0.018252.62s99.2%1.05M
CoreWeave2.5%$0.15$0.50$0.051370.62s99.7%1.05M
StreamLake2.4%$0.084$0.28$0.017271.16s98.8%1.02M
Friendli2.3%$0.15$0.50$0.031280.54s99.0%1.05M
Baseten1.8%$0.15$0.50$0.031220.59s99.1%1.05M
Modal1.8%$0.15$0.50$0.03910.30s100.0%1.05M
Wafer1.4%$0.07$0.50$0.03430.64s100.0%1.05M
SiliconFlow1.3%$0.15$0.50$0.03282.79s97.2%1.05M
Reka AI1.0%$0.06$1.60$0.04871.44s99.3%262K
DeepInfra1.0%$0.075$0.25$0.015441.51s99.2%1.05M
Sail Research0.8%$0.045$0.60$0.029322.12s99.9%1.05M
Morph0.8%$0.126$0.75$0.015451.19s92.9%1.05M
Phala0.6%$0.113$0.375$0.022214.71s100.0%1.05M
DigitalOcean0.6%$0.15$0.50$0.03470.62s99.3%1.05M
NEAR AI0.6%$0.105$0.35$0.025121.70s99.9%1.05M
Venice0.5%$0.15$0.50$0.03891.47s99.0%1.05M
Crusoe0.3%$0.15$0.50$0.0372.51.22s99.2%1.05M
AtlasCloud0.3%$0.15$0.50$0.03457.42s93.3%1.05M
Inceptron0.2%$0.12$0.55$0.0991090.58s95.9%1.05M
DekaLLM0.1%$0.10$1.00$0.041051.38s97.2%1.05M

List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.

GLM 5.3: prices and speed by provider

Output prices run from $2.20 to $8.80 per million tokens across 32 providers, median $4.40.

ProviderShareInput $/MOutput $/MCache read $/MTokens/sLatencyUptime 3dContext
Wafer21.3%$0.039$4.80$0.0381000.73s100.0%1.05M
Mistral11.5%$1.40$4.40$0.141081.50s99.5%1.05M
SiliconFlow9.4%$1.12$3.52$0.208483.46s99.7%1.05M
Z.ai9.2%$1.40$4.40$0.26403.72s100.0%1.05M
inference.net8.8%$0.08$4.40$0.0551080.52s100.0%1.05M
Decart5.7%$0.842$2.65$0.1961230.77s99.9%1.05M
Baseten3.7%$1.40$4.40$0.14661.11s99.8%1.05M
Sail Research3.5%$0.20$3.40$0.15691.85s100.0%1.05M
Together3.5%$1.40$4.40$0.261340.34s96.0%1.05M
Fireworks3.0%$1.40$4.40$0.26571.00s99.7%1.05M
Crusoe2.7%$1.40$4.40$0.261470.54s99.9%1.05M
Prime Intellect2.4%$1.40$4.40$0.261470.83s100.0%1.05M
Modal2.3%$1.40$4.40$0.261100.33s99.9%1.05M
Makora1.7%$0.14$4.40$0.101070.84s98.5%1.05M
Morph0.8%$0.407$6.00$0.108301.02s97.3%1.05M
DeepInfra0.7%$0.563$2.50$0.125801.27s98.3%1.05M
Inceptron0.7%$0.60$2.20$0.201040.51s99.2%1.05M
Reka AI0.6%$0.14$4.20$0.105630.73s98.2%1.05M
NovitaAI0.4%$0.70$2.20$0.13512.32s98.4%1.05M
Phala0.4%$0.84$2.64$0.156611.47s99.9%1.05M
Baidu Qianfan0.3%$1.40$4.40$0.2653.51.33s99.4%1.05M
Friendli0.3%$1.26$3.96$0.2341100.52s99.8%1.05M
Alibaba Cloud Int.0.3%$1.19$3.74$0.238541.52s100.0%1M
AkashML0.2%$0.19$4.40$0.19791.35s100.0%1.05M
DigitalOcean0.1%$0.91$2.86$0.169470.96s98.9%1.05M
Venice<0.1%$1.40$4.40$0.26731.06s96.9%1M
Cloudflare<0.1%$1.40$4.40$0.26356.87s98.8%1.05M
io.net<0.1%$0.75$3.40$0.20711.00s97.5%262K
Nebius Token Factory<0.1%$1.40$4.40–1680.47s89.4%1.02M
GMICloud-$0.98$3.08$0.182405.34s100.0%1.05M
Parasail-$1.40$4.40$0.26740.62s99.7%1.05M
AtlasCloud-$1.40$4.40$0.26451.17s98.8%1.05M

List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.

All GLM models

ModelTokens, 7dProvidersOutput $/MContextReleased
GLM 5.3 Flash11.2T29$0.25 to $1.601.05M2026-08-26
GLM 5.33.42T32$2.20 to $8.801.05M2026-08-18
GLM 5.21.63T25$1.68 to $10.001.05M2026-06-16
GLM 4.756B6$1.75 to $2.50205K2025-12-22
GLM 550B8$1.92 to $3.20205K2026-02-11
GLM 5.129B13$3.04 to $4.40205K2026-04-07
GLM 4.621B4$1.75 to $2.20205K2025-09-30
GLM 4.7 Flash15B3$0.40200K2026-01-19
GLM 4.5 Air12B3$0.85 to $1.10131K2025-07-25
GLM 4.52B1$2.20131K2025-07-25
GLM 4.6V1B2$0.90131K2025-12-08
GLM 4.5V240M2$1.8066K2025-08-11

Finance your GPUs

Serving GLM on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.

Prefer email? hello@amcompute.com