GLM API Providers
Who serves GLM, at what price and speed.
GLM is Z.ai's (formerly Zhipu AI) family of open-weight models for coding and agents, from GLM-4.5 to GLM-5.3 and its Flash variants.
Together served 16% of the 16.4T GLM tokens that went through a large public LLM router over Oct 3 to Oct 9, 2026. Z.ai's own API served 8%; 42 providers list the family.
Who serves GLM tokens
| Provider | Tokens, 7d | Share |
|---|---|---|
| Together | 2.55T | 15.9% |
| Relace | 2.01T | 12.5% |
| Z.ai (lab) | 1.30T | 8.1% |
| Parasail | 1.15T | 7.1% |
| Wafer | 1.14T | 7.1% |
| inference.net | 955B | 5.9% |
| Decart | 842B | 5.2% |
| NovitaAI | 738B | 4.6% |
| Mistral | 576B | 3.6% |
| Fireworks | 571B | 3.5% |
| SiliconFlow | 467B | 2.9% |
| Open Inference | 467B | 2.9% |
| Baseten | 360B | 2.2% |
| StreamLake | 350B | 2.2% |
| CoreWeave | 329B | 2.0% |
| 26 others | 2.28T | 14.2% |
All GLM models combined, Oct 3 to Oct 9, 2026.
GLM 5.3 Flash: prices and speed by provider
Output prices run from $0.25 to $1.60 per million tokens across 29 providers, median $0.50.
| Provider | Share | Input $/M | Output $/M | Cache read $/M | Tokens/s | Latency | Uptime 3d | Context |
|---|---|---|---|---|---|---|---|---|
| Together | 20.8% | $0.15 | $0.50 | $0.03 | 97 | 0.68s | 99.8% | 1.05M |
| Relace | 13.8% | $0.04 | $0.50 | $0.013 | 46 | 1.42s | 100.0% | 1.05M |
| Parasail | 10.3% | $0.15 | $0.50 | $0.03 | 95 | 1.11s | 99.9% | 1.05M |
| Z.ai | 8.0% | $0.15 | $0.50 | $0.03 | 34 | 3.50s | 100.0% | 1.05M |
| NovitaAI | 6.5% | $0.084 | $0.28 | $0.017 | 30 | 1.69s | 95.6% | 1.05M |
| inference.net | 4.5% | $0.07 | $0.50 | $0.04 | 83 | 0.66s | 99.8% | 1.05M |
| Decart | 4.3% | $0.093 | $0.31 | $0.019 | 48 | 1.79s | 99.9% | 1.05M |
| Fireworks | 4.2% | $0.15 | $0.50 | $0.03 | 75 | 1.26s | 98.2% | 1.05M |
| Open Inference | 4.2% | $0.044 | $0.45 | $0.01 | 23 | 4.16s | 99.6% | 1.05M |
| GMICloud | 2.6% | $0.09 | $0.30 | $0.018 | 25 | 2.62s | 99.2% | 1.05M |
| CoreWeave | 2.5% | $0.15 | $0.50 | $0.05 | 137 | 0.62s | 99.7% | 1.05M |
| StreamLake | 2.4% | $0.084 | $0.28 | $0.017 | 27 | 1.16s | 98.8% | 1.02M |
| Friendli | 2.3% | $0.15 | $0.50 | $0.03 | 128 | 0.54s | 99.0% | 1.05M |
| Baseten | 1.8% | $0.15 | $0.50 | $0.03 | 122 | 0.59s | 99.1% | 1.05M |
| Modal | 1.8% | $0.15 | $0.50 | $0.03 | 91 | 0.30s | 100.0% | 1.05M |
| Wafer | 1.4% | $0.07 | $0.50 | $0.03 | 43 | 0.64s | 100.0% | 1.05M |
| SiliconFlow | 1.3% | $0.15 | $0.50 | $0.03 | 28 | 2.79s | 97.2% | 1.05M |
| Reka AI | 1.0% | $0.06 | $1.60 | $0.04 | 87 | 1.44s | 99.3% | 262K |
| DeepInfra | 1.0% | $0.075 | $0.25 | $0.015 | 44 | 1.51s | 99.2% | 1.05M |
| Sail Research | 0.8% | $0.045 | $0.60 | $0.029 | 32 | 2.12s | 99.9% | 1.05M |
| Morph | 0.8% | $0.126 | $0.75 | $0.015 | 45 | 1.19s | 92.9% | 1.05M |
| Phala | 0.6% | $0.113 | $0.375 | $0.022 | 21 | 4.71s | 100.0% | 1.05M |
| DigitalOcean | 0.6% | $0.15 | $0.50 | $0.03 | 47 | 0.62s | 99.3% | 1.05M |
| NEAR AI | 0.6% | $0.105 | $0.35 | $0.025 | 12 | 1.70s | 99.9% | 1.05M |
| Venice | 0.5% | $0.15 | $0.50 | $0.03 | 89 | 1.47s | 99.0% | 1.05M |
| Crusoe | 0.3% | $0.15 | $0.50 | $0.03 | 72.5 | 1.22s | 99.2% | 1.05M |
| AtlasCloud | 0.3% | $0.15 | $0.50 | $0.03 | 45 | 7.42s | 93.3% | 1.05M |
| Inceptron | 0.2% | $0.12 | $0.55 | $0.099 | 109 | 0.58s | 95.9% | 1.05M |
| DekaLLM | 0.1% | $0.10 | $1.00 | $0.04 | 105 | 1.38s | 97.2% | 1.05M |
List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.
GLM 5.3: prices and speed by provider
Output prices run from $2.20 to $8.80 per million tokens across 32 providers, median $4.40.
| Provider | Share | Input $/M | Output $/M | Cache read $/M | Tokens/s | Latency | Uptime 3d | Context |
|---|---|---|---|---|---|---|---|---|
| Wafer | 21.3% | $0.039 | $4.80 | $0.038 | 100 | 0.73s | 100.0% | 1.05M |
| Mistral | 11.5% | $1.40 | $4.40 | $0.14 | 108 | 1.50s | 99.5% | 1.05M |
| SiliconFlow | 9.4% | $1.12 | $3.52 | $0.208 | 48 | 3.46s | 99.7% | 1.05M |
| Z.ai | 9.2% | $1.40 | $4.40 | $0.26 | 40 | 3.72s | 100.0% | 1.05M |
| inference.net | 8.8% | $0.08 | $4.40 | $0.055 | 108 | 0.52s | 100.0% | 1.05M |
| Decart | 5.7% | $0.842 | $2.65 | $0.196 | 123 | 0.77s | 99.9% | 1.05M |
| Baseten | 3.7% | $1.40 | $4.40 | $0.14 | 66 | 1.11s | 99.8% | 1.05M |
| Sail Research | 3.5% | $0.20 | $3.40 | $0.15 | 69 | 1.85s | 100.0% | 1.05M |
| Together | 3.5% | $1.40 | $4.40 | $0.26 | 134 | 0.34s | 96.0% | 1.05M |
| Fireworks | 3.0% | $1.40 | $4.40 | $0.26 | 57 | 1.00s | 99.7% | 1.05M |
| Crusoe | 2.7% | $1.40 | $4.40 | $0.26 | 147 | 0.54s | 99.9% | 1.05M |
| Prime Intellect | 2.4% | $1.40 | $4.40 | $0.26 | 147 | 0.83s | 100.0% | 1.05M |
| Modal | 2.3% | $1.40 | $4.40 | $0.26 | 110 | 0.33s | 99.9% | 1.05M |
| Makora | 1.7% | $0.14 | $4.40 | $0.10 | 107 | 0.84s | 98.5% | 1.05M |
| Morph | 0.8% | $0.407 | $6.00 | $0.108 | 30 | 1.02s | 97.3% | 1.05M |
| DeepInfra | 0.7% | $0.563 | $2.50 | $0.125 | 80 | 1.27s | 98.3% | 1.05M |
| Inceptron | 0.7% | $0.60 | $2.20 | $0.20 | 104 | 0.51s | 99.2% | 1.05M |
| Reka AI | 0.6% | $0.14 | $4.20 | $0.105 | 63 | 0.73s | 98.2% | 1.05M |
| NovitaAI | 0.4% | $0.70 | $2.20 | $0.13 | 51 | 2.32s | 98.4% | 1.05M |
| Phala | 0.4% | $0.84 | $2.64 | $0.156 | 61 | 1.47s | 99.9% | 1.05M |
| Baidu Qianfan | 0.3% | $1.40 | $4.40 | $0.26 | 53.5 | 1.33s | 99.4% | 1.05M |
| Friendli | 0.3% | $1.26 | $3.96 | $0.234 | 110 | 0.52s | 99.8% | 1.05M |
| Alibaba Cloud Int. | 0.3% | $1.19 | $3.74 | $0.238 | 54 | 1.52s | 100.0% | 1M |
| AkashML | 0.2% | $0.19 | $4.40 | $0.19 | 79 | 1.35s | 100.0% | 1.05M |
| DigitalOcean | 0.1% | $0.91 | $2.86 | $0.169 | 47 | 0.96s | 98.9% | 1.05M |
| Venice | <0.1% | $1.40 | $4.40 | $0.26 | 73 | 1.06s | 96.9% | 1M |
| Cloudflare | <0.1% | $1.40 | $4.40 | $0.26 | 35 | 6.87s | 98.8% | 1.05M |
| io.net | <0.1% | $0.75 | $3.40 | $0.20 | 71 | 1.00s | 97.5% | 262K |
| Nebius Token Factory | <0.1% | $1.40 | $4.40 | – | 168 | 0.47s | 89.4% | 1.02M |
| GMICloud | - | $0.98 | $3.08 | $0.182 | 40 | 5.34s | 100.0% | 1.05M |
| Parasail | - | $1.40 | $4.40 | $0.26 | 74 | 0.62s | 99.7% | 1.05M |
| AtlasCloud | - | $1.40 | $4.40 | $0.26 | 45 | 1.17s | 98.8% | 1.05M |
List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.
All GLM models
| Model | Tokens, 7d | Providers | Output $/M | Context | Released |
|---|---|---|---|---|---|
| GLM 5.3 Flash | 11.2T | 29 | $0.25 to $1.60 | 1.05M | 2026-08-26 |
| GLM 5.3 | 3.42T | 32 | $2.20 to $8.80 | 1.05M | 2026-08-18 |
| GLM 5.2 | 1.63T | 25 | $1.68 to $10.00 | 1.05M | 2026-06-16 |
| GLM 4.7 | 56B | 6 | $1.75 to $2.50 | 205K | 2025-12-22 |
| GLM 5 | 50B | 8 | $1.92 to $3.20 | 205K | 2026-02-11 |
| GLM 5.1 | 29B | 13 | $3.04 to $4.40 | 205K | 2026-04-07 |
| GLM 4.6 | 21B | 4 | $1.75 to $2.20 | 205K | 2025-09-30 |
| GLM 4.7 Flash | 15B | 3 | $0.40 | 200K | 2026-01-19 |
| GLM 4.5 Air | 12B | 3 | $0.85 to $1.10 | 131K | 2025-07-25 |
| GLM 4.5 | 2B | 1 | $2.20 | 131K | 2025-07-25 |
| GLM 4.6V | 1B | 2 | $0.90 | 131K | 2025-12-08 |
| GLM 4.5V | 240M | 2 | $1.80 | 66K | 2025-08-11 |
Finance your GPUs
Serving GLM on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.
Prefer email? hello@amcompute.com