GMICloud: Models and Prices
GPU cloud renting NVIDIA servers, with an inference engine on top.
GMICloud served 1.45T tokens of other labs' open-weight models over Oct 3 to Oct 9, 2026 through a large public LLM router, #12 among providers.
Open-weight models
| Model | Family | Share of model | Tokens, 7d | Input $/M | Output $/M | Tokens/s | Uptime 3d | Context |
|---|---|---|---|---|---|---|---|---|
| DeepSeek V4 Flash 0423 | DeepSeek | 15.5% | 474B | $0.091 | $0.182 | 44 | 99.8% | 1.05M |
| GLM 5.3 Flash | GLM | 2.6% | 286B | $0.09 | $0.30 | 25 | 99.2% | 1.05M |
| MiMo-V2.6-Pro | MiMo | 17.3% | 211B | $0.435 | $0.87 | 40 | 88.9% | 1.05M |
| DeepSeek V4 Flash 0731 | DeepSeek | 2.8% | 143B | $0.286 | $0.858 | 54 | 100.0% | 1.05M |
| DeepSeek V4.1 Flash | DeepSeek | 0.3% | 106B | $0.18 | $0.72 | 112 | 99.9% | 1.05M |
| MiniMax M3 | MiniMax | 6.8% | 82B | $0.24 | $0.96 | 59 | 99.8% | 1.05M |
| MiMo-V2.6-Flash | MiMo | 0.5% | 52B | $0.14 | $0.28 | 21 | 79.7% | 1.05M |
| MiMo-V2.5 | MiMo | 4.5% | 34B | $0.119 | $0.238 | 25 | 92.9% | 1.05M |
| DeepSeek V3.2 | DeepSeek | 10.3% | 29B | $0.209 | $0.31 | 37 | 98.9% | 164K |
| MiniMax M2.7 | MiniMax | 51.9% | 22B | $0.21 | $0.84 | 41 | 99.1% | 197K |
| DeepSeek V4 Flash Vision Exp | DeepSeek | 9.5% | 3B | $0.44 | $1.32 | 55 | 100.0% | 1.05M |
| DeepSeek V4 Pro 0423 | DeepSeek | 0.4% | 2B | $0.957 | $1.91 | 23 | 97.3% | 1.05M |
| DeepSeek V3 0324 | DeepSeek | 46.2% | 2B | $0.29 | $1.14 | 30 | 100.0% | 128K |
| GLM 5.3 | GLM | - | - | $0.98 | $3.08 | 40 | 100.0% | 1.05M |
| Hy3 | Hunyuan | - | - | $0.14 | $0.58 | 66 | 99.8% | 262K |
| GLM 5.2 | GLM | - | - | $1.40 | $4.40 | 27 | 100.0% | 1.05M |
| DeepSeek V4 Pro 0813 | DeepSeek | - | - | $1.06 | $3.17 | 31 | 100.0% | 1.05M |
| MiMo-V2.5-Pro | MiMo | - | - | $0.304 | $0.609 | 23 | 96.0% | 1.05M |
| Kimi K2.6 | Kimi | - | - | $0.855 | $3.60 | - | - | 262K |
| Qwen3 235B A22B Instruct 2507 | Qwen | - | - | $0.087 | $0.35 | - | 98.8% | 262K |
| GLM 5 | GLM | - | - | $0.60 | $1.92 | 58 | 98.6% | 203K |
| Kimi K2.7 Code | Kimi | - | - | $0.95 | $4.00 | 31 | 100.0% | 262K |
| Qwen3.5 397B A17B | Qwen | - | - | $0.60 | $3.60 | 59 | 58.0% | 262K |
| GLM 5.1 | GLM | - | - | $1.40 | $4.40 | 59.5 | 100.0% | 203K |
| MiniMax M2.5 | MiniMax | - | - | $0.30 | $1.20 | 51 | 97.7% | 197K |
| Hy3 preview | Hunyuan | - | - | $0.18 | $0.60 | 92 | 100.0% | 262K |
List prices per million tokens as of October 9, 2026. Share: of all tokens served on that model; a dash where the volume is not reported.
Company
- Website
- gmicloud.ai
- Headquarters
- Mountain View, CA
- Compute
- Operates data centers in the U.S., Taiwan and Asia; building a 16 MW, roughly 7,000-GPU GB300 site in Taoyuan, Taiwan.
- Latest financing
- $668M: $223M Series B equity led by ARCHIV and a $445M credit facility led by CTBC (2026).
Sources: DCD.
Finance your GPUs
Serving tokens on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.
Prefer email? hello@amcompute.com