GMICloud: Models and Prices

GPU cloud renting NVIDIA servers, with an inference engine on top.

GMICloud served 1.45T tokens of other labs' open-weight models over Oct 3 to Oct 9, 2026 through a large public LLM router, #12 among providers.

Open-weight models

ModelFamilyShare of modelTokens, 7dInput $/MOutput $/MTokens/sUptime 3dContext
DeepSeek V4 Flash 0423DeepSeek15.5%474B$0.091$0.1824499.8%1.05M
GLM 5.3 FlashGLM2.6%286B$0.09$0.302599.2%1.05M
MiMo-V2.6-ProMiMo17.3%211B$0.435$0.874088.9%1.05M
DeepSeek V4 Flash 0731DeepSeek2.8%143B$0.286$0.85854100.0%1.05M
DeepSeek V4.1 FlashDeepSeek0.3%106B$0.18$0.7211299.9%1.05M
MiniMax M3MiniMax6.8%82B$0.24$0.965999.8%1.05M
MiMo-V2.6-FlashMiMo0.5%52B$0.14$0.282179.7%1.05M
MiMo-V2.5MiMo4.5%34B$0.119$0.2382592.9%1.05M
DeepSeek V3.2DeepSeek10.3%29B$0.209$0.313798.9%164K
MiniMax M2.7MiniMax51.9%22B$0.21$0.844199.1%197K
DeepSeek V4 Flash Vision ExpDeepSeek9.5%3B$0.44$1.3255100.0%1.05M
DeepSeek V4 Pro 0423DeepSeek0.4%2B$0.957$1.912397.3%1.05M
DeepSeek V3 0324DeepSeek46.2%2B$0.29$1.1430100.0%128K
GLM 5.3GLM--$0.98$3.0840100.0%1.05M
Hy3Hunyuan--$0.14$0.586699.8%262K
GLM 5.2GLM--$1.40$4.4027100.0%1.05M
DeepSeek V4 Pro 0813DeepSeek--$1.06$3.1731100.0%1.05M
MiMo-V2.5-ProMiMo--$0.304$0.6092396.0%1.05M
Kimi K2.6Kimi--$0.855$3.60--262K
Qwen3 235B A22B Instruct 2507Qwen--$0.087$0.35-98.8%262K
GLM 5GLM--$0.60$1.925898.6%203K
Kimi K2.7 CodeKimi--$0.95$4.0031100.0%262K
Qwen3.5 397B A17BQwen--$0.60$3.605958.0%262K
GLM 5.1GLM--$1.40$4.4059.5100.0%203K
MiniMax M2.5MiniMax--$0.30$1.205197.7%197K
Hy3 previewHunyuan--$0.18$0.6092100.0%262K

List prices per million tokens as of October 9, 2026. Share: of all tokens served on that model; a dash where the volume is not reported.

Company

Headquarters
Mountain View, CA
Compute
Operates data centers in the U.S., Taiwan and Asia; building a 16 MW, roughly 7,000-GPU GB300 site in Taoyuan, Taiwan.
Latest financing
$668M: $223M Series B equity led by ARCHIV and a $445M credit facility led by CTBC (2026).

Sources: DCD.

Finance your GPUs

Serving tokens on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.

Prefer email? hello@amcompute.com