Hunyuan API Providers
Who serves Hunyuan, at what price and speed.
Hunyuan (Hy) is Tencent's model family. Tencent publishes open weights and serves almost all Hunyuan tokens itself through Tencent Cloud.
Tencent Cloud served 97% of the 10.1T Hunyuan tokens that went through a large public LLM router over Oct 3 to Oct 9, 2026. Tencent's own API served 97%; 7 providers list the family.
Who serves Hunyuan tokens
| Provider | Tokens, 7d | Share |
|---|---|---|
| Tencent Cloud (lab) | 9.71T | 96.8% |
| NovitaAI | 168B | 1.7% |
| SiliconFlow | 152B | 1.5% |
| Phala | 650M | <0.1% |
| AtlasCloud | 0 | 0.0% |
| inference.net | 0 | 0.0% |
| Reka AI | 0 | 0.0% |
All Hunyuan models combined, Oct 3 to Oct 9, 2026.
Hy4 preview: prices and speed by provider
Output prices run from $2.50 to $2.50 per million tokens across 4 providers, median $2.50.
| Provider | Share | Input $/M | Output $/M | Cache read $/M | Tokens/s | Latency | Uptime 3d | Context |
|---|---|---|---|---|---|---|---|---|
| Tencent Cloud | 96.0% | $0.834 | $2.50 | $0.042 | 43 | 4.46s | 99.8% | 1.05M |
| NovitaAI | 2.1% | $0.834 | $2.50 | $0.042 | 29 | 5.61s | 100.0% | 1M |
| SiliconFlow | 1.9% | $0.834 | $2.50 | $0.042 | 40 | 1.76s | 97.5% | 1.05M |
| DeepInfra | - | $0.834 | $2.50 | $0.042 | 12 | 11.59s | 97.8% | 1.05M |
List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.
Hy3: prices and speed by provider
Output prices run from $0.528 to $0.80 per million tokens across 5 providers, median $0.58.
| Provider | Share | Input $/M | Output $/M | Cache read $/M | Tokens/s | Latency | Uptime 3d | Context |
|---|---|---|---|---|---|---|---|---|
| Tencent Cloud | 100.0% | $0.132 | $0.528 | $0.033 | 36 | 4.38s | 100.0% | 262K |
| Phala | <0.1% | $0.15 | $0.64 | $0.04 | 63 | 2.70s | 100.0% | 262K |
| NovitaAI | - | $0.14 | $0.58 | $0.035 | 41 | 3.86s | 99.0% | 262K |
| GMICloud | - | $0.14 | $0.58 | $0.035 | 66 | 2.62s | 99.8% | 262K |
| AtlasCloud | - | $0.20 | $0.80 | $0.05 | 95 | 2.28s | 100.0% | 262K |
List prices as of October 9, 2026, cheapest endpoint per provider. Tokens/s and latency are medians over 30 minutes.
All Hunyuan models
| Model | Tokens, 7d | Providers | Output $/M | Context | Released |
|---|---|---|---|---|---|
| Hy4 preview | 7.96T | 4 | $2.50 | 1.05M | 2026-08-28 |
| Hy3 | 2.10T | 5 | $0.528 to $0.80 | 262K | 2026-07-06 |
| Hy3 preview | 5B | 1 | $0.60 | 262K | 2026-04-22 |
| Hy-MT2-30B-A3B | 1B | 1 | $0.295 | 8K | 2026-08-20 |
| Hunyuan A13B Instruct | 50M | 1 | $0.57 | 131K | 2025-07-08 |
| Hy-MT2-7B | 47M | 1 | $0.295 | 8K | 2026-08-19 |
| Hy-MT2-1.8B | 46M | 1 | $0.177 | 8K | 2026-08-20 |
Finance your GPUs
Serving Hunyuan on your own GPUs? Tell us the cluster, the models you serve and the contracts behind them, and we will share it with partner funders that fit.
Prefer email? hello@amcompute.com