glm-5.2 API — compare providers
The same model (glm-5.2) is served by 22 providers; OpenRouter offers a free endpoint — the priciest (Alibaba) costs $9.57/1M blended. On quality, glm-5.2 scores 1472 on the LMArena (Chatbot Arena) — so the cheapest endpoint is also the best value per dollar here.
Quality and intelligence scores are attached to the model itself — identical whichever provider serves it. Verified 2026-09-04.
| Rank | Provider | Input /1M | Output /1M | Blended /1M | Value (Elo/$) |
|---|---|---|---|---|---|
| 1 | OpenRouter | $0 | $0 | $0 ← cheapest | — ← best value |
| 2 | Baidu | $0.357 | $1.122 | $1.479 | 995 |
| 3 | Novita AI | $0.413 | $1.298 | $1.711 | 860 |
| 4 | DeepInfra | $0.4875 | $1.56 | $2.047 | 719 |
| 5 | DigitalOcean | $0.7 | $2.2 | $2.9 | 508 |
| 6 | CoreWeave | $0.76 | $2.42 | $3.18 | 463 |
| 7 | AtlasCloud | $0.938 | $2.948 | $3.886 | 379 |
| 8 | Phala | $1.26 | $3 | $4.26 | 346 |
| 9 | GMICloud | $1.05 | $3.3 | $4.35 | 338 |
| 10 | Reka | $1.1 | $3.75 | $4.85 | 304 |
| 11 | SiliconFlow | $1.19 | $3.74 | $4.93 | 299 |
| 12 | Together AI | $1.4 | $4.4 | $5.8 | 254 |
| 13 | Mistral AI | $1.4 | $4.4 | $5.8 | 254 |
| 14 | Crusoe | $1.4 | $4.4 | $5.8 | 254 |
| 15 | Venice | $1.4 | $4.4 | $5.8 | 254 |
| 16 | Parasail | $1.4 | $4.4 | $5.8 | 254 |
| 17 | Friendli | $1.4 | $4.4 | $5.8 | 254 |
| 18 | Cloudflare | $1.4 | $4.4 | $5.8 | 254 |
| 19 | Z.AI | $1.4 | $4.4 | $5.8 | 254 |
| 20 | BaseTen | $2.1 | $6.6 | $8.7 | 169 |
| 21 | Fireworks AI | $2.1 | $6.6 | $8.7 | 169 |
| 22 | Alibaba | $2.31 | $7.26 | $9.57 | 154 |
Value = model Elo ÷ blended price per 1M tokens. Since the model's quality is identical across providers, the cheapest endpoint is also the best value per dollar.
Prices are per 1M tokens (USD). "Blended" = input + output, for coarse ranking. Same underlying model, 22 serving providers (a free endpoint is available). Context window: 131,072 tokens (32,768 max output) — see how it ranks in biggest context windows.
Cost per 1,000 requests by workload
What each provider actually bills for a representative job, not just the sticker price.
| Provider | Chatbot 1000 in / 500 out | RAG / long context 8000 in / 500 out | Batch summarize 4000 in / 1000 out |
|---|---|---|---|
| OpenRouter | $0.00 | $0.00 | $0.00 |
| Baidu | $0.92 | $3.42 | $2.55 |
| Novita AI | $1.06 | $3.95 | $2.95 |
| DeepInfra | $1.27 | $4.68 | $3.51 |
| DigitalOcean | $1.80 | $6.70 | $5.00 |
| CoreWeave | $1.97 | $7.29 | $5.46 |
| AtlasCloud | $2.41 | $8.98 | $6.70 |
| Phala | $2.76 | $11.58 | $8.04 |
| GMICloud | $2.70 | $10.05 | $7.50 |
| Reka | $2.98 | $10.68 | $8.15 |
| SiliconFlow | $3.06 | $11.39 | $8.50 |
| Together AI | $3.60 | $13.40 | $10.00 |
| Mistral AI | $3.60 | $13.40 | $10.00 |
| Crusoe | $3.60 | $13.40 | $10.00 |
| Venice | $3.60 | $13.40 | $10.00 |
| Parasail | $3.60 | $13.40 | $10.00 |
| Friendli | $3.60 | $13.40 | $10.00 |
| Cloudflare | $3.60 | $13.40 | $10.00 |
| Z.AI | $3.60 | $13.40 | $10.00 |
| BaseTen | $5.40 | $20.10 | $15.00 |
| Fireworks AI | $5.40 | $20.10 | $15.00 |
| Alibaba | $5.94 | $22.11 | $16.50 |
Estimate your own workload
Input tokens/request: Output tokens/request: Requests:
| Provider | Estimated cost (USD) |
|---|
Self-host / quantized builds
Open weights under MIT — you can run glm-5.2 on your own hardware instead of paying per token. Quantized builds lower the memory footprint; whether it fits your GPU or Mac depends on the model's size. Formats: GGUF for llama.cpp / Ollama / LM Studio, AWQ and GPTQ for GPU serving, MLX for Apple silicon.
- GGUF builds llama.cpp · Ollama · LM Studio
- AWQ builds GPU serving (vLLM)
- GPTQ builds GPU serving
- MLX builds Apple silicon
- Ollama library one-command local run
Links search the live model hubs (Hugging Face, Ollama) so they track new builds as they appear — we don't host weights. Which formats exist depends on what the model owner and community have published.