glm-5.2 API — compare providers

The same model (glm-5.2) is served by 22 providers; OpenRouter offers a free endpoint — the priciest (Alibaba) costs $9.57/1M blended. On quality, glm-5.2 scores 1472 on the LMArena (Chatbot Arena) — so the cheapest endpoint is also the best value per dollar here.

LicenseMITopen weights · permissive · terms
Knowledge cutoffJanuary 2026source · ranked
Quality (Elo)1472LMArena (Chatbot Arena) · source
Intelligence52.6Artificial Analysis Intelligence Index · source · ranked

Quality and intelligence scores are attached to the model itself — identical whichever provider serves it. Verified 2026-09-04.

RankProviderInput /1MOutput /1MBlended /1MValue (Elo/$)
1OpenRouter$0$0$0 ← cheapest ← best value
2Baidu$0.357$1.122$1.479995
3Novita AI$0.413$1.298$1.711860
4DeepInfra$0.4875$1.56$2.047719
5DigitalOcean$0.7$2.2$2.9508
6CoreWeave$0.76$2.42$3.18463
7AtlasCloud$0.938$2.948$3.886379
8Phala$1.26$3$4.26346
9GMICloud$1.05$3.3$4.35338
10Reka$1.1$3.75$4.85304
11SiliconFlow$1.19$3.74$4.93299
12Together AI$1.4$4.4$5.8254
13Mistral AI$1.4$4.4$5.8254
14Crusoe$1.4$4.4$5.8254
15Venice$1.4$4.4$5.8254
16Parasail$1.4$4.4$5.8254
17Friendli$1.4$4.4$5.8254
18Cloudflare$1.4$4.4$5.8254
19Z.AI$1.4$4.4$5.8254
20BaseTen$2.1$6.6$8.7169
21Fireworks AI$2.1$6.6$8.7169
22Alibaba$2.31$7.26$9.57154

Value = model Elo ÷ blended price per 1M tokens. Since the model's quality is identical across providers, the cheapest endpoint is also the best value per dollar.

Prices are per 1M tokens (USD). "Blended" = input + output, for coarse ranking. Same underlying model, 22 serving providers (a free endpoint is available). Context window: 131,072 tokens (32,768 max output) — see how it ranks in biggest context windows.

Cost per 1,000 requests by workload

What each provider actually bills for a representative job, not just the sticker price.

ProviderChatbot
1000 in / 500 out
RAG / long context
8000 in / 500 out
Batch summarize
4000 in / 1000 out
OpenRouter$0.00$0.00$0.00
Baidu$0.92$3.42$2.55
Novita AI$1.06$3.95$2.95
DeepInfra$1.27$4.68$3.51
DigitalOcean$1.80$6.70$5.00
CoreWeave$1.97$7.29$5.46
AtlasCloud$2.41$8.98$6.70
Phala$2.76$11.58$8.04
GMICloud$2.70$10.05$7.50
Reka$2.98$10.68$8.15
SiliconFlow$3.06$11.39$8.50
Together AI$3.60$13.40$10.00
Mistral AI$3.60$13.40$10.00
Crusoe$3.60$13.40$10.00
Venice$3.60$13.40$10.00
Parasail$3.60$13.40$10.00
Friendli$3.60$13.40$10.00
Cloudflare$3.60$13.40$10.00
Z.AI$3.60$13.40$10.00
BaseTen$5.40$20.10$15.00
Fireworks AI$5.40$20.10$15.00
Alibaba$5.94$22.11$16.50

Estimate your own workload

Input tokens/request:   Output tokens/request:   Requests:

ProviderEstimated cost (USD)

Self-host / quantized builds

Open weights under MIT — you can run glm-5.2 on your own hardware instead of paying per token. Quantized builds lower the memory footprint; whether it fits your GPU or Mac depends on the model's size. Formats: GGUF for llama.cpp / Ollama / LM Studio, AWQ and GPTQ for GPU serving, MLX for Apple silicon.

Links search the live model hubs (Hugging Face, Ollama) so they track new builds as they appear — we don't host weights. Which formats exist depends on what the model owner and community have published.