Chinese open-weight flagship models: price and measured speed
Open weights mean the same model is sold by several hosts at different prices and different speeds. This page puts 20 measured endpoints — 4 flagship models from DeepSeek, Z.ai, Moonshot AI and Alibaba, served by 7 providers — in a single table, priced and timed. Blended price runs from$2.64 to $20.7 per 1M tokens; throughput from 9 to 264 tokens/s. Latency is our own, measured from eu-paris.
Every measured endpoint, cheapest first
| Model | Provider | In /1M | Out /1M | Blended | TTFT p50 | Tok/s |
|---|---|---|---|---|---|---|
| deepseek-v4 | DeepSeek | $0.66 | $1.98 | $2.64 ← cheapest | 850 ms | 55 |
| deepseek-v4 | DeepInfra | $1.3 | $2.6 | $3.9 | 1208 ms | 36 |
| deepseek-v4 | Novita AI | $1.1088 | $3.3264 | $4.435 | 928 ms | 56 |
| deepseek-v4 | Alibaba | $1.122 | $3.366 | $4.488 | 1412 ms | 67 |
| glm-5.3 | DeepInfra | $1.2 | $4 | $5.2 | 800 ms | 116 |
| deepseek-v4 | Together AI | $1.32 | $3.96 | $5.28 | 693 ms | 98 |
| deepseek-v4 | Fireworks AI | $1.32 | $3.96 | $5.28 | 428 ms | 85 |
| glm-5.3 | Novita AI | $1.4 | $4.4 | $5.8 | 1392 ms | 45 |
| glm-5.3 | Together AI | $1.4 | $4.4 | $5.8 | 189 ms | 118 |
| glm-5.3 | Fireworks AI | $1.4 | $4.4 | $5.8 | 931 ms | 73 |
| qwen3.8-2.4t-a95b | Novita AI | $2 | $6 | $8 | 1330 ms | 41 |
| qwen3.8-2.4t-a95b | Alibaba | $2 | $6 | $8 | 1499 ms | 42 |
| qwen3.8-2.4t-a95b | DeepInfra | $2 | $6 | $8 | 801 ms | 72 |
| qwen3.8-2.4t-a95b | Together AI | $2 | $6 | $8 | 558 ms | 264 ← fastest |
| kimi-k3 | DeepInfra | $2.85 | $14.25 | $17.1 | 1746 ms | 9 |
| kimi-k3 | Together AI | $3 | $15 | $18 | 708 ms | 138 |
| kimi-k3 | Fireworks AI | $3 | $15 | $18 | 1832 ms | 47 |
| kimi-k3 | Moonshot AI | $3 | $15 | $18 | 2040 ms | 44 |
| kimi-k3 | Novita AI | $3 | $15 | $18 | 842 ms | 71 |
| kimi-k3 | Alibaba | $3.45 | $17.25 | $20.7 | 1850 ms | 31 |
Blended = input + output per 1M tokens, a coarse ranking aid rather than a bill — size your own workload on each endpoint page. TTFT = time to first token, p50 over 12 samples, measured from eu-paris on 2026-09-06.
The model sets the price, the host sets the speed
On price, the model still dominates here: switching model moves the bill by 6.5× between the cheapest host of each, against at most 2× between hosts of identical weights. On throughput the host dominates: 15.9× between hosts of one model, against 2.7× between the fastest host of each model — so the same weights, bought from two hosts, do not run at the same speed.
Read per model on the family pages, where the same figures are broken out by region:deepseek-v4 · glm-5.3 · kimi-k3 · qwen3.8-2.4t-a95b.
What is in this tier, and what is missing
| Model | Maker | Hosts measured | Hosts priced | Intelligence |
|---|---|---|---|---|
| deepseek-v4 | DeepSeek | 6 | 18 | — |
| glm-5.3 | Z.ai | 4 | 20 | 59.5 |
| kimi-k3 | Moonshot AI | 6 | 13 | 59.7 |
| qwen3.8-2.4t-a95b | Alibaba | 4 | 7 | — |
20 of the 58 priced endpoints in this tier carry a measured speed. The rest are in the price index but not in the benchmark — which says nothing about how fast they are, only that we have not called them. Intelligence, where shown, is the Artificial Analysis Intelligence Index, a third-party score we report but do not produce; several of these models have no published score, which is why the tier is not defined by it.
How the tier is defined
Each maker's main open-weight model (DeepSeek V4, Kimi K3, GLM-5.3, Qwen3.8 2.4T-A95B) — the maker's own top of range, not our ranking. That line is the makers' own: where a maker ships both a full model and a light one, it is the maker who names the light one “flash”. Moonshot publishes no light variant of Kimi K3, so that family appears in the flagship tier only. We do not decide which models are comparable on quality — this index publishes price, specifications and measured latency, and leaves capability judgments to the benchmarks that run them. The other half of the range is on chinese open-weight flash models.
Prices carry a source and a verification date on each endpoint page. Latency verified 2026-09-06. See also the speed-per-dollar board, why the same model has several prices, and the methodology.
FAQ
Which flagship Chinese model API is cheapest?
As of 2026-09-06, deepseek-v4 on DeepSeek at $2.64 per 1M tokens blended (input + output). It is not necessarily the fastest: measured from eu-paris it streams 55 tokens/s, against 264 for the fastest endpoint in this tier.
Does the provider matter more than the model?
On price, no. Switching model moves the bill by 6.5× between the cheapest host of each, while hosts of the same weights differ by at most 2×.
How is the speed measured?
First-party. Each endpoint is called with the same prompt, streamed, and its time to first token and output throughput recorded — 12 samples per endpoint per run, from eu-paris, and separately from us-east. All regions run at the same hour, because figures taken at different hours are not comparable. Method: stream, prompt=fixe, out=256 tok, samples=5/région, region=eu-paris.
What counts as a "flagship" model here?
Each maker's main open-weight model (DeepSeek V4, Kimi K3, GLM-5.3, Qwen3.8 2.4T-A95B) — the maker's own top of range, not our ranking. We do not rank these models on quality — this page compares price and measured speed only.