Chinese open-weight flash models: price and measured speed
Open weights mean the same model is sold by several hosts at different prices and different speeds. This page puts 13 measured endpoints — 3 flash models from DeepSeek, Z.ai and Alibaba, served by 6 providers — in a single table, priced and timed. Blended price runs from$0.26 to $1.637 per 1M tokens; throughput from 28 to 135 tokens/s. Latency is our own, measured from eu-paris.
Every measured endpoint, cheapest first
| Model | Provider | In /1M | Out /1M | Blended | TTFT p50 | Tok/s |
|---|---|---|---|---|---|---|
| deepseek-v4-flash | DeepInfra | $0.08 | $0.18 | $0.26 ← cheapest | 1296 ms | 28 |
| glm-5.3-flash | Novita AI | $0.075 | $0.25 | $0.325 | 941 ms | 66 |
| glm-5.3-flash | DeepInfra | $0.075 | $0.25 | $0.325 | 1160 ms | 39 |
| deepseek-v4-flash | Together AI | $0.14 | $0.28 | $0.42 | 635 ms | 82 |
| qwen3.8-flash | Alibaba | $0.15 | $0.47 | $0.62 | 953 ms | 103 |
| qwen3.8-flash | Together AI | $0.15 | $0.47 | $0.62 | 2199 ms | 135 ← fastest |
| qwen3.8-flash | Novita AI | $0.15 | $0.47 | $0.62 | 857 ms | 100 |
| glm-5.3-flash | Fireworks AI | $0.15 | $0.5 | $0.65 | 799 ms | 95 |
| glm-5.3-flash | Together AI | $0.15 | $0.5 | $0.65 | 606 ms | 85 |
| deepseek-v4-flash | Fireworks AI | $0.22 | $0.66 | $0.88 | 370 ms | 117 |
| deepseek-v4-flash | DeepSeek | $0.22 | $0.66 | $0.88 | 431 ms | 115 |
| deepseek-v4-flash | Alibaba | $0.352 | $1.056 | $1.408 | 960 ms | 111 |
| deepseek-v4-flash | Novita AI | $0.4092 | $1.2276 | $1.637 | 1079 ms | 115 |
Blended = input + output per 1M tokens, a coarse ranking aid rather than a bill — size your own workload on each endpoint page. TTFT = time to first token, p50 over 12 samples, measured from eu-paris on 2026-09-06.
The host moves the numbers more than the model
Hold the model constant and change only the host: the price of the very same weights varies by up to 6.3×. Change the model instead, comparing each one at its own cheapest host, and you move the bill by 2.4×. The decision that costs you money is which endpoint you call, not which model you pick. On throughput the host dominates: 4.2× between hosts of one model, against 1.4× between the fastest host of each model — so the same weights, bought from two hosts, do not run at the same speed.
Read per model on the family pages, where the same figures are broken out by region:deepseek-v4-flash · glm-5.3-flash · qwen3.8-flash.
What is in this tier, and what is missing
| Model | Maker | Hosts measured | Hosts priced | Intelligence |
|---|---|---|---|---|
| deepseek-v4-flash | DeepSeek | 6 | 22 | 51.8 |
| glm-5.3-flash | Z.ai | 4 | 19 | 57.5 |
| qwen3.8-flash | Alibaba | 3 | 3 | — |
13 of the 44 priced endpoints in this tier carry a measured speed. The rest are in the price index but not in the benchmark — which says nothing about how fast they are, only that we have not called them. Intelligence, where shown, is the Artificial Analysis Intelligence Index, a third-party score we report but do not produce; several of these models have no published score, which is why the tier is not defined by it.
How the tier is defined
The light variant each maker publishes under its own “flash” name (DeepSeek V4 Flash, GLM-5.3 Flash, Qwen3.8 Flash) — the maker's own naming, not our ranking. That line is the makers' own: where a maker ships both a full model and a light one, it is the maker who names the light one “flash”. Moonshot publishes no light variant of Kimi K3, so that family appears in the flagship tier only. We do not decide which models are comparable on quality — this index publishes price, specifications and measured latency, and leaves capability judgments to the benchmarks that run them. The other half of the range is on chinese open-weight flagship models.
Prices carry a source and a verification date on each endpoint page. Latency verified 2026-09-06. See also the speed-per-dollar board, why the same model has several prices, and the methodology.
FAQ
Which flash Chinese model API is cheapest?
As of 2026-09-06, deepseek-v4-flash on DeepInfra at $0.26 per 1M tokens blended (input + output). It is not necessarily the fastest: measured from eu-paris it streams 28 tokens/s, against 135 for the fastest endpoint in this tier.
Does the provider matter more than the model?
On price, yes. Holding the model constant, hosts of the same weights differ by up to 6.3×, while switching model — comparing each model at its cheapest host — moves the bill by 2.4×. The endpoint you pick matters more than the model you pick.
How is the speed measured?
First-party. Each endpoint is called with the same prompt, streamed, and its time to first token and output throughput recorded — 12 samples per endpoint per run, from eu-paris, and separately from us-east. All regions run at the same hour, because figures taken at different hours are not comparable. Method: stream, prompt=fixe, out=256 tok, samples=5/région, region=eu-paris.
What counts as a "flash" model here?
The light variant each maker publishes under its own “flash” name (DeepSeek V4 Flash, GLM-5.3 Flash, Qwen3.8 Flash) — the maker's own naming, not our ranking. We do not rank these models on quality — this page compares price and measured speed only.