Best LLM APIs by speed per dollar

A third take on value — this one measured, not borrowed. This ranks11 endpoints by output throughput(tokens per second, streamed from our own EU benchmark) divided by each endpoint's blended price per 1M tokens. Where the Elo andintelligence boards use public leaderboards, the speed figure here is measured directly — each row carries the region and date it was recorded.

Ranked by speed per dollar

#EndpointTokens/sBlended /1MTok/s per $
1gpt-oss-20b · Groq · eu-paris 2026-09-041,987$0.3755,298.7
2gpt-oss-120b · Cerebras · eu-paris 2026-09-042,240.1$1.12,036.5
3gemini-3.6-flash · Google · eu-paris 2026-09-047,729$4.51,717.6
4gpt-oss-120b · Groq · eu-paris 2026-09-04776$0.751,034.7
5gemini-3.8-flash · Google · eu-paris 2026-09-043,845.2$4.5854.5
6mistral-small · Mistral AI · eu-paris 2026-09-04142.1$0.75189.5
7qwen3.8-27b · Groq · eu-paris 2026-09-04496.6$4.8103.5
8command-r · Cohere · eu-paris 2026-09-0439.6$0.7552.8
9mistral-large · Mistral AI · eu-paris 2026-09-0470.3$235.1
10mistral-medium · Mistral AI · eu-paris 2026-09-04118.2$913.1
11command-a · Cohere · eu-paris 2026-09-0449.7$12.54

Speed per dollar = measured output tokens/s ÷ (input + output price per 1M tokens). Throughput is provider-specific, so — unlike Elo or the intelligence index — the same model can rank differently across hosts. Blended is a coarse ranking aid, not a workload estimate; size your own on each model's page.

For contrast: raw throughput

The fastest endpoints by measured tokens/s alone — price aside.

Speed rankEndpointTokens/sBlended /1MTok/s per $
1gemini-3.6-flash · Google7,729$4.51,717.6
2gemini-3.8-flash · Google3,845.2$4.5854.5
3gpt-oss-120b · Cerebras2,240.1$1.12,036.5
4gpt-oss-20b · Groq1,987$0.3755,298.7
5gpt-oss-120b · Groq776$0.751,034.7
6qwen3.8-27b · Groq496.6$4.8103.5
7mistral-small · Mistral AI142.1$0.75189.5
8mistral-medium · Mistral AI118.2$913.1
9mistral-large · Mistral AI70.3$235.1
10command-a · Cohere49.7$12.54

First-party throughput benchmark · latency data verified 2026-09-05. Coverage is limited to endpoints we can call directly and grows as the benchmark runs across more models. Cross-check value from a different angle on theElo and intelligence boards, or read how it's all measured in the methodology.

FAQ

Which LLM API gives the most speed per dollar?

As of 2026-09-05, gpt-oss-20b on Groq leads: 5,298.7 tokens/s per dollar — 1,987 tokens/s measured at $0.375 per 1M tokens (blended input + output).

How is throughput measured?

First-party: each endpoint is streamed from an EU host and its output tokens/s recorded, with the region and date shown per row. Unlike the Elo and intelligence rankings — which borrow public benchmarks — the speed number here is measured directly. Method: stream, prompt=fixe, out=256 tok, samples=5/région, region=eu-paris.

Is the fastest model the best value on speed?

Not always. The raw throughput leader is gemini-3.6-flash on Google at 7,729 tokens/s, but speed per dollar also weighs price — so a slightly slower, much cheaper endpoint can rank higher. The leader by speed per dollar is gpt-oss-20b.