The efficient frontier: when the big model is worth it

Every scored model on one chart: price againstmeasured intelligence. Only 7 of57 models sit on the Pareto frontier — for all the rest, something else is both cheaper and smarter. The frontier answers the real question: the last few intelligence points cost about 12781× more per point than the first ones. Pay for them when your task needs them; below the curve, you pay for nothing.

$0.1$0.3$1$3$10$30$60102030405060gpt-oss-20b (AkashML) — score 15.2 at $0.12/1M · Pareto-optimalDeepSeek V4 Flash 0731 (Baidu) — score 51.8 at $0.15/1M · Pareto-optimalgpt-oss-120b (AkashML) — score 24.1 at $0.2/1M · dominatedMinistral 3 3B 2512 (Mistral) — score 7.1 at $0.2/1M · dominatedGemma 3 27B (DeepInfra) — score 7.4 at $0.24/1M · dominatedGranite 4.2 8B (CoreWeave) — score 19.6 at $0.25/1M · dominatedNemotron 3.5 Lightning (DeepInfra) — score 23.6 at $0.28/1M · dominatedMinistral 3 8B 2512 (Mistral) — score 9 at $0.3/1M · dominatedGLM 5.3 Flash (Relace) — score 57.5 at $0.309/1M · Pareto-optimalGemma 4 26B A4B (Cloudflare) — score 26.1 at $0.4/1M · dominatedMinistral 3 14B 2512 (Mistral) — score 11.2 at $0.4/1M · dominatedLlama 4 Scout (DeepInfra) — score 10.3 at $0.4/1M · dominatedGemma 4 31B (CoreWeave) — score 29.7 at $0.44/1M · dominatedGLM 4.7 Flash (Venice) — score 23.3 at $0.46/1M · dominatedNemotron 3 Super (DeepInfra) — score 25.7 at $0.485/1M · dominatedDeepSeek V3.2 (GMICloud) — score 25.1 at $0.518/1M · dominatedGPT-5.6 Luna (OpenAI) — score 51.2 at $0.7/1M · dominatedGPT-4o mini — score 6.7 at $0.75/1M · dominatedLlama 4 Maverick (DigitalOcean) — score 14.5 at $0.896/1M · dominatedGLM 4.5 Air (Novita) — score 16.7 at $0.98/1M · dominatedQwen3.6 35B A3B (AkashML) — score 32.1 at $1/1M · dominatedQwen3.5-35B-A3B (DeepInfra) — score 29.9 at $1.14/1M · dominatedMiniMax M3 (CoreWeave) — score 45.4 at $1.19/1M · dominatedDeepSeek V3.1 (DeepInfra) — score 21.4 at $1.2/1M · dominatedDeepSeek V3 (DeepInfra) — score 14.2 at $1.21/1M · dominatedMiniMax M2.5 (Venice) — score 34.5 at $1.22/1M · dominatedMiniMax M2.7 (Novita) — score 38.9 at $1.35/1M · dominatedGLM 5.2 (Baidu) — score 52.6 at $1.479/1M · dominatedQwen3.5-27B (Alibaba) — score 34.6 at $1.755/1M · dominatedGPT-4.1 mini — score 14.8 at $2/1M · dominatedGLM 4.7 (DeepInfra) — score 34.5 at $2.15/1M · dominatedGLM 4.6 (Venice) — score 23.4 at $2.18/1M · dominatedQwen3.6 27B (Chutes) — score 37.7 at $2.3/1M · dominatedQwen3.5-122B-A10B (SiliconFlow) — score 32.9 at $2.34/1M · dominatedMistral Medium 3 (Mistral) — score 12.5 at $2.4/1M · dominatedQwen3.8 27B (Parasail) — score 52 at $2.44/1M · dominatedGLM 5 (GMICloud) — score 40.5 at $2.52/1M · dominatedMiniMax M1 (Minimax) — score 17.9 at $2.6/1M · dominatedDeepSeek R1 (DeepInfra) — score 20.4 at $2.65/1M · dominatedNemotron 3 Ultra (DeepInfra) — score 38.3 at $2.7/1M · dominatedKimi K2.5 (SiliconFlow) — score 36 at $2.7/1M · dominatedQwen3.5 397B A17B (Alibaba) — score 34.3 at $2.73/1M · dominatedKimi K2.6 (Baidu) — score 45.1 at $2.856/1M · dominatedGLM 5.1 (Baidu) — score 41 at $3.764/1M · dominatedKimi K2.7 Code (DeepInfra) — score 43 at $4.08/1M · dominatedGemini 3.8 Flash (Google) — score 58.7 at $4.5/1M · Pareto-optimalGemini 3.6 Flash — score 51.6 at $4.5/1M · dominatedGLM 5.3 (Reka) — score 59.5 at $4.65/1M · Pareto-optimalGPT-5.6 Sol (OpenAI) — score 58.9 at $6/1M · dominatedGPT-5.6 Terra (OpenAI) — score 55 at $7/1M · dominatedGrok 4.6 (xAI) — score 60.9 at $8/1M · Pareto-optimalMistral Medium 3.5 (Mistral) — score 30.4 at $9/1M · dominatedClaude Sonnet 5 (Anthropic) — score 55.3 at $12/1M · dominatedGPT-4o — score 11.1 at $12.5/1M · dominatedKimi K3 (DeepInfra) — score 59.7 at $17.1/1M · dominatedClaude Opus 5 (Anthropic) — score 63 at $30/1M · Pareto-optimalClaude Fable 5 (Anthropic) — score 62.1 at $60/1M · dominatedintelligence score ↑ · blended $/1M (log) →

Ringed dots are Pareto-optimal; small dots are dominated (a cheaper, smarter alternative exists). Hover for details, click for the model's page.

The optimal picks, and the price of each extra point

Walking up the frontier from cheapest to smartest — what each step costs per benchmark point gained.

ModelLabScore/1M$ per extra point
gpt-oss-20b (AkashML)OpenAI15.2$0.12
DeepSeek V4 Flash 0731 (Baidu)DeepSeek51.8$0.15$0.001
GLM 5.3 Flash (Relace)Z.ai57.5$0.309$0.028
Gemini 3.8 Flash (Google)Google58.7$4.5$3.49
GLM 5.3 (Reka)Z.ai59.5$4.65$0.19
Grok 4.6 (xAI)xAI60.9$8$2.39
Claude Opus 5 (Anthropic)Anthropic63$30$10.48

Intelligence = Artificial Analysis Intelligence Index (0–100); price = cheapest live blended offer per model, verified 2026-09-05. The benchmark is an aggregate — a dominated model can still win on your specific task, latency, or context needs. See alsothe value ranking,the price of the frontier over time andthe age discount ·methodology.

FAQ

When is a frontier model worth paying for?

When your task needs the last few benchmark points — and only then. On the current frontier, going from DeepSeek V4 Flash 0731 (Baidu) to Claude Opus 5 (Anthropic) costs about $10.48 per extra intelligence point versus $0.001 at the cheap end — roughly 12781× more per point. For tasks a mid model already handles, the premium buys nothing.

What does Pareto-optimal mean for an LLM API?

A model is Pareto-optimal when no other model is both cheaper and higher-scoring. Only 7 of the 57 scored models qualify — every other model is dominated: something else offers more intelligence for less money.

Should I always pick a model on the frontier?

On this metric, yes — a dominated model gives you less score per dollar by definition. But the benchmark doesn't capture everything: latency, context window, modalities, licensing, or strength on your specific task can justify an off-frontier pick. Use the frontier as the default, and pay off it only for a reason you can name.