The efficient frontier: when the big model is worth it
Every scored model on one chart: price againstmeasured intelligence. Only 7 of57 models sit on the Pareto frontier — for all the rest, something else is both cheaper and smarter. The frontier answers the real question: the last few intelligence points cost about 12781× more per point than the first ones. Pay for them when your task needs them; below the curve, you pay for nothing.
Ringed dots are Pareto-optimal; small dots are dominated (a cheaper, smarter alternative exists). Hover for details, click for the model's page.
The optimal picks, and the price of each extra point
Walking up the frontier from cheapest to smartest — what each step costs per benchmark point gained.
| Model | Lab | Score | /1M | $ per extra point |
|---|---|---|---|---|
| gpt-oss-20b (AkashML) | OpenAI | 15.2 | $0.12 | — |
| DeepSeek V4 Flash 0731 (Baidu) | DeepSeek | 51.8 | $0.15 | $0.001 |
| GLM 5.3 Flash (Relace) | Z.ai | 57.5 | $0.309 | $0.028 |
| Gemini 3.8 Flash (Google) | 58.7 | $4.5 | $3.49 | |
| GLM 5.3 (Reka) | Z.ai | 59.5 | $4.65 | $0.19 |
| Grok 4.6 (xAI) | xAI | 60.9 | $8 | $2.39 |
| Claude Opus 5 (Anthropic) | Anthropic | 63 | $30 | $10.48 |
Intelligence = Artificial Analysis Intelligence Index (0–100); price = cheapest live blended offer per model, verified 2026-09-05. The benchmark is an aggregate — a dominated model can still win on your specific task, latency, or context needs. See alsothe value ranking,the price of the frontier over time andthe age discount ·methodology.
FAQ
When is a frontier model worth paying for?
When your task needs the last few benchmark points — and only then. On the current frontier, going from DeepSeek V4 Flash 0731 (Baidu) to Claude Opus 5 (Anthropic) costs about $10.48 per extra intelligence point versus $0.001 at the cheap end — roughly 12781× more per point. For tasks a mid model already handles, the premium buys nothing.
What does Pareto-optimal mean for an LLM API?
A model is Pareto-optimal when no other model is both cheaper and higher-scoring. Only 7 of the 57 scored models qualify — every other model is dominated: something else offers more intelligence for less money.
Should I always pick a model on the frontier?
On this metric, yes — a dominated model gives you less score per dollar by definition. But the benchmark doesn't capture everything: latency, context window, modalities, licensing, or strength on your specific task can justify an off-frontier pick. Use the frontier as the default, and pay off it only for a reason you can name.