LLM API cost guides
Short, data-derived reads on where an LLM API bill actually goes — each one computed live from the price index, not hand-written opinion.
The falling price of intelligence
How much measured reasoning $1 buys, generation after generation — one to two orders of magnitude gained in two years, drawn on one chart.
The price trend: every model by release date
One dot per model — release date against current cheapest price, log scale, colored by lab. Each generation ships cheaper; the whole market in one look.
The efficient frontier: when the big model is worth it
Price against intelligence on one chart, the Pareto frontier drawn through it, and what each extra benchmark point costs — when to pay for frontier, when you're paying for nothing.
The price of the frontier
Every model that set a new intelligence record, and what each costs today — the record has been held at twenty cents and at sixty dollars per million tokens.
The age discount: what last season's model costs
Models under six months cost several times more than models past two years — how the discount accrues, and when an older model is the smarter buy.
How fast models get replaced
Each lab's release timeline, its median days between two models, and the accelerating market-wide release rate — how long a model really stays current.
Best value LLM APIs: quality per dollar
Which model gives the most measured quality for each dollar — cheap and good, not just cheap.
Best intelligence per dollar (2nd benchmark)
A second value ranking, from the Artificial Analysis Intelligence Index — aggregate reasoning score per dollar, cross-checking the Elo board.
Best embedding model APIs (MTEB per dollar)
Embedding models ranked on a dated MTEB snapshot, then by quality per dollar for the ones with a sourced API price — the rest ship open weights.
Same model, different price
The same open model is hosted by many providers at very different prices. Where the choice matters most.
Output tokens are the expensive ones
Generated tokens cost several times more than supplied ones, so trimming responses beats trimming prompts.
Prompt caching: who offers it, what it saves
Reused input tokens billed at a fraction of the normal price — coverage, real savings, worked example.
Biggest context windows, and their price
How much you can feed each model in one call — the largest windows and what the cheapest host charges.
Truly open weights vs. open-but-restricted
"Open weight" ranges from Apache-2.0 to non-commercial clauses — which models you can actually self-host and ship.
Knowledge cutoff dates by model
How current each model's training data is — the freshest knowledge and what the cheapest host charges for it.
Figures refresh whenever the index rebuilds. Every number traces back to a provider's published price.