Methodology

One rule governs the whole index: every number is traceable to a source URL and a date it was verified. Quality is measured, never asserted; prices are cross-checked across independent sources and arbitrated by majority; anything we can't source, we don't print. This page explains exactly how each figure is built, so you can decide whether to trust it. It currently covers 539endpoints across 42 providers and 108distinct models.

Pricing: sources and majority arbitration

Each endpoint's headline price comes from a primary source — the provider's own pricing page, or the OpenRouter catalogue where it aggregates that provider. That figure is then cross-checked against independent secondary sources (LiteLLM, models.dev). Prices are per 1M tokens in USD. "Blended" means input + output added together — a coarse aid for ranking, not a workload estimate; each model page lets you size your own mix.

We never invent agreement: a secondary source only enters the comparison if it actually lists that model. When sources are present, the price carries one of four confidence levels, arbitrated on a tolerance of 20% blended:

Every model page shows the per-provider price, the source link, the confidence tag and the date it was verified — so a "diverges" flag is visible to you, not swept away.

Freshness and staleness

Each price stores the date it was last checked against its source. A compact freshness signal sits under the site banner on every page, recomputed on each rebuild — currently 539 prices, 100% verified within seven days. Category and ranking tables carry a per-row "Verified" column with the relative age of each price. Two thresholds keep stale numbers honest: past 45 days a price shows a "may be stale" badge, and past 90 days it is dropped from indexing entirely (noindex) — a doubtful number does more harm than a missing one.

Quality benchmarks

Quality is borrowed from primary leaderboards, each read on a single dated snapshot to avoid mixing noise from different sources. Scores attach to the model, not the provider — the same model scores the same wherever it's served, which is why the cheapest host of a given model is also its best value.

The human-preference score is LMArena (Chatbot Arena) Elo, read directly from the primary LMArena text leaderboard. A second, benchmark-derived measure is the Artificial Analysis Intelligence Index(0–100), from a public snapshot of artificialanalysis.ai. Model matches are curated by hand from a single leaderboard; ambiguous variants are dropped rather than guessed. Embeddings use the MTEB multilingual mean task score on one dated snapshot. The value board ranks 52 models on Elo; the intelligence board 56 on the AA index; the embeddings board 15 models (4 priced).

Speed: first-party measurement

Unlike the borrowed benchmarks above, throughput is measured here directly: each reachable endpoint is streamed from a European host with a fixed prompt, and its output tokens per second recorded, with the region and date shown per row. This is the one figure we generate ourselves rather than cite.

Scope: every figure is a real first-party measurement (five streamed samples per endpoint from a Paris host, p50/p95 reported), but coverage is still limited to the directly-reachable endpoints we can measure (11 on the speed board). The speed-per-dollar ranking grows more complete as the benchmark reaches more models; where an endpoint isn't yet measured, lean on the sourced price and quality figures.

Derived metrics

The rankings combine the figures above by simple, stated arithmetic — no black box.Value per dollar = model Elo ÷ blended price. Intelligence per dollar = AA Index ÷ blended price. Speed per dollar = measured tokens/s ÷ blended price. Because quality is identical across providers, the cheapest endpoint of a model is also its best value — which is exactly why top-of-leaderboard models are rarely top value. A model is labelled frontier when its AA Intelligence Index is ≥ 55.

Licenses, context and cutoffs

Open-weight vs. proprietary status, context-window size, maximum output and knowledge-cutoff dates are curated per model, each with its own source link, and cross-checked on the web before publishing. For open-weight models we also estimate self-host memory: VRAM ≈ parameters × bytes-per-weight (0.5 at int4, 1.0 at int8, 2.0 at fp16) × 1.2 overhead for KV cache and activations, which feeds the self-host break-even calculator.

What this index does not do

It invents nothing: a figure without a source and a date does not get printed. It is an index, not a gateway — we don't route your traffic or resell tokens, so a provider's ranking is never for sale. Affiliate links, where present, are disclosed and never change the order of any table; rankings are computed purely from the sourced numbers. All data on the site is free to reuse under CC BY 4.0with attribution.