Biggest context windows, and the price of using them
A model's context window caps how much you can feed it in one call — the whole point of long-context work like RAG, big-codebase reasoning, or long documents. Across 108 models here, the median window is259k tokens and the largest reaches1311k. Window size is a property of the model, identical whichever provider serves it — so the price shown is each model's cheapest host.
Cheapest long context: deepseek-v4-flash gives you a1,311k-token window at just$0.15/1M blended (Baidu).
Long windows aren't free to use: cost scales with tokens sent, and the output premium still applies — see output vs inputand prompt caching for how to keep a big-context bill down.
How many models reach each size
Share of models whose context window is at least this large.
Does a bigger window cost more?
Every model plotted by window size against its cheapest blended price (both log scales). Less correlated than you'd expect: million-token windows exist at budget prices, and some small-window flagships are the dearest dots here. Click a brand to isolate it.
Largest context windows
Ranked by window size; where two match, the cheaper host is listed first.
| Model | Context | Max output | Cheapest host | Blended /1M |
|---|---|---|---|---|
| deepseek-v4-flash | 1,311k | 131,072 | Baidu | $0.15 |
| glm-5.3-flash | 1,311k | 131,072 | Relace | $0.309 |
| glm-5.3 | 1,311k | 235,929 | Reka | $4.65 |
| deepseek-v4-flash-vision-exp | 1,049k | 131,072 | DeepInfra | $0.862 |
| llama-4-maverick | 1,049k | 115,200 | DigitalOcean | $0.896 |
| minimax-m3 | 1,049k | 235,929 | CoreWeave | $1.19 |
| deepseek-v4 | 1,049k | 393,216 | Alibaba | $2.323 |
| gemini-3.5-flash-lite | 1,049k | 65,536 | $2.8 | |
| gemini-3.6-flash | 1,049k | 65,536 | $4.5 | |
| gemini-3.8-flash | 1,049k | 65,536 | $4.5 | |
| gemini-3.7-flash | 1,049k | 65,536 | $4.5 | |
| qwen3.8-2.4t-a95b | 1,049k | 131,072 | Novita AI | $8 |
| kimi-k3 | 1,049k | 16,384 | DeepInfra | $17.1 |
| gpt-4.1-mini | 1,048k | 32,768 | OpenAI | $2 |
| nemotron-3-super-120b-a12b | 1,000k | 16,384 | DeepInfra | $0.485 |
| qwen3.8-27b | 1,000k | 235,929 | Parasail | $2.44 |
| minimax-m1 | 1,000k | 40,000 | Minimax | $2.6 |
| claude-fable-5-1 | 1,000k | 128,000 | Anthropic | $60 |
| gpt-6-astra | 1,000k | 128,000 | OpenAI | $60 |
| grok-4.5 | 500k | 32,768 | xAI | $8 |
Context in tokens; prices per 1M tokens (USD), from each model's cheapest published rate. Verified 2026-09-05.
FAQ
What is the largest LLM context window available?
The largest window in this index reaches 1311k tokens, against a median of 259k across 108 models. Window size caps how much you can send in a single call.
What's the cheapest way to get a long context window?
deepseek-v4-flash offers a 1,311k-token window at just $0.15 per 1M tokens blended via Baidu — the cheapest long-context option here.
Does the provider change a model's context window?
No — window size is a property of the model, identical whichever provider serves it. So you can route to the cheapest host without losing context length. What does vary with usage is cost, which scales with tokens sent.