Biggest context windows, and the price of using them

A model's context window caps how much you can feed it in one call — the whole point of long-context work like RAG, big-codebase reasoning, or long documents. Across 108 models here, the median window is259k tokens and the largest reaches1311k. Window size is a property of the model, identical whichever provider serves it — so the price shown is each model's cheapest host.

Cheapest long context: deepseek-v4-flash gives you a1,311k-token window at just$0.15/1M blended (Baidu).

Long windows aren't free to use: cost scales with tokens sent, and the output premium still applies — see output vs inputand prompt caching for how to keep a big-context bill down.

How many models reach each size

≥ 32k100% of models
≥ 128k95% of models
≥ 200k64% of models
≥ 1,000k18% of models

Share of models whose context window is at least this large.

Does a bigger window cost more?

Every model plotted by window size against its cheapest blended price (both log scales). Less correlated than you'd expect: million-token windows exist at budget prices, and some small-window flagships are the dearest dots here. Click a brand to isolate it.

$0.1$0.3$1$3$10$30$6032k64k128k256k512k1M$ / 1M tokens (blended, log)Qwen 2.5 72B (DeepInfra) — 131k context at $0.76/1M · cheapest via DeepInfraQwen3 Next 80B A3B Instruct (Alibaba) — 262k context at $0.878/1M · cheapest via AlibabaQwen3 30B A3B Instruct 2507 (SiliconFlow) — 262k context at $0.39/1M · cheapest via SiliconFlowQwen3 14B (DeepInfra) — 131k context at $0.36/1M · cheapest via DeepInfraQwen3 32B (DeepInfra) — 131k context at $0.36/1M · cheapest via DeepInfraQwen3 235B A22B Instruct 2507 (GMICloud) — 262k context at $0.438/1M · cheapest via GMICloudQwen3 Coder (DeepInfra) — 262k context at $1.3/1M · cheapest via DeepInfraQwen3 235B A22B Thinking 2507 (Alibaba) — 131k context at $2.53/1M · cheapest via AlibabaQwen3 Coder 30B A3B Instruct (Novita) — 262k context at $0.34/1M · cheapest via Novita AIQwen3 Next 80B A3B Thinking (Google) — 262k context at $1.35/1M · cheapest via GoogleQwen3 VL 235B A22B Instruct (DeepInfra) — 262k context at $1.08/1M · cheapest via DeepInfraQwen3 VL 235B A22B Thinking (Alibaba) — 131k context at $4.4/1M · cheapest via AlibabaQwen3 VL 30B A3B Instruct (Alibaba) — 262k context at $0.65/1M · cheapest via AlibabaQwen3 VL 30B A3B Thinking (SiliconFlow) — 262k context at $1.29/1M · cheapest via SiliconFlowQwen3 VL 8B Instruct (Alibaba) — 262k context at $0.572/1M · cheapest via AlibabaQwen3 Coder Next (Parasail) — 262k context at $0.92/1M · cheapest via ParasailQwen3.5 397B A17B (Alibaba) — 262k context at $2.73/1M · cheapest via AlibabaQwen3.5-35B-A3B (DeepInfra) — 262k context at $1.14/1M · cheapest via DeepInfraQwen3.5-27B (Alibaba) — 262k context at $1.755/1M · cheapest via AlibabaQwen3.5-9B (SiliconFlow) — 262k context at $0.25/1M · cheapest via SiliconFlowQwen3.5-122B-A10B (SiliconFlow) — 262k context at $2.34/1M · cheapest via SiliconFlowQwen3.6 35B A3B (AkashML) — 262k context at $1/1M · cheapest via AkashMLQwen3.6 27B (Chutes) — 262k context at $2.3/1M · cheapest via ChutesQwen3.8 2.4T A95B (Novita) — 1049k context at $8/1M · cheapest via Novita AIQwen3.8 27B (Parasail) — 1000k context at $2.44/1M · cheapest via ParasailMixtral 8x22B Instruct (Mistral) — 66k context at $8/1M · cheapest via Mistral AIMistral Nemo (DeepInfra) — 131k context at $0.049/1M · cheapest via DeepInfraMistral Large 3 (Mistral) — 131k context at $2/1M · cheapest via Mistral AISaba (Mistral) — 33k context at $0.8/1M · cheapest via Mistral AIMistral Medium (Mistral) — 131k context at $9/1M · cheapest via Mistral AIMistral Medium 3 (Mistral) — 131k context at $2.4/1M · cheapest via Mistral AIMistral Small 3.2 24B (DeepInfra) — 131k context at $0.275/1M · cheapest via DeepInfraCodestral 2508 (Mistral) — 256k context at $1.2/1M · cheapest via Mistral AIMistral Medium 3.1 (Mistral) — 131k context at $2.4/1M · cheapest via Mistral AIVoxtral Small 24B 2507 (Mistral) — 33k context at $0.4/1M · cheapest via Mistral AIMinistral 3 14B 2512 (Mistral) — 262k context at $0.4/1M · cheapest via Mistral AIMinistral 3 8B 2512 (Mistral) — 262k context at $0.3/1M · cheapest via Mistral AIMinistral 3 3B 2512 (Mistral) — 131k context at $0.2/1M · cheapest via Mistral AIMistral Small (Mistral) — 131k context at $0.75/1M · cheapest via Mistral AIMistral Medium 3.5 (Mistral) — 262k context at $9/1M · cheapest via Mistral AIDeepSeek Chat (DeepInfra) — 64k context at $1.21/1M · cheapest via DeepInfraDeepSeek V3 (DeepInfra) — 164k context at $1.21/1M · cheapest via DeepInfraDeepSeek V3 0324 (DeepInfra) — 164k context at $1.14/1M · cheapest via DeepInfraDeepSeek R1 (DeepInfra) — 128k context at $2.65/1M · cheapest via DeepInfraDeepSeek V3.1 (DeepInfra) — 164k context at $1.2/1M · cheapest via DeepInfraDeepSeek V3.1 Terminus (AtlasCloud) — 164k context at $1.25/1M · cheapest via AtlasCloudDeepSeek V3.2 Exp (SiliconFlow) — 164k context at $0.68/1M · cheapest via SiliconFlowDeepSeek V3.2 (GMICloud) — 164k context at $0.518/1M · cheapest via GMICloudDeepSeek V4 Pro 0813 (Alibaba) — 1049k context at $2.323/1M · cheapest via AlibabaDeepSeek V4 Flash 0731 (Baidu) — 1311k context at $0.15/1M · cheapest via BaiduDeepSeek V4 Flash Vision Exp (DeepInfra) — 1049k context at $0.862/1M · cheapest via DeepInfraGLM 4.5 Air (Novita) — 131k context at $0.98/1M · cheapest via Novita AIGLM 4.5V (Novita) — 66k context at $2.4/1M · cheapest via Novita AIGLM 4.6 (Venice) — 205k context at $2.18/1M · cheapest via VeniceGLM 4.6V (Novita) — 131k context at $1.2/1M · cheapest via Novita AIGLM 4.7 (DeepInfra) — 205k context at $2.15/1M · cheapest via DeepInfraGLM 4.7 Flash (Venice) — 203k context at $0.46/1M · cheapest via VeniceGLM 5 (GMICloud) — 205k context at $2.52/1M · cheapest via GMICloudGLM 5.1 (Baidu) — 205k context at $3.764/1M · cheapest via BaiduGLM 5.2 (Baidu) — 1049k context at $1.479/1M · cheapest via BaiduGLM 5.3 (Reka) — 1311k context at $4.65/1M · cheapest via RekaGLM 5.3 Flash (Relace) — 1311k context at $0.309/1M · cheapest via RelaceGPT-4o — 128k context at $12.5/1M · cheapest via OpenAIGPT-4o mini — 128k context at $0.75/1M · cheapest via OpenAIGPT-4.1 mini — 1048k context at $2/1M · cheapest via OpenAIgpt-oss-20b (AkashML) — 131k context at $0.12/1M · cheapest via AkashMLgpt-oss-120b (AkashML) — 131k context at $0.2/1M · cheapest via AkashMLGPT-5.6 Sol (OpenAI) — 400k context at $6/1M · cheapest via OpenAIGPT-5.6 Terra (OpenAI) — 400k context at $7/1M · cheapest via OpenAIGPT-5.6 Luna (OpenAI) — 400k context at $0.7/1M · cheapest via OpenAIGPT-6 Astra (OpenAI) — 1000k context at $60/1M · cheapest via OpenAIGemma 3 27B (DeepInfra) — 131k context at $0.24/1M · cheapest via DeepInfraGemma 4 31B (CoreWeave) — 262k context at $0.44/1M · cheapest via CoreWeaveGemma 4 26B A4B (Cloudflare) — 262k context at $0.4/1M · cheapest via CloudflareGemini 3.6 Flash — 1049k context at $4.5/1M · cheapest via GoogleGemini 3.5 Flash Lite (Google) — 1049k context at $2.8/1M · cheapest via GoogleGemini 3.7 Flash (Google) — 1049k context at $4.5/1M · cheapest via GoogleGemini 3.8 Flash (Google) — 1049k context at $4.5/1M · cheapest via GoogleLlama 3.1 8B (DeepInfra) — 131k context at $0.06/1M · cheapest via DeepInfraLlama 3.1 70B Instruct (DeepInfra) — 131k context at $0.8/1M · cheapest via DeepInfraLlama 3.2 3B Instruct (Parasail) — 131k context at $0.38/1M · cheapest via ParasailLlama 3.3 70B (DeepInfra) — 131k context at $0.42/1M · cheapest via DeepInfraLlama 4 Scout (DeepInfra) — 328k context at $0.4/1M · cheapest via DeepInfraLlama 4 Maverick (DigitalOcean) — 1049k context at $0.896/1M · cheapest via DigitalOceanMiniMax M1 (Minimax) — 1000k context at $2.6/1M · cheapest via MinimaxMiniMax M2 (Minimax) — 205k context at $1.275/1M · cheapest via MinimaxMiniMax M2.1 (Novita) — 205k context at $1.5/1M · cheapest via Novita AIMiniMax M2.5 (Venice) — 205k context at $1.22/1M · cheapest via VeniceMiniMax M2.7 (Novita) — 205k context at $1.35/1M · cheapest via Novita AIMiniMax M3 (CoreWeave) — 1049k context at $1.19/1M · cheapest via CoreWeaveKimi K2 Thinking (Google) — 262k context at $3.1/1M · cheapest via GoogleKimi K2.5 (SiliconFlow) — 262k context at $2.7/1M · cheapest via SiliconFlowKimi K2.6 (Baidu) — 262k context at $2.856/1M · cheapest via BaiduKimi K2.7 Code (DeepInfra) — 262k context at $4.08/1M · cheapest via DeepInfraKimi K3 (DeepInfra) — 1049k context at $17.1/1M · cheapest via DeepInfraClaude Fable 5 (Anthropic) — 200k context at $60/1M · cheapest via AnthropicClaude Sonnet 5 (Anthropic) — 200k context at $12/1M · cheapest via AnthropicClaude Opus 5 (Anthropic) — 200k context at $30/1M · cheapest via AnthropicClaude Fable 5.1 (Anthropic) — 1000k context at $60/1M · cheapest via AnthropicNemotron 3 Nano 30B A3B (Crusoe) — 262k context at $0.25/1M · cheapest via CrusoeNemotron 3 Super (DeepInfra) — 1000k context at $0.485/1M · cheapest via DeepInfraNemotron 3 Ultra (DeepInfra) — 262k context at $2.7/1M · cheapest via DeepInfraNemotron 3.5 Lightning (DeepInfra) — 262k context at $0.28/1M · cheapest via DeepInfraCommand R (Cohere) — 128k context at $0.75/1M · cheapest via CohereCommand A (Cohere) — 256k context at $12.5/1M · cheapest via CohereGrok 4.5 (xAI) — 500k context at $8/1M · cheapest via xAIGrok 4.6 (xAI) — 256k context at $8/1M · cheapest via xAIGranite 4.2 8B (CoreWeave) — 131k context at $0.25/1M · cheapest via CoreWeave

Largest context windows

Ranked by window size; where two match, the cheaper host is listed first.

ModelContextMax outputCheapest hostBlended /1M
deepseek-v4-flash1,311k131,072Baidu$0.15
glm-5.3-flash1,311k131,072Relace$0.309
glm-5.31,311k235,929Reka$4.65
deepseek-v4-flash-vision-exp1,049k131,072DeepInfra$0.862
llama-4-maverick1,049k115,200DigitalOcean$0.896
minimax-m31,049k235,929CoreWeave$1.19
deepseek-v41,049k393,216Alibaba$2.323
gemini-3.5-flash-lite1,049k65,536Google$2.8
gemini-3.6-flash1,049k65,536Google$4.5
gemini-3.8-flash1,049k65,536Google$4.5
gemini-3.7-flash1,049k65,536Google$4.5
qwen3.8-2.4t-a95b1,049k131,072Novita AI$8
kimi-k31,049k16,384DeepInfra$17.1
gpt-4.1-mini1,048k32,768OpenAI$2
nemotron-3-super-120b-a12b1,000k16,384DeepInfra$0.485
qwen3.8-27b1,000k235,929Parasail$2.44
minimax-m11,000k40,000Minimax$2.6
claude-fable-5-11,000k128,000Anthropic$60
gpt-6-astra1,000k128,000OpenAI$60
grok-4.5500k32,768xAI$8

Context in tokens; prices per 1M tokens (USD), from each model's cheapest published rate. Verified 2026-09-05.

FAQ

What is the largest LLM context window available?

The largest window in this index reaches 1311k tokens, against a median of 259k across 108 models. Window size caps how much you can send in a single call.

What's the cheapest way to get a long context window?

deepseek-v4-flash offers a 1,311k-token window at just $0.15 per 1M tokens blended via Baidu — the cheapest long-context option here.

Does the provider change a model's context window?

No — window size is a property of the model, identical whichever provider serves it. So you can route to the cheapest host without losing context length. What does vary with usage is cost, which scales with tokens sent.