Which AI model has the longest context window?
The short answer
As of July 26, 2026, Gemini 2.5 Pro has the longest production context window at 2,000,000 tokens (Google).
The next tier down is the 1M-token class, and the tier below that clusters around 128K-200K. The full ranking below is generated from live model data, so it reflects the current lineup as models are added and retired.
Ranked by context window size
Every non-deprecated model we track, largest window first. Data as of July 26, 2026.
| Model | Context (tokens) | Provider |
|---|---|---|
| Gemini 2.5 Pro | 2,000,000 | |
| GPT-5.6 Sol | 1,050,000 | OpenAI |
| GPT-5.6 Terra | 1,050,000 | OpenAI |
| GPT-5.6 Luna | 1,050,000 | OpenAI |
| Claude Opus 5 | 1,000,000 | Anthropic |
| Claude Opus 5 (Fast Mode) | 1,000,000 | Anthropic |
| Claude Sonnet 5 | 1,000,000 | Anthropic |
| Claude Fable 5 | 1,000,000 | Anthropic |
| GPT-4.1 | 1,000,000 | OpenAI |
| GPT-4.1 Mini | 1,000,000 | OpenAI |
| GPT-4.1 Nano | 1,000,000 | OpenAI |
| Gemini 3.1 Pro | 1,000,000 | |
| Gemini 3 Flash | 1,000,000 | |
| Gemini 3.1 Flash-Lite | 1,000,000 | |
| Gemini 2.5 Flash | 1,000,000 | |
| Gemini 2.5 Flash-Lite | 1,000,000 | |
| DeepSeek V4 Pro | 512,000 | DeepSeek |
| GPT-5.5 | 400,000 | OpenAI |
| GPT-5.5 Pro | 400,000 | OpenAI |
| GPT-5.4 | 400,000 | OpenAI |
| GPT-5.4 Mini | 400,000 | OpenAI |
| GPT-5.4 Nano | 400,000 | OpenAI |
| GPT-5.4 Pro | 400,000 | OpenAI |
| GPT-5.3 | 400,000 | OpenAI |
| GPT-5.2 | 400,000 | OpenAI |
| GPT-5.2 Pro | 400,000 | OpenAI |
| GPT-5.1 | 400,000 | OpenAI |
| GPT-5 | 400,000 | OpenAI |
| GPT-5 Mini | 400,000 | OpenAI |
| GPT-5 Nano | 400,000 | OpenAI |
| GPT-5 Pro | 400,000 | OpenAI |
| Qwen3.5 397B | 262,144 | Alibaba |
| Claude Haiku 4.5 | 200,000 | Anthropic |
| o3 | 200,000 | OpenAI |
| o3-mini | 200,000 | OpenAI |
| o3-pro | 200,000 | OpenAI |
| o4-mini | 200,000 | OpenAI |
| Qwen 2.5 72B | 131,072 | Alibaba |
| Qwen 2.5 Coder 32B | 131,072 | Alibaba |
| Qwen3 Coder 480B | 131,072 | Alibaba |
| GPT-4o | 128,000 | OpenAI |
| GPT-4o mini | 128,000 | OpenAI |
| GPT-4 Turbo | 128,000 | OpenAI |
| Llama 3.3 70B | 128,000 | Meta |
| Llama 3.1 405B | 128,000 | Meta |
| Llama 3.1 70B | 128,000 | Meta |
| Llama 3.1 8B | 128,000 | Meta |
| Mistral Large | 128,000 | Mistral |
| DeepSeek V3 | 128,000 | DeepSeek |
| DeepSeek V3.1 | 128,000 | DeepSeek |
| DeepSeek R1 | 128,000 | DeepSeek |
| GLM-5.1 | 128,000 | zhipu |
The gap between "context length" and "useful context length"
A 1M or 2M token window doesn't mean the model uses every token equally well. Independent evaluations consistently show:
- Quality degrades on retrieval tasks as context grows past roughly 32K for most models.
- The largest-window Gemini models hold recall quality across long context better than most, currently the strongest at this.
- Claude Sonnet and Opus maintain quality well to around 100K, then drift.
- Open-weights models in particular often degrade sharply past 32K despite a larger advertised window.
If your workload depends on the model finding a specific fact buried deep in a long context, test it with your actual prompts before choosing on advertised window size alone.
When you actually need a long context window
- Loading entire codebases for refactoring or audit.
- Long-document Q&A without chunking and retrieval.
- Multi-document synthesis where retrieval would lose cross-document relationships.
- Multi-turn conversations with extensive history you don't want to summarize.
For everything else, most chat, RAG, classification, and extraction, 32K-128K is plenty, and shorter context is cheaper to run. If cost is the deciding factor, cross-reference the cheapest model ranking.
A note on very large advertised windows
Some open-weights models advertise windows in the multi-million-token range. Advertised size and independently-validated recall quality at that length are different things, and availability via a given provider varies. Where a model's window isn't confirmed on its provider's current pricing page, we don't list it here. Treat any headline window well beyond 1M as unvalidated until you've tested recall on your own prompts.
Get cost at your context length
Paste your full context into the counter. It shows exact token counts and per-call cost across every model, so you can see which fit your workload and what they'd cost.
Try this on every model
- Claude Opus 5 $5.00/$25.00
- Claude Opus 5 (Fast Mode) $10.00/$50.00
- Claude Sonnet 5 $2.00/$10.00
- Claude Haiku 4.5 $1.00/$5.00
- Claude Fable 5 $10.00/$50.00
- GPT-5.6 Sol $5.00/$30.00
- GPT-5.6 Terra $2.50/$15.00
- GPT-5.6 Luna $1.00/$6.00
- GPT-5.5 $5.00/$30.00
- GPT-5.5 Pro $30.00/$180.00
- GPT-5.4 $2.50/$15.00
- GPT-5.4 Mini $0.75/$4.50
- GPT-5.4 Nano $0.20/$1.25
- GPT-5.4 Pro $30.00/$180.00
- GPT-5.3 $1.75/$14.00
- GPT-5.2 $1.75/$14.00
- GPT-5.2 Pro $21.00/$168.00
- GPT-5.1 $1.25/$10.00
- GPT-5 $1.25/$10.00
- GPT-5 Mini $0.25/$2.00
- GPT-5 Nano $0.05/$0.40
- GPT-5 Pro $15.00/$120.00
- GPT-4.1 $2.00/$8.00
- GPT-4.1 Mini $0.40/$1.60
- GPT-4.1 Nano $0.10/$0.40
- o3 $2.00/$8.00
- o3-mini $1.10/$4.40
- o3-pro $20.00/$80.00
- o4-mini $1.10/$4.40
- GPT-4o $2.50/$10.00
- GPT-4o mini $0.15/$0.60
- GPT-4 Turbo $10.00/$30.00
- Gemini 3.1 Pro $2.00/$12.00
- Gemini 3 Flash $0.50/$3.00
- Gemini 3.1 Flash-Lite $0.25/$1.50
- Gemini 2.5 Pro $1.25/$10.00
- Gemini 2.5 Flash $0.30/$2.50
- Gemini 2.5 Flash-Lite $0.10/$0.40
- Llama 3.3 70B $1.04/$1.04
- Llama 3.1 405B $3.50/$3.50
- Llama 3.1 70B $0.59/$0.79
- Llama 3.1 8B $0.18/$0.18
- Mistral Large $2.00/$6.00
- DeepSeek V3 $0.27/$1.10
- DeepSeek V3.1 $0.60/$1.70
- DeepSeek R1 $3.00/$7.00
- DeepSeek V4 Pro $1.74/$3.48
- Qwen 2.5 72B $0.90/$0.90
- Qwen 2.5 Coder 32B $0.80/$0.80
- Qwen3 Coder 480B $2.00/$2.00
- Qwen3.5 397B $0.60/$3.60
- GLM-5.1 $1.40/$4.40