Which AI model has the longest context window?
The short answer
As of September 7, 2026, GPT-6 Astra has the longest production context window at 1,050,000 tokens (OpenAI).
The next tier down is the 1M-token class, and the tier below that clusters around 128K-200K. The full ranking below is generated from live model data, so it reflects the current lineup as models are added and retired.
Ranked by context window size
Every non-deprecated model we track, largest window first. Data as of September 7, 2026.
| Model | Context (tokens) | Provider |
|---|---|---|
| GPT-6 Astra | 1,050,000 | OpenAI |
| GPT-5.6 Sol | 1,050,000 | OpenAI |
| GPT-5.6 Terra | 1,050,000 | OpenAI |
| GPT-5.6 Luna | 1,050,000 | OpenAI |
| Gemini 3.8 Flash | 1,048,576 | |
| Gemini 3.7 Flash | 1,048,576 | |
| Gemini 3.6 Flash | 1,048,576 | |
| Gemini 3.5 Flash | 1,048,576 | |
| Gemini 3.5 Flash-Lite | 1,048,576 | |
| Gemini 3.1 Pro | 1,048,576 | |
| Gemini 3 Flash | 1,048,576 | |
| Gemini 3.1 Flash-Lite | 1,048,576 | |
| Gemini 2.5 Pro | 1,048,576 | |
| Gemini 2.5 Flash | 1,048,576 | |
| Gemini 2.5 Flash-Lite | 1,048,576 | |
| DeepSeek V4 Pro 0813 | 1,048,576 | DeepSeek |
| Kimi K3 | 1,048,576 | Moonshot AI |
| Claude Opus 5 | 1,000,000 | Anthropic |
| Claude Opus 5 (Fast Mode) | 1,000,000 | Anthropic |
| Claude Sonnet 5 | 1,000,000 | Anthropic |
| Claude Fable 5.1 | 1,000,000 | Anthropic |
| Claude Mythos 5.1 | 1,000,000 | Anthropic |
| Claude Fable 5 | 1,000,000 | Anthropic |
| Claude Mythos 5 | 1,000,000 | Anthropic |
| GPT-4.1 | 1,000,000 | OpenAI |
| GPT-4.1 Mini | 1,000,000 | OpenAI |
| GPT-4.1 Nano | 1,000,000 | OpenAI |
| DeepSeek V4 Flash | 1,000,000 | DeepSeek |
| Qwen3.6-Plus | 1,000,000 | Alibaba |
| Qwen3.7-Plus | 1,000,000 | Alibaba |
| Qwen3.8 Flash | 1,000,000 | Alibaba |
| GLM-5.3 | 1,000,000 | Zhipu AI |
| GLM-5.3 Flash | 1,000,000 | Zhipu AI |
| GLM-5.2 | 1,000,000 | Zhipu AI |
| MiniMax M3 | 524,288 | MiniMax |
| DeepSeek V4 Pro | 512,000 | DeepSeek |
| GPT-5.5 | 400,000 | OpenAI |
| GPT-5.5 Pro | 400,000 | OpenAI |
| GPT-5.4 | 400,000 | OpenAI |
| GPT-5.4 Mini | 400,000 | OpenAI |
| GPT-5.4 Nano | 400,000 | OpenAI |
| GPT-5.4 Pro | 400,000 | OpenAI |
| GPT-5.2 | 400,000 | OpenAI |
| GPT-5.2 Pro | 400,000 | OpenAI |
| GPT-5.1 | 400,000 | OpenAI |
| GPT-5 | 400,000 | OpenAI |
| GPT-5 Mini | 400,000 | OpenAI |
| GPT-5 Nano | 400,000 | OpenAI |
| GPT-5 Pro | 400,000 | OpenAI |
| Qwen3.5 397B | 262,144 | Alibaba |
| Qwen3.5 9B | 262,144 | Alibaba |
| Kimi K2.7 Code | 262,144 | Moonshot AI |
| Kimi K2.6 | 262,144 | Moonshot AI |
| Claude Haiku 4.5 | 200,000 | Anthropic |
| o3 | 200,000 | OpenAI |
| o3-mini | 200,000 | OpenAI |
| o3-pro | 200,000 | OpenAI |
| o4-mini | 200,000 | OpenAI |
| o1 | 200,000 | OpenAI |
| o1-pro | 200,000 | OpenAI |
| GPT-OSS 120B | 131,072 | OpenAI |
| Llama 3.3 70B | 131,072 | Meta |
| Qwen 2.5 72B | 131,072 | Alibaba |
| Qwen 2.5 Coder 32B | 131,072 | Alibaba |
| Qwen3 Coder 480B | 131,072 | Alibaba |
| GPT-4o | 128,000 | OpenAI |
| GPT-4o mini | 128,000 | OpenAI |
| GPT-4 Turbo | 128,000 | OpenAI |
| Llama 3.1 405B | 128,000 | Meta |
| Llama 3.1 70B | 128,000 | Meta |
| Llama 3.1 8B | 128,000 | Meta |
| Mistral Large | 128,000 | Mistral |
| DeepSeek V3 | 128,000 | DeepSeek |
| DeepSeek V3.1 | 128,000 | DeepSeek |
| DeepSeek R1 | 128,000 | DeepSeek |
| GLM-5.1 | 128,000 | Zhipu AI |
| GPT-3.5 Turbo | 16,385 | OpenAI |
The gap between "context length" and "useful context length"
A 1M or 2M token window doesn't mean the model uses every token equally well. Independent evaluations consistently show:
- Quality degrades on retrieval tasks as context grows past roughly 32K for most models.
- The largest-window Gemini models hold recall quality across long context better than most, currently the strongest at this.
- Claude Sonnet and Opus maintain quality well to around 100K, then drift.
- Open-weights models in particular often degrade sharply past 32K despite a larger advertised window.
If your workload depends on the model finding a specific fact buried deep in a long context, test it with your actual prompts before choosing on advertised window size alone.
When you actually need a long context window
- Loading entire codebases for refactoring or audit.
- Long-document Q&A without chunking and retrieval.
- Multi-document synthesis where retrieval would lose cross-document relationships.
- Multi-turn conversations with extensive history you don't want to summarize.
For everything else, most chat, RAG, classification, and extraction, 32K-128K is plenty, and shorter context is cheaper to run. If cost is the deciding factor, cross-reference the cheapest model ranking.
A note on very large advertised windows
Some open-weights models advertise windows in the multi-million-token range. Advertised size and independently-validated recall quality at that length are different things, and availability via a given provider varies. Where a model's window isn't confirmed on its provider's current pricing page, we don't list it here. Treat any headline window well beyond 1M as unvalidated until you've tested recall on your own prompts.
Get cost at your context length
Paste your full context into the counter. It shows exact token counts and per-call cost across every model, so you can see which fit your workload and what they'd cost.
Try this on every model
- Claude Opus 5 $5.00/$25.00
- Claude Opus 5 (Fast Mode) $10.00/$50.00
- Claude Sonnet 5 $2.00/$10.00
- Claude Haiku 4.5 $1.00/$5.00
- Claude Fable 5.1 $10.00/$50.00
- Claude Mythos 5.1 $10.00/$50.00
- Claude Fable 5 $10.00/$50.00
- Claude Mythos 5 $10.00/$50.00
- GPT-6 Astra $10.00/$50.00
- GPT-5.6 Sol $4.00/$20.00
- GPT-5.6 Terra $2.00/$12.00
- GPT-5.6 Luna $0.20/$1.20
- GPT-5.5 $5.00/$30.00
- GPT-5.5 Pro $30.00/$180.00
- GPT-5.4 $2.50/$15.00
- GPT-5.4 Mini $0.75/$4.50
- GPT-5.4 Nano $0.20/$1.25
- GPT-5.4 Pro $30.00/$180.00
- GPT-5.2 $1.75/$14.00
- GPT-5.2 Pro $21.00/$168.00
- GPT-5.1 $1.25/$10.00
- GPT-5 $1.25/$10.00
- GPT-5 Mini $0.25/$2.00
- GPT-5 Nano $0.05/$0.40
- GPT-5 Pro $15.00/$120.00
- GPT-4.1 $2.00/$8.00
- GPT-4.1 Mini $0.40/$1.60
- GPT-4.1 Nano $0.10/$0.40
- o3 $2.00/$8.00
- o3-mini $1.10/$4.40
- o3-pro $20.00/$80.00
- o4-mini $1.10/$4.40
- o1 $15.00/$60.00
- o1-pro $150.00/$600.00
- GPT-4o $2.50/$10.00
- GPT-4o mini $0.15/$0.60
- GPT-4 Turbo $10.00/$30.00
- GPT-3.5 Turbo $0.50/$1.50
- GPT-OSS 120B $0.15/$0.60
- Gemini 3.8 Flash $0.75/$3.75
- Gemini 3.7 Flash $0.75/$3.75
- Gemini 3.6 Flash $0.75/$3.75
- Gemini 3.5 Flash $1.50/$9.00
- Gemini 3.5 Flash-Lite $0.30/$2.50
- Gemini 3.1 Pro $2.00/$12.00
- Gemini 3 Flash $0.50/$3.00
- Gemini 3.1 Flash-Lite $0.25/$1.50
- Gemini 2.5 Pro $1.25/$10.00
- Gemini 2.5 Flash $0.30/$2.50
- Gemini 2.5 Flash-Lite $0.10/$0.40
- Llama 3.3 70B $1.04/$1.04
- Llama 3.1 405B $3.50/$3.50
- Llama 3.1 70B $0.59/$0.79
- Llama 3.1 8B $0.18/$0.18
- Mistral Large $2.00/$6.00
- DeepSeek V3 $0.27/$1.10
- DeepSeek V3.1 $0.60/$1.70
- DeepSeek R1 $3.00/$7.00
- DeepSeek V4 Pro 0813 $1.32/$3.96
- DeepSeek V4 Pro $1.74/$3.48
- DeepSeek V4 Flash $0.14/$0.28
- Qwen 2.5 72B $0.90/$0.90
- Qwen 2.5 Coder 32B $0.80/$0.80
- Qwen3 Coder 480B $2.00/$2.00
- Qwen3.5 397B $0.60/$3.60
- Qwen3.5 9B $0.17/$0.25
- Qwen3.6-Plus $0.50/$3.00
- Qwen3.7-Plus $0.32/$1.28
- Qwen3.8 Flash $0.15/$0.47
- GLM-5.3 $1.40/$4.40
- GLM-5.3 Flash $0.15/$0.50
- GLM-5.2 $1.40/$4.40
- GLM-5.1 $1.40/$4.40
- Kimi K3 $3.00/$15.00
- Kimi K2.7 Code $0.95/$4.00
- Kimi K2.6 $1.20/$4.50
- MiniMax M3 $0.30/$1.20