#tHow Many Tokens?

← Back to counter

Which AI model has the longest context window?

The short answer

As of July 26, 2026, Gemini 2.5 Pro has the longest production context window at 2,000,000 tokens (Google).

The next tier down is the 1M-token class, and the tier below that clusters around 128K-200K. The full ranking below is generated from live model data, so it reflects the current lineup as models are added and retired.

Ranked by context window size

Every non-deprecated model we track, largest window first. Data as of July 26, 2026.

ModelContext (tokens)Provider
Gemini 2.5 Pro 2,000,000 Google
GPT-5.6 Sol 1,050,000 OpenAI
GPT-5.6 Terra 1,050,000 OpenAI
GPT-5.6 Luna 1,050,000 OpenAI
Claude Opus 5 1,000,000 Anthropic
Claude Opus 5 (Fast Mode) 1,000,000 Anthropic
Claude Sonnet 5 1,000,000 Anthropic
Claude Fable 5 1,000,000 Anthropic
GPT-4.1 1,000,000 OpenAI
GPT-4.1 Mini 1,000,000 OpenAI
GPT-4.1 Nano 1,000,000 OpenAI
Gemini 3.1 Pro 1,000,000 Google
Gemini 3 Flash 1,000,000 Google
Gemini 3.1 Flash-Lite 1,000,000 Google
Gemini 2.5 Flash 1,000,000 Google
Gemini 2.5 Flash-Lite 1,000,000 Google
DeepSeek V4 Pro 512,000 DeepSeek
GPT-5.5 400,000 OpenAI
GPT-5.5 Pro 400,000 OpenAI
GPT-5.4 400,000 OpenAI
GPT-5.4 Mini 400,000 OpenAI
GPT-5.4 Nano 400,000 OpenAI
GPT-5.4 Pro 400,000 OpenAI
GPT-5.3 400,000 OpenAI
GPT-5.2 400,000 OpenAI
GPT-5.2 Pro 400,000 OpenAI
GPT-5.1 400,000 OpenAI
GPT-5 400,000 OpenAI
GPT-5 Mini 400,000 OpenAI
GPT-5 Nano 400,000 OpenAI
GPT-5 Pro 400,000 OpenAI
Qwen3.5 397B 262,144 Alibaba
Claude Haiku 4.5 200,000 Anthropic
o3 200,000 OpenAI
o3-mini 200,000 OpenAI
o3-pro 200,000 OpenAI
o4-mini 200,000 OpenAI
Qwen 2.5 72B 131,072 Alibaba
Qwen 2.5 Coder 32B 131,072 Alibaba
Qwen3 Coder 480B 131,072 Alibaba
GPT-4o 128,000 OpenAI
GPT-4o mini 128,000 OpenAI
GPT-4 Turbo 128,000 OpenAI
Llama 3.3 70B 128,000 Meta
Llama 3.1 405B 128,000 Meta
Llama 3.1 70B 128,000 Meta
Llama 3.1 8B 128,000 Meta
Mistral Large 128,000 Mistral
DeepSeek V3 128,000 DeepSeek
DeepSeek V3.1 128,000 DeepSeek
DeepSeek R1 128,000 DeepSeek
GLM-5.1 128,000 zhipu

The gap between "context length" and "useful context length"

A 1M or 2M token window doesn't mean the model uses every token equally well. Independent evaluations consistently show:

If your workload depends on the model finding a specific fact buried deep in a long context, test it with your actual prompts before choosing on advertised window size alone.

When you actually need a long context window

For everything else, most chat, RAG, classification, and extraction, 32K-128K is plenty, and shorter context is cheaper to run. If cost is the deciding factor, cross-reference the cheapest model ranking.

A note on very large advertised windows

Some open-weights models advertise windows in the multi-million-token range. Advertised size and independently-validated recall quality at that length are different things, and availability via a given provider varies. Where a model's window isn't confirmed on its provider's current pricing page, we don't list it here. Treat any headline window well beyond 1M as unvalidated until you've tested recall on your own prompts.

Get cost at your context length

Paste your full context into the counter. It shows exact token counts and per-call cost across every model, so you can see which fit your workload and what they'd cost.

Try this on every model

Try the live counter →