What is the cheapest AI model?
The short answer
As of July 26, 2026, GPT-5 Nano at $0.05 per million input tokens and $0.40 per million output is the cheapest model in the catalog on a typical 1,000-input / 200-output workload (OpenAI).
That ranking is generated from the same live pricing data that powers the counter, so it stays current as providers change rates and new models launch. For high-volume routing, classification, and extraction the cheapest tier is usually the right call; for harder reasoning you will normally step up a tier or two.
Cheapest models, ranked
Sorted by per-1M-call cost on a typical 1,000-token input + 200-token output workload. Prices as of July 26, 2026.
| # | Model | Input $/M | Output $/M | 1M calls | Accuracy |
|---|---|---|---|---|---|
| #1 | GPT-5 Nano OpenAI |
$0.05 | $0.40 | $130 | exact |
| #2 | GPT-4.1 Nano OpenAI |
$0.10 | $0.40 | $180 | exact |
| #3 | Gemini 2.5 Flash-Lite |
$0.10 | $0.40 | $180 | exact |
| #4 | Llama 3.1 8B Meta |
$0.18 | $0.18 | $216 | exact |
| #5 | GPT-4o mini OpenAI |
$0.15 | $0.60 | $270 | exact |
| #6 | GPT-5.4 Nano OpenAI |
$0.20 | $1.25 | $450 | exact |
| #7 | DeepSeek V3 DeepSeek |
$0.27 | $1.10 | $490 | ≈±3% |
| #8 | Gemini 3.1 Flash-Lite |
$0.25 | $1.50 | $550 | exact |
| #9 | GPT-5 Mini OpenAI |
$0.25 | $2.00 | $650 | exact |
| #10 | GPT-4.1 Mini OpenAI |
$0.40 | $1.60 | $720 | exact |
| #11 | Llama 3.1 70B Meta |
$0.59 | $0.79 | $748 | exact |
| #12 | Gemini 2.5 Flash |
$0.30 | $2.50 | $800 | exact |
The "1M calls" column is the number that actually matters at scale: it multiplies each model's input and output rate by that workload and projects it to a million calls, so a cheap headline input rate with an expensive output rate lands where it really belongs.
"Cheap" depends on what you need
The cheapest model isn't always the right answer. Weigh these before committing:
- Tokenizer accuracy. OpenAI, Anthropic, Google, and Llama counts in the table are exact (official tokenizers or real BPE). Mistral, DeepSeek, Qwen, and GLM are estimated within about ±3%. The accuracy column flags which is which.
- Capability gap. The cheapest models are built for high-volume routing, classification, and extraction. They fall behind on hard reasoning. For mid-tier tasks, a small "mini" model or Claude Haiku is usually the right step up.
- Output-heavy vs input-heavy. Most models charge 3-5x more for output than input, so a workload that generates a lot (translation, code generation) is priced very differently from one that reads a lot (RAG, summarization). The per-1M-calls figure already accounts for a 5:1 input:output shape; your real ratio moves the ranking.
- Context length needs. Cheap high-volume models often cap context lower than the flagships. If you routinely send long inputs, check the longest context window ranking alongside price.
- Preview vs GA. Preview and early-access tiers can shift pricing or behavior. GA models give pricing stability.
Cheaper than this list
If you need to go below even the cheapest hosted rate:
- Self-host a small open-weights model (Llama, Qwen) on your own GPU. Hardware cost only, no per-token bill.
- Prompt caching drops cached input to roughly 10% of the standard rate on OpenAI, Anthropic, Google, and DeepSeek. See how much prompt caching saves.
- Batch API gives a flat 50% discount for asynchronous workloads on OpenAI, Anthropic, and Google. See which providers offer batch discounts.
Get a real cost comparison
Paste your prompt into the counter. It shows the actual token count and per-call cost across every model, so you can choose by total cost on your workload instead of by per-million headline.
Try this on every model
- Claude Opus 5 $5.00/$25.00
- Claude Opus 5 (Fast Mode) $10.00/$50.00
- Claude Sonnet 5 $2.00/$10.00
- Claude Haiku 4.5 $1.00/$5.00
- Claude Fable 5 $10.00/$50.00
- GPT-5.6 Sol $5.00/$30.00
- GPT-5.6 Terra $2.50/$15.00
- GPT-5.6 Luna $1.00/$6.00
- GPT-5.5 $5.00/$30.00
- GPT-5.5 Pro $30.00/$180.00
- GPT-5.4 $2.50/$15.00
- GPT-5.4 Mini $0.75/$4.50
- GPT-5.4 Nano $0.20/$1.25
- GPT-5.4 Pro $30.00/$180.00
- GPT-5.3 $1.75/$14.00
- GPT-5.2 $1.75/$14.00
- GPT-5.2 Pro $21.00/$168.00
- GPT-5.1 $1.25/$10.00
- GPT-5 $1.25/$10.00
- GPT-5 Mini $0.25/$2.00
- GPT-5 Nano $0.05/$0.40
- GPT-5 Pro $15.00/$120.00
- GPT-4.1 $2.00/$8.00
- GPT-4.1 Mini $0.40/$1.60
- GPT-4.1 Nano $0.10/$0.40
- o3 $2.00/$8.00
- o3-mini $1.10/$4.40
- o3-pro $20.00/$80.00
- o4-mini $1.10/$4.40
- GPT-4o $2.50/$10.00
- GPT-4o mini $0.15/$0.60
- GPT-4 Turbo $10.00/$30.00
- Gemini 3.1 Pro $2.00/$12.00
- Gemini 3 Flash $0.50/$3.00
- Gemini 3.1 Flash-Lite $0.25/$1.50
- Gemini 2.5 Pro $1.25/$10.00
- Gemini 2.5 Flash $0.30/$2.50
- Gemini 2.5 Flash-Lite $0.10/$0.40
- Llama 3.3 70B $1.04/$1.04
- Llama 3.1 405B $3.50/$3.50
- Llama 3.1 70B $0.59/$0.79
- Llama 3.1 8B $0.18/$0.18
- Mistral Large $2.00/$6.00
- DeepSeek V3 $0.27/$1.10
- DeepSeek V3.1 $0.60/$1.70
- DeepSeek R1 $3.00/$7.00
- DeepSeek V4 Pro $1.74/$3.48
- Qwen 2.5 72B $0.90/$0.90
- Qwen 2.5 Coder 32B $0.80/$0.80
- Qwen3 Coder 480B $2.00/$2.00
- Qwen3.5 397B $0.60/$3.60
- GLM-5.1 $1.40/$4.40