#tHow Many Tokens?

← Back to counter

What is the cheapest AI model?

The short answer

As of September 7, 2026, GPT-5 Nano at $0.05 per million input tokens and $0.40 per million output is the cheapest model in the catalog on a typical 1,000-input / 200-output workload (OpenAI).

That ranking is generated from the same live pricing data that powers the counter, so it stays current as providers change rates and new models launch. For high-volume routing, classification, and extraction the cheapest tier is usually the right call; for harder reasoning you will normally step up a tier or two.

Cheapest models, ranked

Sorted by per-1M-call cost on a typical 1,000-token input + 200-token output workload. Prices as of September 7, 2026.

#ModelInput $/MOutput $/M1M callsAccuracy
#1 GPT-5 Nano
OpenAI
$0.05 $0.40 $130 exact
#2 GPT-4.1 Nano
OpenAI
$0.10 $0.40 $180 exact
#3 Gemini 2.5 Flash-Lite
Google
$0.10 $0.40 $180 exact
#4 DeepSeek V4 Flash
DeepSeek
$0.14 $0.28 $196 ≈±3%
#5 Llama 3.1 8B
Meta
$0.18 $0.18 $216 exact
#6 Qwen3.5 9B
Alibaba
$0.17 $0.25 $220 ≈±3%
#7 Qwen3.8 Flash
Alibaba
$0.15 $0.47 $244 ≈±3%
#8 GLM-5.3 Flash
Zhipu AI
$0.15 $0.50 $250 ≈±3%
#9 GPT-4o mini
OpenAI
$0.15 $0.60 $270 exact
#10 GPT-OSS 120B
OpenAI
$0.15 $0.60 $270 ≈±3%
#11 GPT-5.6 Luna
OpenAI
$0.20 $1.20 $440 exact
#12 GPT-5.4 Nano
OpenAI
$0.20 $1.25 $450 exact

The "1M calls" column is the number that actually matters at scale: it multiplies each model's input and output rate by that workload and projects it to a million calls, so a cheap headline input rate with an expensive output rate lands where it really belongs.

"Cheap" depends on what you need

The cheapest model isn't always the right answer. Weigh these before committing:

Cheaper than this list

If you need to go below even the cheapest hosted rate:

Get a real cost comparison

Paste your prompt into the counter. It shows the actual token count and per-call cost across every model, so you can choose by total cost on your workload instead of by per-million headline.

Try this on every model

Try the live counter →