Auto-generated from each provider's pricing page. Pricing changes, model launches, and tokenizer-accuracy upgrades, in reverse chronological order. The leaderboard at the top updates with every build.
Live ranking
AI cost leaderboard, ranked cheapest first
Cost to run a 1,000-token-input / 200-token-output prompt 1,000 times, across the eight cheapest non-deprecated models we track. Updates automatically with every pricing snapshot.
Added. Together AI lists openai/gpt-oss-120b at $0.15 input / $0.60 output per 1M tokens and the serverless catalogue publishes a 131,072-token context window, so price, context and identity are all documented. This is the first OpenAI-authored model tracked here that is billed by a third-party host rather than OpenAI, and it lands well below every hosted GPT tier on output while carrying roughly a third of the GPT-5 line's context window. Token counts use the o200k vocabulary the model is built on, labelled approximate because the harmony variant layers additional special tokens over that base. Added noindex, the default for new entries.
Nine models appeared on Together AI's pricing page without qualifying for an entry. Four are held on identity rather than numbers: Muse Glimmer 30B at $0.35 / $1.50 now has a published 131,072-token context window, which clears the blocker recorded on 2026-09-05, but its model string still reads meta-models/Muse-Glimmer-30B against the meta-llama convention used for Meta's models, so provider, family and tokenizer remain unassignable. Inkling at $1.00 / $4.05 publishes a 524,288-token window but comes from Thinking Machines, a provider not yet tracked here, with no documented tokenizer to map it to. Ternary Bonsai 27B publishes a 262,144-token window at $0.00 / $0.00, and free pricing would take the top of /cheapest-ai-model/ on a rate that needs human confirmation before it ranks. Cogito v2.1 671B at $1.25 / $1.25, Rnj-1 Instruct at $0.15 / $0.15, LFM2.5-8B-A1B at $0.03 / $0.12 and Gemma 4 31B at $0.39 / $0.97 are absent from the serverless catalogue entirely. Qwen3.8-2.4T-A95B at $2.00 / $6.00 and Qwen3.7-Max at $1.25 / $3.75 remain listed with no context length for a fourth consecutive reading. Context window drives the /longest-context-window/ ranking, so each stays out rather than being seeded with an inferred value.
Flagged for manual review: contextWindow scraped 272,000 vs stored 400,000, not applied. The pricing page renders a context value of "<272K" against four rows - GPT-5.5, GPT-5.5 Pro, GPT-5.4 and GPT-5.4 Pro - while leaving the column blank for the other thirty-odd models in the same table. That pattern reads as an input-tier threshold on the rate card rather than a model context window, which matches the GPT-5 line's documented 400,000-token total of 272,000 input plus 128,000 output. The 32% gap sits under the automatic swing guard, so it would have applied silently and pushed four rows down /longest-context-window/. Held at 400,000 pending a human read of the rate card.
Flagged for manual review: contextWindow scraped 1,048,575 vs stored 1,000,000, not applied. The same discrepancy appears on GLM-5.3 Flash and GLM-5.2 at 1,048,575 and on DeepSeek V4 Flash at 1,048,576, all stored here as a round 1,000,000. Earlier readings of this catalogue recorded 1,000,000 for models in this group, so the source has either been restated or is being read differently between runs, and the difference is under 5%. Applying it would lift four rows above the other million-token entries on /longest-context-window/ on a change that is cosmetic if the round number was right. Left unchanged pending a direct read of the catalogue.
Indexation candidate, verify - left noindex, no swap made, carried forward from 2026-09-04. OpenAI's model catalogue now presents GPT-6 Astra at the head of the lineup as its most capable model with no availability restriction stated on the page, which is a step toward the general availability that would justify taking GPT-5.6 Sol's indexed slot. It is not confirmation: the Trusted Access Program limitation recorded at launch is simply not mentioned here, and an unmentioned restriction is not a lifted one. Revisit when OpenAI states general availability directly.
No rate changes. All four providers returned readable pricing and every stored rate matched exactly. Anthropic again required the documented fallback, anthropic.com/pricing redirecting to claude.com/pricing, and confirmed Opus 5, Sonnet 5, Haiku 4.5, both Fable tiers, both Mythos tiers and the $10 / $50 fast mode rate unchanged, along with the note that the Sonnet 5 increase to $3 / $15 will not occur. OpenAI's pricing page redirected to developers.openai.com and returned the full catalogue down to the legacy davinci-002 and babbage-002 tiers with every tracked rate matching. Google's listing covered the Gemini 3.8, 3.7, 3.6 and 3.5 Flash tiers and the 2.5 line at stored rates but omitted Gemini 3.1 Pro, Gemini 3 Flash and Gemini 3.1 Flash-Lite, and Together AI's listing again omitted DeepSeek V3, V3.1 and R1, Qwen3 Coder 480B, both Kimi K2 tiers, GLM-5.1, Mistral Large and the Llama 3.1 tiers, so no deprecation weight is drawn from either gap.
Added. Gemini 3.8 Flash opens Google's 3.8 generation at $0.75 input / $3.75 output per 1M tokens — the same rate as Gemini 3.7 Flash and Gemini 3.6 Flash, and half the rate of Gemini 3.5 Flash, so the generation advances without a price move. It carries the 1,048,576-token context window shared by the rest of the Flash line, which places it alongside the other Gemini entries rather than reordering the top of /longest-context-window/. Introductory pricing runs through 2026-12-31 and increases on 2027-01-01. The model was carried forward unadded on 2026-09-04 because no input token limit was published; Google now documents one, so the entry is created here. Added noindex, the default for new entries.
Not added, flagged for manual review. Together AI lists Muse Glimmer 30B at $0.35 input / $1.50 output per 1M tokens, and its serverless catalogue gives a 131,072-token context length, so price and window are both published. What is missing is the identity behind it: the model string reads meta-models/Muse-Glimmer-30B, a prefix that does not match the meta-llama convention Together AI uses for Meta's models, and the model has no page of its own. Provider, family, and tokenizer therefore cannot be assigned without guessing, and the tokenizer choice drives every count we render. Held out until the provider documents the model.
Three models carried forward unadded, all for the same reason: no published context length. Together AI lists Gemma 4 31B at $0.39 / $0.97, Qwen3.8-2.4T-A95B at $2.00 / $6.00, and Qwen3.7-Max at $1.25 / $3.75, and none of the three states an input token limit in the serverless model table. Context window drives the /longest-context-window/ ranking, so each stays out until its provider publishes one rather than being seeded with an inferred value.
Added. OpenAI now lists GPT-6 Astra at $10 input / $50 output per 1M tokens on a 1.05M-token context window, opening the GPT-6 generation above the GPT-5.6 line. It is described as OpenAI's most capable model, built for the hardest end-to-end work, and it accepts a higher maximum reasoning effort setting than GPT-5.6 Sol. The context window matches Sol exactly, so the two sit together in the same tier on /longest-context-window/, and the move is entirely on price: 2.5x Sol in both directions. Availability is limited at launch, rolling out to enterprises through OpenAI's Trusted Access Program with broader API and subscription access described as coming soon. Added noindex, the default for new entries.
Indexation candidate, verify — left noindex, no swap made. GPT-6 Astra is described as OpenAI's most capable model and reads as the successor to GPT-5.6 Sol at the top of the lineup, which would make it the candidate for Sol's indexed slot. It was not swapped in because access is still limited to enterprises in the Trusted Access Program, and GPT-5.6 Sol remains listed and generally available as the flagship customers can actually call. Worth revisiting once Astra reaches general availability.
Four models carried forward unadded, all for the same reason: no published context length. Google lists Gemini 3.8 Flash at $0.75 / $3.75, matching the 3.7 and 3.6 Flash rates, but neither the pricing page nor the model overview states an input token limit. Together AI lists Qwen3.8-2.4T-A95B at $2.00 / $6.00 and Qwen3.7-Max at $1.25 / $3.75, both showing a dash for context length in the serverless model table, and MiniMax M2.7 at $0.30 / $1.20 with no serverless documentation entry at all. Context window drives the /longest-context-window/ ranking, so each stays out until its provider publishes one rather than being seeded with an inferred value.
Marked retired, closing the review opened on 2026-08-27. OpenAI's deprecations page documents the removal directly: developers using the gpt-5.3-chat-latest snapshot were notified on 2026-05-08 and the model was shut down on 2026-08-10, with GPT-5.6 Sol named as the recommended replacement. That snapshot is the exact api id this entry is pinned to, which settles the question the five previous readings could not. What kept the flag off until now was the ambiguity of the gap rather than doubt about the absence — gpt-5.3-codex is still listed at the same $1.75 / $14 rates, so the 5.3 generation itself was never withdrawn, only the chat tier. A dated shutdown notice is the confirmation that was missing. The entry is retained with its final rates rather than deleted, and its page stays live; it now drops out of the /cheapest-ai-model/ and /longest-context-window/ rankings and the live model count moves from 75 to 74.
Five models appeared across the provider pages without qualifying for an entry, none of them added. Together AI still lists Qwen3.8-2.4T-A95B at $2.00 / $6.00 and Qwen3.7-Max at $1.25 / $3.75 with the context length rendered as a dash, and MiniMax M2.7 at $0.30 / $1.20 while the model remains absent from the serverless catalogue, so all three are held back on the same grounds as before: contextWindow feeds the /longest-context-window/ ranking directly and a guessed window would reorder that page. OpenAI's catalogue adds GPT-5.6 Cyber at $12.50 / $75.00, a specialised tier for authorised vulnerability research, again with no published context window, and GPT-5.5 Cyber appears on the pricing page but not in the model catalogue at all. Gemini Omni 1.1 Flash at $1.50 / $9.00 is excluded on different grounds and needs no further review: Google documents it as a video generation and editing model rather than a text model, so it falls outside what this site tracks. Add the four pending tiers once their providers publish context lengths.
No rate changes. All four providers returned readable pricing and every stored rate matched exactly across the models tracked here. OpenAI's read was the most complete yet, reaching down to the legacy davinci-002 and babbage-002 tiers, which is what made the GPT-5.3 absence worth pursuing to the deprecations page. Anthropic required the documented fallback: anthropic.com/pricing now redirects to claude.com/pricing, which carries only consumer plan tiers and no API rates, while platform.claude.com/docs/en/about-claude/pricing returned the full table and confirmed Opus 5, Sonnet 5, Haiku 4.5, both Fable tiers and both Mythos tiers unchanged, along with the $10 / $50 fast mode rate. The Sonnet 5 note recorded on 2026-08-12 is confirmed accurate: the increase to $3 / $15 scheduled for 2026-09-01 did not take effect and $2 / $10 stands as the standard rate. Together AI's listing was partial again, omitting DeepSeek V3.1 and R1, Qwen3 Coder 480B, both Kimi K2 tiers, GLM-5.1 and Mistral Large, so no deprecation weight is drawn from those gaps and the open questions on deepseek-v4-pro and the two Kimi K2 rows still await a complete provider listing.
Added. Anthropic now lists Claude Fable 5.1 at $10 input / $50 output per 1M tokens on a full 1M-token context window, taking over from Claude Fable 5 at the top of the lineup. Headline rates and context window are identical to Fable 5, so the generation advances without a price move, and both sit together in the 1M tier on /longest-context-window/. The real change is in prompt caching: cache hits and refreshes on the 5.1 tier are billed at 0.025x the base input rate ($0.25 per 1M tokens) against the 0.1x multiplier every other Claude model uses. Model id claude-fable-5-1 and the 1M context window are confirmed on the public model overview. Added noindex, the default for new entries.
Added alongside Claude Fable 5.1 and priced identically at $10 input / $50 output per 1M tokens on a 1M-token context window, including the same 0.025x cache-hit multiplier. Remains limited availability through Project Glasswing for defensive cybersecurity workflows, with no self-serve sign-up, exactly as Claude Mythos 5 is. The model id is the one field not directly published: limited-availability models are absent from the public model overview table, so claude-mythos-5-1 follows the convention confirmed there for Claude Fable 5.1 and matches the existing claude-mythos-5 entry. Added noindex, the default for new entries, and no indexation swap was considered — the rules exclude the Fable and Mythos tiers.
Added. Together AI now lists Qwen/Qwen3.8-Flash at $0.15 input / $0.47 output, opening Alibaba's Qwen3.8 generation, and the serverless catalogue publishes a 1,000,000-token context window. It is the cheapest million-token Qwen model tracked here, undercutting Qwen3.7-Plus by more than half on input and roughly two thirds on output at the same window, which places it on /cheapest-ai-model/ alongside the GLM-5.3 Flash tier it is priced against. Token counts use the Qwen BPE (approximately +/-3%). Added noindex, the default for new entries.
Not added, needs a published context length. Together AI lists Qwen3.8-2.4T-A95B at $2.00 / $6.00 and Qwen3.7-Max at $1.25 / $3.75, but the serverless catalogue shows no context length for either — the field renders as a dash. Both are held back rather than entered with a guessed window, because contextWindow feeds the /longest-context-window/ ranking directly and an invented figure would reorder that page. This repeats the treatment Qwen3.7-Max already had when it first appeared without a published length. Add them once Together AI publishes the windows.
Not added, needs a published context length. Together AI's pricing page lists MiniMax M2.7 at $0.30 / $1.20, the same rates as the newer MiniMax M3 already tracked here, but the model does not appear in the serverless catalogue at all, so there is no published context window to record. Held back on the same basis as the two Qwen tiers flagged today.
Note updated, no rate change. Rates hold at $4.00 / $20.00 with cached input at $0.40, but OpenAI's pricing page now states that the GPT-5.6 Sol pricing is promotional and available at least through 2026-11-21. That end date is recorded on the model so the current rate is not read as permanent — the same care taken with the Claude Sonnet 5 introductory rate, which by contrast was confirmed permanent on 2026-08-12.
Possible deprecation, verify — first absence, no action taken. GPT-5.3 did not appear in today's read of OpenAI's pricing page, which listed GPT-5.2, GPT-5.1 and GPT-5 as a group and skipped 5.3. Every other OpenAI model held its stored rate exactly, so this is more likely a gap in one reading than a withdrawal. Left active at its stored $1.75 / $14.00, matching the policy of never acting on a single absence.
Possible deprecation, still unresolved — carried again, so the previous entry's expectation that it would not be is not met. Today's read of Together AI's pricing page returned a partial listing: it omitted DeepSeek V3.1, DeepSeek R1, Qwen3 Coder 480B, both Kimi K2 tiers, GLM-5.1 and Mistral Large, all of which are known to exist. Absence in a reading that incomplete is not evidence of removal, so it carries no weight against this entry or the two Kimi K2 rows flagged on 2026-08-30. One blocker has cleared since then: the page builder now keeps a model page live when a model is retired, so setting the flag would no longer turn a live URL into a 404. What is still needed is a complete provider listing, or a direct check against DeepSeek, before the flag can be set. Left active at its stored $1.74 / $3.48.
Added. Together AI now lists zai-org/GLM-5.3 at $1.40 input / $4.40 output, completing the GLM-5.3 generation above the GLM-5.3 Flash tier added two days ago. Both the pricing page and the serverless catalogue publish a 1,000,000-token context window. The rates and the window are identical to the GLM-5.2 flagship, so the generation advances without a price move and the two sit together in the 1M tier on /longest-context-window/. Token counts use the Llama-family BPE proxy applied to the other GLM entries (approximately +/-10%), as we do not yet ship a real ChatGLM tokenizer. Added noindex, the default for new entries.
Possible deprecation, verify — second consecutive absence, and a deprecation was prepared this run and then withdrawn. The row is missing from both Together AI sources, the pricing page and the serverless catalogue, on 2026-08-28 and again today: four readings across two days, with Kimi K3 still listed and unchanged at $3.00 / $15.00. That evidence points to the K2 generation being retired. The flag was nonetheless reverted, because the page builder omits deprecated models entirely, so setting it removes /kimi-k2-7-code/ and turns a live, linkable URL into a 404 — the outcome the no-delete rule protects against. Resolving this needs a human decision on how deprecated models should be published (a retained page marked deprecated, or a redirect) before the flag can be set safely. Left active at its stored $0.95 / $4.00.
Possible deprecation, verify — second consecutive absence, on the same evidence and with the same withdrawn deprecation as Kimi K2.7 Code. Absent from both Together AI sources on 2026-08-28 and again today while Kimi K3 remains listed. Left active at its stored $1.20 / $4.50 pending the same decision about how deprecated models should be published.
Possible deprecation, verify — third consecutive absence. The undated DeepSeek V4 Pro row is again missing from both Together AI sources, while the dated DeepSeek V4 Pro 0813 remains in both at an unchanged $1.32 / $3.96 on a 1,048,576-token window. The evidence matches that behind the two Kimi K2 flags this run, and none of the three were deprecated: the page builder omits deprecated models, so the flag would remove the model page and 404 a live URL. That risk is highest here, because this entry is indexed. Left active at its stored $1.74 / $3.48 pending a direct check against DeepSeek rather than the reseller listing. This is the last run it should be carried on absence alone.
Possible deprecation, verify — first recorded absence. GLM-5.1 does not appear on Together AI’s pricing page or in the serverless catalogue today, while GLM-5.2 holds at $1.40 / $4.40 and the GLM-5.3 generation is now listed in full. A single day of absence is not enough to act on, so the entry is left active at its stored $1.40 / $4.40 on a 128,000-token window pending a second reading.
Possible deprecation, verify — third consecutive absence. The standalone GPT-5.3 chat tier is again missing from OpenAI’s pricing page, while gpt-5.3-codex is still listed at exactly the stored rates, $1.75 / $14.00. Every neighbouring tier remains present and unchanged. The 5.3 generation has therefore not been withdrawn — only the chat entry this row is pinned to — which keeps the supersession ambiguous. Left active at its stored rates pending confirmation.
No change. Together AI’s serverless catalogue prints Qwen3.7-Plus’s context length as 1,000,000 again, after showing a dash on 2026-08-28. This matches the stored value, which was held through the gap rather than changed. Rates re-confirmed at $0.32 / $1.28.
No rate changes. Every stored price re-confirmed across all four providers. OpenAI holds at GPT-5.6 Sol $4.00 / $20.00, Terra $2.00 / $12.00 and Luna $0.20 / $1.20, with GPT-5.5 and 5.5 Pro, the GPT-5.4 tiers, GPT-5.2, 5.1, 5, 5 Mini, 5 Nano, the GPT-4.1 tiers, GPT-4o, GPT-4o Mini, GPT-4 Turbo, o1, o3, o4-mini and GPT-3.5 Turbo all unchanged. Anthropic holds at Claude Fable 5 and Mythos 5 $10 / $50, Opus 5 $5 / $25, Opus 5 fast mode $10 / $50, Sonnet 5 $2 / $10 and Haiku 4.5 $1 / $5, the page repeating that Sonnet 5’s $2 / $10 is now the standard rate and that the increase once scheduled for 2026-09-01 will not take place. Google holds across all ten Gemini entries, the introductory rate on 3.7 and 3.6 Flash still running through 2026-12-31. Together AI re-confirms GLM-5.2 at $1.40 / $4.40, Kimi K3 at $3.00 / $15.00, MiniMax M3 at $0.30 / $1.20, DeepSeek V4 Pro 0813 at $1.32 / $3.96, V4 Flash at $0.14 / $0.28, Qwen3.5 397B at $0.60 / $3.60, Qwen3.5 9B at $0.17 / $0.25, Qwen3.6-Plus at $0.50 / $3.00, Qwen3.7-Plus at $0.32 / $1.28 and Llama 3.3 70B at $1.04 / $1.04. The only data change this run is the addition of GLM-5.3. Deprecations were prepared for the two Kimi K2 rows and withdrawn before shipping, because the page builder drops deprecated models and the flag would have removed two live URLs; both are left active and flagged.
Added. Together AI now lists zai-org/GLM-5.3-Flash at $0.15 input / $0.50 output, opening Zhipu AI's GLM-5.3 generation above the existing GLM-5.2 and GLM-5.1 entries. The serverless catalogue publishes a 1,000,000-token context window, so the entry pairs a full million-token window with one of the lowest rates in the set — roughly a ninth of GLM-5.2 at the same context size. It enters near the top of the /cheapest-ai-model/ ranking among million-token models and joins the 1M tier on /longest-context-window/. Token counts use the Llama-family BPE proxy applied to the other GLM entries (≈±10%), as we do not yet ship a real ChatGLM tokenizer. Added noindex, the default for new entries.
Possible deprecation, verify — second consecutive absence. The standalone GPT-5.3 chat tier is again missing from OpenAI's pricing page, one day after the same gap was first recorded. Its sibling gpt-5.3-codex is still listed at exactly the stored rates, $1.75 / $14.00, so the 5.3 generation itself has not been withdrawn — only the chat entry this row is pinned to. Every neighbouring tier remains listed and unchanged. Two readings in as many days make the absence look real rather than a one-off parse miss, but with the codex sibling still present at the same price the supersession is ambiguous, so the entry is left active at its stored rates pending confirmation.
Possible deprecation, verify. The undated DeepSeek V4 Pro row did not appear in either of today's two reads of Together AI — the pricing page or the serverless catalogue — although the dated DeepSeek V4 Pro 0813 remains in both at an unchanged $1.32 / $3.96. Against that, this entry was re-confirmed at $1.74 / $3.48 as recently as 2026-08-27, so a removal one day later would be unusually abrupt, and Together AI's pricing table is long enough that a row can be missed on a single retrieval. Left active at its stored rates pending confirmation. This is an indexed entry, so it should be checked directly against the provider before any change.
Possible deprecation, verify. Not seen in either of today's reads of Together AI, the pricing page or the serverless catalogue, having been re-confirmed at $0.95 / $4.00 on 2026-08-27. Kimi K2.6 is absent from the same two reads, while Kimi K3 remains listed and unchanged at $3.00 / $15.00 on a 1,048,576-token window. Two sibling rows disappearing together one day after confirmation reads more like an incomplete retrieval than a withdrawal, so both are left active at their stored rates pending confirmation.
Possible deprecation, verify. Not seen in either of today's reads of Together AI, alongside Kimi K2.7 Code and on the same reasoning — the row was re-confirmed at $1.20 / $4.50 within the past week and Kimi K3 is still listed. Left active at its stored rates pending confirmation.
Still not added — no context length published. Together AI continues to list Qwen3.8-2.4T-A95B at an unchanged $2.50 input / $6.25 output, but the serverless catalogue still prints its context length as a dash. Qwen3.7-Max is in the same position at $1.25 / $3.75, also with no published window. Neither entry can be completed without a sourced context window, so both remain unlisted. Revisit once Together AI publishes the figures.
No change. Together AI's serverless catalogue now prints a dash for Qwen3.7-Plus's context length, where it published 1,000,000 on 2026-08-12. A value disappearing from a catalogue is not evidence that the window shrank, so the stored 1,000,000 is unchanged. Rates re-confirmed at $0.32 / $1.28.
No rate changes. Every stored price re-confirmed across all four providers. OpenAI holds at GPT-5.6 Sol $4.00 / $20.00, Terra $2.00 / $12.00 and Luna $0.20 / $1.20, with GPT-5.5, 5.4, 5.2, 5.1, 5, 4.1, 4o, 4 Turbo, o1, o1-pro, o3, o3-mini, o3-pro, o4-mini and GPT-3.5 Turbo all unchanged. Anthropic holds at Claude Fable 5 and Mythos 5 $10 / $50, Opus 5 $5 / $25, Opus 5 fast mode $10 / $50, Sonnet 5 $2 / $10 and Haiku 4.5 $1 / $5, with the page repeating that Sonnet 5's $2 / $10 is now the standard rate and the increase once scheduled for 2026-09-01 will not take place. Google holds across all ten Gemini entries, the introductory rate on 3.7 and 3.6 Flash still running through 2026-12-31. Together AI re-confirms GLM-5.2 at $1.40 / $4.40 on a 1,000,000-token window, Kimi K3 at $3.00 / $15.00, MiniMax M3 at $0.30 / $1.20, DeepSeek V4 Pro 0813 at $1.32 / $3.96, V4 Flash at $0.14 / $0.28, Qwen3.5 397B at $0.60 / $3.60, Qwen3.5 9B at $0.17 / $0.25, Qwen3.7-Plus at $0.32 / $1.28 and Llama 3.3 70B at $1.04 / $1.04. The only data change this run is the addition of GLM-5.3 Flash.
Possible deprecation, verify. GPT-5.3 no longer appears on OpenAI's pricing page. Every neighbouring tier is still listed — GPT-5.4 at $2.50 / $15.00, GPT-5.2 at $1.75 / $14.00 and GPT-5.1 at $1.25 / $10.00 — so the gap sits in an otherwise complete run of the GPT-5 series. The tier was confirmed present at $1.75 / $14.00 as recently as 2026-08-26, which makes this a same-week disappearance rather than a long-standing absence. Against that, the reading rests on a single retrieval of the pricing page, and OpenAI's model catalogue lists only the GPT-5.6 variants, so it cannot corroborate either way. The entry is left active at its stored rates pending confirmation; a model is only marked deprecated here once its removal is clearly established.
No change. Every stored rate re-confirmed across all four providers. OpenAI holds at GPT-5.6 Sol $4.00 / $20.00, Terra $2.00 / $12.00 and Luna $0.20 / $1.20, with the GPT-5.5, 5.4, 5.2, 5.1, 5, 4.1, 4o, 4 Turbo, o1, o3, o4-mini and GPT-3.5 Turbo tiers all unchanged. Anthropic holds at Claude Fable 5 and Mythos 5 $10 / $50, Opus 5 $5 / $25, Opus 5 fast mode $10 / $50, Sonnet 5 $2 / $10 and Haiku 4.5 $1 / $5; Anthropic's pricing page now states that Sonnet 5's $2 / $10 is the standard rate and that the increase to $3 / $15 previously scheduled for 2026-09-01 will not take place. Google holds across all ten Gemini entries, with the introductory rate on 3.7 and 3.6 Flash still running through 2026-12-31. Together AI re-confirms GLM-5.2 at $1.40 / $4.40 on a 1,000,000-token window, Kimi K3 at $3.00 / $15.00, Kimi K2.7 Code at $0.95 / $4.00, MiniMax M3 at $0.30 / $1.20, DeepSeek V4 Pro at $1.74 / $3.48, V4 Pro 0813 at $1.32 / $3.96 and V4 Flash at $0.14 / $0.28, along with the Qwen and Llama entries. No prices were applied this run.
Added to expand rate coverage. OpenAI's first-generation reasoning tier, listed at $15.00 input / $60.00 output with a 200,000-token context window and a 100,000-token maximum output. It has been on OpenAI's pricing page throughout, but was outside the set tracked here until now. Superseded by the o3 and o4 lines on price and capability, so it is carried for rate comparison rather than as a current recommendation, and the entry is noindex.
Added to expand rate coverage. The most expensive model on OpenAI's pricing page at $150.00 input / $600.00 output — ten times the o1 rate in both directions, and roughly thirty-seven times GPT-5.6 Sol on input. Same 200,000-token context window and 100,000-token maximum output as o1. This becomes the costliest entry in the set and sits at the far end of the /cheapest-ai-model/ ranking. Noindex.
Added to expand rate coverage. OpenAI's long-running budget tier at $0.50 input / $1.50 output, pinned to the gpt-3.5-turbo-0125 snapshot. Its 16,385-token context window is the smallest of any model tracked here and places it last on the /longest-context-window/ ranking. It uses the older cl100k tokenizer rather than the o200k vocabulary of the GPT-4o generation and later, so token counts differ from the rest of the OpenAI lineup for the same text. Noindex.
Applied after manual verification. The same-day flag on this entry — scraped 1,000,000 against a stored 512,000, a 95% jump held back by the implausible-swing guard — has been checked against the provider and confirmed. Together AI's serverless catalogue prints 1000000 for zai-org/GLM-5.2, and Zhipu AI's own documentation independently describes a 1M-token context window, so the stored 512,000 was the incorrect value rather than the new reading being a parse error. Rates are unchanged at $1.40 / $4.40, and the catalogue re-confirms them. GLM-5.2 moves up the /longest-context-window/ ranking as a result.
Flagged for manual review: contextWindow scraped 1,000,000 vs stored 512,000. Together AI's serverless catalogue now prints 1000000 for zai-org/GLM-5.2, against the 512,000 recorded here on 2026-08-12. That is a 95% jump, well past the plausible-swing threshold, so the stored value is unchanged pending human verification. Read consistently across two separate fetches of the catalogue, so this looks like a real listing rather than a parse error — but contextWindow feeds the /longest-context-window/ ranking directly, and at 1M this entry would move several places up an indexed page, so it needs confirming against the provider page before it is applied. Rates unchanged at $1.40 / $4.40.
Not added — no context length published. Together AI lists a new Qwen tier, Qwen3.8-2.4T-A95B (Qwen/Qwen3.8-2.4T-A95B), at $2.50 input / $6.25 output, above the Qwen3.7 line in both directions. The serverless catalogue prints its context length as a dash, so the entry cannot be completed yet. Qwen3.7-Max ($1.25 / $3.75) remains in the same position, still with no published window. Flagged for manual review; revisit once Together AI publishes the context lengths.
No change. Every stored rate re-confirmed across all four providers. OpenAI holds at GPT-5.6 Sol $4 / $20, Terra $2 / $12 and Luna $0.20 / $1.20, with the GPT-5.5, 5.4, 5.3, 5.2, 5.1, 5, 4.1, 4o, o3 and o4-mini tiers all unchanged — the Sol cut applied on 2026-08-24 has held. Anthropic holds at Fable 5 and Mythos 5 $10 / $50, Opus 5 $5 / $25, Opus 5 fast mode $10 / $50, Sonnet 5 $2 / $10 and Haiku 4.5 $1 / $5, with Sonnet 5's $2 / $10 still confirmed as the standard price rather than introductory. Google holds across all ten Gemini entries. Together AI holds across the DeepSeek, Qwen, Kimi, GLM, MiniMax and Llama entries. No prices were applied this run.
OpenAI cut its flagship. Input drops 20% from $5.00 to $4.00 per 1M tokens, and cached input moves in step from $0.50 to $0.40, holding the usual 10% ratio. The rate was re-confirmed at $5.00 as recently as 2026-08-21, so this is a fresh change rather than a stale reading. Both legs of the cut are inside the plausible-swing threshold and move consistently, so they were applied.
Output drops 33% from $30.00 to $20.00 per 1M tokens alongside the input cut. GPT-5.6 Sol now undercuts GPT-5.5 ($5.00 / $30.00) in both directions while carrying roughly 2.6x the context, so the flagship is no longer the more expensive of the two. This reorders the /cheapest-ai-model/ ranking among the premium tiers.
Precision correction, not a capacity change. We had recorded the rounded 1,000,000; Together AI's serverless catalogue publishes the exact figure as 1,048,576 tokens. A 4.9% adjustment, the same rounding fix applied to Llama 3.3 70B on 2026-08-13 and across the Gemini lineup on 2026-08-21. Rates unchanged at $3.00 / $15.00.
Not added — no context length published. OpenAI lists two cybersecurity tiers, GPT-5.6 Cyber and GPT-5.5 Cyber, both at $12.50 / $75.00 with cached input at $1.25. Neither carries a context window on OpenAI's pricing page or its model catalogue, so neither can be given a complete entry yet. Flagged for manual review; revisit once OpenAI publishes the window.
No change. Claude Fable 5 and Mythos 5 ($10 / $50), Opus 5 ($5 / $25), Opus 5 fast mode ($10 / $50), Sonnet 5 ($2 / $10) and Haiku 4.5 ($1 / $5) all re-confirmed at their stored rates. Sonnet 5's $2 / $10 remains the standard price.
No change. All nine Gemini entries re-confirmed, including the 3.7 and 3.6 Flash pair at $0.75 / $3.75, 3.5 Flash at $1.50 / $9.00, and the 3.5 and 3.1 Flash-Lite tiers at $0.30 / $2.50 and $0.25 / $1.50. The introductory pricing on 3.7 and 3.6 Flash still runs through 2026-12-31.
Added. claude-mythos-5 shares Claude Fable 5's specs and pricing — $10 / $50 with a full 1M-token context window — and appears on Anthropic's published pricing table. It is not generally available: access is limited to approved customers in Project Glasswing, for defensive cybersecurity workflows, with no self-serve sign-up. The entry is included for rate comparison and labelled accordingly.
Added. Google's newest and most capable Flash tier, at $0.75 / $3.75 with a 1,048,576-token context window. Introductory pricing runs through 2026-12-31 and increases on 2027-01-01, so this rate has a known expiry.
Added. The Flash tier directly below Gemini 3.7 Flash, at the same rate and the same 1,048,576-token context window. Same introductory pricing window through 2026-12-31.
Added. The most expensive Flash tier Google currently lists, at double the rate of the newer 3.6 and 3.7 Flash models above it. Output is 6x input. 1,048,576-token context window.
Added. This is the model whose rates were misread against the Gemini 3.1 Flash-Lite row on 2026-08-12 and correctly held back by the swing guard. It now has its own entry, which closes that ambiguity: 3.5 Flash-Lite is $0.30 / $2.50 and 3.1 Flash-Lite is $0.25 / $1.50.
Added. A dated V4 Pro release that Together AI lists alongside the undated DeepSeek V4 Pro at its own rate and its own context length — 1,048,576 tokens against 512,000. Both are live listings, so this is a second entry rather than a change to the existing one.
Correction of a stale stored value, not a capacity cut. Google's own spec page for gemini-2.5-pro publishes an input token limit of 1,048,576; the 2,000,000 we held came from a larger window that was announced but never shipped. A 48% reduction, inside the plausible-swing threshold and read from a single-model spec page rather than a multi-row table, so applied. This reorders the /longest-context-window/ ranking, where Gemini 2.5 Pro previously stood alone at the top. Rates unchanged at $1.25 / $10.00.
Precision correction across five entries — Gemini 3.1 Pro, Gemini 3 Flash, Gemini 3.1 Flash-Lite, Gemini 2.5 Flash and Gemini 2.5 Flash-Lite. We had recorded the rounded 1,000,000; every Gemini spec page publishes the exact figure as 1,048,576 tokens. A 4.9% adjustment, the same rounding fix applied to Llama 3.3 70B on 2026-08-13. No rate changes.
Not added — no context length published. Together AI now lists two Qwen tiers above Qwen3.7-Plus: Qwen3.8-2.4T-A95B at $2.50 / $6.25 and Qwen3.7-Max at $1.25 / $3.75. Both show a blank context length in Together AI's model catalogue, so neither can be given a complete entry yet. Flagged for manual review; revisit once Alibaba or Together AI publishes the window.
Possible deprecation, verify. Qwen3 Coder 480B and DeepSeek R1 no longer appear on Together AI's pricing page or in its serverless model catalogue. One absent fetch is not proof of removal, so both entries are left active and unchanged pending a human check.
No change. Gemini 3 Flash dropped off Google's pricing table this run, but the model catalogue still lists gemini-3-flash-preview as available and its spec page still resolves, so the entry stays active at $0.50 / $3.00. Its context window was corrected to 1,048,576 along with the rest of the Gemini lineup.
Scope note. Together AI's pricing page carries a number of listings outside the families we track — among them Muse Glimmer 30B, Inkling, Cogito v2.1 671B, NVIDIA Nemotron 3 Ultra, Gemma 4 31B, LFM2.5-8B-A1B, gpt-oss-120B and MiniMax M2.7. None were added. Expanding the tracked set is a deliberate editorial decision rather than an automatic one, so it is left for a human.
No change. Every OpenAI entry re-confirmed at its stored rate, including the full GPT-5.6 line (Sol $5 / $30, Terra $2 / $12, Luna $0.20 / $1.20), GPT-5.5 and GPT-5.5 Pro, the GPT-5.4 tiers, GPT-5.3, GPT-5.2, GPT-5.1, GPT-5, GPT-4.1, o3, o4-mini and GPT-4o.
No change. Claude Opus 5 ($5 / $25), Opus 5 fast mode ($10 / $50), Sonnet 5 ($2 / $10), Haiku 4.5 ($1 / $5) and Fable 5 ($10 / $50) all re-confirmed. Sonnet 5's $2 / $10 remains the standard rate — the increase to $3 / $15 once scheduled for 2026-09-01 is still cancelled.
Added. moonshotai/Kimi-K3, Moonshot AI's flagship on Together AI and the first Kimi generation to carry a full 1M-token context window, against the 256K of the K2 line. Moonshot AI is a new provider for us. Its tokenizer is a Moonshot BPE we do not implement, so counts use the Llama-family BPE heuristic as a proxy and the entry is labelled approximate — the same treatment given to the GLM entries.
Added. moonshotai/Kimi-K2.7-Code, Moonshot AI's coding-specialised tier, at roughly a third the rate of Kimi K3 with a 256K context window (262,144 tokens). Tokenizer proxied as with Kimi K3.
Added. moonshotai/Kimi-K2.6, the general-purpose model of Moonshot AI's K2 generation, sharing the 256K context window of K2.7 Code but priced above it in both directions. Tokenizer proxied as with Kimi K3.
Added. MiniMaxAI/MiniMax-M3, pairing a 512K context window (524,288 tokens) with one of the lowest rates of any long-context open-weight model we track. MiniMax is a new provider for us, and its tokenizer is likewise proxied by the Llama-family heuristic, so the entry is labelled approximate.
Precision correction, not a capacity change. We had recorded the rounded 128,000; Together AI publishes the exact figure as 131,072 tokens, which is how we already store the Qwen entries. A 2.4% adjustment, well inside the plausible-swing threshold, so applied. Rates are unchanged at $1.04 in both directions.
No change — this closes the review opened on 2026-08-12. Google's pricing page now returns $0.25 input / $1.50 output for Gemini 3.1 Flash-Lite, matching our stored rate exactly, and separately lists Gemini 3.5 Flash-Lite at $0.30 / $2.50. That confirms yesterday's reading of $0.30 / $2.50 against this row was a misread of the adjacent model rather than a real increase, and the swing guard was right to hold it. Every other Gemini entry re-confirmed unchanged.
No change — this closes the possible-deprecation review opened on 2026-08-02. That run's fetch returned a partial lineup and omitted GPT-5.3, o3-pro and GPT-4 Turbo; this run returned the full table and all three are present at their stored rates ($1.75/$14, $20/$80 and $10/$30), so no deprecation was warranted. All thirty OpenAI entries re-confirmed exactly, including the GPT-5.6 Luna cut applied by hand on 2026-08-11.
Deferred — no change applied. OpenAI's pricing table has added a security-focused tier, gpt-5.6-cyber at $12.50 input / $75 output (cached $1.25), alongside a gpt-5.5-cyber at the same rate. Both are the most expensive OpenAI models on the page. Prices and model ids are clear, but neither the pricing page nor the models reference publishes a context window for either, while it does publish 1.05M for Sol, Terra and Luna. Since each entry feeds the /longest-context-window/ ranking, they stay out until the figure is published.
Still deferred — no change applied. Three models on Together AI's pricing page have clear rates but no published context length in the serverless-models reference: Qwen3.8-2.4T-A95B ($2.50/$6.25, new to this run and the largest Qwen tier listed), Qwen3.7-Max ($1.25/$3.75, unchanged and deferred since 2026-08-02) and MiniMax M2.7 ($0.30/$1.20, priced identically to the M3 added this run). Its sibling MiniMax M3 was added because its context length is published. Also untracked by choice: Qwen3 235B and the small legacy tiers Qwen2.5 7B Instruct Turbo and Llama 3 8B Instruct Lite.
Still deferred — no change applied, blocker unchanged. Claude Mythos 5 remains listed at $10 / $50, matching Fable 5, and the long-context section still confirms the full 1M-token window at standard pricing. It is still limited availability behind a waitlist, and the page still names it inconsistently (Claude Mythos 5 in the pricing table, Claude Mythos Preview elsewhere) without publishing an API model id. Adding it would mean inventing the string users copy into their code. Every other Anthropic entry re-confirmed unchanged, including the Sonnet 5 rate of $2 / $10 now recorded as permanent.
Still deferred for a fifth run — no change applied, blocker unchanged. Gemini 3.6 Flash ($1.50/$7.50), Gemini 3.5 Flash ($1.50/$9.00) and Gemini 3.5 Flash-Lite ($0.30/$2.50) all hold their prices, and Google still publishes no input token limit for any of them. A fourth model, Gemini 3.5 Live Translate ($3.50/$21.00), is new to this run and has the same gap.
No rate change — $2 input / $10 output is unchanged. What changed is that the increase we had been warning about is cancelled. Anthropic now states the $2/$10 rate announced as introductory pricing through 2026-08-31 is the standard price, and the scheduled move to $3/$15 on 2026-09-01 will not occur. The model note has been rewritten, since it previously told readers to budget for a 50% increase in under three weeks.
Added, resolving the deferral opened on 2026-08-02. Zhipu AI’s newest flagship on Together AI, priced identically to the GLM-5.1 we already carry but with a 512K-token context window against GLM-5.1’s 128K. The earlier deferral was blocked on the missing API string and context window; both come from Together’s serverless-models reference, which lists zai-org/GLM-5.2 at 512,000 tokens. Tokenizer remains the Llama-family BPE proxy used for GLM-5.1 (approx-10pct).
Added, resolving the deferral opened on 2026-08-02. Alibaba’s newest general-purpose tier on Together AI, Qwen/Qwen3.7-Plus at a 1M-token context window. Cheaper in both directions than the Qwen3.6-Plus it succeeds. The sibling Qwen3.7-Max stays deferred — see its own entry.
Added, resolving the deferral opened on 2026-08-02. Qwen/Qwen3.6-Plus, the first Qwen generation on Together AI to carry a full 1M-token context window. Already superseded on price by Qwen3.7-Plus, but still listed and callable.
Added. Qwen/Qwen3.5-9B, the small dense model of the Qwen3.5 generation, carrying the same 256K native context (262,144 tokens) as its 397B sibling at a fraction of the rate.
Added. deepseek-ai/DeepSeek-V4-Flash-0731, the small fast tier of DeepSeek’s V4 generation, at roughly one twelfth the rate of V4 Pro. Notable for a 1M-token context window — twice that of the larger V4 Pro.
Model id only — no rate change. Gemini 3.1 Flash-Lite has moved from preview to generally available and the id has dropped its -preview suffix. Callers pinned to the old preview id should update.
Flagged for manual review — no change applied, and the stored rate is very likely correct. A read of Google’s pricing page returned $0.30 input / $2.50 output for this row against our stored $0.25 / $1.50. The output figure is a 67% jump, past the 60% threshold at which an automated read is not trusted. It is almost certainly a row misread rather than a real increase: $0.30 / $2.50 are the rates recorded for the separate Gemini 3.5 Flash-Lite model, which sits adjacent in the same table and was seen at exactly those figures in the 2026-08-02 run. The stored $0.25 / $1.50 stands.
Still deferred — no change applied. Qwen3.7-Max is listed at $1.25 / $3.75, unchanged from the 2026-08-02 read, and the API string Qwen/Qwen3.7-Max is now confirmed. The remaining blocker is the context length, which Together’s serverless-models reference leaves blank for this model while publishing it for every other Qwen3.x tier. Since each entry feeds the /longest-context-window/ ranking, it stays out until the figure is published. Its siblings Qwen3.7-Plus and Qwen3.6-Plus were added this run now that their context lengths are available.
Deferred — no change applied. Anthropic’s pricing table has added Claude Mythos 5 at $10 / $50, matching Fable 5, and the long-context section confirms it carries the full 1M-token context window at standard pricing. Two things hold it back: access is limited availability behind a waitlist rather than general release, and the page names it inconsistently (Claude Mythos 5 in the pricing table, Claude Mythos Preview in the long-context section) without publishing an API model id anywhere. Adding it would mean inventing the string users copy into their code. It stays out until the id is published.
Still deferred for a fourth run — no change applied, and the blocker is unchanged. Gemini 3.6 Flash ($1.50/$7.50) and Gemini 3.5 Flash ($1.50/$9.00) are both confirmed Stable with ids gemini-3.6-flash and gemini-3.5-flash, but Google still publishes no input token limit for either, on the pricing page, the models overview, or the per-model pages (which returned 404 on the slug patterns tried). Adding them would mean inventing the one figure /longest-context-window/ is built on.
No changes. Every OpenAI rate that appeared in this fetch matched our stored values exactly, including the GPT-5.6 Luna cut applied by hand on 2026-08-11 ($0.20 / $1.20), which the page now confirms — closing out the review opened on 2026-08-02. Sol ($5/$30), Terra ($2/$12), GPT-5.5, 5.4, 5.4-mini, 5.2, 5.1, 5, 4.1, o3, o4-mini and 4o all re-confirmed. The fetch returned a partial lineup — roughly a dozen models against the thirty we track, omitting the Pro and Nano tiers and GPT-5.3 — so no deprecation conclusions were drawn from the gaps.
HUMAN-VERIFIED, now applied. The 80% cut flagged on 2026-08-02 (see the flagged entry below) was confirmed against OpenAI's pricing page and corroborated by multiple third-party trackers — OpenAI cut Luna 80% on 2026-07-30, from $1.00/$6.00 (cached $0.10) to $0.20/$1.20 (cached $0.02). The swing exceeded the automated 60% guard, so it was correctly held for manual review; this entry records the verified application. Luna now leads the /cheapest-ai-model/ ranking.
OpenAI cut GPT-5.6 Terra by 20% in both directions one week after launch. Cached input drops in step, from $0.25 to $0.20 per 1M. Terra now undercuts GPT-5.4 ($2.50/$15) on price while carrying roughly 2.6x the context window, which makes the older model hard to justify for new work. The other two GPT-5.6 tiers were not cut in the same proportion: Sol held at $5/$30, and the figure shown for Luna was too large a swing to apply automatically (see the Luna entry below).
FLAGGED FOR MANUAL REVIEW — no change applied. OpenAI's pricing page listed GPT-5.6 Luna at $0.20 input / $1.20 output (cached $0.02) against our stored $1.00 / $6.00 (cached $0.10). That is an 80% reduction, well past the 60% threshold at which we stop trusting an automated read and require a human to confirm. Two things argue it may be genuine rather than a parse error: it is an exact 5x cut applied consistently across input, output and cached input, and its sibling Terra was independently confirmed cut 20% in the same fetch. But an 80% swing is also exactly what a misread pricing table looks like, and Luna feeds the /cheapest-ai-model/ ranking, where a wrong figure would reorder the page. The stored rate stands at $1.00 / $6.00 until verified by hand.
STILL DEFERRED for a third run — no change applied. Gemini 3.6 Flash ($1.50/$7.50), Gemini 3.5 Flash ($1.50/$9.00) and Gemini 3.5 Flash-Lite ($0.30/$2.50, new to this run) all have confirmed, stable prices, and the models overview now lists all three as Stable with model ids gemini-3.6-flash, gemini-3.5-flash and gemini-3.5-flash-lite. The blocker is unchanged: Google still does not publish an input token limit for any of them, on either the models overview or the Vertex AI reference. Since every model entry feeds the /longest-context-window/ ranking, adding them would mean inventing the one figure that page is built on. They stay out until Google publishes the limit.
DEFERRED — no change applied. Together's pricing page has picked up a batch of models we do not yet track, most notably GLM-5.2 ($1.40/$4.40, priced identically to the GLM-5.1 we already carry) and a Qwen3.7 line (Plus at $0.32/$1.28, Max at $1.25/$3.75), alongside Kimi K3, MiniMax M3 and Qwen3.6-Plus. Prices are clear, but the pricing table alone gives neither a context window nor an exact API model string, and Together's per-model pages for GLM-5.2 returned 404 on the slug patterns tried. Deferred rather than guessed. Existing entries were all re-confirmed against this fetch: Llama 3.3 70B ($1.04/$1.04), DeepSeek V4 Pro ($1.74/$3.48), GLM-5.1 ($1.40/$4.40) and Qwen3.5 397B ($0.60/$3.60) are unchanged.
POSSIBLE DEPRECATION, VERIFY — no change applied, all entries left active. GPT-5.3 ($1.75/$14), o3-pro ($20/$80) and GPT-4 Turbo ($10/$30) did not appear in this run's OpenAI pricing fetch. A single absent fetch is not evidence of removal — pricing pages routinely drop older tiers into collapsed or secondary tables — so per policy the entries stay live and unchanged. If they are still missing on the next two runs, that is worth a manual check against OpenAI's deprecations page. Every other OpenAI model we track was re-confirmed at its stored rate.
Added Qwen3.5 397B at $0.60 input / $3.60 output per 1M tokens (cached input $0.35) with a 262,144-token native context window. This clears one of the models deferred earlier today: the price was confirmed on Together's pricing page, and the context window and exact API id (Qwen/Qwen3.5-397B-A17B) have now been sourced from Together's own model page, so no figure is being guessed. A 397B-parameter MoE with 17B active, it leads the Qwen3.5 generation that has displaced Qwen3 Coder 480B and Qwen 2.5 72B on Together's main pricing table. Note that it breaks the older Qwen convention of charging the same rate in both directions — output is 6x input here.
STILL DEFERRED — no change applied. Re-checked the Gemini 3.6 Flash ($1.50/$7.50) and Gemini 3.5 Flash ($1.50/$9.00) entries held back earlier today. Prices re-confirmed on the pricing page, but the context windows are still unpublished: the Gemini models overview lists both as Stable without token limits, and the Vertex AI model reference does not state them either. Held back for a second run rather than shipping a guessed context window that would feed the /longest-context-window/ ranking. Separately, Gemini 3 Flash — which did not appear in this run's pricing fetch — is confirmed still available on the models overview under Preview status, not in the deprecated section, so its entry stays active and unchanged at $0.50/$3.00.
Anthropic shipped Claude Opus 5. The claude-opus entry now resolves to Opus 5 (apiId: claude-opus-5), following the same convention used for the 4.7 to 4.8 repoint. Standard-mode pricing is unchanged at $5 input / $25 output per 1M tokens — Anthropic has now held the Opus rate flat across 4.5, 4.6, 4.7, 4.8 and 5. Opus 4.8 and 4.7 remain available at the same rate.
Fast mode (still a research preview) now covers Claude Opus 5 as well as Opus 4.8, at an unchanged $10 input / $50 output per 1M tokens. Fast-mode pricing applies across the full context window including requests over 200k input tokens. Claude API first-party only — not on Bedrock, Google Cloud, or the Batch API, and not available on Opus 4.7 or 4.6.
The claude-sonnet entry now resolves to Claude Sonnet 5 (apiId: claude-sonnet-5). Sonnet 5 is on INTRODUCTORY pricing of $2 input / $10 output through 2026-08-31; standard pricing of $3 / $15 takes effect 2026-09-01. This is a scheduled, dated increase published by Anthropic, not a promotion of unknown duration — anyone budgeting past August should plan on $3 / $15. Sonnet 4.6 remains available at $3 / $15.
Corrected the context window for claude-opus, claude-opus-fast and claude-sonnet from 200K to 1M. Anthropic's pricing docs now state that Claude 4.6 and later models include the full 1M-token context window at standard pricing, with no long-context surcharge — a 900k-token request bills at the same per-token rate as a 9k one. Prompt caching and batch discounts apply at standard rates across the whole window. Claude Haiku 4.5 is below the 4.6 cutoff and stays at 200K.
Added Claude Fable 5 at $10 input / $50 output per 1M tokens with a 1M-token context window — a new tier priced above Opus. Note that Fable 5 uses the newer Claude 4.7+ tokenizer, which produces roughly 30% more tokens for the same text than Sonnet 4.6 and earlier, so the real cost gap versus older Claude models is wider than the headline rate. Claude Mythos 5 is priced identically but was NOT added: it is limited-availability (via anthropic.com/glasswing) rather than generally purchasable.
Added the GPT-5.6 line: Sol ($5/$30, cached $0.50), Terra ($2.50/$15, cached $0.25) and Luna ($1/$6, cached $0.10). All three carry a 1.05M-token context window, up from 400K across the GPT-5.0 to 5.5 generations. Sol and Terra match the headline rates of GPT-5.5 and GPT-5.4 respectively while offering ~2.6x the context, so there is little reason to stay on the older models at the same price.
Together.ai raised Llama 3.3 70B from $0.88 to $1.04 per 1M tokens, an 18% increase. Within the plausible-swing threshold and confirmed on the live pricing page, so applied.
Added DeepSeek V4 Pro at $1.74 input / $3.48 output per 1M tokens (cached input $0.20) with a 512K-token context window — the largest context of any open-weight model we track, and roughly 4x the 128K that the V3 generation offered. It has replaced V3.1 as the DeepSeek entry on Together's main pricing page.
NOT ADDED PENDING VERIFICATION. Three new Gemini models appeared with confirmed prices — Gemini 3.6 Flash ($1.50/$7.50), Gemini 3.5 Flash ($1.50/$9.00) and Gemini 3.5 Flash-Lite ($0.30/$2.50) — but neither the pricing page nor the models overview page publishes their context windows or exact API model ids, and the per-model doc pages 404. Rather than guess a context window (which would feed the /longest-context-window/ ranking directly), these are held back until the figures can be sourced. All six existing Gemini entries were re-verified against the live page and are unchanged.
NOT ADDED PENDING VERIFICATION. Together listed several new models whose context windows could not be sourced (their model pages 404): GLM-5.2 ($1.40/$4.40, identical to GLM-5.1), Qwen3.7-Plus ($0.32/$1.28), Qwen3.6-Plus ($0.50/$3.00), Qwen3.5-397B-A17B ($0.60/$3.60) and Qwen3.5 9B ($0.17/$0.25). Prices recorded here for the next run; entries held back rather than shipped with a guessed context window.
POSSIBLE DEPRECATION, VERIFY — no change applied. The GPT-5.0 through 5.2 tiers, the GPT-4.1 family, the o3/o4 reasoning series and the GPT-4o family did not appear in this run's fetch of either the pricing page or the models page, which now foreground the GPT-5.6 line. That is suggestive but not proof of removal — the fetch may simply have captured the flagship section. All entries left active per the never-remove-on-one-fetch rule. Worth a manual check of the OpenAI deprecations page before flipping any flags.
POSSIBLE DEPRECATION, VERIFY — no change applied. DeepSeek V3.1, DeepSeek R1, Qwen3 Coder 480B and Qwen 2.5 72B are no longer visible on Together's serverless pricing table, which now leads with DeepSeek V4 Pro and the Qwen3.5+ generation. These entries were already flagged as indicative pricing and are left active. Anyone billing against them should confirm their provider's current rate.
Anthropic's modest Opus 4.8 upgrade lands at the same standard price but quietly drops the fast-tier 3x. Worth the model-string swap; worth a serious look if latency is your bottleneck.
Anthropic released Claude Opus 4.8 on 2026-05-28. Standard mode pricing unchanged at $5 input / $25 output per 1M tokens. The claude-opus entry now resolves to 4.8 (apiId: claude-opus-4-8). Inherits the Opus 4.7 tokenizer behavior (up to 35% more tokens than legacy Claude models for the same text). Anthropic describes 4.8 as 'a modest but tangible improvement' with gains in agentic coding, reasoning, knowledge work, and honesty.
Added Claude Opus 4.8 Fast Mode as a separate entry: $10 input / $50 output per 1M tokens, producing tokens at ~2.5x normal speed. Anthropic dropped the fast-tier price 3x vs Opus 4.7 (was $30/$150). Useful for latency-sensitive workloads where Opus quality is needed but the standard tier's throughput is the bottleneck.
Shipped real BPE tokenization for the Llama family via llama-tokenizer-js (lazy-loaded ~2MB chunk on first Llama count). All llama-* models now labeled 'exact' instead of '≈±3%'. Mistral / Qwen / DeepSeek / GLM still use heuristic (now character-class-aware — buckets text into ASCII / digit / CJK / whitespace and applies per-class ratios — more accurate than the prior constant-ratio version, still labeled ≈±3%).
Added 5 OSS models: Llama 3.3 70B ($0.88/$0.88, current Together flagship Meta), DeepSeek V3.1 ($0.60/$1.70 Together listing), DeepSeek R1 ($3/$7 reasoning), Qwen3 Coder 480B ($2/$2 current Alibaba coding flagship), GLM-5.1 ($1.40/$4.40 new Zhipu provider). Llama 3.1 entries kept but flagged 'no longer on Together's main page — verify provider'. Llama 4 NOT added — not on Together's current pricing page.