← all articles

ai model pricing in september 2026: what a month of work really costs

ai model pricing as of september 2026 splits into three bands. flagships like claude fable 5.1 and gpt-6 astra cost $10 per million input tokens and $50 per million output. mid priced models (claude sonnet 5, gpt-6 sol and google's gemini 3.1 pro) sit at $2 in and $10 to $12 out. fast models run from $0.10 to $1 in.

for a realistic month of agent heavy work, that is the difference between $1.74 and $174. our verdict: default to a mid tier model with caching on, and only pay flagship rates for the hard 10% of tasks.

every api price below comes from the vendor's own pricing page, checked on september 24, 2026. the tier labels are our own grouping by price.

what does each model cost per million tokens?

this is the core llm api pricing table. prices are standard, pay as you go, in usd per million tokens, for prompts under 200k tokens.

vendormodeltierinputcached inputoutput
anthropicclaude fable 5.1flagship$10$0.25$50
anthropicclaude opus 5.5flagship$4$0.20$20
anthropicclaude sonnet 5mid$2$0.20$10
anthropicclaude haiku 4.5fast$1$0.10$5
openaigpt-6 astraflagship$10$1$50
openaigpt-6 solmid$2$0.20$10
openaigpt-6 lunafast$0.10$0.01$0.50
googlegemini 3.1 pro previewflagship$2$0.20$12
googlegemini 3.8 flashmid$0.75$0.075$3.75
googlegemini 3.5 flash-litefast$0.30$0.03$2.50
xaigrok-4.7flagship$2$0.50$6
xaigrok-4.3mid$1.25$0.20$2.50
deepseekdeepseek v4 pro (peak)open weight$1.32$0.044$3.96
deepseekdeepseek-flash (peak)open weight$0.30$0.006$1.20

sources: claude api pricing, openai api pricing, gemini api pricing, xai pricing and deepseek pricing.

a few things the table hides:

  • claude sonnet 5 stayed cheap. its $2 / $10 launch price was meant to be introductory until august 31, 2026. anthropic cancelled the planned september 1 rise to $3 / $15, so $2 / $10 is now the standard rate.
  • gemini flash has an expiry date. gemini 3.8 flash is $0.75 / $3.75 through december 31, 2026, then doubles to $1.50 / $7.50 on january 1, 2027.
  • long prompts cost more on google, xai and openai. gemini 3.1 pro goes to $4 / $18 above 200k tokens, grok-4.7 doubles to $4 / $12, and gpt-6 models charge 2x input and 1.5x output above 272k. claude 4.6 and later bill the full 1m window at the standard rate.
  • deepseek halves off peak. the peak window is 01:00 to 04:00 and 06:00 to 10:00 utc, monday to friday, excluding chinese public holidays. outside it, every price above is cut in half.
  • tokens are not the same size. anthropic says the tokenizer in claude 4.7 and later produces roughly 30% more tokens for the same text. haiku 4.5 uses the older one.

how much do caching and batch discounts save?

caching is the biggest lever on any bill, because agents resend the same context on every turn. a cached input token costs 10% of the base input price on most claude models, 5% on opus 5.5 and 2.5% on fable 5.1. openai and google also price cached input at roughly a tenth of fresh input. anthropic does charge 1.25x to write a 5 minute cache entry, or 2x for 1 hour, and openai charges 1.25x for cache writes on gpt-5.6 and later.

batch is the second lever. anthropic, openai and google all take 50% off for asynchronous batch jobs. openai's flex tier charges batch rates for regular requests that accept slower responses. xai offers 20% off batch, and only on grok-4.3 and the grok-4.20 models.

cheap models also run on third party hosts. on together ai, openai's open weight gpt-oss 120b is $0.15 in and $0.60 out, and alibaba's qwen3.8 flash is $0.09 in and $0.28 out.

what does a realistic month of work cost on each model?

here is a worked example, so the cost per million tokens turns into a bill. picture one person running a coding or research agent most working days in a month. that is about 20 million input tokens and 2 million output tokens. agents reread their context constantly, so assume 14 million of those input tokens are cache hits.

modelno cachingwith caching (14m cached)
claude fable 5.1$300.00$163.50
gpt-6 astra$300.00$174.00
claude opus 5.5$120.00$66.80
gemini 3.1 pro preview$64.00$38.80
claude sonnet 5$60.00$34.80
gpt-6 sol$60.00$34.80
grok-4.7$52.00$31.00
claude haiku 4.5$30.00$17.40
grok-4.3$30.00$15.30
deepseek v4 pro (peak)$34.32$16.46
gemini 3.8 flash$22.50$13.05
gemini 3.5 flash-lite$11.00$7.22
deepseek-flash (peak)$8.40$4.28
gpt-6 luna$3.00$1.74

the cached column ignores the cache write surcharges at anthropic and openai and google's hourly cache storage fee. it also ignores the tokenizer gap, which could add up to 30% to the newer claude rows for the same work.

two readings jump out. first, caching cuts most bills by 35 to 50%. second, the priciest flagship month costs 100 times the cheapest fast model month. if your agent spends its day grepping files and summarising logs, you are paying flagship rates for fast model work. we covered why agents burn so much context in what is a context window.

are the $20 and $100 to $200 plans a better deal?

for heavy users on flagship models, a subscription is often the better deal in our view. it caps your spend at a known number, while our worked flagship month runs $163.50 to $174 on the api.

vendorentry planhigh usage plans
anthropicclaude pro, $20 a month ($17 on annual billing)max 5x at $100, max 20x at $200
openaichatgpt plus, $20 a monthpro at $100 or $200
googlegoogle ai pro, $19.99 a monthgoogle ai ultra at $99.99 (5x) or $199.99 (20x) a month
xaisupergrok, $30 a month (app store)supergrok plus, $100 a month (app store)

sources: claude plans, max plan details, chatgpt plus, chatgpt pro tiers, google ai plans and the grok app store listing.

one caveat on openai. since september 10, 2026, openai has temporarily paused new sign ups and upgrades to the $200 pro tier. the $100 tier is still open.

on chatgpt plus vs claude pro, the price is identical at $20 monthly, so choose on the model and the tools. on the web, claude pro is the only one of the two with an annual discount. google ai pro is the value pick if you also want its 5 tb of storage.

which model should you use for what?

stop picking one model for everything and pick one per job instead.

  1. hard reasoning, gnarly refactors, final drafts: claude opus 5.5. it is 60% cheaper than fable 5.1 on input and output, and its $0.20 cache reads tie gemini 3.1 pro for the cheapest of any flagship in our table.
  2. daily agent and coding work: claude sonnet 5 or gpt-6 sol. both are $2 / $10 and about $35 for our cached month. this is the default for most people.
  3. long documents over 200k tokens: claude sonnet 5, because it keeps its standard rate across the full 1m window, while gemini 3.1 pro and grok-4.7 charge more past 200k and gpt-6 sol past 272k.
  4. high volume classification, extraction, routing: gpt-6 luna. at $1.74 for our cached month it is in a class of its own.
  5. cheap and open weight: deepseek-flash, run off peak, or gpt-oss 120b on a host like together ai.
  6. you chat all day and never touch an api: a $20 plan. move to a $100 tier only once you hit limits weekly.

if you are wiring these into an agent, what is an ai agent explains where the tokens go.

your context should not be locked to one model

look at the dates in this post. a price that doubles on january 1. a plan tier paused on september 10. an introductory rate that became the standard price. the best model for your budget will change again next quarter.

so the thing worth protecting is your context, not your choice of vendor. provider memory is tied to one product, which is the trap we laid out in ai memory compared.

remynd is a mac app that captures the focused window, runs ocr locally with apple vision and keeps the searchable index on your mac. you can exclude apps and sites, and recordings are kept for 30 days by default. its agent window runs your own claude code or codex cli against that history, with read only access, so you can switch models without losing context. more in ai coding agent memory of your work.

if you want a cheaper model, settings let you add a custom api endpoint compatible with the openai responses api. call transcription runs on device with mlx. some features do reach the network, so the honest claim is that the archive lives on your mac. details on our security page.

download remynd for mac and keep your context when you switch models.

common questions

what is the cheapest capable ai model api in september 2026? +
for everyday work the cheap end is gpt-6 luna at $0.10 in and $0.50 out per million tokens, deepseek-flash at $0.30 in and $1.20 out at peak, and gemini 3.5 flash-lite at $0.30 in and $2.50 out. with caching on, all three cost under $10 for our worked month of heavy agent use.
is claude haiku 4.5 cheaper than gemini flash? +
no. haiku 4.5 is $1 in and $5 out per million tokens, while gemini 3.8 flash is $0.75 in and $3.75 out through december 31, 2026. that flash price doubles on january 1, 2027, at which point haiku becomes the cheaper of the two.
should i pay for chatgpt plus or claude pro? +
both cost $20 a month billed monthly, and claude pro drops to $17 a month on annual billing. if you mostly chat, pick the model you prefer. if you run agents all day on flagship models, a $100 tier can come in under the api bill, since our worked month costs $163.50 to $174 on the flagships even with caching. on a mid tier model the same month is about $35, so the api is cheaper.
how much do prompt caching and batch discounts actually save? +
on claude, openai and gemini a cached input token costs about a tenth of a fresh one, or less, and batch jobs are half price. in our worked example caching alone cut the monthly bill for most models by roughly 35 to 50 percent.
do token prices compare directly across vendors? +
not quite. anthropic says its newer tokenizer, used from claude 4.7 onward, produces about 30% more tokens for the same text. so a claude model with the same sticker price as a rival can cost more for identical work.