claude vs chatgpt vs gemini vs grok: which ai to use in september 2026
as of september 2026, the short answer to claude vs chatgpt vs gemini is this. use claude opus 5.5 for writing, analysis and coding agents. use gpt-6 astra for computer use and the longest documents. use gemini 3.1 pro if you live in google workspace and want the best $19.99 plan. grok 4.7 is the cheapest flagship api of the four, but in our view the weakest consumer pick.
all four labs shipped a new model inside the last 30 days, so treat this as a snapshot, not a ranking for the year.
what shipped in september 2026?
four labs shipped new models in the same month. anthropic released claude fable 5.1 on september 1, then claude opus 5.5 on september 22, which it says performs at the level of fable 5.1 on most work, at 40% of fable's api price. openai launched gpt-6 astra on september 3, rolling out to chatgpt plans over the following days, and added gpt-6 sol and luna on september 22, about an hour after opus 5.5.
google shipped gemini 3.8 flash on september 2. its top pro model is still gemini 3.1 pro, released february 19, 2026. xai released grok 4.7 on september 21, at the same price as grok 4.6.
| model | released | context window | api price per 1m tokens (in / out) |
|---|---|---|---|
| claude opus 5.5 | sep 22, 2026 | 1m | $4 / $20 |
| claude fable 5.1 | sep 1, 2026 | 1m | $10 / $50 |
| gpt-6 astra | sep 3, 2026 | 1,050,000 | $10 / $50 |
| gemini 3.1 pro | feb 19, 2026 | 1m | $2 / $12 |
| gemini 3.8 flash | sep 2, 2026 | 1,048,576 | $0.75 / $3.75 until dec 31, 2026 |
| grok 4.7 | sep 21, 2026 | 500,000 | $2 / $6 |
api prices come from anthropic, openai's model page, google and xai's docs. gemini 3.1 pro and grok 4.7 cost more above 200,000 prompt tokens, and gemini 3.8 flash's introductory price doubles to $1.50 / $7.50 on january 1, 2027.
which ai is best for writing?
claude, and in our view it is not close for everyday prose. anthropic's pitch for opus 5.5 is that testers found its writing clearer and that it puts the most important information up front, which is exactly what a status update or a client email needs.
the crowd leans the same way. on the arena text leaderboard, updated september 13, 2026 from over 8.1 million votes, anthropic models held 7 of the top 12 places. gemini 3.8 flash was the best non anthropic, non meta entry at 1493.
gpt-6 astra and grok 4.7 are both capable writers. neither has a top 12 spot on that board as of this writing.
verdict: writing goes to claude opus 5.5.
which is strongest at reasoning and analysis?
the independent number that matters is the artificial analysis intelligence index. as of september 24, 2026, claude opus 5.5 leads at 58 on its max setting. claude fable 5.1 and gpt-6 astra tie at 53, and gpt-6 sol scores 48, each at max.
on professional work, anthropic reports opus 5.5 at 1846 elo on gdpval-aa v2.1, a test of real work across 44 occupations, against 1735 for fable 5.1. that is the vendor's own measurement, so read it as direction rather than gospel.
openai's headline claims for astra are about math and puzzles: 98% on frontiermath tier 4 and 99.9% on arc agi 3. impressive, and not very relevant to reading a spreadsheet and telling you what changed.
verdict: for analysis you will act on, opus 5.5. for hard math, astra.
which is best for research and long documents?
research splits on where your sources live. gemini wins if they sit in gmail, docs and drive, because google ai pro bundles deep research and a 1 million token window for $19.99. google's model card puts gemini 3.1 pro at 51.4% on humanity's last exam with search and code tools. anthropic's fable 5.1 launch post reports 65.0% on the same exam with tools.
for sheer length, gpt-6 astra holds 1,050,000 tokens in the api, with a 128,000 token output cap. prompts over 272,000 input tokens are billed at 2x input and 1.5x output for the whole request. claude opus 5.5 reads 1 million tokens and writes up to 128,000. gemini 3.1 pro reads 1 million and writes up to 64k. grok 4.7 stops at 500,000.
a big window is not the same as good recall at the end of it. we explain why in what is a context window.
verdict: web and inbox research to gemini, a 900 page data room to astra.
which ai is best at agentic tasks?
this is the category that moved most in 2026, and the one that matters most if you want an ai agent to finish work rather than suggest it.
- coding agents: anthropic reports opus 5.5 at 66.4% on terminal-bench 4.0 against 57.9% for gpt-6 astra, and says one tester finished a 680,000 line code migration in less than a day.
- computer use: openai calls astra "the world's best computer use model", able to fill out online forms, update crm records and organize your calendar. it reports 72.6% on osworld 2.0 (offline set, partial score). anthropic reports opus 5.5 at 81.8% partial on osworld 2.0, so the vendor numbers do not settle it.
- cheap volume: grok 4.7 at $2 / $6 and gemini 3.8 flash at $0.75 / $3.75 are the models we would run a thousand small jobs on.
if you want to see how desktop agents drive a mac, read computer use agents on a mac.
verdict: claude for code, astra if you want computer use inside chatgpt work and codex.
how much do the consumer plans cost?
the entry tiers have converged on about $20 a month, and on the web the chatgpt plus price matches claude pro exactly. the power tiers now start at about $100.
this covers the consumer plans. for api prices per million tokens across every tier, see ai model pricing compared.
| plan | price per month | flagship you get |
|---|---|---|
| claude pro | $20, or $17 billed yearly | opus 5.5; fable only through usage credits |
| claude max | from $100 | 5x or 20x pro usage, fable up to 50% of weekly limits |
| chatgpt plus | $20 | gpt-6 models, including astra |
| chatgpt pro | from $100 ($200 for pro 20x) | 5x usage, pro reasoning on astra |
| google ai plus | $4.99 | gemini, 2x free limits |
| google ai pro | $19.99 | gemini 3.1 pro, deep research, 5 tb |
| google ai ultra | $99.99 or $199.99 | 5x or 20x pro usage |
| supergrok | $30 | higher limits, multi agent reasoning |
| supergrok heavy | $300 | highest usage at the fastest speed |
claude, chatgpt and google prices come from the vendors' own pages (claude, chatgpt, google). the chatgpt pro 20x and supergrok prices come from the vendors' chatgpt and grok app store listings, where chatgpt plus shows as $19.99.
the verdict: pick this one if...
here is the best ai model for work by job, as of september 24, 2026.
| pick | if you mostly... |
|---|---|
| claude opus 5.5 | write, analyze documents, or run coding agents |
| gpt-6 astra | need an agent to operate apps, or feed it the longest files |
| gemini 3.1 pro | live in gmail and docs and want the best value $20 plan |
| gemini 3.8 flash | want fast, cheap answers at high volume |
| grok 4.7 | want the cheapest flagship api and follow news on x |
if you can only pay for one, pay for claude pro. if money is tight, google ai pro is the best value on the list.
the model will change. your context shouldn't.
look at the dates above. opus 5 lasted two months, from july 24 to september 22, before opus 5.5 replaced it. google shipped three flash models in six weeks, by its own count. the best answer to claude vs chatgpt vs gemini in december will not be the answer today.
that is an argument for keeping your working context outside any one vendor. each assistant's memory is scoped to its own product, so switching means starting over.
remynd keeps that context on your side. it is a mac app that captures the focused window, runs ocr locally with apple vision, and keeps the searchable index on your mac. you can exclude apps and sites, and recordings are kept for 30 days by default. call transcription runs on device.
the agent window runs your own claude code or codex cli against that history with read only access, so you can swap models without losing what you were working on. settings also take a custom api endpoint compatible with the openai responses api. sign in and cloud agents do reach the network, so the honest claim is that the archive lives on your mac. the details are in private ai on your mac.
download remynd for mac and let the models take turns.