For most business work in September 2026, Claude Opus 5.5 is the best value of the four: it matches or beats the two flagships on most published benchmarks at well under half their price. Claude Fable 5.1 is for the hardest, longest-running agent jobs. GPT-6 Astra leads on a few specialist benchmarks, including business workflow automation and science, and costs the same as Fable 5.1. Jev is not a chatbot at all. It makes fast, cheap yes/no and multiple-choice decisions inside software, and it is the one to watch if you run high-volume automations.
All four came out between 1 and 22 September 2026. Here is how they compare, with every number sourced at the bottom.
Benchmark scores are as published by the model makers. The GPT-6 Astra scores below come from Anthropic's comparison table, because OpenAI's own launch materials did not publish matching figures. Treat vendor benchmarks as a guide, not a guarantee.
The four models in one table
| Claude Opus 5.5 | Claude Fable 5.1 | GPT-6 Astra | Jev | |
|---|---|---|---|---|
| Made by | Anthropic | Anthropic | OpenAI | TypeSafe AI |
| Released | 22 Sep 2026 | 1 Sep 2026 | 3 Sep 2026 (preview), 4 Sep (public) | 15 Sep 2026 (early access) |
| What it is | General-purpose LLM | General-purpose LLM, Anthropic's largest | General-purpose LLM, OpenAI's flagship | "System One" decision model, not an LLM |
| Output | Text, code | Text, code | Text, code | Typed answers with probabilities |
| Input price per million tokens | $4 | $10 | $10 | $0.042 |
| Output price per million tokens | $20 | $50 | $50 | Free |
| Context window | 1M tokens | 1M tokens | 1.05M tokens | Not published |
| Max output | 128K tokens | 128K tokens | 128K tokens | n/a |
| Knowledge cutoff | June 2026 | June 2026 | April 2026 | Not published |
| Available to | Everyone | Everyone | Everyone (phased rollout) | Waitlist only |
Claude Opus 5.5: the new default
Opus 5.5 is Anthropic's newest model and, unusually, it outscores Anthropic's own larger model. On Anthropic's figures it beats Fable 5.1 on agentic coding, knowledge work and reasoning, generates output more than 30% faster than Opus 5, and costs around 40% less to run than Opus 5 on typical workloads. Anthropic's documentation now tells developers to start with Opus 5.5 for most work.
It runs at medium effort by default, which keeps answers quick and cheap, and can be turned up to high or max for harder jobs.
Best for: most business AI work. Website chat assistants, drafting and summarising, research, coding, and AI agents that do multi-step tasks. We covered the release in full in Claude Opus 5.5 is out.
Claude Fable 5.1: for the longest, hardest jobs
Fable 5.1 is Anthropic's top-tier model, positioned for ambitious coding projects, long-horizon agents and multi-day autonomous sessions. Anthropic describes it as the same underlying model as Claude Mythos 5.1, which is only available to vetted organisations through trusted access programmes. Fable 5.1 is the generally available version with standard safeguards.
It costs two and a half times as much as Opus 5.5 per token. Three weeks after launch, Anthropic's own guidance is to reach for it only when Opus 5.5 at higher effort still falls short.
Best for: very long autonomous jobs, large codebases and anything where Opus 5.5 has been tested and genuinely is not enough.
GPT-6 Astra: OpenAI's flagship
GPT-6 Astra is OpenAI's newest model, described by OpenAI as state of the art on computer use, browsing, professional work, software engineering, cybersecurity and science. It was trained on more than 100,000 GPUs at OpenAI's Stargate site in Texas and uses a reasoning technique OpenAI calls "recurrent depth", also described as looped transformers. OpenAI president Greg Brockman called it a "generational leap", according to Axios.
Two points worth knowing. The public version is restricted and rejects some cybersecurity prompts, with fuller access through a programme called Daybreak. And some critics have raised concerns that recurrent depth makes the model's reasoning harder to inspect, which matters if you need to explain why an AI made a decision.
Best for: teams already on ChatGPT or OpenAI's API, computer-use and browsing tasks, and scientific or research work, where it leads the published comparison.
Jev: a different kind of model
Jev, from San Francisco start-up TypeSafe AI, does not write anything. You give it some information and a set of questions with fixed answer options, and it returns an answer to each with a calibrated confidence score. TypeSafe calls this a "System One" model, after the fast, intuitive mode of thinking in psychology.
That makes it useless for writing an email and very good for deciding which inbox an email belongs in. TypeSafe claims responses in 70 to 500 milliseconds and says it is 40 to 200 times faster than frontier LLMs for these tasks. Because the answer can only be one of the options you defined, it cannot make one up.
Best for: high-volume classification and routing inside automations. Sorting enquiries, scoring leads, flagging risky messages, checking another AI's output. We wrote a full guide in What is Jev?
Benchmarks side by side
These are from Anthropic's Opus 5.5 launch page. Jev is not included because it does not do these kinds of tasks.
| Benchmark (what it tests) | Opus 5.5 | Fable 5.1 | GPT-6 Astra |
|---|---|---|---|
| Terminal-Bench 4.0 (agentic coding) | 66.4% | 55.8% | 57.9% |
| FrontierCode v1.1 (agentic coding) | 54.4% | 50.3% | 53.3% |
| GDPval-AA v2.1 (real knowledge work, Elo) | 1846 | 1735 | 1542 |
| AutomationBench (business workflows) | 40.0% | 31.4% | 41.4% |
| Humanity's Last Exam, with tools (reasoning) | 67.7% | 65.6% | 57.2% |
| Terminal-Bench-Science 0.1 (research tasks) | 58.7% | 52.6% | 64.6% |
Opus 5.5 leads four of the six. GPT-6 Astra leads the other two. Fable 5.1 does not lead any, which says more about how good Opus 5.5 is than how weak Fable is.
What each one costs for the same job
Take one AI-handled customer enquiry: 2,000 tokens in, 1,000 tokens out. Multiplied by a thousand enquiries:
| Model | Cost per 1,000 enquiries |
|---|---|
| Claude Opus 5.5 | about $28 |
| Claude Fable 5.1 | about $70 |
| GPT-6 Astra | about $70 |
Jev cannot write the reply, so it is not a like-for-like comparison. But if the job is only to decide which kind of enquiry it is, at $0.042 per million input tokens with free output, 2,000 tokens in costs under a hundredth of a cent. A thousand of them costs about 8 cents.
That is the real point of Jev. You would not replace your chatbot with it. You would put it in front, so the expensive model only gets called when it is actually needed.
Which one should you use?
If you are a business owner choosing an AI subscription: Claude with Opus 5.5 or ChatGPT with GPT-6 Astra will both handle day-to-day writing, research and admin well. Pick the one your team already knows. The difference between them matters less than using either one properly.
If you are building a customer-facing AI assistant: Opus 5.5. It is cheaper than the flagships, gives shorter answers, and Anthropic reports it tied for the lowest prompt injection rate in its testing, which matters when strangers are typing into it.
If you are automating a high-volume process: use two models. Something like Jev to sort and score, and Opus 5.5 to write whatever needs writing. Jev is waitlist only for now, so for most businesses the practical version today is Opus 5.5 at low effort doing the sorting, with Jev as the cost-cutting upgrade once it opens up.
If you are running long autonomous coding or research jobs: test Opus 5.5 at high effort first, then Fable 5.1 or GPT-6 Astra if it falls short.
This is the kind of decision we make when we build AI agents and automations for clients. The model is one part of it. What it is connected to, and what happens when it gets something wrong, matters more.
Sources
- Introducing Claude Opus 5.5, Anthropic
- Introducing Claude Fable 5.1 and Claude Mythos 5.1, Anthropic
- Models overview, Claude Platform documentation
- Introducing GPT-6 Astra, OpenAI Developer Community
- GPT-6 Astra, Wikipedia
- OpenAI releases new model GPT-6 Astra, says it may represent AGI, Axios
- Jev (AI model), Wikipedia
- Jev: TypeSafe's System One Model Explained, DataCamp
- A new kind of AI model from a ChatGPT inventor is thrilling developers, TechCrunch
Last checked against these sources on 22 September 2026.
Frequently asked questions
Is Claude Opus 5.5 better than GPT-6 Astra?
On the benchmarks Anthropic published, Opus 5.5 leads on agentic coding, knowledge work and reasoning, while GPT-6 Astra leads on AutomationBench and Terminal-Bench-Science. Opus 5.5 is also cheaper, at $4 input and $20 output per million tokens against $10 and $50 for GPT-6 Astra.
What is the difference between Claude Opus 5.5 and Claude Fable 5.1?
Fable 5.1 is Anthropic's largest model, built for long-running agent work, and costs $10 input and $50 output per million tokens. Opus 5.5 costs $4 and $20, and on Anthropic's published benchmarks it matches or beats Fable 5.1 on most tasks. Anthropic recommends starting with Opus 5.5 and moving to Fable 5.1 only if it falls short.
What is Jev and how is it different from ChatGPT or Claude?
Jev is a model from TypeSafe AI that returns structured decisions with confidence scores instead of writing text. You define the questions and the possible answers, and Jev picks one. It is built for fast, cheap classification inside software, not for conversation.
What is the cheapest of these AI models?
Jev, at $0.042 per million input tokens with free output, but it only makes decisions and cannot write. Of the three models that write text, Claude Opus 5.5 is the cheapest at $4 input and $20 output per million tokens.
Which AI model is best for a small business?
For most small businesses, Claude Opus 5.5 is the best balance of quality and cost in September 2026. For simple day-to-day use, the one your team already knows, Claude or ChatGPT, matters more than the benchmark differences.
Is GPT-6 Astra the same as Astra-6?
Yes. The model is officially called GPT-6 Astra. It was released by OpenAI in limited preview on 3 September 2026 and publicly on 4 September 2026.