Anthropic, OpenAI and Google DeepMind released four major models in a single month. Claude Fable 5.1 launched on September 1. GPT-6 Astra followed on September 3. GPT-6.1 Sol and Gemini 4 Argon arrived in the final days of September. While launch announcements suggested distinct strengths, the benchmark scores overlap significantly. The differences lie in price, access rules, and cost per task.
In this article
One change altered the lineup. OpenAI cancelled GPT-6.1 Astra on September 28 after it failed internal scope and authorization tests. GPT-6 Astra remains OpenAI’s top model for now.
Specs, Pricing and Access
Astra and Fable 5.1 share the same $10 input and $50 output list price. Sol and Argon list at one-fifth of that. Argon’s price is introductory and doubles later.
| Feature | GPT-6 Astra | GPT-6.1 Sol | Gemini 4 Argon | Claude Fable 5.1 |
|---|---|---|---|---|
| Developer | OpenAI | OpenAI | Google DeepMind | Anthropic |
| Released | Sep 3, 2026 | Sep 29, 2026 | Announced Sep 30, 2026 | Sep 1, 2026 |
| Access | OpenAI API, ChatGPT, Codex | OpenAI API, ChatGPT Work, Codex | Fairwind Program only | Claude API, Bedrock, Google Cloud, Microsoft Foundry |
| Input / output (per 1M) | $10 / $50 | $2 / $10 | $2 / $10 intro, then $4 / $20 | $10 / $50 |
| Cached input (per 1M) | $1.00 | $0.10 | $0.10 intro | $0.25 |
| Long-prompt pricing | Above 272K: $20 / $75 | Above 272K: 2x input, 1.5x output | Not disclosed | Flat to 1M |
| Context window | 1.05M | 1.05M | Not disclosed | 1M |
| Max output per response | 128K | 128K | 1M | 128K |
| Reasoning control | 5 effort levels, low to max | 5 effort levels, low to max | Not disclosed | Effort levels incl. low, medium, high |
| Open weights | No | No | No | No |
Sources: OpenAI GPT-6.1 Sol coverage, Gemini 4 Argon coverage, GPT-6 Astra long-context pricing, Claude Fable 5.1 specs. Standard first-party list prices, short-context tier.
The cached-input row matters most for agents. Agents resend system prompts, tool schemas and history on every step. Astra’s $1.00 cache read is 4x Fable 5.1’s and 10x Sol’s.
Argon’s 1M output cap is the only structural outlier. The other 3 stop at 128K tokens per response.
Benchmarks: Where Each Model Leads
No model sweeps the board. Argon leads the knowledge-work and long-horizon coding rows. Astra leads frontier software engineering and computer use. Opus 5.5, not in this lineup, leads Terminal-Bench 4.0.
| Benchmark | What it tests | Gemini 4 Argon | GPT-6 Astra | Claude Fable 5.1 |
|---|---|---|---|---|
| DeepSWE v1.1 | Long-horizon software engineering | 77.9% | 74.1% | 67.4% |
| Vals Index | Finance, coding, legal and tax work | 68.9% | 63.1% | 65.8% |
| FrontierSWE v2 | Frontier software engineering | 55.0% | 65.5% | 56.3% |
| Terminal-Bench 4.0 | Agentic work in a terminal | 57.4% | 58.2% | 57.9% |
| CWE-bench v1 | Vulnerability remediation | 68% (tie) | 68% (tie) | 58% |
| OSWorld-2.0 | Computer use | 69.2% | 72.6% | Not in Google’s table |
Source: Google DeepMind’s published comparison, as reported in our Gemini 4 Argon coverage. Vendor-reported.
GPT-6.1 Sol is not in Google’s table. OpenAI’s own numbers place it close to Astra:
- DeepSWE v1.1: Sol matches Astra at roughly one-fifth of the cost.
- OSWorld 2.0 offline set: Sol lands within 2.1 points of Astra at about one-seventh the cost per task.
- AutomationBench 1.0.6: Sol scores 2.2 points above Claude Opus 5.5 at medium effort.
- Terminal-Bench Science 0.1: Astra still leads at 68.1%. OpenAI recommends Astra for the hardest research.
Independent signals point the other way on raw intelligence. On the Artificial Analysis Intelligence Index, Astra scores 61. Fable 5.1 scores 5 points higher. On its coding-agent index, Fable 5.1 in Claude Code scores 70 against Astra’s 67. Artificial Analysis also reports that Argon equals Astra on the Intelligence Index.
On ARC-AGI-2, Astra scores 95% and Fable 5.1 scores 90%.
Cost per Task: Same List Price, Different Bill
Artificial Analysis puts Claude Fable 5.1 at $9.18 per task, against $4.72 for GPT-6 Astra. That is about 1.9x, at identical list prices.
The gap comes from token volume, not rates. Cost per task multiplies price by tokens spent. With equal rates, the gap implies Fable 5.1 spent more tokens per task in that run. Anthropic also notes its newer tokenizer produces roughly 30% more tokens for the same text.
The two cheaper models change the picture further:
- Gemini 4 Argon: Artificial Analysis reports Argon equals Astra’s Intelligence Index at 60% of Astra’s cost per task, using introductory prices.
- GPT-6.1 Sol: On Terminal-Bench Science, OpenAI reports $5.47 per task for Sol against $23.80 for Astra.
Caching can reverse the ranking for agents. The per-task figures above do not model heavy cache reuse. A long-running agent rereads the same context on every step. Here is the arithmetic for a 200K-token cached context, before output tokens:
| Model | Cached rate (per 1M) | Per step | Per 100 steps |
|---|---|---|---|
| GPT-6 Astra | $1.00 | $0.20 | $20.00 |
| Claude Fable 5.1 | $0.25 | $0.05 | $5.00 |
| GPT-6.1 Sol | $0.10 | $0.02 | $2.00 |
| Gemini 4 Argon (intro) | $0.10 | $0.02 | $2.00 |
Illustrative math from list cache rates. Astra’s 200



