Four frontier models in one month
Sound familiar?
Your team picked a model six months ago. Now every newsletter says a newer one is better, the benchmark charts all disagree, and nobody can tell you whether switching would save money or just create work.
September 2026 was the busiest month for AI model releases so far. Anthropic released Claude Fable 5.1 on September 1. Google followed with Gemini 3.8 Flash on September 2. OpenAI released GPT-6 Astra on September 3, then the smaller GPT-6 Sol and GPT-6 Luna on September 22. Meta's current model, Muse Spark 1.3, is also competing for the same customers.
This comparison is written for people choosing a model for a product, not for a leaderboard. It covers what each model costs, what its maker built it for, and how to decide using your own data.
The spec sheet: flagship models side by side
| Model | Maker | Released | Context window | Max output | Price per 1M tokens |
|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | Sep 3, 2026 | 1.05M tokens (922K input) | 128K | $10 / $50 |
| Claude Fable 5.1 | Anthropic | Sep 1, 2026 | 1M tokens | 128K | $10 / $50 |
| Gemini 3.8 Flash | Sep 2, 2026 | 1M tokens | 65K | $0.75 / $3.75 (introductory) | |
| Muse Spark 1.3 | Meta | 2026 | 1M tokens | Not published | $1.25 / $4.25 |
Two price details change the maths. OpenAI charges 2x the input rate and 1.5x the output rate for GPT-6 Astra requests over 272,000 input tokens. Google's Gemini 3.8 Flash price is introductory until December 31, 2026, and doubles to $1.50 / $7.50 from January 1, 2027.
Prices as of September 28, 2026. Check each provider before you budget: OpenAI pricing, Anthropic pricing, Gemini pricing, Muse Spark.
You are probably not choosing between flagships
The headline models are each maker's most expensive tier. Most production features run on a cheaper model from the same family, and the gap between tiers is far larger than the gap between vendors.
| Tier | OpenAI | Anthropic | Meta | |
|---|---|---|---|---|
| Top: hardest, longest tasks | GPT-6 Astra $10 / $50 | Claude Fable 5.1 $10 / $50 | — | — |
| Workhorse: most product features | GPT-6 Sol $2 / $10 | Claude Opus 5.5 $4 / $20; Claude Sonnet 5 $2 / $10 | Gemini 3.8 Flash $0.75 / $3.75 | Muse Spark 1.3 $1.25 / $4.25 |
| Budget: high volume, simple tasks | GPT-6 Luna $0.10 / $0.50 | Claude Haiku 4.5 $1 / $5 | Gemini 3.5 Flash-Lite $0.30 / $2.50 | — |
The makers say the same thing. Anthropic's own documentation tells developers to start with its Opus tier for most workloads, and to move to Fable 5.1 only for demanding, long-running agent work or when tests on Opus fall short. OpenAI describes Astra as built for "the hardest end-to-end work".
What each model costs at real volume
Take a typical feature: 10,000 requests a day, each sending about 2,000 tokens of context and instructions and getting about 500 tokens back. Over 30 days, that is 600 million input tokens and 150 million output tokens. At list prices, with no caching or batch discounts, the monthly bill looks like this.
| GPT-6 Astra or Claude Fable 5.1 | 13500 USD |
|---|---|
| Claude Opus 5.5 | 5400 USD |
| GPT-6 Sol or Claude Sonnet 5 | 2700 USD |
| Muse Spark 1.3 | 1388 USD |
| Gemini 3.8 Flash (intro price) | 1013 USD |
| GPT-6 Luna | 135 USD |
Source: Tricolens calculation from list prices, Sep 28, 2026. 2,000 input and 500 output tokens per request, 30 days, no caching or batch discounts.
The top tier costs about 100 times the budget tier for the same traffic. That does not make the budget tier the right answer: a cheap model that gets 1 in 10 answers wrong can cost more in support tickets than it saves. It does mean the top tier should be reserved for the steps that genuinely need it.
Discounts narrow these gaps. OpenAI, Anthropic and Google all offer 50% off for batch processing when you can wait for results, and prompt caching cuts the cost of any context you reuse across requests.
Which model for which job
IfYour task is long, multi-step agent work where a failure is expensive, such as changes across a codebase or a full research report
ThenShortlist the top tier: GPT-6 Astra and Claude Fable 5.1. Test both on your own cases.
IfYou run customer-facing chat or support at scale
ThenStart in the workhorse tier (GPT-6 Sol, Claude Sonnet 5, Gemini 3.8 Flash) and step up only where quality falls short.
IfYou process large volumes of documents in the background
ThenTest Gemini 3.8 Flash and GPT-6 Luna with batch pricing. Budget for Gemini's price doubling in January 2027.
IfA single request needs a whole contract, case file or codebase
ThenAll four flagships accept about 1M tokens. Watch the GPT-6 Astra surcharge above 272K input tokens.
IfYou want a lower-cost option from a fourth vendor
ThenAdd Muse Spark 1.3 to the shortlist. It is priced between Gemini Flash and the workhorse tiers of OpenAI and Anthropic.
Want a model recommendation for your use case?
Get a free consultationWhy benchmark tables mislead buyers
Every maker publishes the benchmarks that flatter its model, and the versions rarely match. Anthropic reports 65.0% for Claude Fable 5.1 on Humanity's Last Exam with tools. Google reports 54.9% for Gemini 3.8 Flash on "HLE-Verified", a different variant of the same test. Meta reports 90.3 for Muse Spark 1.3 on OSWorld 2.0, while Anthropic reports 77.9% for Fable 5.1 on "OSWorld 2.0 (partial)".
None of these pairs can be compared directly. Even when the tests match, a public benchmark measures general ability, not how well a model handles your customers, your documents or your tone of voice. The only comparison that answers your question is one run on your own work:
- Collect 50 to 100 real examples of the task, including the awkward ones that go wrong today
- Write down what a good answer looks like for each, so scoring is not a matter of taste
- Run every shortlisted model on the same examples with the same instructions
- Score quality, response time and cost per task, not quality alone
- Keep the examples and re-run them whenever a provider ships a new version
Build so switching is cheap
The ranking will change again before the end of the year. The teams that benefit from each new release are the ones that can switch models in a day, not a quarter. That takes three things: a thin layer in your code between your product and any one provider, prompts stored and versioned like code, and the saved test set above.
We cover the engineering side in our guide to integrating LLMs into web applications. If you would rather have it built for you, see our AI integration services.
Key takeaway
Pick the cheapest model that passes your own tests, reserve the top tier for the steps that need it, and build so you can switch when the next release lands.
Frequently Asked Questions
Written by
Hiren Patel
AI & Engineering
