Why AI bills surprise teams
Sound familiar?
The prototype cost a few dollars a day. Three months after launch, the AI line on the cloud bill is bigger than the servers, and nobody can say which feature is responsible.
AI APIs charge per token, a short chunk of text that is usually a word or part of a word. You pay for every token you send (your instructions, the user's message, any documents or conversation history) and every token the model writes back. Output tokens cost 3 to 8 times more than input tokens across the major providers.
Per-token prices have fallen, but three things push bills up anyway. Conversations resend their history on every turn. AI agents make dozens of model calls to finish one task. And teams default to the most capable model for everything, including steps a model costing a tenth as much would handle.
What the main models cost per token
| Model | Maker | Input | Output | Discounts available |
|---|---|---|---|---|
| GPT-6 Astra | OpenAI | $10.00 | $50.00 | Cached input $1; batch 50% off |
| Claude Fable 5.1 | Anthropic | $10.00 | $50.00 | Cache reads $0.25; batch 50% off |
| Claude Opus 5.5 | Anthropic | $4.00 | $20.00 | Prompt caching; batch 50% off |
| GPT-6 Sol | OpenAI | $2.00 | $10.00 | Cached input $0.20; batch 50% off |
| Claude Sonnet 5 | Anthropic | $2.00 | $10.00 | Prompt caching; batch 50% off |
| Muse Spark 1.3 | Meta | $1.25 | $4.25 | Cached input |
| DeepSeek-V4-Pro | DeepSeek | $1.32 | $3.96 | 50% off outside peak hours |
| Gemini 3.8 Flash | $0.75 | $3.75 | Introductory until Dec 31, 2026; batch 50% off | |
| DeepSeek-V4-Flash | DeepSeek | $0.30 | $1.20 | 50% off outside peak hours |
| GPT-6 Luna | OpenAI | $0.10 | $0.50 | Cached input $0.01; batch 50% off |
DeepSeek prices shown are peak rates; its off-peak rates are half. Prices as of September 28, 2026, from each provider: OpenAI, Anthropic, Google, Meta, DeepSeek. Check them before you budget.
For a side-by-side of what these models are good at, see our AI model comparison.
Three worked examples
The examples below use list prices and 30-day months. The token counts are typical for each kind of feature. Replace them with counts from your own requests to get a real estimate.
| Feature | Assumed volume | Monthly tokens | Cost by model |
|---|---|---|---|
| Customer support chat | 5,000 conversations a day, 6 turns each, about 3,000 tokens in and 250 out per turn | 2.7 billion in, 225 million out | GPT-6 Sol: $7,650. Gemini 3.8 Flash: $2,869 |
| Document extraction, run overnight in batches | 200,000 pages a month, about 1,500 tokens in and 300 out per page | 300 million in, 60 million out | Claude Sonnet 5 batch: $600. Gemini 3.8 Flash batch: $225. GPT-6 Luna batch: $30 |
| Coding or research agent | 200 tasks a day, about 40 model calls per task, 30,000 tokens in and 1,500 out per call | 7.2 billion in, 360 million out | Claude Fable 5.1 with no caching: $90,000 |
The agent is the outlier. Each task rereads the same files and instructions dozens of times, so input tokens dominate. That is exactly the cost prompt caching is designed to remove.
| Claude Fable 5.1, no caching | 90000 USD |
|---|---|
| GPT-6 Astra, 90% of input cached | 31680 USD |
| Claude Fable 5.1, 90% of input cached | 26820 USD |
Source: Tricolens calculation from OpenAI and Anthropic list prices, Sep 28, 2026. Excludes cache-write charges.
Want these numbers worked out for your feature?
Get a free consultationSix levers that cut an AI bill
- Right-size each step. Use a budget model for classification, routing and extraction, and save the top tier for the steps that need it
- Cache repeated context. Instructions, documents and conversation history that repeat across requests can cost 90% less as cached input
- Batch anything that can wait. OpenAI, Anthropic and Google all take 50% off for batch jobs
- Cap output length. Output tokens cost 3 to 8 times more, so ask for concise answers and set a maximum
- Schedule flexible work off-peak. DeepSeek charges half price outside its peak hours
- Watch long-context surcharges. GPT-6 Astra doubles the input rate for prompts over 272,000 tokens
Routing different requests to different models is the lever with the biggest effect, and we cover how to build it in model routing: why one AI model is no longer enough.
Budget for next year's prices, too
Prices move in both directions. Gemini 3.8 Flash's introductory rate ends on December 31, 2026, when input rises from $0.75 to $1.50 per million tokens and output from $3.75 to $7.50. Any feature budgeted on today's Gemini price should be re-costed at the January rate.
Set a monthly spending alert with every provider you use, track cost per feature rather than one AI line on the bill, and review your model choices every quarter. If you would like help designing a cost-controlled AI feature, see our AI integration services.
Key takeaway
Estimate from real token counts, cache and batch wherever you can, and match the model tier to each step. The cheapest bill comes from architecture, not from haggling over the per-token price.
Frequently Asked Questions
Written by
Bhumin Patel
AI & Engineering
