LLM cost optimization:When switching modelsactually pays off
A frontier model became a fifth cheaper in September. What that means for the applications you already run, and the one number you need to check it.
Get in touch- Why price per million tokens misleads you
- How to work out what one task costs you
- Which tier is good enough for which job
- Which levers actually reduce your cost per task
- How to tell that quality holds after the switch
- When migration is worth it and when to skip it
- Where the cost calculation leads you astray
- Frequently asked questions
- How to check your own numbers
- Where the information on this page comes from

Let's talk about your project.
First we check whether the project fits your business model. Then you get a proposal with phases and effort.
Have your AI costs reviewed or call: +49 151 1576 5566LLM cost optimization starts with the right metric, and the price per million tokens is not it. What matters is the cost per completed task: per quote, per support ticket, per processed invoice. The two numbers move apart. When Anthropic released Claude Opus 5.5 on 22 September 2026, the token price fell by 20 percent, while the cost of a typical task fell by roughly 40 percent according to the vendor, because the model needs fewer tokens for the same work. The practical consequence: log what a typical task costs you for one or two weeks, recalculate it at the new prices, and only switch when the saving exceeds the cost of testing. At small volumes it often stays below.
- The metric that holds up is cost per task, not price per million tokens.
- Claude Opus 5.5 has cost four US dollars per million input tokens and twenty per million output tokens since 22 September 2026, a fifth below its predecessor.
- Reused input served from cache costs 0.20 US dollars per million tokens there, which is the biggest lever in long running workloads.
- A switch only pays off once the monthly saving clears the one off cost of migration and quality testing.
On this page
- Why price per million tokens misleads you
- How to work out what one task costs you
- Which tier is good enough for which job
- Which levers actually reduce your cost per task
- How to tell that quality holds after the switch
- When migration is worth it and when to skip it
- Where the cost calculation leads you astray
- Frequently asked questions
- How to check your own numbers
- Where the information on this page comes from
Why price per million tokens misleads you
Model providers bill in tokens. A token is a chunk of text, roughly three quarters of an English word. Your invoice has two parts: the text you send in and the text the model writes back, with output priced several times higher than input.
The token price is a unit price, not a consumption figure. It says nothing about how many tokens a model needs to finish your task, and that is where the bill is decided. Anthropic quantified the gap at its own launch: the price per token dropped by a fifth against the previous model, while the cost of a typical workload dropped by around two fifths according to the vendor. The difference comes entirely from the new model needing fewer steps and fewer tokens for the same job.
The effect also runs the other way, which is what catches teams out. The same pricing documentation notes that the newer models in this family use a different tokenizer that produces about 30 percent more tokens for identical text. Compare pricing tables alone and a model can look cheaper while your monthly bill goes up. So never compare token prices against each other. Compare finished tasks.
is the gap between the token price cut and the drop in cost per task at the September 2026 launch: a fifth against roughly two fifths.
How to work out what one task costs you
You do not need new tooling for this. Every API response reports usage: input tokens, output tokens and, if you use caching, cache reads and cache writes separately. Log those four values for one or two weeks alongside the type of task, and you have the number that matters.
The formula is a multiplication. Cost per task equals input tokens times input price plus output tokens times output price, divided by one million. Use the mean across many real tasks rather than one convenient example, and calculate the most expensive decile as well, because outliers are what blow up a month.
How much the mix matters shows in a worked example the vendor published. For a session with 2.2 million input tokens, 91 percent of them served from cache, and 60,000 output tokens, the same session cost 3.50 US dollars on the previous model and 2.40 on the new one. The more instructive figure sits inside that total: those 60,000 output tokens account for 1.20 US dollars, as much as six million cache reads.
Simpler workloads land far below that. The pricing documentation runs 10,000 support conversations at an average of 3,700 tokens each through the smallest model in the family and arrives at roughly 37 US dollars for the whole batch. Running that same work on the frontier model costs a multiple with no visible return.
| Line item in a sample session | Volume | Cost on the new model |
|---|---|---|
| Input served from cache | about 2 million tokens | about 0.40 US dollars |
| Fresh input | about 0.2 million tokens | about 0.80 US dollars |
| Model output | 60,000 tokens | 1.20 US dollars |
| Same session on the previous model | same volumes | 3.50 US dollars |
| Same session on the new model | same volumes | 2.40 US dollars |
Which tier is good enough for which job
The largest saving rarely comes from moving to the newest version of the same frontier model. It comes from noticing how many tasks never needed a frontier model at all. Within one provider family, the smallest and largest tiers can differ by a factor of ten. A request that splits an invoice into fields, or files an email into one of five categories, is usually just as reliable on the smallest tier as on the largest.
The vendor recommends exactly that laddering in its own documentation: the small tier for simple work, the mid tier for most production workloads, the large tier only for demanding reasoning. In practice a mixed design works best, where a cheap model does the groundwork and only the hard remainder reaches the expensive one. Our article on AI for SMEs covers how to put such workflows on a permanent footing.
If data protection rules keep you out of the cloud, the arithmetic changes completely, because the line items become hardware and electricity rather than tokens. We weigh that up in our article on local AI for business.
| Tier | Price per million tokens, input and output | What it is good enough for |
|---|---|---|
| Small tier | 1 and 5 US dollars | Classifying, summarising, extracting fields, lookups |
| Mid tier | 2 and 10 US dollars | Production work: draft replies, analysis, writing tasks |
| Large tier, new version | 4 and 20 US dollars | Long workflows, multi step reasoning, hard domain questions |
| Large tier, previous version | 5 and 25 US dollars | Same use cases, more expensive than the new version since 22 Sep 2026 |
is what 10,000 support conversations averaging 3,700 tokens each cost on the smallest model tier, as calculated in the vendor pricing documentation.
Which levers actually reduce your cost per task
The strongest lever is caching repeated input. If your application sends the same instructions, the same product catalogue or the same conversation history with every request, you pay for that block in full every single time without caching. With caching it costs 0.20 US dollars per million tokens on the new model instead of four. An independent analysis puts numbers on it: 2.8 million input tokens cost 11.20 US dollars uncached and 1.62 US dollars at a nine in ten hit rate.
Caching is not free. The first write costs more than a normal input, specifically 1.25 times for the short cache and double for the one hour cache according to the pricing documentation. That yields a usable rule of thumb from the vendor: the short cache pays for itself from the first read, the one hour cache from the second. Enable caching for requests that never repeat and you pay a premium for nothing.
The second lever is batching. Anything that does not need to be ready within seconds, such as overnight analysis, product descriptions or data enrichment, can run through a separate endpoint at half price on both input and output. That discount stacks with caching.
The third lever is effort level. Newer models can be told to think longer before answering. The vendor puts the difference between high and medium effort at roughly 20,000 extra tokens, around 0.40 US dollars per task, which is about what one retry costs. The rule follows directly: high effort is worth it when it prevents a retry, and not otherwise.
| Lever | Effect | When it pays off |
|---|---|---|
| Prompt caching | Repeated text costs a twentieth of the normal input price | As soon as the same block is sent more than once |
| Batch processing | Half price on both input and output | When results can take hours rather than seconds |
| Smaller model tier | Up to ten times cheaper than the large tier | For classifying, extracting and summarising |
| Lower effort level | Saves around 0.40 US dollars per task | When results hold up without extended thinking |
| Shorter system instructions | Permanently reduces input volume | When prompts have grown over months |
How to tell that quality holds after the switch
A cheaper model only saves money if nobody has to fix its output afterwards. As soon as somebody does, the cents you saved disappear in the first quarter hour of staff time. Quality is therefore not a footnote to the decision, it is the decision.
Do not lean on published leaderboards for this. The vendor writes in its own announcement that at this level of capability, benchmark margins have become a less reliable guide to real world differences, and reports a standard error of plus or minus 2.6 points on one of the headline scores. A two point lead in a table tells you very little about your quotes and your tickets.
What does tell you something takes half a day to build: an evaluation set of twenty to thirty real cases from your own operation where you already know the correct answer, deliberately including hard ones. Run that set through both models and compare the results side by side. Only when the new model matches quality at lower cost is the saving real. Keep the set afterwards, because the next model launch is already scheduled and the work will be done.
- 01
Same quality, clearly cheaper
SwitchMigrate, rerun the evaluation set after four weeks and hold the invoice against the previous month.
- 02
Same quality, barely cheaper
WaitKeep the evaluation set and pull it out at the next price move. The effort does not pay for small amounts.
- 03
Worse quality
StayCheck whether better instructions or caching close the gap before you drop a model tier.
When migration is worth it and when to skip it
Put a third number next to the other two. On one side sits the monthly saving: your current invoice times the difference you measured. On the other sits the one off effort of building an evaluation set, migrating and watching the results for two weeks. For a well documented application that is typically a handful of hours.
The vendor's published example gives you the order of magnitude. At ten tasks a day, the monthly bill there falls from 770 to 528 US dollars. A saving of that size covers the effort within weeks. If your current bill is in the low double digits, the switch saves less than the testing costs. In that case the right call is to build the evaluation set anyway and fold the migration into the next piece of work you touch.
Run the calculation in the other direction as well. When the same work gets cheaper, you can either bank the money or get more done for the same budget, for instance by adding a verification step or serving a second language. For teams currently building their first assistant, the second option is usually worth more. Our article on the AI assistant for business sets out the available designs.
US dollars per month is what a sample application with ten tasks a day costs before and after the switch, calculated at list prices.
Where the cost calculation leads you astray
List prices are not your prices. Booking through a cloud platform adds a ten percent premium for guaranteed regional routing according to the vendor, and restricting processing to one country applies a factor of 1.1 to every line item on newer models. Both can be the right choice, but they belong in the calculation. In the other direction, volume discounts are negotiated individually above certain thresholds.
Second, most spreadsheets miss the side costs. Tools the model calls itself are sometimes billed separately, web search for example at ten US dollars per thousand searches. Workflows that spin up several parallel workers consume far more; the vendor cites roughly seven times the tokens of an ordinary session for such multi agent runs. If you are building in that direction, our article on the AI agent is worth a look first.
Third, any snapshot dates quickly. The vendor has announced two further models in the same line for the coming weeks, and competitors will respond. Build a pricing table today and you own an outdated pricing table next quarter. What lasts is not the price list but the method: your own measurement, your own evaluation set, and a fixed slot each quarter where you hold both against current prices.
Finally, price is not the only variable. The same new model is more than thirty percent faster at generating output according to the vendor, and more resistant to instructions smuggled into third party text. If your application processes documents from outside your organisation, that second point may well outweigh the price.
Frequently asked questions
How much is one million tokens?
It depends on the model tier. On current models it ranges from one US dollar per million input tokens on the small tier to ten on the largest, with output priced roughly five times higher than input. As a rough guide, a million tokens is around 750,000 English words.
How expensive are LLMs to run in production?
That depends on your volume, not on a plan. Multiply cost per task by tasks per month. The vendor quotes 528 US dollars a month at list prices for a sample application with ten substantial tasks a day, while a batch of 10,000 simple support conversations on the smallest tier costs around 37 US dollars.
Are LLM costs going down?
The past trend is documented: the frontier model released in September 2026 costs a fifth less than its predecessor, and three fifths less for cached input. That is not a promise about the future. Plan with your current invoice and recheck once a quarter.
How do I measure cost per task?
Every API response reports input, output and cache tokens. Log those values for one or two weeks alongside the task type, then take the mean. Calculate the most expensive decile too, because outliers drive the monthly total.
Should I always pick the cheapest model?
No. A cheap model whose output needs human correction costs more than an expensive one that runs clean. Your own evaluation set on real cases decides it. For classifying, extracting and summarising the small tier is usually enough; for long multi step workflows it rarely is.
Does prompt caching really save money?
Substantially, for repeated blocks of text. Cached input costs 0.20 instead of four US dollars per million tokens on the new frontier model. The first write costs more than a normal input, so the short cache pays off from the first read and the one hour cache from the second.
Why do LLM costs creep up without anyone deciding to spend more?
Because input volume grows on its own. Instructions get extended over months, conversation histories get longer, documents get attached. The price per token stays flat while tokens per task climb. Caching and an annual prompt cleanup often beat a model switch here.
How to check your own numbers
- 01
Log your usage
Record input, output and cache tokens per task for one or two weeks. The values are already in every API response.
- 02
Build the cost per task
Take the mean and the most expensive decile, then multiply by your monthly volume. That figure is your benchmark, not the token price.
- 03
Create an evaluation set
Collect twenty to thirty real cases with known correct answers and run them through both models before you change anything.
- 04
Pull the levers first
Caching, batch processing and a smaller tier for sub tasks often deliver more than the newest version of the same model.
Model prices change faster than anyone can maintain a spreadsheet. What lasts is your own measurement: what a task costs you, and whether the output holds up.
Google spam updateGoogle Spam Update: What Can Your Numbers Really Tell You?
Detect AI-written textDetect AI-written text: what AI detectors get wrong
AI agentsAgentic AI Explained: What It Means for Your Business
Google AI ModeGoogle AI Mode: What It Means for Your Website
AI agentsWhat Is an AI Agent, and How Is It Different From a Chatbot?
Structured DataStructured Data: What It Really Does for AI Search
AI AssistantAI Assistant for Business: Types, Uses and Data Protection
GEOE-E-A-T: Trust Signals for Google and AI Search
Local AILocal AI for Business: What “Local” Really Means
GEOAI Crawlers in robots.txt: Managing GPTBot and Co.
AI for SMEsAI for SMEs: How to Introduce AI Step by Step
Perplexity SEOPerplexity SEO: How to Get Your Site Cited as a Source
ChatGPT SEOChatGPT SEO: How to Get Your Business Found in ChatGPT
llms.txtllms.txt: What the File Does and When It Pays Off
AI SEOAI SEO: What Changes Compared to Traditional SEO
AI OverviewsGoogle AI Overviews: How Google Picks Its Sources
GEO vs AEO vs LLMOGEO vs AEO vs LLMO: The AI Search Terms Explained
Measure AI VisibilityMeasure AI Visibility: Method, Metrics and Limits
ChatGPT AdsChatGPT Ads: How to Advertise on ChatGPT in Germany
AI Literacy ObligationAI Literacy Obligation: What Article 4 Requires Since 2026
AI DisclosureAI Chatbot Disclosure: Article 50 in Practice
AI Text WatermarkAI Text Watermark: What Claude Marks and What It Doesn't
ShopwareShopware Plugin Development: Buy or Build? A Guide
Where the information on this page comes from
- Anthropic: Introducing Claude Opus 5.5accessed 26 Sep 2026
- Anthropic: What a task costs on Opus 5.5accessed 26 Sep 2026
- Claude Platform Docs: Pricingaccessed 26 Sep 2026
- TechCrunch: Anthropic releases Opus 5.5 with lower prices and Fable-level performanceaccessed 26 Sep 2026
- Finout: Claude Opus 5.5 Pricing 2026, what Anthropic's new flagship actually costsaccessed 26 Sep 2026


