Blog · LLM costs · 26 Sep 2026
Blog
LLM costs

LLM cost optimization:When switching modelsactually pays off

A frontier model became a fifth cheaper in September. What that means for the applications you already run, and the one number you need to check it.

Get in touch
Enquiry
Nikolai Schöbel und Jeremias Burger, Co-Founder Scalableloops

Let's talk about your project.

First we check whether the project fits your business model. Then you get a proposal with phases and effort.

Have your AI costs reviewed or call: +49 151 1576 5566
Blog · LLM costs

LLM cost optimization starts with the right metric, and the price per million tokens is not it. What matters is the cost per completed task: per quote, per support ticket, per processed invoice. The two numbers move apart. When Anthropic released Claude Opus 5.5 on 22 September 2026, the token price fell by 20 percent, while the cost of a typical task fell by roughly 40 percent according to the vendor, because the model needs fewer tokens for the same work. The practical consequence: log what a typical task costs you for one or two weeks, recalculate it at the new prices, and only switch when the saving exceeds the cost of testing. At small volumes it often stays below.

In brief
  • The metric that holds up is cost per task, not price per million tokens.
  • Claude Opus 5.5 has cost four US dollars per million input tokens and twenty per million output tokens since 22 September 2026, a fifth below its predecessor.
  • Reused input served from cache costs 0.20 US dollars per million tokens there, which is the biggest lever in long running workloads.
  • A switch only pays off once the monthly saving clears the one off cost of migration and quality testing.
Published 26 Sep 2026Nikolai Schöbel and Jeremias Burger10 min read
Nikolai SchöbelJeremias Burger

Nikolai Schöbel and Jeremias Burger

Co-founders of Scalableloops GmbH. Nikolai Schöbel leads online marketing and AI strategy, Jeremias Burger the AI architecture. Both build AI systems and train teams on them in their own agency work.

On this page
  1. Why price per million tokens misleads you
  2. How to work out what one task costs you
  3. Which tier is good enough for which job
  4. Which levers actually reduce your cost per task
  5. How to tell that quality holds after the switch
  6. When migration is worth it and when to skip it
  7. Where the cost calculation leads you astray
  8. Frequently asked questions
  9. How to check your own numbers
  10. Where the information on this page comes from
Metric

Why price per million tokens misleads you

Model providers bill in tokens. A token is a chunk of text, roughly three quarters of an English word. Your invoice has two parts: the text you send in and the text the model writes back, with output priced several times higher than input.

The token price is a unit price, not a consumption figure. It says nothing about how many tokens a model needs to finish your task, and that is where the bill is decided. Anthropic quantified the gap at its own launch: the price per token dropped by a fifth against the previous model, while the cost of a typical workload dropped by around two fifths according to the vendor. The difference comes entirely from the new model needing fewer steps and fewer tokens for the same job.

The effect also runs the other way, which is what catches teams out. The same pricing documentation notes that the newer models in this family use a different tokenizer that produces about 30 percent more tokens for identical text. Compare pricing tables alone and a model can look cheaper while your monthly bill goes up. So never compare token prices against each other. Compare finished tasks.

2x

is the gap between the token price cut and the drop in cost per task at the September 2026 launch: a fifth against roughly two fifths.

Anthropic launch announcement, 22 Sep 2026
Measurement

How to work out what one task costs you

You do not need new tooling for this. Every API response reports usage: input tokens, output tokens and, if you use caching, cache reads and cache writes separately. Log those four values for one or two weeks alongside the type of task, and you have the number that matters.

The formula is a multiplication. Cost per task equals input tokens times input price plus output tokens times output price, divided by one million. Use the mean across many real tasks rather than one convenient example, and calculate the most expensive decile as well, because outliers are what blow up a month.

How much the mix matters shows in a worked example the vendor published. For a session with 2.2 million input tokens, 91 percent of them served from cache, and 60,000 output tokens, the same session cost 3.50 US dollars on the previous model and 2.40 on the new one. The more instructive figure sits inside that total: those 60,000 output tokens account for 1.20 US dollars, as much as six million cache reads.

Simpler workloads land far below that. The pricing documentation runs 10,000 support conversations at an average of 3,700 tokens each through the smallest model in the family and arrives at roughly 37 US dollars for the whole batch. Running that same work on the frontier model costs a multiple with no visible return.

Line item in a sample sessionVolumeCost on the new model
Input served from cacheabout 2 million tokensabout 0.40 US dollars
Fresh inputabout 0.2 million tokensabout 0.80 US dollars
Model output60,000 tokens1.20 US dollars
Same session on the previous modelsame volumes3.50 US dollars
Same session on the new modelsame volumes2.40 US dollars
Model choice

Which tier is good enough for which job

The largest saving rarely comes from moving to the newest version of the same frontier model. It comes from noticing how many tasks never needed a frontier model at all. Within one provider family, the smallest and largest tiers can differ by a factor of ten. A request that splits an invoice into fields, or files an email into one of five categories, is usually just as reliable on the smallest tier as on the largest.

The vendor recommends exactly that laddering in its own documentation: the small tier for simple work, the mid tier for most production workloads, the large tier only for demanding reasoning. In practice a mixed design works best, where a cheap model does the groundwork and only the hard remainder reaches the expensive one. Our article on AI for SMEs covers how to put such workflows on a permanent footing.

If data protection rules keep you out of the cloud, the arithmetic changes completely, because the line items become hardware and electricity rather than tokens. We weigh that up in our article on local AI for business.

TierPrice per million tokens, input and outputWhat it is good enough for
Small tier1 and 5 US dollarsClassifying, summarising, extracting fields, lookups
Mid tier2 and 10 US dollarsProduction work: draft replies, analysis, writing tasks
Large tier, new version4 and 20 US dollarsLong workflows, multi step reasoning, hard domain questions
Large tier, previous version5 and 25 US dollarsSame use cases, more expensive than the new version since 22 Sep 2026
37 USD

is what 10,000 support conversations averaging 3,700 tokens each cost on the smallest model tier, as calculated in the vendor pricing documentation.

Claude Platform Docs, Pricing
Levers

Which levers actually reduce your cost per task

The strongest lever is caching repeated input. If your application sends the same instructions, the same product catalogue or the same conversation history with every request, you pay for that block in full every single time without caching. With caching it costs 0.20 US dollars per million tokens on the new model instead of four. An independent analysis puts numbers on it: 2.8 million input tokens cost 11.20 US dollars uncached and 1.62 US dollars at a nine in ten hit rate.

Caching is not free. The first write costs more than a normal input, specifically 1.25 times for the short cache and double for the one hour cache according to the pricing documentation. That yields a usable rule of thumb from the vendor: the short cache pays for itself from the first read, the one hour cache from the second. Enable caching for requests that never repeat and you pay a premium for nothing.

The second lever is batching. Anything that does not need to be ready within seconds, such as overnight analysis, product descriptions or data enrichment, can run through a separate endpoint at half price on both input and output. That discount stacks with caching.

The third lever is effort level. Newer models can be told to think longer before answering. The vendor puts the difference between high and medium effort at roughly 20,000 extra tokens, around 0.40 US dollars per task, which is about what one retry costs. The rule follows directly: high effort is worth it when it prevents a retry, and not otherwise.

LeverEffectWhen it pays off
Prompt cachingRepeated text costs a twentieth of the normal input priceAs soon as the same block is sent more than once
Batch processingHalf price on both input and outputWhen results can take hours rather than seconds
Smaller model tierUp to ten times cheaper than the large tierFor classifying, extracting and summarising
Lower effort levelSaves around 0.40 US dollars per taskWhen results hold up without extended thinking
Shorter system instructionsPermanently reduces input volumeWhen prompts have grown over months
Quality

How to tell that quality holds after the switch

A cheaper model only saves money if nobody has to fix its output afterwards. As soon as somebody does, the cents you saved disappear in the first quarter hour of staff time. Quality is therefore not a footnote to the decision, it is the decision.

Do not lean on published leaderboards for this. The vendor writes in its own announcement that at this level of capability, benchmark margins have become a less reliable guide to real world differences, and reports a standard error of plus or minus 2.6 points on one of the headline scores. A two point lead in a table tells you very little about your quotes and your tickets.

What does tell you something takes half a day to build: an evaluation set of twenty to thirty real cases from your own operation where you already know the correct answer, deliberately including hard ones. Run that set through both models and compare the results side by side. Only when the new model matches quality at lower cost is the saving real. Keep the set afterwards, because the next model launch is already scheduled and the work will be done.

  1. 01

    Same quality, clearly cheaper

    Switch

    Migrate, rerun the evaluation set after four weeks and hold the invoice against the previous month.

  2. 02

    Same quality, barely cheaper

    Wait

    Keep the evaluation set and pull it out at the next price move. The effort does not pay for small amounts.

  3. 03

    Worse quality

    Stay

    Check whether better instructions or caching close the gap before you drop a model tier.

Decision

When migration is worth it and when to skip it

Put a third number next to the other two. On one side sits the monthly saving: your current invoice times the difference you measured. On the other sits the one off effort of building an evaluation set, migrating and watching the results for two weeks. For a well documented application that is typically a handful of hours.

The vendor's published example gives you the order of magnitude. At ten tasks a day, the monthly bill there falls from 770 to 528 US dollars. A saving of that size covers the effort within weeks. If your current bill is in the low double digits, the switch saves less than the testing costs. In that case the right call is to build the evaluation set anyway and fold the migration into the next piece of work you touch.

Run the calculation in the other direction as well. When the same work gets cheaper, you can either bank the money or get more done for the same budget, for instance by adding a verification step or serving a second language. For teams currently building their first assistant, the second option is usually worth more. Our article on the AI assistant for business sets out the available designs.

528 vs 770

US dollars per month is what a sample application with ten tasks a day costs before and after the switch, calculated at list prices.

Anthropic, cost breakdown at model launch
Limits

Where the cost calculation leads you astray

List prices are not your prices. Booking through a cloud platform adds a ten percent premium for guaranteed regional routing according to the vendor, and restricting processing to one country applies a factor of 1.1 to every line item on newer models. Both can be the right choice, but they belong in the calculation. In the other direction, volume discounts are negotiated individually above certain thresholds.

Second, most spreadsheets miss the side costs. Tools the model calls itself are sometimes billed separately, web search for example at ten US dollars per thousand searches. Workflows that spin up several parallel workers consume far more; the vendor cites roughly seven times the tokens of an ordinary session for such multi agent runs. If you are building in that direction, our article on the AI agent is worth a look first.

Third, any snapshot dates quickly. The vendor has announced two further models in the same line for the coming weeks, and competitors will respond. Build a pricing table today and you own an outdated pricing table next quarter. What lasts is not the price list but the method: your own measurement, your own evaluation set, and a fixed slot each quarter where you hold both against current prices.

Finally, price is not the only variable. The same new model is more than thirty percent faster at generating output according to the vendor, and more resistant to instructions smuggled into third party text. If your application processes documents from outside your organisation, that second point may well outweigh the price.

Frequently asked questions

Frequently asked questions

How much is one million tokens?

It depends on the model tier. On current models it ranges from one US dollar per million input tokens on the small tier to ten on the largest, with output priced roughly five times higher than input. As a rough guide, a million tokens is around 750,000 English words.

How expensive are LLMs to run in production?

That depends on your volume, not on a plan. Multiply cost per task by tasks per month. The vendor quotes 528 US dollars a month at list prices for a sample application with ten substantial tasks a day, while a batch of 10,000 simple support conversations on the smallest tier costs around 37 US dollars.

Are LLM costs going down?

The past trend is documented: the frontier model released in September 2026 costs a fifth less than its predecessor, and three fifths less for cached input. That is not a promise about the future. Plan with your current invoice and recheck once a quarter.

How do I measure cost per task?

Every API response reports input, output and cache tokens. Log those values for one or two weeks alongside the task type, then take the mean. Calculate the most expensive decile too, because outliers drive the monthly total.

Should I always pick the cheapest model?

No. A cheap model whose output needs human correction costs more than an expensive one that runs clean. Your own evaluation set on real cases decides it. For classifying, extracting and summarising the small tier is usually enough; for long multi step workflows it rarely is.

Does prompt caching really save money?

Substantially, for repeated blocks of text. Cached input costs 0.20 instead of four US dollars per million tokens on the new frontier model. The first write costs more than a normal input, so the short cache pays off from the first read and the one hour cache from the second.

Why do LLM costs creep up without anyone deciding to spend more?

Because input volume grows on its own. Instructions get extended over months, conversation histories get longer, documents get attached. The price per token stays flat while tokens per task climb. Caching and an annual prompt cleanup often beat a model switch here.

Next steps

How to check your own numbers

  1. 01

    Log your usage

    Record input, output and cache tokens per task for one or two weeks. The values are already in every API response.

  2. 02

    Build the cost per task

    Take the mean and the most expensive decile, then multiply by your monthly volume. That figure is your benchmark, not the token price.

  3. 03

    Create an evaluation set

    Collect twenty to thirty real cases with known correct answers and run them through both models before you change anything.

  4. 04

    Pull the levers first

    Caching, batch processing and a smaller tier for sub tasks often deliver more than the newest version of the same model.

Model prices change faster than anyone can maintain a spreadsheet. What lasts is your own measurement: what a task costs you, and whether the output holds up.

or call: +49 151 1576 5566

Further reading

Projekt-Detail

    Got a project in mind?

    We reply personally. First a use-case check, then an architecture proposal.

    Start your inquiry