Blog · AI voice agent · 8 Oct 2026
Blog
AI voice agent

AI voice agent pricing:what a call costsand what you must say

Transcription and speech synthesis have become markedly cheaper. What that means for your phone line, and the three questions worth settling before the first call.

Get in touch
Enquiry
Nikolai Schöbel und Jeremias Burger, Co-Founder Scalableloops

Let's talk about your project.

First we check whether the project fits your business model. Then you get a proposal with phases and effort.

Have your voice agent plan reviewed or call: +49 151 1576 5566
Blog · AI voice agent

AI voice agent pricing comes down to three building blocks: a transcript of what the caller says, a language model that works out the reply, and speech synthesis. The agent answers the phone, understands spoken language, replies in a synthetic voice and hands over to a person when it should. The price of two of them has fallen sharply. On 1 October 2026 Microsoft released three models that put transcription at 0.54 US dollars per audio hour and speech output at 15 or 22 US dollars per million characters. Worked through for a three minute call, that comes to roughly four cents, with the language model in the middle still to be added. Which means the hard part is no longer the price. It is the three questions in front of it: where the calls are processed, what you have to tell callers about the AI, and when recording a conversation becomes a criminal offence.

In brief
  • On 1 October 2026 Microsoft released MAI-Transcribe-2-Streaming for live transcription in 60 languages, plus MAI-Voice-2.1 and MAI-Voice-2.1-Flash for speech output in 23 languages.
  • Transcription is priced at 0.54 US dollars per audio hour as an introductory rate through the end of the year; speech output costs 22 and 15 US dollars per million characters.
  • A three minute call therefore costs around four cents for transcription and speech combined, before the language model that writes the reply.
  • Since 2 August 2026 callers must be told they are talking to an AI system, and in Germany recording a call without permission is a criminal offence.
Published 8 Oct 2026Nikolai Schöbel and Jeremias Burger8 min read
Nikolai SchöbelJeremias Burger

Nikolai Schöbel and Jeremias Burger

Co-founders of Scalableloops GmbH. Nikolai Schöbel leads online marketing and AI strategy, Jeremias Burger the AI architecture. Both build AI systems and train teams on them in their own agency work.

On this page
  1. What is an AI voice agent?
  2. Which building blocks drive AI voice agent pricing
  3. AI voice agent pricing: what one call costs
  4. Where an AI voice agent processes your calls
  5. EU AI Act: what you have to tell callers about the AI
  6. When missing call recording consent becomes a criminal offence
  7. Who benefits from AI phone answering
  8. Frequently asked questions
  9. How to approach your voice AI agent setup
  10. Where the information on this page comes from
Basics

What is an AI voice agent?

An AI voice agent is software that picks up the phone instead of letting it ring out. It listens, recognises speech, decides what to do based on the rules and data you give it, and answers in a synthetically generated voice. Depending on the job it may simply take a message, state opening hours, book an appointment or route the caller to the right person.

What separates it from a classic menu with keypad options is that the caller just talks, with no options to work through. What separates it from a website chatbot is the channel. On a phone line every pause is audible, because people in conversation expect an answer within a fraction of a second. That is exactly why providers are currently competing so hard on speed.

One thing is worth keeping in view: an agent like this decides nothing you have not allowed in advance. It is a switchboard that understands speech, not a colleague. What it may say, which cases it takes and when it hands over to a person are all settings. For the wider picture of what such systems do inside a business, see our article on the AI assistant in business.

Architecture

Which building blocks drive AI voice agent pricing

AI phone answering runs in three steps that follow one another and are billed separately. First a real-time transcription model turns speech into text, continuously, while the caller is still talking. Then a language model turns that text into a reply. Finally a speech synthesis model turns the reply into audible speech (and only at that point does the caller hear anything, which is why delays from all three steps pile up here).

On 1 October 2026 Microsoft introduced new models for the first and third of those steps. MAI-Transcribe-2-Streaming transcribes as the caller speaks and, according to the vendor, covers 60 languages with automatic, continuous language detection, so it works out by itself which language is being spoken. For output there is MAI-Voice-2.1 and the faster MAI-Voice-2.1-Flash, both covering 23 languages and 26 locales. Microsoft states an end to end latency of 150 milliseconds for 45 seconds of audio for the Flash variant and describes it as roughly 60 percent cheaper than comparable models. The models are available through Microsoft Foundry, the MAI Playground, Vercel, Azure Voice Live and, in part, OpenRouter.

In practice the step that matters most is the second one, and it is the one the coverage says least about. The language model in the middle determines how good the answers are, and it costs extra. Anyone who mistakes the transcription and speech prices for the total will budget too low. We collected the ways to bring those model costs down under LLM cost optimization.

StepWhat happensHow it is billed
TranscriptionSpeech is turned into text as the caller talksPer hour of audio
ReplyA language model works out what should be saidPer volume of text processed
Speech outputThe text is rendered as an audible voicePer million characters
TelephonyThe call is routed into the systemPer minute or as a flat rate, depending on provider
60

languages are covered by the new transcription model according to the vendor, with continuous automatic language detection.

Microsoft AI, 1 Oct 2026
Cost

AI voice agent pricing: what one call costs

AI voice agent cost can be calculated exactly for two of the three building blocks, because speech synthesis pricing and transcription pricing are both published. Transcription is 0.54 US dollars per audio hour, and Microsoft explicitly calls that an introductory price through the end of the year. Whether it holds after that is not stated. Speech output is 22 US dollars per million characters, or 15 in the faster and cheaper Flash variant. Converted at the European Central Bank reference rate of 7 October 2026 (1 euro bought 1.1177 US dollars), you get the figures in the table.

Work a three minute call through. Transcription runs for the whole call, which is 0.05 audio hours and therefore about 2.4 cents. For speech output, assume the agent is speaking for roughly half the time. A minute of spoken text runs to about 750 to 900 characters including spaces, so a minute and a half is somewhere around 1,100 to 1,350 characters. With the Flash variant that costs about 1.5 to 1.8 cents. Together you land at roughly four cents per call, or just under five with the more expensive voice.

Scaled up to 100 calls across 20 working days, so 2,000 calls a month, that is about 79 to 85 euros for transcription and speech. This is a calculation with stated assumptions rather than a measurement of your business: change the call length or the share of talking and the number moves. Two items are missing from that total entirely, namely the language model for the replies and the telephony. And buying a finished service rather than building one costs more than the raw model prices, because dialogue logic, call routing and operations are bundled into it.

Building blockList priceConvertedUnit
Transcription (MAI-Transcribe-2-Streaming)0.54 USDabout 0.48 EURper hour of audio, introductory price through year end
Speech output (MAI-Voice-2.1)22 USDabout 19.68 EURper million characters
Speech output (MAI-Voice-2.1-Flash)15 USDabout 13.42 EURper million characters
Language model for the replynot included in these pricesdepends on the model chosenper volume of text processed
4 cents

is what transcription and speech output together come to on a three minute call, calculated from the published prices.

Own calculation, ECB rate 7 Oct 2026
Data

Where an AI voice agent processes your calls

An AI voice agent starts from a different place than a chatbot under data protection law, because a voice is more than text. Microsoft writes in its own documentation for the speech services that audio of people speaking, and the transcripts that go with it, may count as personal or sensitive data under various privacy laws, since they carry not only the voice but, depending on context, personal information in what is said. The same document states that you, as the operator, are responsible for obtaining every permission needed to process that data.

On storage the statement is refreshingly clear. For real time transcription, Microsoft says audio is processed only in the server's memory and that no data is stored at rest. For the question of which country that happens in, the EU Data Boundary applies: Microsoft commits to storing and processing customer data and personal data for its enterprise online services within the EU and EFTA. That commitment carries explicitly named exceptions in which data does leave the boundary, and for Azure it only applies if you deploy the service in a region inside it.

From that follows a check nobody will run for you. Which region does your service run in, is the data processing agreement in place, and is the specific model even available in that region. With freshly released models the last point is not a formality, because regional availability tends to arrive in stages. If the data is not to leave your premises at all, the route is to run the speech services in your own environment, which Microsoft offers as containers and where, according to the vendor, audio and transcript do not go to the cloud. We set out what that means more generally under local AI.

Disclosure

EU AI Act: what you have to tell callers about the AI

Under the EU AI Act, telling callers they are speaking to an AI has not been a matter of taste since 2 August 2026. That is the date Article 50 took effect, and paragraph 1 requires AI systems intended to interact directly with natural persons to be designed so that those persons are informed they are interacting with an AI system. The German Federal Network Agency, which supervises this, puts it just as plainly in its overview: affected persons must be informed about the interaction with the AI system.

There is an exception, and on a phone line it carries less weight than many hope. The information can be dropped where the fact is obvious anyway to a reasonably well informed and observant person. With a synthetic voice that sounds good, that is precisely what is in doubt, and the better the voices get, the weaker the argument becomes. We treat the spoken disclosure as the safe route. It has to come at the first interaction at the latest, which means in the opening line of the call rather than in a policy page.

In practice that means one sentence in the greeting saying what the caller is dealing with, and a route to a human that actually works. Article 99 sets out what non compliance can cost: fines of up to 15 million euros or 3 percent of worldwide annual turnover, with the lower of the two applying to small and medium sized companies. We wrote up the detail in our article on the chatbot disclosure duty under Article 50, and the separate training obligation in Article 4.

QuestionWhat applies
Do I have to announce the AI?Yes, unless it is obvious that an AI system is speaking
When does the disclosure have to come?At the first interaction at the latest, so in the greeting
Who is responsible?The provider of the system under paragraph 1; deployer duties apply in further cases
Since when does this apply?Article 50 of the EU AI Act has applied since 2 August 2026
What does a breach cost?Up to 15 million euros or 3 percent of annual turnover, the lower figure for SMEs
Decision

Who benefits from AI phone answering

Whether AI phone answering pays off depends less on the technology than on what gets dropped when the phone rings. In the trades it is often the work itself, because the same person orders materials, advises customers and answers calls. In a medical practice it is the patients at the desk while the line is busy. In both cases the gain is not a saved salary but the calls getting answered at all. What such an agent looks like in practice is shown on our AI voice agent page.

One argument against cannot be calculated away. If your calls are mostly delicate, meaning complaints, emergencies or negotiations, then a machine at the front of the line is the wrong call however cheap it is. The sensible dividing line, in our experience, runs between recurring and one off enquiries.

  1. 01

    Trades and field service businesses

    Usually pays off

    Many short, similar calls about appointments and availability while nobody is in the office. Start with taking messages and callback requests, not with advice.

  2. 02

    Practices, law firms, administration

    Worth checking

    High call volume on fixed topics, but health or client data often comes up in the conversation. Settle data protection and retention first.

  3. 03

    Retail and customer service

    Useful for standard cases

    Delivery status, opening hours and returns work well. Complaints belong with a person from the start.

  4. 04

    Consulting and high value sales

    Probably not

    Every call is different and the first impression decides. Good call forwarding beats an agent here.

Frequently asked questions

Frequently asked questions

What is an AI voice agent, and where does it fit in customer service?

It is software that answers calls, understands spoken language, replies in a synthetically generated voice and hands over to a person when needed. In customer service it works best on recurring enquiries such as delivery status or opening hours, with complaints going straight to a person.

What does an AI voice agent cost per call?

Around four cents for transcription and speech output combined on a three minute call, calculated from the Microsoft prices published in October 2026. The language model for the replies and the telephony come on top.

Do I have to tell callers they are speaking to an AI?

Yes, unless it is obvious. Article 50 of the EU AI Act has applied since 2 August 2026 and requires that people are informed they are interacting with an AI system, at the first interaction at the latest.

Can I record calls handled by an AI voice agent?

Only with consent. In Germany, recording the non publicly spoken word without authorisation is a criminal offence under section 201 of the Criminal Code, carrying up to three years imprisonment or a fine.

Where is the audio processed?

That depends on the provider and the region you choose. Microsoft states that real time transcription is processed only in server memory with no data stored at rest, and commits under the EU Data Boundary to processing within the EU and EFTA, subject to named exceptions.

How many languages do these systems handle?

Microsoft's new transcription model covers 60 languages according to the vendor and detects continuously which one is being spoken. Speech output covers 23 languages across 26 locales.

Is an AI voice agent worth it for a small business?

If calls regularly go unanswered, usually yes, because the gain sits in the calls you now take rather than in saved wages. If your calls are mostly delicate or advice heavy, probably not.

Next steps

How to approach your voice AI agent setup

  1. 01

    Count the calls

    Log for two weeks how many calls come in, how many go unanswered and what they were about. Without that list, any calculation is guesswork.

  2. 02

    Pick one case

    Take the most frequent recurring enquiry, usually an appointment or a callback request, and let the agent handle only that to begin with.

  3. 03

    Settle disclosure and consent

    Write the greeting that names the AI, decide whether you need to record at all, and build in the route to a person.

  4. 04

    Check region and contract

    Before launch, confirm which region the service runs in, that the data processing agreement is signed, and how long transcripts are kept.

The technology is no longer the expensive part, and it is no longer the difficult part either. What stays difficult is deciding which conversation you are willing to hand to a machine and which one you are not.

or call: +49 151 1576 5566

Further reading

Projekt-Detail

    Got a project in mind?

    We reply personally. First a use-case check, then an architecture proposal.

    Start your inquiry