AI voice agent pricing:what a call costsand what you must say
Transcription and speech synthesis have become markedly cheaper. What that means for your phone line, and the three questions worth settling before the first call.
Get in touch- What is an AI voice agent?
- Which building blocks drive AI voice agent pricing
- AI voice agent pricing: what one call costs
- Where an AI voice agent processes your calls
- EU AI Act: what you have to tell callers about the AI
- When missing call recording consent becomes a criminal offence
- Who benefits from AI phone answering
- Frequently asked questions
- How to approach your voice AI agent setup
- Where the information on this page comes from

Let's talk about your project.
First we check whether the project fits your business model. Then you get a proposal with phases and effort.
Have your voice agent plan reviewed or call: +49 151 1576 5566AI voice agent pricing comes down to three building blocks: a transcript of what the caller says, a language model that works out the reply, and speech synthesis. The agent answers the phone, understands spoken language, replies in a synthetic voice and hands over to a person when it should. The price of two of them has fallen sharply. On 1 October 2026 Microsoft released three models that put transcription at 0.54 US dollars per audio hour and speech output at 15 or 22 US dollars per million characters. Worked through for a three minute call, that comes to roughly four cents, with the language model in the middle still to be added. Which means the hard part is no longer the price. It is the three questions in front of it: where the calls are processed, what you have to tell callers about the AI, and when recording a conversation becomes a criminal offence.
- On 1 October 2026 Microsoft released MAI-Transcribe-2-Streaming for live transcription in 60 languages, plus MAI-Voice-2.1 and MAI-Voice-2.1-Flash for speech output in 23 languages.
- Transcription is priced at 0.54 US dollars per audio hour as an introductory rate through the end of the year; speech output costs 22 and 15 US dollars per million characters.
- A three minute call therefore costs around four cents for transcription and speech combined, before the language model that writes the reply.
- Since 2 August 2026 callers must be told they are talking to an AI system, and in Germany recording a call without permission is a criminal offence.
On this page
- What is an AI voice agent?
- Which building blocks drive AI voice agent pricing
- AI voice agent pricing: what one call costs
- Where an AI voice agent processes your calls
- EU AI Act: what you have to tell callers about the AI
- When missing call recording consent becomes a criminal offence
- Who benefits from AI phone answering
- Frequently asked questions
- How to approach your voice AI agent setup
- Where the information on this page comes from
What is an AI voice agent?
An AI voice agent is software that picks up the phone instead of letting it ring out. It listens, recognises speech, decides what to do based on the rules and data you give it, and answers in a synthetically generated voice. Depending on the job it may simply take a message, state opening hours, book an appointment or route the caller to the right person.
What separates it from a classic menu with keypad options is that the caller just talks, with no options to work through. What separates it from a website chatbot is the channel. On a phone line every pause is audible, because people in conversation expect an answer within a fraction of a second. That is exactly why providers are currently competing so hard on speed.
One thing is worth keeping in view: an agent like this decides nothing you have not allowed in advance. It is a switchboard that understands speech, not a colleague. What it may say, which cases it takes and when it hands over to a person are all settings. For the wider picture of what such systems do inside a business, see our article on the AI assistant in business.
Which building blocks drive AI voice agent pricing
AI phone answering runs in three steps that follow one another and are billed separately. First a real-time transcription model turns speech into text, continuously, while the caller is still talking. Then a language model turns that text into a reply. Finally a speech synthesis model turns the reply into audible speech (and only at that point does the caller hear anything, which is why delays from all three steps pile up here).
On 1 October 2026 Microsoft introduced new models for the first and third of those steps. MAI-Transcribe-2-Streaming transcribes as the caller speaks and, according to the vendor, covers 60 languages with automatic, continuous language detection, so it works out by itself which language is being spoken. For output there is MAI-Voice-2.1 and the faster MAI-Voice-2.1-Flash, both covering 23 languages and 26 locales. Microsoft states an end to end latency of 150 milliseconds for 45 seconds of audio for the Flash variant and describes it as roughly 60 percent cheaper than comparable models. The models are available through Microsoft Foundry, the MAI Playground, Vercel, Azure Voice Live and, in part, OpenRouter.
In practice the step that matters most is the second one, and it is the one the coverage says least about. The language model in the middle determines how good the answers are, and it costs extra. Anyone who mistakes the transcription and speech prices for the total will budget too low. We collected the ways to bring those model costs down under LLM cost optimization.
| Step | What happens | How it is billed |
|---|---|---|
| Transcription | Speech is turned into text as the caller talks | Per hour of audio |
| Reply | A language model works out what should be said | Per volume of text processed |
| Speech output | The text is rendered as an audible voice | Per million characters |
| Telephony | The call is routed into the system | Per minute or as a flat rate, depending on provider |
languages are covered by the new transcription model according to the vendor, with continuous automatic language detection.
AI voice agent pricing: what one call costs
AI voice agent cost can be calculated exactly for two of the three building blocks, because speech synthesis pricing and transcription pricing are both published. Transcription is 0.54 US dollars per audio hour, and Microsoft explicitly calls that an introductory price through the end of the year. Whether it holds after that is not stated. Speech output is 22 US dollars per million characters, or 15 in the faster and cheaper Flash variant. Converted at the European Central Bank reference rate of 7 October 2026 (1 euro bought 1.1177 US dollars), you get the figures in the table.
Work a three minute call through. Transcription runs for the whole call, which is 0.05 audio hours and therefore about 2.4 cents. For speech output, assume the agent is speaking for roughly half the time. A minute of spoken text runs to about 750 to 900 characters including spaces, so a minute and a half is somewhere around 1,100 to 1,350 characters. With the Flash variant that costs about 1.5 to 1.8 cents. Together you land at roughly four cents per call, or just under five with the more expensive voice.
Scaled up to 100 calls across 20 working days, so 2,000 calls a month, that is about 79 to 85 euros for transcription and speech. This is a calculation with stated assumptions rather than a measurement of your business: change the call length or the share of talking and the number moves. Two items are missing from that total entirely, namely the language model for the replies and the telephony. And buying a finished service rather than building one costs more than the raw model prices, because dialogue logic, call routing and operations are bundled into it.
| Building block | List price | Converted | Unit |
|---|---|---|---|
| Transcription (MAI-Transcribe-2-Streaming) | 0.54 USD | about 0.48 EUR | per hour of audio, introductory price through year end |
| Speech output (MAI-Voice-2.1) | 22 USD | about 19.68 EUR | per million characters |
| Speech output (MAI-Voice-2.1-Flash) | 15 USD | about 13.42 EUR | per million characters |
| Language model for the reply | not included in these prices | depends on the model chosen | per volume of text processed |
is what transcription and speech output together come to on a three minute call, calculated from the published prices.
Where an AI voice agent processes your calls
An AI voice agent starts from a different place than a chatbot under data protection law, because a voice is more than text. Microsoft writes in its own documentation for the speech services that audio of people speaking, and the transcripts that go with it, may count as personal or sensitive data under various privacy laws, since they carry not only the voice but, depending on context, personal information in what is said. The same document states that you, as the operator, are responsible for obtaining every permission needed to process that data.
On storage the statement is refreshingly clear. For real time transcription, Microsoft says audio is processed only in the server's memory and that no data is stored at rest. For the question of which country that happens in, the EU Data Boundary applies: Microsoft commits to storing and processing customer data and personal data for its enterprise online services within the EU and EFTA. That commitment carries explicitly named exceptions in which data does leave the boundary, and for Azure it only applies if you deploy the service in a region inside it.
From that follows a check nobody will run for you. Which region does your service run in, is the data processing agreement in place, and is the specific model even available in that region. With freshly released models the last point is not a formality, because regional availability tends to arrive in stages. If the data is not to leave your premises at all, the route is to run the speech services in your own environment, which Microsoft offers as containers and where, according to the vendor, audio and transcript do not go to the cloud. We set out what that means more generally under local AI.
EU AI Act: what you have to tell callers about the AI
Under the EU AI Act, telling callers they are speaking to an AI has not been a matter of taste since 2 August 2026. That is the date Article 50 took effect, and paragraph 1 requires AI systems intended to interact directly with natural persons to be designed so that those persons are informed they are interacting with an AI system. The German Federal Network Agency, which supervises this, puts it just as plainly in its overview: affected persons must be informed about the interaction with the AI system.
There is an exception, and on a phone line it carries less weight than many hope. The information can be dropped where the fact is obvious anyway to a reasonably well informed and observant person. With a synthetic voice that sounds good, that is precisely what is in doubt, and the better the voices get, the weaker the argument becomes. We treat the spoken disclosure as the safe route. It has to come at the first interaction at the latest, which means in the opening line of the call rather than in a policy page.
In practice that means one sentence in the greeting saying what the caller is dealing with, and a route to a human that actually works. Article 99 sets out what non compliance can cost: fines of up to 15 million euros or 3 percent of worldwide annual turnover, with the lower of the two applying to small and medium sized companies. We wrote up the detail in our article on the chatbot disclosure duty under Article 50, and the separate training obligation in Article 4.
| Question | What applies |
|---|---|
| Do I have to announce the AI? | Yes, unless it is obvious that an AI system is speaking |
| When does the disclosure have to come? | At the first interaction at the latest, so in the greeting |
| Who is responsible? | The provider of the system under paragraph 1; deployer duties apply in further cases |
| Since when does this apply? | Article 50 of the EU AI Act has applied since 2 August 2026 |
| What does a breach cost? | Up to 15 million euros or 3 percent of annual turnover, the lower figure for SMEs |
When missing call recording consent becomes a criminal offence
Call recording consent is the point where a data protection question turns into a criminal law question, and it is routinely overlooked. Under section 201 paragraph 1 of the German Criminal Code, anyone who records the non publicly spoken word of another person without authorisation faces up to three years imprisonment or a fine. A phone call is non publicly spoken word. That holds regardless of whether a person or a machine is listening at the other end, and it holds for recording meetings too, which is being promoted as one of the headline uses for these models.
The decisive word is unauthorised. Ask first and get a yes, and the recording is authorised. Record because the software offers the button, and it is not. For a voice agent that means consent belongs at the start of the call, alongside the AI disclosure, and anyone who declines needs a route that still gets them where they were going.
A distinction here saves a great deal of effort. A running transcript that exists only in memory so the system can form its reply, and disappears afterwards, is a different thing from a stored recording kept for later review. The less you retain, the smaller this whole topic becomes. So the question of what you genuinely need to store is worth answering before you choose a provider rather than after. For the legal assessment of your own situation, take legal advice rather than this article.
of the German Criminal Code makes unauthorised recording of the non publicly spoken word punishable by up to three years imprisonment or a fine.
Who benefits from AI phone answering
Whether AI phone answering pays off depends less on the technology than on what gets dropped when the phone rings. In the trades it is often the work itself, because the same person orders materials, advises customers and answers calls. In a medical practice it is the patients at the desk while the line is busy. In both cases the gain is not a saved salary but the calls getting answered at all. What such an agent looks like in practice is shown on our AI voice agent page.
One argument against cannot be calculated away. If your calls are mostly delicate, meaning complaints, emergencies or negotiations, then a machine at the front of the line is the wrong call however cheap it is. The sensible dividing line, in our experience, runs between recurring and one off enquiries.
- 01
Trades and field service businesses
Usually pays offMany short, similar calls about appointments and availability while nobody is in the office. Start with taking messages and callback requests, not with advice.
- 02
Practices, law firms, administration
Worth checkingHigh call volume on fixed topics, but health or client data often comes up in the conversation. Settle data protection and retention first.
- 03
Retail and customer service
Useful for standard casesDelivery status, opening hours and returns work well. Complaints belong with a person from the start.
- 04
Consulting and high value sales
Probably notEvery call is different and the first impression decides. Good call forwarding beats an agent here.
Frequently asked questions
What is an AI voice agent, and where does it fit in customer service?
It is software that answers calls, understands spoken language, replies in a synthetically generated voice and hands over to a person when needed. In customer service it works best on recurring enquiries such as delivery status or opening hours, with complaints going straight to a person.
What does an AI voice agent cost per call?
Around four cents for transcription and speech output combined on a three minute call, calculated from the Microsoft prices published in October 2026. The language model for the replies and the telephony come on top.
Do I have to tell callers they are speaking to an AI?
Yes, unless it is obvious. Article 50 of the EU AI Act has applied since 2 August 2026 and requires that people are informed they are interacting with an AI system, at the first interaction at the latest.
Can I record calls handled by an AI voice agent?
Only with consent. In Germany, recording the non publicly spoken word without authorisation is a criminal offence under section 201 of the Criminal Code, carrying up to three years imprisonment or a fine.
Where is the audio processed?
That depends on the provider and the region you choose. Microsoft states that real time transcription is processed only in server memory with no data stored at rest, and commits under the EU Data Boundary to processing within the EU and EFTA, subject to named exceptions.
How many languages do these systems handle?
Microsoft's new transcription model covers 60 languages according to the vendor and detects continuously which one is being spoken. Speech output covers 23 languages across 26 locales.
Is an AI voice agent worth it for a small business?
If calls regularly go unanswered, usually yes, because the gain sits in the calls you now take rather than in saved wages. If your calls are mostly delicate or advice heavy, probably not.
How to approach your voice AI agent setup
- 01
Count the calls
Log for two weeks how many calls come in, how many go unanswered and what they were about. Without that list, any calculation is guesswork.
- 02
Pick one case
Take the most frequent recurring enquiry, usually an appointment or a callback request, and let the agent handle only that to begin with.
- 03
Settle disclosure and consent
Write the greeting that names the AI, decide whether you need to record at all, and build in the route to a person.
- 04
Check region and contract
Before launch, confirm which region the service runs in, that the data processing agreement is signed, and how long transcripts are kept.
The technology is no longer the expensive part, and it is no longer the difficult part either. What stays difficult is deciding which conversation you are willing to hand to a machine and which one you are not.
AI agentsAgentic AI Explained: What It Means for Your Business
Gemini 4 ArgonGemini 4 Argon: What the Model Is Actually Built For
Computer useComputer Use Agent: What It Operates and What It May Do
Google spam updateGoogle Spam Update: What Can Your Numbers Really Tell You?
LLM costsLLM Cost Optimization: When a Model Switch Pays Off
Detect AI-written textDetect AI-written text: what AI detectors get wrong
Google AI ModeGoogle AI Mode: What It Means for Your Website
AI agentsWhat Is an AI Agent, and How Is It Different From a Chatbot?
Structured DataStructured Data: What It Really Does for AI Search
AI AssistantAI Assistant for Business: Types, Uses and Data Protection
GEOE-E-A-T: Trust Signals for Google and AI Search
Local AILocal AI for Business: What “Local” Really Means
GEOAI Crawlers in robots.txt: Managing GPTBot and Co.
AI for SMEsAI for SMEs: How to Introduce AI Step by Step
Perplexity SEOPerplexity SEO: How to Get Your Site Cited as a Source
ChatGPT SEOChatGPT SEO: How to Get Your Business Found in ChatGPT
llms.txtllms.txt: What the File Does and When It Pays Off
AI SEOAI SEO: What Changes Compared to Traditional SEO
AI OverviewsGoogle AI Overviews: How Google Picks Its Sources
GEO vs AEO vs LLMOGEO vs AEO vs LLMO: The AI Search Terms Explained
Measure AI VisibilityMeasure AI Visibility: Method, Metrics and Limits
ChatGPT AdsChatGPT Ads: How to Advertise on ChatGPT in Germany
AI Literacy ObligationAI Literacy Obligation: What Article 4 Requires Since 2026
AI DisclosureAI Chatbot Disclosure: Article 50 in Practice
AI Text WatermarkAI Text Watermark: What Claude Marks and What It Doesn't
Where the information on this page comes from
- Microsoft AI: Our first streaming transcription model debuts at no. 1 on Artificial Analysisaccessed 8 Oct 2026
- TestingCatalog: Microsoft launches MAI-Transcribe-2-Streamingaccessed 8 Oct 2026
- SiliconANGLE: Microsoft targets ultra-realistic voice agents with its first streaming transcription modelaccessed 8 Oct 2026
- Microsoft Learn: Data, privacy, and security for Speech to textaccessed 8 Oct 2026
- Microsoft Learn: What is the EU Data Boundary?accessed 8 Oct 2026
- Regulation (EU) 2024/1689 (EU AI Act), Art. 50, 99, 113accessed 8 Oct 2026
- Bundesnetzagentur: Transparency obligations under the AI Actaccessed 8 Oct 2026
- German Criminal Code, section 201: Violation of the privacy of the spoken wordaccessed 8 Oct 2026
- dejure.org: section 201 StGBaccessed 8 Oct 2026
- European Central Bank: euro reference rates of 7 Oct 2026accessed 8 Oct 2026
- fiumu: speaking time calculator, length of voiceover scriptsaccessed 8 Oct 2026
- sinusaudio: calculating text length against audio runtimeaccessed 8 Oct 2026


