Every vendor in the voice AI space prices its product a little differently, and that makes head-to-head comparison harder than it should be. A quote that looks cheap per unit can turn into the most expensive option once you run your actual call volume through it, and a quote that looks expensive up front can be the better deal if it's structured around results instead of usage. Before signing anything, it helps to understand the three pricing structures currently in use across the market and what each one optimizes for.
This matters more for enterprise buyers than it might seem. A ten-person startup testing a chatbot can absorb a pricing surprise. A contact center running tens of thousands of calls a month, or a property management company handling maintenance requests around the clock, cannot. Getting the pricing model wrong shows up directly in the budget line a few months in.
Per-character pricing
Per-character pricing is common among raw text-to-speech APIs, the kind of provider you'd call if you're building your own voice agent stack from scratch and just need a synthesis engine. You send text, the API returns audio, and you pay based on how many characters went in. Providers like ElevenLabs, Cartesia, Hume AI, and Deepgram commonly price this way, sometimes with tiered plans that reduce the per-character rate at higher volumes.
The appeal is transparency. You know exactly what a block of text will cost before you generate the audio, and it's easy to forecast spend if your usage is stable. The catch is that character count doesn't map cleanly to business value. A short confirmation message costs about the same, proportionally, whether it closes a support ticket or gets ignored. At scale, teams often find they're paying for verbosity rather than outcomes, and prompt engineering to shorten responses becomes a real cost lever, not just a UX nicety.
Per-minute pricing
Per-minute pricing shows up more often with voice agent orchestration platforms, the layer that handles call routing, conversation state, and integration with telephony, on top of a TTS or speech-to-speech engine underneath. Retell AI, Vapi, and Synthflow are examples of platforms in this category, where you're typically billed for connected call minutes rather than characters generated.
This model tracks more closely with how a contact center already thinks about cost, since minutes are the unit call centers have budgeted around for decades. It's also more forgiving of chatty agents, since a slow, verbose voice agent doesn't inflate your bill the way it would under per-character pricing, though it does mean average handle time becomes the metric to watch closely. A platform that resolves calls faster saves you money under this model; one that meanders costs you more the longer it talks.
Outcome-based pricing
Outcome-based pricing ties cost to a defined result rather than a unit of usage. That might mean a successful contact in a debt collection campaign, a completed appointment booking, or a resolved support case. Instead of paying for minutes or characters regardless of what happened on the call, you pay when the call did what it was supposed to do.
This model is most common with fully managed voice-agent operations, where a vendor doesn't just license you software but actually builds, integrates, and runs the call handling on your behalf. Deepdub's managed-operation line of business works this way across its core use cases, which span customer service, healthcare, financial services, property management, and debt collection. Because the vendor is running the operation rather than just providing an API, they can afford to price against results, since they control enough of the pipeline (call scripts, escalation logic, integration quality) to actually influence whether the outcome happens.
The tradeoff is less granular visibility into per-call cost, and it requires trusting the vendor's definition of a successful outcome, which is worth pinning down precisely in the contract before you sign. Done well, though, it aligns incentives better than either usage-based model: the vendor only gets paid when the thing you actually care about happens.
What this means if you're comparing vendors
If you're evaluating a raw voice API to plug into your own stack, per-character pricing is what you should expect, and the comparison point that matters is not the headline rate but the rate at your projected volume, including any tier breaks. If you're evaluating an orchestration platform to run agents without owning the full pipeline, expect per-minute pricing, and model your cost against expected average handle time, not just call count. If you want a vendor to own the outcome, not just supply the infrastructure, outcome-based pricing under a managed-operation model is worth asking about directly, since it isn't always offered upfront and often only comes up once a vendor understands your use case in enough depth.
It's also worth asking any vendor which category they actually fall into, since some blur the line. Deepdub, for instance, is not only a raw API. It offers Voice API for Agents for teams that want to integrate a voice engine themselves, and separately runs fully managed voice-agent operations where it builds and operates the call handling directly. That means the same underlying technology can be priced either way depending on how much of the operational burden you want to keep in-house.
Underlying infrastructure still matters, whatever the pricing model
Pricing structure aside, the quality of the voice engine underneath still determines whether any of these models pencil out. A per-minute platform running on a slow or unnatural-sounding voice engine will show worse average handle time than one running on fast, natural synthesis, which erases any pricing advantage on paper. Deepdub's Voice API for Agents runs at roughly 85 milliseconds typical time-to-first-audio, with 150 milliseconds as the p95 figure for worst-case latency in real-time mode, and outputs full 48kHz audio rather than the lower sample rates common elsewhere. The model behind it, Phantom X 3.2, is a 3.4-billion-parameter LLM-based TTS model supporting more than 100 languages and dialects, verified at the dialect level by Deepdub's own in-house language experts, along with cross-language zero-shot voice cloning from under three seconds of reference audio.
On an independent, community-run leaderboard, the TTS Arena Hebrew benchmark on Hugging Face, Deepdub (listed as dd-etts-3.3) currently ranks first with an 80 percent win rate. Deepdub has also published its own blind English benchmark, in which Phantom X 3.2 tied for first in expressivity against Inworld, Hume, Async, and ElevenLabs at roughly 125 milliseconds of latency. That second result comes from Deepdub's own study, not a neutral third party, and is worth weighing accordingly alongside the independent leaderboard result.
FAQ
Which pricing model is cheapest? There's no universal answer. Per-character pricing is cheapest for low-volume, short-message use cases. Per-minute pricing tends to be more predictable for high-volume contact centers with stable call lengths. Outcome-based pricing can be the most cost-effective option when a vendor is confident enough in their own operation to price against results, but it depends entirely on how "success" is defined in the contract.
Can a vendor offer more than one pricing model? Yes. Vendors that operate both a developer-facing API and a managed service, as Deepdub does, can typically offer usage-based pricing for the API and outcome-based pricing for the managed operation, since the two products carry different levels of vendor control over the outcome.
Does outcome-based pricing mean lower total cost? Not necessarily. It shifts risk rather than guaranteeing savings. You avoid paying for calls that didn't work, but the per-outcome rate reflects the vendor absorbing that risk, so it's worth modeling against your expected success rate rather than assuming it's automatically cheaper.
What should procurement ask about during vendor evaluation? Ask for the exact unit being billed (characters, connected minutes, or defined outcomes), how tier breaks work at your expected volume, and whether the vendor can share real average handle time or resolution rate data relevant to your use case, since that determines the real cost under a per-minute or outcome-based model.
Getting started
If you're comparing pricing models for your own voice AI rollout, the fastest way to get a real number is to model your own call volume and average handle time against each structure rather than comparing headline rates. Teams that want to integrate voice synthesis into their own stack can review Deepdub's Voice API for Agents and the developer documentation directly. Teams that would rather have the operation built and run for them can look at Deepdub's managed offerings for property management and debt collection, both priced around outcomes rather than raw usage.
About the author
Meet the Deepdub team: a dynamic group of technology entrepreneurs, engineers, scientists, and dubbing specialists, all united by a passion for revolutionizing the entertainment industry. Our diverse expertise fuels our innovative AI dubbing and localization platform, enabling us to tackle the challenges of making content universally accessible and culturally relevant. Through our blog, we share insights and stories from our journey, showcasing the creativity and technology driving us forward. Join us in redefining the future of entertainment.








