Text to speech

Converting written text into spoken audio, the final stage of a pipeline voice agent's turn, or one channel of a realtime model.

Text-to-speech turns a model's written reply into audio a caller actually hears, the last stage of a pipeline voice agent's turn. Models differ on how quickly they produce the first audio a caller hears, how natural the voice sounds, and whether a specific cloned or custom voice can be used instead of a stock one.

A regulated contact center needs its voice to sound consistent and professional on every single call, and the wait before the agent starts speaking is exactly what a caller experiences as the agent being slow, regardless of how fast the reasoning behind it actually was.

A legal disclosure that must be read in full without interruption calls for a different balance of speed and naturalness than a quick conversational reply where responsiveness matters more than a fully natural voice.

In Dynamiq, choose from a curated catalog spanning ElevenLabs, Cartesia, Deepgram, OpenAI, Groq and self-hosted options for the text-to-speech stage, each with measured latency and a curated voice lineup, plus a custom voice ID field for a cloned or fine-tuned voice from your own provider account.

See an agent on your own workflow.

Bring a process and its documents. Our engineers will show you how Dynamiq runs it, in your environment or ours.