Speech to text
Speech-to-text transcribes a caller's spoken words into text a language model can reason over, the first of three stages in a pipeline voice agent's turn, alongside the language model and text-to-speech. Models differ on latency, accuracy, language coverage, and whether they can detect the end of a turn themselves as part of transcription.
A misheard word in a regulated call, an account number, a medication name, a dollar figure, can propagate into a wrong action downstream in a way that matters far more than it would in a text chat, where a person can reread what they typed.
A phone caller's narrowband audio is transcribed by a model tuned for telephony, while a clearer browser call uses a different model tuned for that audio instead.
In Dynamiq, choose from a curated catalog spanning Deepgram, AssemblyAI, Cartesia, ElevenLabs, Speechmatics, OpenAI, Groq and self-hosted options for the speech-to-text stage of a pipeline voice agent, each listed with its measured latency, price and word error rate, or skip the stage entirely with a single realtime speech-to-speech model instead.

See an agent on your own workflow.
Bring a process and its documents. Our engineers will show you how Dynamiq runs it, in your environment or ours.