# Speech to text

URL: https://www.getdynamiq.ai/glossary/speech-to-text

> Converting spoken audio into written text in real time, the first stage of a pipeline voice agent's turn.

Converting spoken audio into written text in real time, the first stage of a pipeline voice agent's turn.

Speech-to-text transcribes a caller's spoken words into text a language model can reason over, the first of three stages in a pipeline voice agent's turn, alongside the language model and text-to-speech. Models differ on latency, accuracy, language coverage, and whether they can detect the end of a turn themselves as part of transcription.

A misheard word in a regulated call, an account number, a medication name, a dollar figure, can propagate into a wrong action downstream in a way that matters far more than it would in a text chat, where a person can reread what they typed.

A phone caller's narrowband audio is transcribed by a model tuned for telephony, while a clearer browser call uses a different model tuned for that audio instead.

**In Dynamiq**, choose from a curated catalog spanning Deepgram, AssemblyAI, Cartesia, ElevenLabs, Speechmatics, OpenAI, Groq and self-hosted options for the speech-to-text stage of a pipeline voice agent, each listed with its measured latency, price and word error rate, or skip the stage entirely with a single realtime speech-to-speech model instead.

## See it in Dynamiq

-   [Voice Agents](https://www.getdynamiq.ai/product/voice-agents)

## Related terms

-   [Voice agent](https://www.getdynamiq.ai/glossary/voice-agent)
-   [Text to speech](https://www.getdynamiq.ai/glossary/text-to-speech)
-   [Barge-in and turn detection](https://www.getdynamiq.ai/glossary/barge-in-and-turn-detection)

## See an agent on your own workflow.

Bring a process and its documents. Our engineers will show you how Dynamiq runs it, in your environment or ours.

[Start free](https://app.getdynamiq.ai/signup) · [Talk to the team](https://www.getdynamiq.ai/book-a-demo)
