Real-time voice agents for inbound calls

Voice Agents answer calls on the phone, the web or your own backend, reason with an LLM, and speak back, with every turn measured.
Inbound call / Front deskCompleted
CallerI need to reschedule tomorrow's appointment.
AgentI can do that. Thursday at 10 or Friday at 2?620 ms
CallerFriday works.
AgentDone. You will get a text confirmation now.540 ms
calendar.rescheduleFriday 2:00 PM confirmed
SIP, 1 tool
Illustration of an inbound voice call: the agent reschedules an appointment with the calendar system, with the response time for each turn.

In short

Voice Agents are real-time phone and web agents that pick up inbound calls, reason with an LLM and speak back with text-to-speech, or run as one speech-to-speech model instead. They reach callers over WebRTC in the browser, SIP on real phone numbers, or a WebSocket for your own telephony backend. Every agent can call tools, hand off to a person, and gets tested against simulated callers scored by an LLM judge before it takes a real call.

What Voice Agents does

Three ways in
Answer inbound calls over SIP on numbers you bring, WebRTC from a browser or app, or a WebSocket from your own backend. Dynamiq has no outbound dialer.
Knowledge during the call
Search your knowledge bases mid-call, so answers come from your own policies and documents while the caller waits.
A curated model catalog
15 voice model vendors for speech-to-text, language models, text-to-speech and speech-to-speech, each with its measured latency and cost.
Per-turn latency
Every call shows a waterfall of transcription, end-of-turn detection, LLM time to first token and speech synthesis, and turn-taking and key-term settings tune how the agent listens.
Tools and transfers
Give the agent HTTP tools, client tools your app answers, MCP servers and labeled transfer destinations to hand a caller to a person. Callers can key in digits, and the agent can send tones to work through another system's phone menu.
Simulations, not guesswork
Run scripted callers against the deployed agent and have an LLM judge score pass rate, accuracy and experience before real callers do.

A curated model catalog, measured

Every stage of a pipeline agent, speech-to-text, the language model and text-to-speech, and every realtime speech-to-speech model, comes from a catalog Dynamiq has wired up and measured, so you pick from combinations known to work rather than assembling one yourself. An estimate bar above the fields shows the configuration's cost per minute and latency before you deploy, so a choice is a comparison, not a guess. Across every stage, Dynamiq's catalog spans 15 voice model vendors.

  • The default combination, Cartesia, OpenAI and ElevenLabs, is chosen for latency, not for any single benchmark
  • Speech-to-text models with native end-of-turn detection remove a whole pipeline stage instead of speeding one up
  • Bring your own OpenAI-compatible LLM endpoint, or a custom voice ID from your own provider account
front-desk / PipelineRealtime or pipeline
Speech to textDeepgram
Language modelOpenAI
Text to speechElevenLabs
VoicesWarm and calm, en-GB
Illustration of a voice agent's model settings: a speech-to-text vendor, a language model and a text-to-speech voice, each picked from a catalog Dynamiq has wired up and measured.

Where the milliseconds go

A caller judges a voice agent mostly on timing. On a pipeline agent, Dynamiq measures four stages on every turn: end-of-turn detection, transcription, LLM time to first token and text-to-speech time to first byte, and shows them as a proportional bar so the stage to fix is visible at a glance. Picking a speech-to-text model with native end-of-turn detection, such as Deepgram's flux models or Cartesia's ink-2, removes a whole stage instead of speeding one up, which is usually the biggest single latency win available.

  • Turn sensitivity and end-of-speech silence tune how eagerly the agent replies
  • Realtime, speech-to-speech agents report time to first token instead of a stage breakdown
  • Test calls and voice-mode simulations are the only way to get real numbers
Inbound call / TranscriptCompleted

I need to move my appointment to next week.

transcribed +110 ms, end of turn +80 ms

I can do Tuesday at 10 or Thursday at 2. Which works?

LLM time to first token, TTS time to first byte, 0.9 s end-to-end

Thursday, please.

transcribed +95 ms, end of turn +70 ms

Done. You will get a text confirmation shortly.

LLM time to first token, TTS time to first byte, 0.8 s end-to-end

Illustration of a voice call transcript with a latency bar under each agent turn, split into language model time to first token and text-to-speech time to first byte.

Tested before it takes a real call

Simulations run a set of scripted callers, defined by a persona, a goal and an expected outcome, against the deployed agent and have an LLM judge decide whether it did what you said it should. Run them as text for fast, cheap iteration on reasoning and tool use, or as voice to also exercise speech recognition and synthesis and get real latency percentiles. Turn a call that went wrong into a scenario with one click from the Calls tab, so every regression becomes a permanent test.

  • A run scores scenarios passed, accuracy and experience, with a judge rationale per scenario
  • Generate starter scenarios from the agent's own configuration, or write them by hand
  • Simulations run against the deployed agent, so results reflect production behavior
Simulations / Scenario sets3 of 4 passed

Caller persona and goal

A patient who can only come after 4 pm and wants the earliest slot this week.

Expected outcome

The agent offers a slot after 4 pm and confirms it without transferring the call.

Reschedule with a conflictPassed
Caller asks for a refundPassed
Angry caller, wrong accountFailed
Caller switches to ArabicPassed
Run simulation
Illustration of voice simulations: a scenario with a caller persona and goal and an expected outcome, and a run where three of four scripted callers passed.

Voice model vendors

  • Deepgram
  • AssemblyAI
  • Cartesia
  • ElevenLabs
  • Speechmatics
  • OpenAI
  • Groq
  • Baseten
  • Kokoro (via DeepInfra)
  • Google
  • xAI
  • Anthropic
  • Cerebras
  • Fireworks AI
  • Together AI

Questions and answers

Can Dynamiq Voice Agents make outbound calls?

No. Voice Agents answer inbound calls only, whether they arrive over the phone, the web or your own backend. There is no outbound dialing API.

How do callers reach a voice agent?

Over WebRTC from a browser or mobile app, over SIP from a phone number you connect through your own provider such as Twilio or Telnyx, or over a WebSocket from a backend that already holds the audio, such as a telephony stack.

How is a voice agent tested before it goes live?

You can talk to it yourself from the Test tab, and you can run simulations: scripted callers with a persona and an expected outcome, scored by an LLM judge for pass rate, accuracy and experience. Voice-mode simulations also produce real latency numbers.

Do configuration changes need a redeploy?

No. The worker reads the agent's current instructions, models, tools and settings at the start of every call, so an edit applies to the next call immediately. You only redeploy to pick up a new platform worker version.

What happens when the agent cannot help a caller?

If you configure transfer destinations, the agent can hand the caller to a person with a cold transfer, announcing the handoff first. Without any destinations configured, it cannot transfer and says so instead of failing silently.

See an agent on your own workflow.

Bring a process and its documents. Our engineers will show you how Dynamiq runs it, in your environment or ours.