Voice agents engage with users in real-time, interactive conversations. Their success depends on what users say, how they sound, and when they speak. Static evaluations provide a task without further user interaction, leaving much of this conversational behavior untested. Human testing captures richer interactions but is costly to scale. Variation across testers and sessions also makes it harder to compare results consistently. A voice user simulator plays the user’s role in a spoken conversation. Microsoft Foundry’s voice user simulator combines goal-driven scenarios, configurable acoustic conditions, and optional conversational dynamics. It helps teams move beyond clean audio and orderly turn-taking to evaluate how voice agents behave during pre-production testing.
Authors: Sheng Liu, Shivank Goel, Morteza Ziyadi
The Same Goal, Different Conversations
A voice user simulator acts as the user in a spoken conversation with a voice agent. You define what the user wants to accomplish, and the simulator pursues that goal through the exchange. It speaks to the voice agent, responds to its questions, and generates each user turn from the exchange so far. Instead of writing every line of a test script, you describe the task and let the interaction unfold.
But the task alone does not capture the full voice experience. The same request may come from a quiet room or a busy street. A user may acknowledge an explanation, interrupt to ask a question, or correct an earlier detail. The voice agent must handle both the request and the conditions surrounding it.
Microsoft Foundry’s voice user simulator lets teams test these variations through three distinct pillars:
- Scenarios define what the simulated user wants to accomplish, along with the facts and constraints that guide the conversation.
- Acoustic conditions shape how the simulated user sounds through voice selection and different listening environments.
- Conversational dynamics shape how the simulated user participates, including acknowledgments, interruptions, and corrections.
Teams can configure these three dimensions independently and combine them to explore the users and interactions their application needs to support. Available options depend on the project and service configuration.
Figure 1. Scenarios, acoustic conditions, and conversational dynamics shape the simulated user. The simulator exchanges speech with the voice agent under test. All figures and dialogue examples are illustrative, not actual UI screenshots, recordings, or measured results.
Scenarios: Define the Goal, Not the Script
A useful test needs a clear user goal, but the path to that goal depends on the conversation. Scenarios give the simulator a goal, relevant facts, and constraints without prescribing a script. The simulator then generates the conversation turn by turn, using the exchange so far to answer questions and respond to suggestions. It is instructed to preserve the supplied facts without coaching the voice agent through its workflow.
In a fictional delivery-date example, the scenario might supply an order number, a preference for Friday, and a requirement to keep the address unchanged. The simulator might provide the order number when asked and respond according to that preference if the voice agent proposes another date. The scenario guides the task while the dialogue develops.
Acoustic Conditions: Different Voices, Different Environments
Voice agents encounter different speakers, surroundings, and call audio quality. Acoustic conditions bring that variety into testing. Teams can select from Foundry supported voices, with male and female voices and different accents among the available options. Voice options depend on service availability. Built-in ambient sounds, such as street traffic, crowd chatter, background TV, and metro-station noise, add a listening environment with adjustable background levels. A telephonic effect applies narrowband, telephone-style audio processing and can be used alone or combined with ambient sound.
Figure 2. Select a user voice and the listening conditions for the delivery-date conversation. The voice options shown are examples.
For the delivery request, a selected voice can speak over telephone-style audio with traffic in the background, representing a user calling from a busy street. Another run could use a different voice in a quiet setting. The same configured voice and audio effects are applied to the user’s turns, including interruptions.
Conversational Dynamics: Test More Than Turn-Taking
Users do more than wait for their turn: they take the floor, signal understanding, and follow up on what they hear. An optional conversational preset adds simulated versions of these behaviors to testing. Barge-in lets the simulated user interrupt while the voice agent is speaking. Backchannels add brief acknowledgments without asking it to stop. Clarification and repair can introduce a question or correction during a longer response.
Figure 3. During the delivery-date conversation, the user can interrupt, acknowledge, or ask a follow-up question. Questions or corrections may also be inserted during a longer response.
In this fictional delivery call, the simulator might interrupt with “Wait, Friday works better,” acknowledge an explanation with “Mm-hm,” or ask “Is the address unchanged?” during a longer explanation. Each contribution has a different purpose and calls for a different response. The aim is to reflect everyday interaction, not simply make the test harder.
From Simulated Conversations to Evaluation
By testing across scenarios, acoustic conditions, and conversational dynamics, teams can explore where a voice agent performs well and where it needs improvement. Reusing these settings across evaluations lets teams compare observed responses when a configuration changes. Voice user simulation makes these varied interactions part of routine testing, alongside other pre-production checks.
Get Started
Open your project in Microsoft Foundry and start with a scenario that reflects a task your voice agent needs to handle. Follow the conversation simulation guide for setup steps, voice and audio configuration, and SDK and REST API examples. Choose the acoustic conditions and conversational behaviors that suit your application, then run the simulation. After the run, review the conversations, listen to available audio, and evaluate how your agent performed. The voice-agent observability guide shows you how.