skip to content

How should a voice agent handle barge-in when a customer talks over its spoken reply?

level: seniorimportance: must knowfreq 50%

answer

  1. stopping playback is only step one
  2. microphone stays open during playback
  3. the model thinks it said everything
  4. echo makes an agent interrupt itself
  5. 'mm-hm' is not a floor grab

basics

~20 s

Stop playback within a couple of hundred milliseconds, keep the microphone open the whole time so nothing is missed, and truncate the conversation history to the words the customer actually heard — not the full reply the model generated.

solid answer

~50 s

Barge-in has three parts and teams usually implement only the first. **Stop fast**: detect incoming speech and cut audio output within roughly 100–300 ms, cancelling any queued synthesis and in-flight generation so nothing plays after the interruption. **Never stop listening**: the recognizer must run continuously, so the interruption's first syllables are already captured; a pipeline that opens the microphone after playback stops loses the beginning of every barge-in. **Repair the state**: the model believes it said its whole sentence, but the customer heard only the first four words, so you must rewrite the assistant turn in history to the truncated text — otherwise it will reference an offer nobody heard. Guard against false triggers with echo cancellation, since without it the agent's own voice re-enters the microphone and it interrupts itself, and treat very short back-channels like "mm-hm" as non-interrupting.

go deeper

for a junior

Know that barge-in means the user can talk over the assistant, that playback must stop quickly, and that the microphone has to stay open while the assistant speaks.

for a middle

Explain the three parts — fast cancellation across every buffer, always-open capture with echo cancellation, and rewriting history to the text actually heard — and why fixed silence timeouts are a poor endpointing rule.

for a senior

Show the production instrumentation: stop latency, false barge-in and missed-onset rates, endpoint accuracy, and how you tune thresholds per acoustic environment rather than shipping one global value.

for a principal

Own the judgment about how much latency and complexity conversational naturalness is worth: a model-based turn detector in the hot path, per-deployment tuning, and the fallback behaviour when it is unavailable.

## Why turn-taking is the hard part of voice Text chat has an explicit floor: the user presses send. Voice has none. The system must continuously decide two things — has the user finished speaking, and is the user trying to take the floor while I am speaking. Get either wrong and the product feels broken in a way no amount of answer quality repairs. In a noisy drive-thru or a multilingual ordering lane, this is the dominant source of complaints. ## The three mechanisms of barge-in ### 1. Fast cancellation When incoming speech is detected during playback, output must stop within a couple of hundred milliseconds. That means cancelling at every layer: the audio device buffer, any queued synthesized chunks, the synthesis request itself, and the language-model generation still streaming. A common bug is stopping the speaker but letting buffered audio drain, so the assistant keeps talking for a second after the customer interrupts — which reads as rudeness. ### 2. Always-open capture The recognizer must be running during playback, not started when playback ends. If capture begins after cancellation, the first 200–400 ms of the interruption — often the whole word that carries the intent, like "no" or "wait" — is gone. Always-open capture is what makes the difference between "it stopped" and "it heard me". This only works with **acoustic echo cancellation**. Without it, the microphone picks up the agent's own voice from the speaker and the system barges in on itself, producing an agent that cannot finish a sentence. In a car lane or on a speakerphone this is the default condition, not an edge case. ### 3. Context repair — the part most teams miss Suppose the agent generated "Your total is fourteen fifty, and would you like to add a drink to that?" and was cut off after "fourteen fifty". The customer never heard the drink offer. If you append the full generated text to the conversation history, the model now believes it made an offer, and its next turn may say "as I mentioned, the drink" — a hallucination from the user's point of view. The fix is to track how much audio actually played, map that back to the corresponding text, and store *that* as the assistant turn, typically with a marker that it was interrupted. Realtime provider stacks expose the played duration for this reason; in a cascaded stack you track it yourself. ## Not every sound is an interruption Humans use back-channels — "mm-hm", "yeah", "right" — to signal listening, not to take the floor. An agent that halts on every one of them is exhausting. Practical policy: require a minimum duration and energy, ignore known short affirmatives, and optionally let a small classifier decide whether the incoming speech is a floor grab or acknowledgement. In genuinely noisy environments, raise the threshold rather than disabling barge-in. ## The mirror problem: deciding the user has finished The other half of turn-taking is endpointing. The old approach is acoustic voice-activity detection with a fixed silence timeout — after N milliseconds of silence, assume the turn ended. It cannot win: a short timeout cuts off a customer who pauses to check a menu or to think mid-order, and a long one makes every exchange feel sluggish. Practice as of mid-2026 has moved to **semantic (model-based) turn detection**: a small model looks at the audio and the partial transcript and predicts whether the utterance is complete. "I'll have a large coffee and…" is syntactically and prosodically unfinished, so the agent waits even through a long pause; "that's everything, thanks" is complete, so it can respond after a very short one. This is an adaptive timeout rather than a fixed one, and it is the single change that most improves how natural a voice agent feels. It is not free — it adds a model call in the hot path, and it degrades when the transcript is noisy or the language is one the detector handles poorly, so keep a generous fixed timeout as a backstop. ## What to measure - **Barge-in stop latency**: speech onset to silence at the speaker. - **False barge-in rate**: interruptions triggered by echo, noise, or back-channels. - **Missed-onset rate**: interruptions where the first words were not captured. - **False endpoint rate**: turns cut off mid-thought, and its opposite, **endpoint delay**. - **Context-repair correctness**: does the stored assistant turn match what was actually heard? Tune stop latency and endpointing thresholds per deployment; the right values in a quiet headset call and in a drive-thru lane are not the same.

  • Why replace a fixed silence timeout with semantic turn detection?
    A fixed timeout has no good value: short enough to feel responsive means cutting off anyone who pauses mid-order; long enough to be safe makes every exchange sluggish. Semantic detection predicts from the partial transcript and prosody whether the utterance is complete, so it waits through a pause after 'I'll have a large coffee and…' but responds quickly after 'that's everything'. Keep a generous fixed timeout as a backstop.
  • What breaks if you append the full generated reply to history after an interruption?
    The model's record diverges from the customer's experience. It believes it delivered an offer or a total that was never heard, and later turns reference it — 'as I said, the drink' — which reads as a hallucination. Track played audio duration, map it back to the corresponding text, and store only that truncated turn, marked as interrupted.
  • How do you stop an agent from barging in on itself?
    Acoustic echo cancellation on the capture path, so the agent's own output is subtracted from what the microphone hears. Without it, speakerphone and open-lane deployments feed the synthesized voice straight back into the recognizer and the agent halts on itself continuously. Add an energy and minimum-duration threshold, and suppress detection on audio that correlates with the current playback.
  • Should every detected sound during playback count as an interruption?
    No. Back-channels such as 'mm-hm' and 'right' signal listening, not a bid for the floor, and halting on them makes the agent feel jumpy. Require a minimum duration and energy, ignore a short list of affirmatives, and optionally classify the incoming speech as floor-grab versus acknowledgement. In noisy environments raise thresholds rather than turning barge-in off.

saying these in an interview costs you the question

  • Only stopping playback and calling barge-in solved
  • Opening the microphone after the assistant finishes speaking
  • Keeping the full generated reply in history after an interruption
  • Relying on a fixed silence timeout to detect end of turn
  • Halting on every back-channel like 'mm-hm'

context