For a tool-heavy voice agent, would you use a cascaded STT-LLM-TTS stack or speech-to-speech?
answer
- text is the auditable artifact
- three hops queue serially
- tone dies at the transcript
- swap one stage without the others
- tool-heavy pushes one way
basics
~20 sCascaded remains the production default for tool-heavy agents: you get a text transcript to log, evaluate and guard, and each stage is independently swappable. Speech-to-speech wins on latency and prosody because nothing is flattened to text, but it is harder to inspect and constrain.
solid answer
~50 sThe cascade — recognize, reason in text, synthesize — buys you an inspectable artifact at every hop. You can log and redact the transcript, run text evals and safety filters on it, swap the recognizer without touching the model, and reuse the same tool-calling and prompt stack as your text product. Its cost is latency, because three stages queue serially, and expressiveness, because tone, hesitation, emotion and emphasis are discarded the moment audio becomes text. A single speech-to-speech model keeps that paralinguistic signal, responds far faster, and handles interruption and prosody natively — but the thing you would evaluate no longer exists as text unless you transcribe it separately, and steering it is less precise. As of mid-2026 the honest answer is that tool-heavy, regulated or auditable agents still ship cascaded, and speech-to-speech is chosen where naturalness dominates and the tool surface is small.
go deeper
Know the two shapes: a cascade converts speech to text, reasons in text and converts back, while speech-to-speech is a single model taking audio in and returning audio.
Be able to name the concrete tradeoffs — serial latency and lost prosody on one side, missing transcript and weaker steering on the other — and say why each stage of a cascade is independently swappable.
Show the decision procedure against a real workload: tool count and consequence, audit obligations, whether tone carries signal, and the latency bar; then describe the mitigations that close most of the gap.
Own it as an architecture bet: vendor coupling, observability and eval maturity, hybrid routing by intent, and an honest statement that this is contested ground in 2026 rather than a solved question.
## The two architectures **Cascaded**: audio → speech recognition → text → language model (with tool calls) → text → speech synthesis → audio. Three or more independent components, each with its own vendor, price and failure mode. **Speech-to-speech (end-to-end)**: audio in, audio out, one model that never materializes a full text intermediate. Turn-taking, interruption and prosody are handled inside the model rather than by your orchestration code. ## What the cascade actually buys **An inspectable artifact at every hop.** The transcript is the single most valuable thing in a regulated voice product: it is what you log, redact, retain, hand to compliance, replay in a dispute and mine for eval sets. Speech-to-speech does not give you one for free. **Independent evolution.** Recognizers improve on a different cadence than reasoning models. In a cascade you can adopt a better recognizer on Monday and a better reasoning model on Friday, and A/B each in isolation. End-to-end couples all three to one vendor's release schedule. **Reuse of the text stack.** Tool definitions, retrieval, prompt templates, guardrails, judge-based evals, prompt caching — everything built for the text product applies unchanged. For an agent whose value is in calling twelve internal tools correctly, that reuse is most of the system. **Precise control.** You can rewrite the user's text before it reaches the model, refuse on a classifier, force a deterministic response for known intents, and pin exactly what gets spoken. ## What the cascade costs **Serial latency.** Recognition must finalize, the model must generate at least a first sentence, synthesis must start. Every hop adds queueing and network time, and the floor is noticeably above human conversational turnaround. Mitigations — streaming recognition, generating and synthesizing sentence-by-sentence, speculative starts on likely intents — narrow the gap but do not close it. **Lost paralinguistics.** Text is a lossy projection of speech. Sarcasm, hesitation, distress, urgency, emphasis on "I said *no* onions" — all gone before the model sees anything. For an empathy-sensitive or emotionally loaded interaction that loss is the product. **Flat output.** Synthesis re-invents prosody from text alone, with no memory of how the user sounded, so replies can be well-worded but tonally mismatched. **Error compounding.** A recognition error becomes the model's premise; the model cannot hear that the word was ambiguous. ## What speech-to-speech buys and costs Buys: sub-second response, natural interruption handling, prosody that carries through the reasoning, and a much simpler client loop. Costs: weaker observability (you must run a parallel transcription pass to get logs, which reintroduces cost and drift), less precise steering, immature tool-calling relative to the text stacks, single-vendor coupling, and evaluation methods that are far less mature than text evals. ## How to decide Ask four questions. 1. **How many tools, and how consequential?** A dozen tools that move money push hard toward cascaded, where you can validate arguments against a text trace. 2. **What is the audit obligation?** If a regulator, a dispute process or an internal review needs a record of what was said and why, you need text at the decision point — which the cascade gives you natively. 3. **Does tone carry information?** Distress lines, sales calls and companionship products lose real signal to text. Order-taking and account lookups mostly do not. 4. **What is the latency bar?** Long-tail agentic work already spends seconds in tools, so a few hundred extra milliseconds are invisible. Rapid back-and-forth chat is where end-to-end feels categorically better. ## The hybrid, and where practice sits in mid-2026 The common compromise is to keep the cascade for the reasoning and tool path but adopt end-to-end characteristics at the edges: a realtime streaming recognizer, model-based turn detection, sentence-level synthesis so the first words play while the rest is still generating, and a separate lightweight audio classifier that extracts sentiment or urgency and passes it to the text model as a tag — recovering some paralinguistic signal without giving up the transcript. Another pattern routes by intent: end-to-end for chit-chat and acknowledgements, cascaded whenever a tool must be called. As of mid-2026, cascaded remains the production default for tool-heavy agents while end-to-end realtime models have matured considerably; the gap is narrowing, and the decision is a judgment call about observability and control versus naturalness, not a settled best practice. State that plainly in an interview rather than claiming one architecture has won.
- Speech-to-speech is faster, so why would a bank still ship a cascade?Because the transcript is the compliance artifact. A cascade produces text at the decision point, which can be redacted, retained, replayed in a dispute and screened by classifiers before a tool that moves money is called. It also lets tool arguments be validated against an inspectable trace. Latency matters less when a turn already spends a second in backend lookups.
- How would you recover some of the tone information a cascade throws away?Run a lightweight audio classifier in parallel with the recognizer and pass its output to the text model as structured metadata — for example an urgency or sentiment tag, or a flag that the speaker was shouting. You keep the transcript as the primary artifact but the model can adapt register. It is coarser than end-to-end prosody, and you should log the tag so its effect on responses is auditable.
- Which parts of the perceived latency gap can a cascade actually close?Most of the avoidable ones: stream recognition rather than waiting for a final transcript, start synthesis on the first complete sentence instead of the whole reply, keep connections warm, and colocate the stages. What remains is the irreducible serial dependency — you cannot synthesize words the model has not generated from a transcript that has not finalized — which is why end-to-end still wins on rapid back-and-forth.
- How would you evaluate a speech-to-speech agent?Transcribe both sides with a separate recognizer to get a text trace for the usual task-completion and tool-correctness evals, then add audio-specific measures the transcript cannot capture: response latency distribution, interruption handling, and human or model ratings of prosody and appropriateness. Be explicit that the transcription pass is itself lossy, so a scoring failure may be the recognizer's, not the agent's.
saying these in an interview costs you the question
- Claiming speech-to-speech has simply replaced cascaded stacks
- Ignoring that the transcript is the audit and eval artifact
- Assuming latency is the only axis of the decision
- Forgetting that tone and emphasis are lost at transcription
- Treating the choice as settled rather than workload-dependent