Why choose streaming speech-to-text over batch transcription for a live dictation product?
answer
- ask who is waiting for the text
- partials arrive early, may be rewritten
- later words disambiguate earlier ones
- endpointing exists only in one mode
- many teams run both passes
basics
~20 sStreaming returns interim text within a few hundred milliseconds so the speaker can see and correct errors as they talk. Batch waits for the whole recording but decodes with full right-hand context, so its final transcript is usually more accurate.
solid answer
~50 sThe choice is about who is waiting. Streaming opens a persistent connection, sends audio in small chunks and returns unstable **interim results** that get revised until a segment is finalized — that is what lets a clinician watch a dictation appear and fix a mis-heard drug name in the moment. Batch uploads finished audio and decodes it with the whole utterance available, so the model can use later words to disambiguate earlier ones; it usually wins on accuracy and costs less per minute of engineering. A common production answer is both: stream for the live view, then re-run a batch pass over the stored audio to produce the transcript of record. What you must measure differs too — streaming cares about time-to-first-partial, how often finalized text flickers, and endpointing delay; batch only cares about final accuracy and turnaround.
go deeper
Know that streaming returns text as the person speaks while batch transcribes a finished recording, and that streaming text can change before it settles.
Be ready to explain why full-audio context makes batch more accurate, what an interim versus final result means, and how you would render partials without flicker.
Show the production shape: a dual-path design, which transcript is authoritative, how user edits survive reconciliation, and the streaming-specific metrics you alert on.
Own the tradeoff at product level — where the extra accuracy is worth double transcription cost, which workflows genuinely need real-time text, and how the answer changes if domain vocabulary drives your error budget.
## The two modes **Batch (file) transcription** takes a complete audio file and returns one transcript. The decoder sees the entire utterance before it has to commit to any word, and turnaround is typically a fraction of real time — a 12-minute dictation may come back in well under a minute. **Streaming transcription** keeps a connection open, pushes audio in chunks of tens to hundreds of milliseconds, and receives hypotheses back continuously. Those early hypotheses are *interim* (partial) results: they are explicitly allowed to change. Once the recognizer is confident that a span will not be revised, it emits a *final* result for that span. ## Why streaming is less accurate Speech is ambiguous locally and disambiguated later. "Hyper tension" versus "hypertension", "ate" versus "eight", "fifteen" versus "fifty" are often only resolvable from the words that follow. A batch decoder has that right-hand context for free. A streaming decoder must either wait (adding latency) or commit and revise. Providers mitigate with a small lookahead window and by rewriting partials, but the accuracy gap on the same audio is real, and it is largest exactly on rare domain vocabulary — the terms a clinical or legal product cares about most. ## What streaming buys The payoff is not accuracy, it is interaction. On a dictation product the clinician sees text form as they speak, notices that a drug name came out wrong, and repeats it immediately instead of proof-reading a wall of text later. Correction-in-the-moment is far cheaper than correction-after-the-fact, and perceived responsiveness is part of the product. Streaming is also mandatory in any conversational setting, because a voice agent cannot begin thinking about a reply until it has words. ## Interim-result handling is a UI problem If you render every partial verbatim, text visibly flickers and rewrites itself, which readers find distracting. Typical mitigations: render partials in a lighter style and finals in normal style; debounce repaints; only re-render the tail segment that is still non-final; and never let downstream logic act on a partial — commands, extraction and storage should key off finalized segments only. ## Endpointing exists only in streaming A batch job knows where the audio ends. A streaming pipeline must *decide* when the speaker has stopped, so it can finalize a segment and hand it on. That decision — endpointing — is its own tuning problem and a common source of complaints ("it cut me off"). ## The hybrid that most teams land on Stream for the live experience and run a batch pass over the same audio for the transcript of record. You pay for the audio twice, which is usually small next to the cost of a wrong clinical term, and you get a better archived transcript plus richer enrichment (word timestamps, diarization) that the batch endpoint often produces more reliably. Reconciliation is the wrinkle: the two transcripts will differ, so decide up front which one the reviewer's edits attach to. ## What to measure - **Time-to-first-partial** — how fast text starts appearing. - **Partial stability** — how much finalized text differs from what was first shown. - **Finalization latency** — audio-end to final segment. - **Endpointing delay** and false-cut rate. - **Final transcript word error rate**, ideally reported separately for general speech and for the domain terms you care about. ## Practical notes as of mid-2026 Providers ship a batch-optimized model and a streaming variant rather than one model for both, and Whisper is no longer the default recommendation — OpenAI's GPT-Transcribe family, Deepgram Nova-3, AssemblyAI Universal and ElevenLabs Scribe are the common production choices, most offering both modes behind similar interfaces. Treat that lineup as changeable; the streaming-versus-batch tradeoff underneath it is stable.
- How would you present interim results without the transcript visibly rewriting itself?Render non-final text in a dimmed or italic style and promote it on finalization, so rewrites read as expected rather than as glitches. Only repaint the trailing non-final segment, debounce updates, and keep the caret and any user edits anchored to finalized spans. Downstream consumers — extraction, commands, storage — should subscribe to finals only, never partials.
- If you run both a streaming and a batch pass, which transcript is authoritative?Pick one and say so in the product. Usually the batch pass is the record, because it decodes with full context and carries richer enrichment; the streaming view is a live aid. The trap is user edits: if a clinician corrects the streaming text and the batch result later overwrites it, you have destroyed their work. Either apply edits after reconciliation, or freeze the transcript at the moment editing begins.
- When is streaming the wrong choice even though the audio is live?When nobody is reading the text in real time. Call recording for later compliance review, meeting archives, or podcast subtitling all arrive as live audio but are consumed afterwards, so you gain nothing from partials and give up accuracy and simplicity. Buffer the audio and transcribe in batch.
saying these in an interview costs you the question
- Claiming streaming and batch produce identical transcripts
- Treating interim results as final and acting on them
- Thinking streaming is chosen for cost, not interactivity
- Ignoring endpointing entirely in a live pipeline
- Assuming batch mode still needs turn or silence detection