skip to content

Whisper transcripts contain repeated phrases over silent audio — how do you diagnose and reduce it?

level: seniorimportance: should knowfreq 44%

answer

  1. Thirty-second windows always decode
  2. No evidence means language prior wins
  3. Subtitle boilerplate is the tell
  4. Three fields expose it
  5. Strip the silence before uploading

basics

~20 s

This is Whisper's known hallucination on non-speech: with no acoustic evidence the decoder falls back on language priors and loops or emits training-set boilerplate. Strip silence with voice-activity detection before transcribing, and drop segments whose no_speech_prob is high or compression_ratio suggests repetition.

solid answer

~50 s

Whisper always decodes a full 30-second window, so a window containing only silence, music or room noise still produces tokens — and with nothing to condition on, the decoder emits whatever its language prior favours, often a repeated phrase or subtitle-style boilerplate absorbed from training data. Diagnose it from `verbose_json`: hallucinated segments typically show a high `no_speech_prob`, a strongly negative `avg_logprob`, and a `compression_ratio` above roughly 2.4 because the text is degenerately repetitive. The fixes stack. Run a voice-activity pass and remove or clip long silences before upload so the model never sees empty windows. Filter the returned segments on those three signals rather than trusting the raw text. Split chunks at silence boundaries so no chunk begins or ends in dead air. And keep the `prompt` short — a long prompt gives the decoder more prior text to echo when the audio says nothing. If you run the weights yourself, disabling conditioning on previous text also stops a loop from propagating forward.

code

python · 22 lines
python
from openai import OpenAI

client = OpenAI()

with open("lecture.mp3", "rb") as audio:
    result = client.audio.transcriptions.create(
        model="whisper-1",
        file=audio,
        response_format="verbose_json",
    )


def suspect(seg):
    return (
        seg.no_speech_prob > 0.6
        or seg.avg_logprob < -1.0
        or seg.compression_ratio > 2.4
    )


clean = [s.text for s in result.segments if not suspect(s)]
print(" ".join(clean))

go deeper

for a junior

Know that Whisper can invent text over silence or music rather than returning nothing, so a transcript should not be trusted line-by-line without some check.

for a middle

Explain the mechanism — fixed 30-second windows always decode, and with no acoustic evidence the language prior takes over — and name the verbose_json fields that expose it.

for a senior

Demonstrate the layered fix in production: voice-activity preprocessing, silence-aware chunking, threshold-based segment filtering, short prompts, and measurement on a labelled sample of your own worst audio.

for a principal

Own the risk framing. Decide where a fabricated sentence is tolerable and where it is not, require human review or confidence surfacing for high-stakes transcripts, and monitor the quality-signal distribution as an input-drift alarm.

## Why it happens at all Whisper's encoder consumes a fixed **30-second window** of audio, padded or trimmed to that length, converted to a log-Mel spectrogram. The decoder then generates text for that window autoregressively. Nothing in the architecture says "produce zero tokens when there is no speech" — the decoder is a language model that will happily continue. When the window contains speech, acoustic evidence dominates and the output tracks it. When the window contains silence, HVAC hum, applause or music, that evidence is absent or uninformative, so the language prior takes over. Two characteristic failure shapes follow: - **Repetition loops.** The decoder emits a phrase, conditions on it, and emits it again — sometimes dozens of times. - **Training-set boilerplate.** Whisper was trained on a large web corpus that included subtitle files, so silent stretches attract stock subtitle phrases: channel sign-offs, thanks-for-watching lines, transcription-service credits. Seeing text no human said is the tell. This is not a bug you configure away; it is a property of running a generative decoder over non-speech, and it is why every serious Whisper pipeline has a filtering stage. ## Diagnosing from the response Request `response_format="verbose_json"` and read the per-segment fields: - **`no_speech_prob`** — the model's own estimate that the window held no speech. Hallucinated segments frequently score high here while still returning confident-looking text; the two are not contradictory, because the no-speech head and the decoder are separate signals. - **`avg_logprob`** — mean token log-probability. Strongly negative values mark segments the model was unsure of. - **`compression_ratio`** — how compressible the emitted text is. Degenerate repetition compresses extremely well, so a high ratio is a direct repetition detector. The reference open-source implementation ships thresholds that are a reasonable starting point for your own filter: treat `compression_ratio` above about **2.4** as repetition, `avg_logprob` below about **-1.0** as low confidence, and `no_speech_prob` above about **0.6** as non-speech. Combine them rather than using one alone — a real segment can trip any single threshold. ## Preventing it upstream Filtering after the fact is a safety net; the stronger move is not to send silence at all. **Voice-activity detection.** Run a VAD over the audio, drop or heavily shorten stretches with no speech, and keep a mapping from the trimmed timeline back to the original so your timestamps stay meaningful. This removes the exact input that triggers the failure. It is the single highest-leverage change in most pipelines. **Chunk on silence.** If you split long audio for the 25 MB limit, cut inside quiet regions rather than on a clock — but trim the quiet rather than leaving a chunk that opens with twenty seconds of nothing. **Normalise levels.** Very quiet recordings push more windows toward the no-evidence regime. Gentle normalisation or gain before transcription helps marginally. **Keep the prompt short.** A long `prompt` hands the decoder more prior text to echo when acoustics are uninformative. If you see glossary terms appearing in quiet stretches, that is prompt bleed — shorten it, or omit it for chunks known to be quiet. ## Extra levers when you run the weights yourself Self-hosting exposes decoding controls the hosted API does not: - **Disable conditioning on previous text.** By default the implementation feeds the previous window's output forward as context; a loop that starts in one window then propagates. Turning that off (`condition_on_previous_text=False`) contains the damage at the cost of some cross-window coherence. - **Temperature fallback.** The reference decoder retries a window at increasing temperature when the compression-ratio or logprob thresholds are breached. Leaving this enabled lets a bad window get a second chance instead of shipping the loop. - **Threshold tuning.** The same three thresholds are exposed as parameters, so you can tighten them for your acoustic conditions rather than accepting defaults. ## Verification, not vibes Build a small labelled sample from your real audio — including its worst cases, the ones with music beds, long pauses and crosstalk — and measure hallucinated-segment rate before and after each change. Log the three quality fields for every segment in production and alert on their distribution shifting; a rise in high-`compression_ratio` segments is an early signal that an upstream capture change (a new microphone, a codec swap, a noisy venue) has degraded your inputs. And document the residual risk. Even a well-tuned pipeline occasionally emits a plausible sentence nobody said, which matters a great deal if the transcript feeds medical notes, legal records or automated actions. Where the stakes are high, surface confidence to reviewers rather than presenting the transcript as ground truth.

  • Why does a high no_speech_prob sometimes accompany confident-sounding hallucinated text?
    Because they come from different parts of the model. The no-speech estimate is its own signal about whether the window contained speech, while the text is produced by the decoder, which will generate fluent output regardless. So the model can simultaneously indicate "probably nothing here" and emit a well-formed sentence — which is exactly why filtering on that field catches hallucinations the text alone never reveals.
  • What does a compression_ratio above roughly 2.4 tell you about a segment?
    That its text is unusually compressible, which in practice means degenerate repetition — the same phrase or token sequence emitted over and over. The reference implementation uses that threshold to trigger a retry at higher temperature. In your own pipeline it is a cheap, reliable detector for loop-shaped hallucination that needs no model access, just the verbose response.
  • What extra control do you gain by running the weights yourself instead of the hosted API?
    Decoding parameters the hosted endpoint does not expose: disabling conditioning on previous text so a loop cannot propagate across windows, tuning the compression-ratio, logprob and no-speech thresholds for your acoustics, and controlling the temperature-fallback retry behaviour. You also decide the chunking algorithm rather than implementing it around a 25 MB upload cap.

saying these in an interview costs you the question

  • Assumes silent audio simply produces empty output
  • Treats every returned segment as something a human said
  • Ignores no_speech_prob and compression_ratio in the response
  • Raises temperature hoping it reduces repetition
  • Blames the audio codec rather than checking for non-speech windows

context