skip to content

What does the prompt parameter do on an OpenAI Whisper transcription request?

level: middleimportance: should knowfreq 40%

answer

  1. Prior context, not an instruction
  2. Only the tail is used
  3. Roughly two hundred tokens survive
  4. Match the prompt to output language
  5. Bias, never a guarantee

basics

~20 s

It supplies prior text as decoder context, biasing spellings, jargon and punctuation style toward what the prompt contains. It is not an instruction field — Whisper does not follow commands in it — and only roughly the last 224 tokens are used.

solid answer

~50 s

`prompt` is a **context prefix fed to the decoder**, not a system instruction. Whisper generates the transcript conditioned on that text, so it nudges the model toward the spellings, proper nouns, acronyms and punctuation style the prompt exhibits — the standard uses are a glossary of domain terms and product names, or the tail of the previous chunk's transcript when you are stitching a long recording. The limits matter. Only about the last 224 tokens are considered, because Whisper's decoder context is 448 tokens with roughly half reserved for the prompt, so a long glossary silently loses its head; put your highest-value terms last. It must be written in the language the endpoint will output. And because it is context rather than an order, instructions like "remove filler words" or "format as bullet points" work at best erratically — the model may transcribe the instruction itself. Treat it as biasing, not control.

code

python · 19 lines
python
from openai import OpenAI

client = OpenAI()

glossary = (
    "Attendees: Priya Raghavan, Tomasz Wiech. "
    "Terms used: SLO, p99 latency, Liquibase, Modulith."
)

with open("standup.mp3", "rb") as audio:
    result = client.audio.transcriptions.create(
        model="whisper-1",
        file=audio,
        language="en",
        prompt=glossary,
        response_format="text",
    )

print(result)

go deeper

for a junior

Know that the prompt biases spelling and style by acting as prior text, and that it is not a place to put instructions about what the transcript should look like.

for a middle

Explain the mechanism — decoder context, roughly 224 usable tokens taken from the tail — and give the two real uses: a curated glossary, and carrying the previous chunk's text forward in a long recording.

for a senior

Show how you validate it: A/B a prompt against a held-out sample, watch for prompt bleed over silence, and back critical spellings with a deterministic post-processing pass rather than trusting a probability nudge.

for a principal

Own the boundary between transcription and downstream transformation — keep the ASR stage faithful and put cleanup, redaction and formatting in an explicit, testable text stage rather than hoping a context prefix enforces policy.

## What it actually is Whisper is an encoder-decoder model. The encoder consumes audio; the decoder emits text tokens autoregressively, conditioned on special task tokens and on whatever prior text it has been given. The `prompt` parameter on the transcription and translation endpoints injects text into that prior-context slot. The model then decodes as though the prompt were the immediately preceding transcript. That single fact explains everything the parameter does well and everything it does badly. ## What it is good at **Proper nouns and domain vocabulary.** Whisper hears an ambiguous acoustic sequence and picks the most probable spelling. If your product is called "KataJob", it will happily write "Kata Job", "Kotajob" or "Cata Job". A prompt containing the correct spelling raises that token sequence's probability, and the transcript stabilises. Same for drug names, ticker symbols, internal acronyms and unusual surnames. **Punctuation and casing style.** A prompt written with full sentence punctuation and capitalisation nudges the output the same way. A prompt written as lowercase fragments without periods tends to produce a less punctuated transcript. This is style transfer by imitation, not a formatting flag. **Continuity across chunks.** When a long recording is split into several requests, each request is stateless. Passing the tail of the previous chunk's transcript as the next chunk's prompt restores the missing context, so terminology and style stay consistent and the decoder is less likely to restart mid-sentence with a mismatched register. ## The 224-token ceiling Whisper's decoder works with a 448-token context, of which roughly half — about 224 tokens — is available for prompt text. The API does not error on a longer prompt; it simply uses the tail. The practical consequences: - A 500-term glossary is mostly wasted. Curate to the terms actually likely in this audio. - Ordering matters. Whatever you most need respected should sit at the **end** of the prompt. - When carrying context between chunks, a few hundred characters of the previous transcript is the right order of magnitude; sending the whole transcript so far is pointless. ## What it is bad at **Instructions.** This is the biggest misconception in interviews. `prompt` is not a system message. "Only transcribe the customer, ignore the agent", "strip ums and ahs", "output JSON" — none of these are reliably honoured, because the model is not instruction-tuned on this field; it is being handed prior text. Sometimes the instruction leaks into the output as transcribed text. If you need transformed output, transcribe faithfully and then post-process with a text model. **Guarantees.** Even for spellings, the prompt is a probability nudge. Loud audio, a strong accent, or a genuinely different word will override it. If a term absolutely must be spelled a certain way, do a deterministic find-and-replace pass on the transcript afterwards. **Hallucination pressure.** A prompt that is long, repetitive or unrelated to the audio can make things worse: the decoder may echo prompt content into the transcript, especially over silence or non-speech. If you see phrases from your glossary appearing in quiet stretches, shorten the prompt or drop it for those chunks. ## Language matching Because the prompt is prior decoder text, it must be in the language the endpoint is going to emit. For transcription, that is the spoken language: a Spanish glossary for Spanish audio. For the translation endpoint, output is English, so the prompt should be English. A mismatch can actively confuse the model, and at the extreme can push a transcription toward the prompt's language. ## Interaction with other parameters Use the `language` parameter for language selection rather than trying to imply it through prompt text — that is what it exists for, and it also skips a detection pass. Keep `temperature` at its default for deterministic-leaning output; raising it while relying on a prompt makes the biasing less predictable. And if you request `verbose_json`, watch the per-segment `avg_logprob` and `compression_ratio` fields when you tune prompts: a prompt that helps should not be pushing segments toward degenerate repetition. ## A practical recipe Build a short prompt from the entities you actually expect — speaker names, product names, the five acronyms this team uses — written as a natural sentence or comma-separated list in the audio's language, kept well under the token ceiling, with the highest-risk terms last. For chunked audio, append the previous chunk's trailing text. Measure with and without it on a held-out sample rather than assuming it helped; on clean audio with common vocabulary, the gain is often negligible and the risk of prompt bleed is not.

  • Why does putting your most important terms at the end of the prompt matter?
    Because only about the last 224 tokens of the prompt reach the decoder — Whisper's 448-token context reserves roughly half for prompt text, and an over-long prompt is silently truncated from the front rather than rejected. Terms sitting in the discarded head have no effect at all, so ordering is a correctness concern, not a style preference.
  • A team puts "remove filler words and format as bullet points" in the prompt. What happens?
    Unreliable behaviour at best, and sometimes the instruction text itself appears in the transcript. The field is prior decoder context, not an instruction channel, and Whisper is not instruction-tuned on it. The correct design is to transcribe faithfully, then run the transcript through a text model that does the cleanup and formatting deterministically.
  • You notice glossary terms appearing during silent stretches of audio. What is going on?
    Prompt bleed. Over non-speech the decoder has little acoustic evidence to condition on, so it leans on prior context and can echo the prompt into the output. Shorten the prompt, drop it for chunks known to be quiet, strip silence with a voice-activity pass first, and filter suspect segments using no_speech_prob from the verbose response.

saying these in an interview costs you the question

  • Treats the prompt as a system instruction Whisper obeys
  • Sends a thousand-term glossary and expects all of it to apply
  • Writes an English prompt for non-English transcription
  • Assumes the prompt guarantees a spelling
  • Blames the model when the prompt text appears in the transcript

context