skip to content

Whisper

Speech-to-text you can either call as an API or run locally, since the weights are open. Most real questions land on file-size limits, timestamp formats, and how you handle audio longer than one request allows.

on this pageshow

questions

6

OpenAI's audio transcription endpoint caps uploads at 25 MB — how do you transcribe a two-hour recording?

level: middleimportance: must knowfreq 72%

answer

  1. The cap is on bytes, not minutes
  2. Sixteen kilohertz mono is enough
  3. Split where nobody is speaking
  4. Feed the previous chunk's tail forward
  5. Chunk timestamps start at zero

basics

~20 s

Re-encode to mono 16 kHz compressed audio first, since Whisper resamples to 16 kHz anyway. If the file is still over the limit, split it at silence boundaries, transcribe each chunk, then add each chunk's start offset to its timestamps before merging.

solid answer

~50 s

The 25 MB ceiling is on **file size per request**, not on duration, so the first move is re-encoding rather than splitting. Whisper's front end resamples everything to 16 kHz mono, so a 48 kHz stereo WAV is paying upload bytes for data the model discards — converting to mono 16 kHz MP3/Opus at a modest bitrate often brings an hour of speech comfortably under the cap. When the recording is still too big, split it, but split on **silence** (or with a small overlap you de-duplicate) so no word is cut in half at a seam. Transcribe chunks independently, pass the tail of the previous chunk's text as the `prompt` so names and terminology stay consistent, and when merging `verbose_json` output add each chunk's start time to every segment's `start` and `end` — chunk timestamps are relative to that chunk, not the original file.

code

python · 23 lines
python
from openai import OpenAI
from pydub import AudioSegment

client = OpenAI()
CHUNK_MS = 10 * 60 * 1000

audio = AudioSegment.from_file("lecture.m4a").set_channels(1).set_frame_rate(16000)
segments = []
context = ""

for start in range(0, len(audio), CHUNK_MS):
    audio[start:start + CHUNK_MS].export("chunk.mp3", format="mp3", bitrate="48k")
    with open("chunk.mp3", "rb") as f:
        part = client.audio.transcriptions.create(
            model="whisper-1",
            file=f,
            response_format="verbose_json",
            prompt=context,
        )
    offset = start / 1000.0
    for s in part.segments:
        segments.append((s.start + offset, s.end + offset, s.text))
    context = part.text[-400:]

go deeper

for a junior

Know that the transcription endpoint rejects files over 25 MB and that the limit is on file size, so compressing the audio is a legitimate first step before any splitting.

for a middle

Be ready to explain the whole pipeline: re-encode to 16 kHz mono, split at silence, transcribe chunks independently, and offset each chunk's timestamps when merging. Say why 16 kHz loses nothing.

for a senior

Show the production details — bounded concurrency against rate limits, retry on 429, seam de-duplication for overlapping chunks, and carrying a glossary or the previous transcript tail forward so terminology stays stable across a long recording.

for a principal

Own the build-vs-call decision. Argue when a chunking pipeline against the hosted API is the right cost and ops profile versus running the open weights, where the upload cap disappears and long-form handling moves inside the runtime you now operate.

## What the limit actually is OpenAI's `/v1/audio/transcriptions` and `/v1/audio/translations` endpoints accept a multipart file upload with a maximum size of **25 MB per request**. It is a size limit, not a duration limit, and there is no server-side chunking: if you post a bigger file the request is rejected outright. Accepted container/codec formats include `flac`, `mp3`, `mp4`, `mpeg`, `mpga`, `m4a`, `ogg`, `wav` and `webm`. So the same two-hour meeting can be either far over the limit or comfortably under it, depending entirely on how it was encoded. ## Step one: re-encode, don't split Whisper's audio front end resamples every input to **16 kHz mono** and turns it into a log-Mel spectrogram before the encoder ever sees it. Anything above that — 48 kHz sampling, stereo channels, 24-bit depth, lossless WAV — is thrown away during preprocessing. Uploading it only burns bandwidth and pushes you over the cap. Practical recipe: downmix to one channel, resample to 16 kHz, and encode with a lossy codec at a speech-appropriate bitrate (roughly 32–64 kbps MP3, or Opus even lower). Two hours of speech at 32 kbps mono is on the order of 30 MB — close enough that a single split is often all you need, versus the ~1.3 GB the raw 48 kHz stereo WAV would be. Do **not** downsample below 16 kHz: 8 kHz telephone-band audio genuinely loses information the model expects, and accuracy drops. ## Step two: split where the audio is quiet If you still exceed the cap, chunk. The naive approach — cut every N minutes on the clock — reliably slices a word in half at each boundary, and Whisper mistranscribes both fragments. Two better approaches: - **Silence-aware splitting.** Detect low-energy regions (a VAD, `ffmpeg`'s `silencedetect`, or a library like `pydub`) and cut inside them, targeting chunks a bit under your byte budget. - **Fixed chunks with overlap.** Cut every N minutes but include a few seconds of overlap on each side, then de-duplicate the repeated text at the seam when merging. Cruder, but it needs no silence detection. ## Step three: carry context forward Each request is stateless — chunk 7 has no idea what chunk 6 heard. That shows up as inconsistent spelling of names, jargon and acronyms across the transcript. The `prompt` parameter accepts free text that is fed to the decoder as prior context, so passing the last few hundred characters of the previous chunk's transcript (or a fixed glossary of domain terms and proper nouns) keeps style and vocabulary stable. Only the tail of that text is used — Whisper's decoder reserves roughly 224 tokens for the prompt — so there is no point sending the whole transcript so far. ## Step four: fix up the timestamps This is the bug most people ship. When you request `response_format="verbose_json"`, every segment (and every word, if you asked for word granularity) carries `start` and `end` in seconds **relative to the file you uploaded**. Chunk 3's first segment starts near 0.0, not near 20:00. Merging without adjustment produces a transcript whose cues all pile up at the beginning and overlap each other. The fix is arithmetic: track each chunk's offset in the original recording and add it to every `start` and `end` before concatenating. If you used overlaps, drop the duplicated span from one side of each seam so the merged timeline is monotonic. ## Operational notes Chunks are independent requests, so they parallelise — but that multiplies your request rate against your account's limits, so bound the concurrency and retry `429`s with backoff. Billing for the hosted model is per minute of audio, so splitting does not itself cost more, though overlaps do add a little duplicated audio. And if you are only after a searchable transcript rather than subtitle timing, `response_format="text"` avoids the whole timestamp-merging problem — you simply concatenate strings. Finally, if this pipeline is a permanent part of your product, weigh it against running the open-weight model yourself, where there is no upload cap at all and the long-form chunking is handled inside the reference implementation instead of by your code.

  • Why does downmixing to 16 kHz mono cost you essentially nothing in accuracy?
    Because Whisper's preprocessing resamples every input to 16 kHz mono and converts it to a log-Mel spectrogram before the encoder runs. Sample rates above 16 kHz and extra channels are discarded at that stage, so shipping them only inflates the upload. The one caveat is codec quality: compress with a bitrate high enough that artifacts don't smear consonants.
  • What goes wrong if you split strictly every ten minutes with no overlap?
    Each boundary has a good chance of landing mid-word or mid-sentence. The truncated audio at the end of one chunk and the start of the next is transcribed as garbage or dropped, and the decoder also loses the sentence context that would have disambiguated nearby words. Cutting inside detected silence, or overlapping a few seconds and de-duplicating, avoids both.
  • How do you keep proper nouns spelled consistently across chunks?
    Pass context in the `prompt` parameter: either the tail of the previous chunk's transcript, or a fixed glossary line listing the names, product terms and acronyms in the recording. It biases the decoder toward those spellings. Only about the last 224 tokens are used, so keep it short and put the highest-value terms last.

saying these in an interview costs you the question

  • Thinks the 25 MB limit is a duration limit
  • Uploads 48 kHz stereo WAV because higher fidelity means better accuracy
  • Splits on a fixed clock interval and cuts words in half
  • Forgets to offset chunk timestamps, producing overlapping cues
  • Assumes the API chunks long files for you server-side

context

open as a page

In OpenAI's audio API, how do /v1/audio/transcriptions and /v1/audio/translations differ?

level: juniorimportance: should knowfreq 48%

basics

~20 s

Transcription writes down what was said in the language it was spoken in. Translation always outputs English, whatever the source language — there is no target-language parameter, because the model was only trained to translate into English.

open as a page

What does the prompt parameter do on an OpenAI Whisper transcription request?

level: middleimportance: should knowfreq 40%

basics

~20 s

It supplies prior text as decoder context, biasing spellings, jargon and punctuation style toward what the prompt contains. It is not an instruction field — Whisper does not follow commands in it — and only roughly the last 224 tokens are used.

open as a page

How do you get word-level timestamps from OpenAI's Whisper transcription API?

level: middleimportance: should knowfreq 56%

basics

~20 s

Send timestamp_granularities: ["word"] on the transcription request and set response_format to verbose_json; any other response format rejects the granularity option. The response then carries a words array with a start and end time per word, alongside the usual segment list.

open as a page

Whisper transcripts contain repeated phrases over silent audio — how do you diagnose and reduce it?

level: seniorimportance: should knowfreq 44%

basics

~20 s

This is Whisper's known hallucination on non-speech: with no acoustic evidence the decoder falls back on language priors and loops or emits training-set boilerplate. Strip silence with voice-activity detection before transcribing, and drop segments whose no_speech_prob is high or compression_ratio suggests repetition.

open as a page

When does running Whisper's open weights yourself beat calling the hosted transcription API?

level: principalimportance: should knowfreq 34%

basics

~20 s

Self-host when data residency forbids sending audio out, when steady high volume makes fixed GPU cost cheaper than per-minute billing, or when you need decoding controls and model sizes the hosted endpoint does not expose. Otherwise the hosted API is far less to operate.

open as a page