OpenAI's audio transcription endpoint caps uploads at 25 MB — how do you transcribe a two-hour recording?
answer
- The cap is on bytes, not minutes
- Sixteen kilohertz mono is enough
- Split where nobody is speaking
- Feed the previous chunk's tail forward
- Chunk timestamps start at zero
basics
~20 sRe-encode to mono 16 kHz compressed audio first, since Whisper resamples to 16 kHz anyway. If the file is still over the limit, split it at silence boundaries, transcribe each chunk, then add each chunk's start offset to its timestamps before merging.
solid answer
~50 sThe 25 MB ceiling is on **file size per request**, not on duration, so the first move is re-encoding rather than splitting. Whisper's front end resamples everything to 16 kHz mono, so a 48 kHz stereo WAV is paying upload bytes for data the model discards — converting to mono 16 kHz MP3/Opus at a modest bitrate often brings an hour of speech comfortably under the cap. When the recording is still too big, split it, but split on **silence** (or with a small overlap you de-duplicate) so no word is cut in half at a seam. Transcribe chunks independently, pass the tail of the previous chunk's text as the `prompt` so names and terminology stay consistent, and when merging `verbose_json` output add each chunk's start time to every segment's `start` and `end` — chunk timestamps are relative to that chunk, not the original file.
code
python · 23 linesfrom openai import OpenAI
from pydub import AudioSegment
client = OpenAI()
CHUNK_MS = 10 * 60 * 1000
audio = AudioSegment.from_file("lecture.m4a").set_channels(1).set_frame_rate(16000)
segments = []
context = ""
for start in range(0, len(audio), CHUNK_MS):
audio[start:start + CHUNK_MS].export("chunk.mp3", format="mp3", bitrate="48k")
with open("chunk.mp3", "rb") as f:
part = client.audio.transcriptions.create(
model="whisper-1",
file=f,
response_format="verbose_json",
prompt=context,
)
offset = start / 1000.0
for s in part.segments:
segments.append((s.start + offset, s.end + offset, s.text))
context = part.text[-400:]go deeper
Know that the transcription endpoint rejects files over 25 MB and that the limit is on file size, so compressing the audio is a legitimate first step before any splitting.
Be ready to explain the whole pipeline: re-encode to 16 kHz mono, split at silence, transcribe chunks independently, and offset each chunk's timestamps when merging. Say why 16 kHz loses nothing.
Show the production details — bounded concurrency against rate limits, retry on 429, seam de-duplication for overlapping chunks, and carrying a glossary or the previous transcript tail forward so terminology stays stable across a long recording.
Own the build-vs-call decision. Argue when a chunking pipeline against the hosted API is the right cost and ops profile versus running the open weights, where the upload cap disappears and long-form handling moves inside the runtime you now operate.
## What the limit actually is OpenAI's `/v1/audio/transcriptions` and `/v1/audio/translations` endpoints accept a multipart file upload with a maximum size of **25 MB per request**. It is a size limit, not a duration limit, and there is no server-side chunking: if you post a bigger file the request is rejected outright. Accepted container/codec formats include `flac`, `mp3`, `mp4`, `mpeg`, `mpga`, `m4a`, `ogg`, `wav` and `webm`. So the same two-hour meeting can be either far over the limit or comfortably under it, depending entirely on how it was encoded. ## Step one: re-encode, don't split Whisper's audio front end resamples every input to **16 kHz mono** and turns it into a log-Mel spectrogram before the encoder ever sees it. Anything above that — 48 kHz sampling, stereo channels, 24-bit depth, lossless WAV — is thrown away during preprocessing. Uploading it only burns bandwidth and pushes you over the cap. Practical recipe: downmix to one channel, resample to 16 kHz, and encode with a lossy codec at a speech-appropriate bitrate (roughly 32–64 kbps MP3, or Opus even lower). Two hours of speech at 32 kbps mono is on the order of 30 MB — close enough that a single split is often all you need, versus the ~1.3 GB the raw 48 kHz stereo WAV would be. Do **not** downsample below 16 kHz: 8 kHz telephone-band audio genuinely loses information the model expects, and accuracy drops. ## Step two: split where the audio is quiet If you still exceed the cap, chunk. The naive approach — cut every N minutes on the clock — reliably slices a word in half at each boundary, and Whisper mistranscribes both fragments. Two better approaches: - **Silence-aware splitting.** Detect low-energy regions (a VAD, `ffmpeg`'s `silencedetect`, or a library like `pydub`) and cut inside them, targeting chunks a bit under your byte budget. - **Fixed chunks with overlap.** Cut every N minutes but include a few seconds of overlap on each side, then de-duplicate the repeated text at the seam when merging. Cruder, but it needs no silence detection. ## Step three: carry context forward Each request is stateless — chunk 7 has no idea what chunk 6 heard. That shows up as inconsistent spelling of names, jargon and acronyms across the transcript. The `prompt` parameter accepts free text that is fed to the decoder as prior context, so passing the last few hundred characters of the previous chunk's transcript (or a fixed glossary of domain terms and proper nouns) keeps style and vocabulary stable. Only the tail of that text is used — Whisper's decoder reserves roughly 224 tokens for the prompt — so there is no point sending the whole transcript so far. ## Step four: fix up the timestamps This is the bug most people ship. When you request `response_format="verbose_json"`, every segment (and every word, if you asked for word granularity) carries `start` and `end` in seconds **relative to the file you uploaded**. Chunk 3's first segment starts near 0.0, not near 20:00. Merging without adjustment produces a transcript whose cues all pile up at the beginning and overlap each other. The fix is arithmetic: track each chunk's offset in the original recording and add it to every `start` and `end` before concatenating. If you used overlaps, drop the duplicated span from one side of each seam so the merged timeline is monotonic. ## Operational notes Chunks are independent requests, so they parallelise — but that multiplies your request rate against your account's limits, so bound the concurrency and retry `429`s with backoff. Billing for the hosted model is per minute of audio, so splitting does not itself cost more, though overlaps do add a little duplicated audio. And if you are only after a searchable transcript rather than subtitle timing, `response_format="text"` avoids the whole timestamp-merging problem — you simply concatenate strings. Finally, if this pipeline is a permanent part of your product, weigh it against running the open-weight model yourself, where there is no upload cap at all and the long-form chunking is handled inside the reference implementation instead of by your code.
- Why does downmixing to 16 kHz mono cost you essentially nothing in accuracy?Because Whisper's preprocessing resamples every input to 16 kHz mono and converts it to a log-Mel spectrogram before the encoder runs. Sample rates above 16 kHz and extra channels are discarded at that stage, so shipping them only inflates the upload. The one caveat is codec quality: compress with a bitrate high enough that artifacts don't smear consonants.
- What goes wrong if you split strictly every ten minutes with no overlap?Each boundary has a good chance of landing mid-word or mid-sentence. The truncated audio at the end of one chunk and the start of the next is transcribed as garbage or dropped, and the decoder also loses the sentence context that would have disambiguated nearby words. Cutting inside detected silence, or overlapping a few seconds and de-duplicating, avoids both.
- How do you keep proper nouns spelled consistently across chunks?Pass context in the `prompt` parameter: either the tail of the previous chunk's transcript, or a fixed glossary line listing the names, product terms and acronyms in the recording. It biases the decoder toward those spellings. Only about the last 224 tokens are used, so keep it short and put the highest-value terms last.
saying these in an interview costs you the question
- Thinks the 25 MB limit is a duration limit
- Uploads 48 kHz stereo WAV because higher fidelity means better accuracy
- Splits on a fixed clock interval and cuts words in half
- Forgets to offset chunk timestamps, producing overlapping cues
- Assumes the API chunks long files for you server-side