In OpenAI's audio API, how do /v1/audio/transcriptions and /v1/audio/translations differ?
answer
- Same audio, two different tasks
- One endpoint has no target choice
- English is the only translation output
- The language field hints the source
- Two-step route for other languages
basics
~20 sTranscription writes down what was said in the language it was spoken in. Translation always outputs English, whatever the source language — there is no target-language parameter, because the model was only trained to translate into English.
solid answer
~40 sBoth endpoints take the same audio upload, but they invoke different tasks. `/v1/audio/transcriptions` returns text in the **spoken** language; you can pass a `language` hint in ISO-639-1 form to skip language detection, which usually improves both accuracy and latency. `/v1/audio/translations` performs speech-to-English translation and has **no target-language parameter at all** — Whisper's training covered X-to-English translation only, so English is the sole possible output. That is the single most-missed fact about the endpoint. If you need Spanish audio rendered as German text, the supported route is two steps: transcribe to Spanish, then translate that text with a text model. The translation endpoint is also the narrower surface generally, so if you need rich metadata or fine-grained timings, prefer transcription plus a separate translation step.
go deeper
Be able to say plainly that transcription keeps the spoken language while translation always returns English, and that there is no way to ask the translation endpoint for a different target.
Explain the language hint's real purpose on the transcription side, the two-step route for non-English targets, and why the prompt must be written in the language the endpoint will output.
Argue for transcribe-then-translate in production: it preserves an auditable source transcript, gives terminology control, and lets one audio pass feed many target languages instead of re-billing audio minutes per language.
Frame it as a pipeline-shape decision — where translation quality is owned, how localisation review fits in, and whether audio-to-English one-shot is ever acceptable for content that carries legal or brand risk.
## Two tasks, one model Whisper was trained multitask: from the same audio encoder it can either **transcribe** (write the speech down in its own language) or **translate** (render the speech as English text). OpenAI exposes these as two endpoints that share a request shape: - `POST /v1/audio/transcriptions` — output is in the source language. - `POST /v1/audio/translations` — output is in English. Same multipart upload, same 25 MB file ceiling, same accepted audio formats, same `prompt` and `temperature` knobs, and overlapping `response_format` choices. ## The asymmetry that trips people up The translation endpoint takes no target language. There is no `target_language` field to set, and passing one does not make it emit French. English is the only output the model can produce for this task, because the multitask training data contained X→English translation pairs and not the reverse or cross-pairs. People search the parameter list for a target-language option, fail to find it, and conclude they have the wrong endpoint — the truth is simpler: the capability does not exist. The `language` parameter, meanwhile, belongs on the **transcription** side and means something different from what its name suggests to newcomers. It is a hint about the language **being spoken**, not a request for the output language. Supplying it (`"es"`, `"ja"`, `"de"`) lets the model skip its language-detection pass, which shaves latency and prevents the classic failure where a few seconds of accented or noisy opening audio cause the whole file to be detected as the wrong language and transcribed as gibberish. ## How to get a non-English target Two steps, no exceptions: 1. Transcribe the audio, ideally with the `language` hint set to the spoken language. 2. Feed that transcript to a text model with a translation instruction, or to a dedicated machine-translation service. This is usually better than the one-shot endpoint even when English *is* your target, for three reasons: you keep the original-language transcript as an artifact, you can review and correct it before translating, and you can translate into several languages from the one transcription pass rather than paying for the audio again per language. ## Prompting differences The `prompt` parameter is decoder context, so its language must match the **output** you expect. On the transcription endpoint, write the prompt in the spoken language — a Spanish glossary for Spanish audio. On the translation endpoint, the output is English, so the prompt should be English. Getting this backwards weakens or confuses the biasing effect rather than helping. ## Response shapes Both endpoints accept the familiar `response_format` values, so you can get a plain string, JSON, or subtitle output. But the transcription endpoint is the richer surface: the `timestamp_granularities` machinery for word-level timing is a transcription-side feature, and it is where the verbose response with per-segment quality fields is most useful. Practically, if your pipeline wants timings, quality filtering, or the detected-language metadata, build it on transcription and treat translation as a separate concern. ## Quality expectations Whisper's translation quality is respectable for gist but is not a substitute for a dedicated translation system on high-value content. It was trained as a speech task, and the English it produces from low-resource languages is noticeably weaker than from high-resource ones. For subtitles a viewer will actually read, the transcribe-then-translate route with a strong text model generally wins on fluency and terminology control — and it gives you a source-language transcript to audit when someone disputes a rendering. ## Billing and limits Both endpoints bill by minutes of audio processed for the hosted Whisper model, and both are bound by the same 25 MB per-request upload cap, so the chunking strategy you build for one applies unchanged to the other. Choosing translation over transcription does not save you anything on either axis — it only changes what language comes back.
- What does the language parameter on the transcription endpoint actually control?It tells the model which language is being **spoken**, in ISO-639-1 form, so it can skip automatic language detection. That trims latency and, more importantly, avoids the failure where noisy or accented opening audio is misdetected and the whole file gets transcribed as the wrong language. It never changes the output language — transcription always writes down the language that was spoken.
- You need Japanese audio rendered as French subtitles. What is the supported approach?Transcribe the audio with the language hint set to Japanese, then translate the resulting text into French with a text model or a machine-translation service. The translation endpoint cannot help: it only ever emits English. The two-step route also lets you keep the Japanese transcript, review it, and fan out to additional target languages without re-processing the audio.
- Which language should the prompt parameter be written in for each endpoint?Match the expected output. For transcription, write the prompt in the spoken language, since the decoder is producing that language and the prompt acts as prior context. For translation, write it in English, because English is what the decoder will emit. Mismatching them dilutes the biasing effect on spellings and terminology instead of strengthening it.
saying these in an interview costs you the question
- Looks for a target-language parameter on the translation endpoint
- Thinks the language parameter selects the output language
- Believes translation can produce any language pair
- Assumes translation output quality matches a dedicated MT system
- Writes a source-language prompt for the translation endpoint