skip to content

How do you get word-level timestamps from OpenAI's Whisper transcription API?

level: middleimportance: should knowfreq 56%

answer

  1. Two parameters must agree
  2. Granularity needs the verbose format
  3. Subtitle formats give cues, not words
  4. Alignment pass costs extra latency
  5. Words array has no speaker field

basics

~20 s

Send timestamp_granularities: ["word"] on the transcription request and set response_format to verbose_json; any other response format rejects the granularity option. The response then carries a words array with a start and end time per word, alongside the usual segment list.

solid answer

~40 s

Word timing is opt-in and format-coupled. On `/v1/audio/transcriptions` you pass `timestamp_granularities[]` with `"word"`, `"segment"`, or both, and you must also set `response_format="verbose_json"` — the plain `json` and `text` formats have no field to put timings in, and `srt`/`vtt` only ever emit cue-level timings, not per-word ones. Segment timestamps come with the verbose response by default and cost nothing extra; word timestamps require an extra forced-alignment pass over the model's cross-attention, so they add measurable latency and you should only ask for them when you need karaoke-style highlighting, precise clipping, or word-accurate search. The verbose response also carries per-segment quality signals — `avg_logprob`, `no_speech_prob`, `compression_ratio` — which are useful for filtering suspect output. This granularity path is a `whisper-1` feature; the newer transcription models on the same endpoint expose a narrower set of response formats.

code

python · 18 lines
python
from openai import OpenAI

client = OpenAI()

with open("interview.mp3", "rb") as audio:
    result = client.audio.transcriptions.create(
        model="whisper-1",
        file=audio,
        response_format="verbose_json",
        timestamp_granularities=["word", "segment"],
    )

for word in result.words:
    print(f"{word.start:6.2f} - {word.end:6.2f}  {word.word}")

for seg in result.segments:
    if seg.no_speech_prob > 0.6:
        print("suspect segment:", seg.text)

go deeper

for a junior

Remember the pairing: to get timestamps at all you need the verbose JSON response format, and word-level timing is an explicit opt-in rather than something you get by default.

for a middle

Explain the parameter coupling and what each response format returns, including that srt and vtt carry only cue-level timings. Name a few fields in the verbose response and say why word granularity costs latency.

for a senior

Show judgment about when the extra alignment pass is worth it, how much boundary accuracy to trust in noisy or overlapping speech, and how you combine these timings with a separate diarization pass since Whisper labels no speakers.

for a principal

Decide where timing precision actually earns its cost across a product: which surfaces need word-accurate alignment, which are served by cue-level subtitles, and whether speaker attribution justifies operating a second diarization model alongside transcription.

## The three knobs Timing detail on OpenAI's transcription endpoint is controlled by two parameters that must agree: - **`response_format`** — one of `json` (default), `text`, `srt`, `verbose_json`, `vtt`. - **`timestamp_granularities[]`** — an array containing `"segment"`, `"word"`, or both. The rule to memorise: **`timestamp_granularities` is only valid when `response_format` is `verbose_json`.** Ask for word timestamps with `response_format="json"` and the request is rejected, because the compact JSON shape is nothing but `{"text": "..."}` — there is no place for timings to live. ## What each format gives you - `text` — a bare transcript string. No timing at all. Cheapest thing to parse when you only want searchable text. - `json` — the same transcript wrapped in a one-field object. - `srt` / `vtt` — ready-made subtitle files. These *do* have timings, but they are **cue-level**: each cue is a phrase or line with a start and end. You cannot recover per-word timing from them, and you cannot request word granularity alongside them. - `verbose_json` — the rich shape: top-level `task`, `language`, `duration` and `text`, plus a `segments` array and, when requested, a `words` array. ## Reading the verbose response Each entry in `segments` describes a chunk of speech the decoder emitted together and carries, alongside `start`, `end` and `text`, several diagnostic fields: - `avg_logprob` — mean token log-probability for the segment; strongly negative means the model was unsure. - `no_speech_prob` — the model's estimate that this window contained no speech at all. - `compression_ratio` — the text's compressibility; a high value flags degenerate repetition. - `temperature`, `tokens`, `seek`, `id` — decoding metadata. Each entry in `words` is much simpler: the `word` itself plus `start` and `end` in seconds. Requesting both granularities gives you both arrays in one response, which is the usual choice — you use segments for display and quality filtering, and words for fine-grained alignment. All times are in seconds relative to **the file you uploaded**. If you chunked a long recording, you must re-base them against each chunk's offset before merging. ## Why word timing costs latency Whisper is an encoder-decoder transformer that emits text tokens; it does not natively output a time for each word. Segment boundaries fall out of the decoding loop almost for free, which is why segment timestamps add no meaningful latency. Word timestamps are produced by a separate alignment step that inspects the decoder's cross-attention over the audio frames and runs dynamic time warping to snap each token to an audio position. That is extra compute after decoding finishes, and OpenAI documents the added latency explicitly. Treat word granularity as a deliberate request, not a default. The alignment is also approximate. Expect boundaries accurate to roughly tens of milliseconds in clean speech, and noticeably worse across overlapping speakers, music, or heavy disfluency. If you are cutting video on these boundaries, pad them. ## What it does not give you Two common misreadings are worth naming. First, **word timestamps are not speaker labels** — Whisper does not do diarization, and nothing in the response tells you who spoke. Speaker attribution needs a separate diarization stage that you align against these timings yourself. Second, **word timing is not a confidence score**; the per-word entries carry no probability. Confidence-shaped signals live at the segment level in `avg_logprob` and `no_speech_prob`. ## Choosing a format in practice Subtitles for a video player: request `srt` or `vtt` and ship the file directly — no post-processing, no timestamp maths. A transcript UI that highlights the current word during playback, a search index that jumps to the exact moment a term was said, or a clip-extraction pipeline: `verbose_json` with word granularity. A bulk archive you only ever full-text search: `text`, and skip the whole question. One version note: this granularity machinery is tied to the Whisper model exposed as `whisper-1`. The same endpoint also serves newer, non-Whisper transcription models whose supported response formats are narrower, so a pipeline that depends on `verbose_json` word arrays is depending on the Whisper path specifically — pin the model rather than assuming any transcription model will answer the same shape.

  • Why do segment timestamps add no latency while word timestamps do?
    Segment boundaries emerge naturally from the decoding loop, so they are already available when generation ends. Word timestamps require an extra pass that reads the decoder's cross-attention over audio frames and warps each token onto an audio position. That alignment runs after decoding, so it is genuine additional compute — and OpenAI documents the added latency for word granularity specifically.
  • Can you use the word timestamps to label who is speaking?
    No. Whisper performs no speaker diarization, and the `words` entries carry only the word text plus start and end times. To attribute speech you run a separate diarization model over the same audio to get speaker-labelled time ranges, then intersect those ranges with Whisper's word timings. Any mismatch at boundaries is yours to reconcile.
  • If you just want a subtitle file, is verbose_json the right choice?
    Usually not. `response_format="srt"` or `"vtt"` returns a finished subtitle file with cue-level timings, so you write it to disk and you are done. Reach for `verbose_json` when you need per-word timing, the language and duration metadata, or the per-segment quality fields for filtering — not when a player just needs cues.

saying these in an interview costs you the question

  • Requests word granularity with response_format json
  • Thinks srt output can be parsed back into word-level timings
  • Believes word timestamps identify the speaker
  • Assumes word timing is free like segment timing
  • Treats the returned times as absolute across a chunked recording

context