How do you add captions to an HTML `<video>` using the `<track>` element, and what distinguishes `kind="captions"` from `kind="subtitles"`?
answer
- a separate text file, not a re-encode
- void child of the media element
- five attributes, one of them boolean
- who can hear decides the kind
- dot before the milliseconds
basics
~20 sAdd a void track child to the video pointing at a WebVTT file, with kind, src, srclang, label and optionally default. Captions convey non-speech audio too and assume the viewer cannot hear; subtitles are a transcription or translation for someone who can hear the audio.
solid answer
~50 sTimed text goes in a `<track>` child of `<video>` or `<audio>`: `<track kind="captions" src="en.vtt" srclang="en" label="English" default>`. The file must be WebVTT — the format whose first line is literally `WEBVTT` — not SRT, which browsers do not parse. `srclang` gives the language of the cues, `label` is the human-readable name the browser puts in its captions menu, and `default` marks the one track enabled without user action; only one track should carry it. The `kind` distinction is about the audience, not the language: `captions` are for a viewer who cannot hear, so they include speaker identification and relevant sounds like "[door slams]", while `subtitles` assume audible dialogue and just render its text, typically translated. Other kinds exist — `descriptions`, `chapters`, `metadata`. A cross-origin `.vtt` file also needs CORS headers plus a `crossorigin` attribute on the media element.
code
html · 5 lines<video controls width="1280" height="720" poster="frame.jpg" crossorigin="anonymous">
<source src="talk.mp4" type="video/mp4">
<track kind="captions" src="talk-en.vtt" srclang="en" label="English" default>
<track kind="subtitles" src="talk-de.vtt" srclang="de" label="Deutsch">
</video>go deeper
Recall the element and its attributes — track with kind, src, srclang, label and default, placed inside the video element — and that the file is WebVTT rather than SRT.
Explain the kind distinction by audience: captions carry non-speech audio for a viewer who cannot hear, subtitles render dialogue for one who can. Mention that label drives the browser's track menu.
Show you have shipped this: the cross-origin CORS plus crossorigin requirement, the text/vtt content type, and why external tracks beat burned-in captions for correction, translation and user restyling.
Own captioning as a delivery obligation — where cue files are produced and reviewed, what quality bar counts as a caption rather than a transcript, and how accessibility requirements shape the media pipeline.
## The markup ```html <video controls width="1280" height="720" poster="frame.jpg"> <source src="talk.webm" type="video/webm"> <source src="talk.mp4" type="video/mp4"> <track kind="captions" src="talk-en.vtt" srclang="en" label="English" default> <track kind="subtitles" src="talk-de.vtt" srclang="de" label="Deutsch"> </video> ``` `<track>` is a void element and must appear after any `<source>` children. Each one references a separate timed-text file; the media file itself is untouched, which is why captions can be added, corrected and translated without re-encoding anything. ## The attributes - **`src`** — required, the URL of the timed-text file. - **`kind`** — `subtitles` (the default when omitted), `captions`, `descriptions`, `chapters` or `metadata`. - **`srclang`** — the BCP 47 language tag of the cues (`en`, `de`, `pt-BR`). Required when `kind="subtitles"`, and worth setting always: it is what tells assistive technology and the browser menu what language the text is in. - **`label`** — the string shown in the browser's track menu. Omit it and the user picks between unlabelled entries, which is a real usability defect on a video with three translations. - **`default`** — boolean; enables this track without user action. Only one track should carry it, or the browser has to arbitrate between conflicting defaults. ## captions versus subtitles This is the part interviews actually probe, and the wrong answer — "captions are the same language, subtitles are translated" — is close enough to sound right. The real distinction is what the viewer can hear: - **`captions`** assume the audio is unavailable to the viewer. They therefore carry everything meaningful on the audio channel: dialogue, speaker identification when it is not obvious on screen, and non-speech sound that matters — `[phone rings]`, `[laughter]`, `[ominous music]`. They exist for deaf and hard-of-hearing viewers, and equally for anyone watching with the sound off in public. - **`subtitles`** assume the audio is audible but not understood. They render the dialogue as text, usually translated, and deliberately leave out sound effects because the viewer can hear them. So a same-language transcript that omits the sound of breaking glass is a subtitle track, not a caption track, and labelling it `kind="captions"` overstates what it provides. The other kinds: `descriptions` are textual descriptions of the *visual* content, meant for audio rendering to a viewer who cannot see the screen; `chapters` provide navigable segment titles; `metadata` cues are never rendered and exist to be read by script. ## WebVTT, not SRT The file format is WebVTT, and it must begin with the token `WEBVTT`: ``` WEBVTT 00:00:01.000 --> 00:00:04.000 [door slams] 00:00:04.500 --> 00:00:07.200 <v Ana>We need to talk about the deploy. ``` Cue times use a dot before the milliseconds — `00:00:04.500`, not the comma SRT uses. A renamed `.srt` will not parse. Serve the file as `text/vtt`; a wrong content type is a common reason a track silently never appears. ## The cross-origin trap Text tracks are subject to CORS. If the `.vtt` lives on a different origin from the page — a CDN or an asset domain — the response needs the appropriate CORS headers *and* the media element needs a `crossorigin` attribute: ```html <video controls crossorigin="anonymous"> <source src="https://cdn.example.com/talk.mp4" type="video/mp4"> <track kind="captions" src="https://cdn.example.com/talk-en.vtt" srclang="en" label="English" default> </video> ``` This is the number one "my captions work locally and not in production" cause, because local development usually serves everything from one origin. ## Why this belongs in markup rather than burned into the video Open captions burned into the pixels cannot be turned off, cannot be restyled for legibility, cannot be searched or translated, and force a re-encode for every correction. A `<track>` is a small text file: swap it and the fix ships. Browsers also give the user control over caption appearance through their own settings, which a burned-in track defeats. ## What the browser does with it When `controls` is present, the browser renders a captions menu listing each track by `label`, and a `default` track renders immediately. Text-track rendering is a browser behaviour rather than something the control bar owns — so a `default` track still displays in a player with custom controls, while the user's ability to switch or disable it disappears with the native menu.
- A same-language transcript of the dialogue is supplied with kind="captions". What is wrong with that?If it contains only spoken words, it is a subtitle track mislabelled as captions. Captions promise everything meaningful on the audio channel — speaker identification and relevant non-speech sound such as `[alarm blaring]`. Declaring `captions` tells a deaf viewer they are getting that completeness, so either enrich the cues or set `kind="subtitles"` honestly.
- Captions work on localhost but never appear once the .vtt is served from a CDN. Why?Text tracks are fetched under CORS. A cross-origin `.vtt` needs the appropriate CORS response headers and the media element needs a `crossorigin` attribute, typically `crossorigin="anonymous"`. Locally everything is same-origin, so the requirement is invisible. Check the content type too: the file should be served as `text/vtt`.
- Why prefer a track element over captions burned into the video frames?A `<track>` is an editable text file: corrections and new languages ship without re-encoding, the user can turn captions off or restyle them through browser settings, and the cues are machine-readable for search and translation. Burned-in captions are permanent pixels — always on, unstylable, and re-encoded for every typo.
saying these in an interview costs you the question
- Saying captions are same-language and subtitles are translations
- Pointing a track at an .srt file
- Omitting srclang or label on translated tracks
- Marking several tracks as default
- Believing captions require re-encoding the video