skip to content

questions

6

In a video-on-demand platform, why is each video cut into short segments listed in a manifest rather than served as one file?

level: juniorimportance: must knowfreq 62%

answer

  1. a menu, not one file
  2. a few seconds per piece
  3. same cut points in every encoding
  4. quality decided for the next piece
  5. small immutable objects over plain HTTP

basics

~20 s

Short segments let a player start after fetching a few seconds, switch bitrate at each segment boundary as bandwidth changes, and seek or retry cheaply, while HTTP servers and caches handle small uniform objects instead of one multi-gigabyte file.

solid answer

~50 s

Segmented delivery (HLS or MPEG-DASH) encodes the video at several bitrates, cuts each **rendition** into segments of a few seconds, and publishes a **manifest** listing the renditions and where their segments live. The player reads the manifest, fetches segments one after another over plain HTTP, and keeps a small buffer ahead of the playhead. Because every rendition is cut at the same time points, the player can take the *next* segment from a different bitrate when its bandwidth changes, which a single file at one bitrate cannot offer. Startup is faster because only the first segment or two must arrive before playback begins, seeking is a jump to the segment covering that time, and a failed request costs a few seconds of video rather than a huge transfer. Segments are also small, immutable files that any ordinary web server or CDN can serve, with no special streaming server in the path.

go deeper

for a junior

Be able to name the three pieces, renditions, segments and the manifest, and say why a player fetching a few seconds at a time can adapt quality and start faster than one downloading a single file.

for a middle

Explain that renditions are cut at identical time points so a switch happens cleanly at a segment boundary, and describe what the manifest declares so the player can filter renditions before fetching media.

for a senior

Show you know the costs: every rendition multiplies encode compute and storage, a keyframe at every segment start costs bits, and segment length bounds live delay. Tie the design to real failure modes such as startup stalls.

for a principal

Frame segmentation as trading up-front encoding and storage for per-viewer adaptability and commodity HTTP delivery, and judge when shared CMAF packaging across protocols pays off against device compatibility.

## The problem with one big file A feature-length video encoded at a single high bitrate can be several gigabytes. Served as **one file**, it forces a hard choice: pick a bitrate that suits a fast home connection and viewers on a congested mobile link stall constantly, or pick a low bitrate and everyone gets a blurry picture. A single file also makes seeking depend on byte-range requests into a file whose internal layout the player must first learn, and a connection dropped late in playback can mean re-requesting a very large object. Video is also unlike a typical web asset. An image or script is fetched once and finished; a video is consumed **progressively over time**, and the network conditions at minute one say little about minute forty. ## How segmented delivery works **Segmented adaptive streaming** — standardised as **HLS** and **MPEG-DASH**, both of which can carry media in the common **CMAF** fragment format — reshapes the video before anyone requests it: 1. A transcode step encodes the source into several **renditions** (also called representations or variants): the same content at different bitrates and resolutions, from a low-resolution rung at a few hundred kbit/s up to a high-resolution rung at several Mbit/s. 2. Each rendition is cut into **segments**, typically a few seconds each, and every rendition is cut at the **same time points**. 3. A **manifest** is published that describes the renditions and how to find their segments. 4. The player downloads the manifest, chooses a starting rendition, and fetches segments sequentially over ordinary HTTP, keeping a **buffer** of several seconds ahead of the playhead. 5. Before each new segment, the player re-evaluates its bandwidth and buffer and may request the next segment from a **different rendition**. ## What the manifest contains | Concept | HLS | MPEG-DASH | |---|---|---| | Top-level document | multivariant (master) playlist | MPD (Media Presentation Description) | | One encoding of the content | variant stream with its own media playlist | `Representation` inside an `AdaptationSet` | | Segment listing | `#EXTINF` entries in the media playlist | an explicit list or a `SegmentTemplate` pattern | | Upper bound on segment length | `#EXT-X-TARGETDURATION` | durations in the segment timeline | | Complete, on-demand asset | `#EXT-X-ENDLIST` present | static presentation type | The manifest carries each rendition's declared bitrate, resolution and codec information, so the player can rule out renditions it cannot decode, or that exceed its screen, before downloading any media. ## Why this beats a single file - **Adaptation**: because renditions share segment boundaries, the player can move up or down in quality at the next boundary without restarting playback. - **Fast startup**: only the first segment or two must arrive before the first frame, and players commonly start on a modest rendition to keep that first download small. - **Cheap seeking**: jumping to minute forty means working out which segment covers that time and requesting it; nothing before it is fetched. - **Resilience**: a failed request costs one segment of a few seconds, which the player retries while it plays from its buffer. - **Plain HTTP infrastructure**: segments are small, immutable files that any web server, object store or CDN can serve, and HTTP passes through firewalls and proxies that often block specialised streaming protocols. - **Shared objects**: the same segment files serve every viewer of that rendition, which is what makes very large audiences affordable. ## What it costs Segmentation is not free: - Every rendition has to be produced up front, which multiplies encoding compute and storage by the number of renditions. - Each segment must begin with a **keyframe** — a frame the decoder can display without any earlier frame — and keyframes are expensive to encode, so very short segments spend extra bits. - Many small requests add HTTP overhead, and manifests for long videos grow long. - For live streams, segment duration puts a floor under how far behind real time viewers sit. Those trade-offs are why segment length and the set of renditions are deliberate design choices rather than defaults. ## Key takeaway A segmented stream turns one large, fixed-quality object into a menu of small, interchangeable pieces. The manifest is the menu, the segments are the pieces, and the player decides — segment by segment — which piece its current connection can afford.

  • Why do players usually start playback on a low or middle rendition rather than the highest one?
    At startup the player has no throughput measurement yet and an empty buffer, so a stall is most likely exactly then. Starting on a modest rendition keeps the first segment small, gets the first frame on screen quickly, and gives the player real download timings for its next choice. It climbs to higher renditions once the buffer has some cushion and the measurements justify it.
  • Can one set of segment files serve both HLS and MPEG-DASH players?
    Often, yes: when media is packaged as CMAF fragments with an encryption scheme both target device groups support, an HLS playlist and a DASH MPD can reference the same segment files, and only the manifests differ. Older setups packaged each protocol separately, doubling stored segments. Device codec support and encryption compatibility still decide whether a single packaging is enough.
  • What happens in a segmented stream when a viewer's connection drops for two seconds?
    The player keeps playing from its buffer while it retries the failed segment request. If the connection returns before the buffer drains, the viewer notices nothing; if the outage outlasts the buffer, playback stalls until a segment arrives, and the player will usually step down a rendition afterwards. Only the interrupted segment is re-requested, not the whole video.

A segmented stream is like a restaurant that offers every course in several portion sizes: the diner chooses the size of each next course based on current appetite instead of committing to one fixed meal up front.

saying these in an interview costs you the question

  • Segments exist so the server can push video over one persistent socket.
  • The player downloads one full file and splits it into segments locally.
  • Each segment is re-encoded on the fly when a player asks for a quality.
  • Segmenting only matters for live streams, not for on-demand video.
  • The server's manifest dictates which quality each player must play.
open as a page

In an adaptive-bitrate video player, how is the bitrate rung for the next segment chosen?

level: middleimportance: must knowfreq 55%

basics

~20 s

The player estimates available throughput from recent segment downloads, checks how many seconds are buffered, and picks the highest rung it can sustain with a safety margin, stepping down quickly when the buffer drains and up cautiously to avoid oscillation.

open as a page

When choosing segment duration for an HLS or MPEG-DASH video service, what do shorter segments gain and cost, given each must start on a keyframe?

level: middleimportance: should knowfreq 46%

basics

~20 s

Shorter segments speed startup, let the player switch bitrate sooner and cut live delay, but usually need more keyframes, which costs compression efficiency, and add requests and manifest size; keyframes must also align across renditions so switches stay seamless.

open as a page

For a live event stream over HLS or MPEG-DASH, how does low-latency chunked delivery cut the delay that whole-segment delivery adds?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Classic live delivery publishes a segment only once it is complete and players hold back several segments, leaving viewers often twenty seconds or more behind; low-latency modes publish sub-second chunks as they are encoded over chunked HTTP, cutting delay to a few seconds.

open as a page

How would you restructure a video-on-demand transcode pipeline where a two-hour upload takes twelve hours to produce six renditions?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Treat transcoding as a queue of small independent tasks: split the source at keyframes into chunks, transcode every chunk of every rendition in parallel on a worker fleet, then assemble, package and publish, with idempotent per-task retries.

open as a page

In progressive-download video playback over HTTP, how does a player seek to minute 40 without downloading everything before it?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

The player uses the file's index to map minute 40 to a byte offset at a nearby keyframe, then sends an HTTP Range request from that offset; a server that supports ranges replies 206 Partial Content with only those bytes.

open as a page