skip to content

questions

12

In a video-on-demand platform, why is each video cut into short segments listed in a manifest rather than served as one file?

level: juniorimportance: must knowfreq 62%

answer

  1. a menu, not one file
  2. a few seconds per piece
  3. same cut points in every encoding
  4. quality decided for the next piece
  5. small immutable objects over plain HTTP

basics

~20 s

Short segments let a player start after fetching a few seconds, switch bitrate at each segment boundary as bandwidth changes, and seek or retry cheaply, while HTTP servers and caches handle small uniform objects instead of one multi-gigabyte file.

solid answer

~50 s

Segmented delivery (HLS or MPEG-DASH) encodes the video at several bitrates, cuts each **rendition** into segments of a few seconds, and publishes a **manifest** listing the renditions and where their segments live. The player reads the manifest, fetches segments one after another over plain HTTP, and keeps a small buffer ahead of the playhead. Because every rendition is cut at the same time points, the player can take the *next* segment from a different bitrate when its bandwidth changes, which a single file at one bitrate cannot offer. Startup is faster because only the first segment or two must arrive before playback begins, seeking is a jump to the segment covering that time, and a failed request costs a few seconds of video rather than a huge transfer. Segments are also small, immutable files that any ordinary web server or CDN can serve, with no special streaming server in the path.

go deeper

for a junior

Be able to name the three pieces, renditions, segments and the manifest, and say why a player fetching a few seconds at a time can adapt quality and start faster than one downloading a single file.

for a middle

Explain that renditions are cut at identical time points so a switch happens cleanly at a segment boundary, and describe what the manifest declares so the player can filter renditions before fetching media.

for a senior

Show you know the costs: every rendition multiplies encode compute and storage, a keyframe at every segment start costs bits, and segment length bounds live delay. Tie the design to real failure modes such as startup stalls.

for a principal

Frame segmentation as trading up-front encoding and storage for per-viewer adaptability and commodity HTTP delivery, and judge when shared CMAF packaging across protocols pays off against device compatibility.

## The problem with one big file A feature-length video encoded at a single high bitrate can be several gigabytes. Served as **one file**, it forces a hard choice: pick a bitrate that suits a fast home connection and viewers on a congested mobile link stall constantly, or pick a low bitrate and everyone gets a blurry picture. A single file also makes seeking depend on byte-range requests into a file whose internal layout the player must first learn, and a connection dropped late in playback can mean re-requesting a very large object. Video is also unlike a typical web asset. An image or script is fetched once and finished; a video is consumed **progressively over time**, and the network conditions at minute one say little about minute forty. ## How segmented delivery works **Segmented adaptive streaming** — standardised as **HLS** and **MPEG-DASH**, both of which can carry media in the common **CMAF** fragment format — reshapes the video before anyone requests it: 1. A transcode step encodes the source into several **renditions** (also called representations or variants): the same content at different bitrates and resolutions, from a low-resolution rung at a few hundred kbit/s up to a high-resolution rung at several Mbit/s. 2. Each rendition is cut into **segments**, typically a few seconds each, and every rendition is cut at the **same time points**. 3. A **manifest** is published that describes the renditions and how to find their segments. 4. The player downloads the manifest, chooses a starting rendition, and fetches segments sequentially over ordinary HTTP, keeping a **buffer** of several seconds ahead of the playhead. 5. Before each new segment, the player re-evaluates its bandwidth and buffer and may request the next segment from a **different rendition**. ## What the manifest contains | Concept | HLS | MPEG-DASH | |---|---|---| | Top-level document | multivariant (master) playlist | MPD (Media Presentation Description) | | One encoding of the content | variant stream with its own media playlist | `Representation` inside an `AdaptationSet` | | Segment listing | `#EXTINF` entries in the media playlist | an explicit list or a `SegmentTemplate` pattern | | Upper bound on segment length | `#EXT-X-TARGETDURATION` | durations in the segment timeline | | Complete, on-demand asset | `#EXT-X-ENDLIST` present | static presentation type | The manifest carries each rendition's declared bitrate, resolution and codec information, so the player can rule out renditions it cannot decode, or that exceed its screen, before downloading any media. ## Why this beats a single file - **Adaptation**: because renditions share segment boundaries, the player can move up or down in quality at the next boundary without restarting playback. - **Fast startup**: only the first segment or two must arrive before the first frame, and players commonly start on a modest rendition to keep that first download small. - **Cheap seeking**: jumping to minute forty means working out which segment covers that time and requesting it; nothing before it is fetched. - **Resilience**: a failed request costs one segment of a few seconds, which the player retries while it plays from its buffer. - **Plain HTTP infrastructure**: segments are small, immutable files that any web server, object store or CDN can serve, and HTTP passes through firewalls and proxies that often block specialised streaming protocols. - **Shared objects**: the same segment files serve every viewer of that rendition, which is what makes very large audiences affordable. ## What it costs Segmentation is not free: - Every rendition has to be produced up front, which multiplies encoding compute and storage by the number of renditions. - Each segment must begin with a **keyframe** — a frame the decoder can display without any earlier frame — and keyframes are expensive to encode, so very short segments spend extra bits. - Many small requests add HTTP overhead, and manifests for long videos grow long. - For live streams, segment duration puts a floor under how far behind real time viewers sit. Those trade-offs are why segment length and the set of renditions are deliberate design choices rather than defaults. ## Key takeaway A segmented stream turns one large, fixed-quality object into a menu of small, interchangeable pieces. The manifest is the menu, the segments are the pieces, and the player decides — segment by segment — which piece its current connection can afford.

  • Why do players usually start playback on a low or middle rendition rather than the highest one?
    At startup the player has no throughput measurement yet and an empty buffer, so a stall is most likely exactly then. Starting on a modest rendition keeps the first segment small, gets the first frame on screen quickly, and gives the player real download timings for its next choice. It climbs to higher renditions once the buffer has some cushion and the measurements justify it.
  • Can one set of segment files serve both HLS and MPEG-DASH players?
    Often, yes: when media is packaged as CMAF fragments with an encryption scheme both target device groups support, an HLS playlist and a DASH MPD can reference the same segment files, and only the manifests differ. Older setups packaged each protocol separately, doubling stored segments. Device codec support and encryption compatibility still decide whether a single packaging is enough.
  • What happens in a segmented stream when a viewer's connection drops for two seconds?
    The player keeps playing from its buffer while it retries the failed segment request. If the connection returns before the buffer drains, the viewer notices nothing; if the outage outlasts the buffer, playback stalls until a segment arrives, and the player will usually step down a rendition afterwards. Only the interrupted segment is re-requested, not the whole video.

A segmented stream is like a restaurant that offers every course in several portion sizes: the diner chooses the size of each next course based on current appetite instead of committing to one fixed meal up front.

saying these in an interview costs you the question

  • Segments exist so the server can push video over one persistent socket.
  • The player downloads one full file and splits it into segments locally.
  • Each segment is re-encoded on the fly when a player asks for a quality.
  • Segmenting only matters for live streams, not for on-demand video.
  • The server's manifest dictates which quality each player must play.
open as a page

A user in Tokyo requests a static image from a website whose origin server sits in Virginia, USA, and the site is served through a CDN with points of presence (PoPs) worldwide. Explain what happens on the very first request for that image versus the hundredth request from a different Tokyo user, and why the CDN makes the site faster.

level: juniorimportance: must knowfreq 75%

basics

~20 s

A CDN stores copies of files on servers (PoPs) close to users everywhere. The first request has to fetch and cache the file from the real origin server, which is slow. Later requests are served from the nearby copy, which is fast because the data travels a much shorter distance.

open as a page

Your team ships a new JS bundle under the same URL /app.js on every deploy, and users keep complaining they're stuck on stale versions after a release even though the file changed on the server. What CDN caching mistake is likely happening, and how would you fix Cache-Control and the deployment strategy to solve it for both this JS bundle and images that rarely change?

level: middleimportance: must knowfreq 80%

basics

~20 s

The file's cache setting is probably telling browsers and the CDN to keep the old copy for too long, and reusing the same filename gives no signal that anything changed. Fix: give changed files new filenames whenever the content changes (like app.abc123.js) and cache those forever, while keeping the URL that points to the latest version very short-lived.

open as a page

In an adaptive-bitrate video player, how is the bitrate rung for the next segment chosen?

level: middleimportance: must knowfreq 55%

basics

~20 s

The player estimates available throughput from recent segment downloads, checks how many seconds are buffered, and picks the highest rung it can sustain with a safety margin, stepping down quickly when the buffer drains and up cautiously to avoid oscillation.

open as a page

A CDN advertises the exact same IP address, say 203.0.113.10, from data centers in Frankfurt, Singapore, and Sao Paulo simultaneously via BGP, and relies on the internet's normal routing to send each user to the nearest one. Explain how this works, and describe a failure scenario where a user's connection to that IP breaks mid-session even though all three data centers are healthy.

level: seniorimportance: must knowfreq 50%

basics

~20 s

Anycast means many servers around the world share one IP address, and the internet's normal routing (BGP) automatically sends each user to whichever one is 'closest' by network path. The problem: if the network path changes mid-connection, a user can suddenly get routed to a different server than before, breaking anything that depended on talking to the same machine.

open as a page

When choosing segment duration for an HLS or MPEG-DASH video service, what do shorter segments gain and cost, given each must start on a keyframe?

level: middleimportance: should knowfreq 46%

basics

~20 s

Shorter segments speed startup, let the player switch bitrate sooner and cut live delay, but usually need more keyframes, which costs compression efficiency, and add requests and manifest size; keyframes must also align across renditions so switches stay seamless.

open as a page

A news website publishes a breaking-news article, then five minutes later the editor fixes a factual error in the headline. Readers in different countries keep seeing the old, wrong headline for varying amounts of time after the fix. What CDN mechanisms would you use to get the correction out quickly and reliably, and what's the trade-off of using them aggressively for every edit?

level: middleimportance: should knowfreq 55%

basics

~20 s

Tell the CDN to throw away its cached copy of that page (purge/invalidate) so the next request fetches the fixed version from the real server. Doing this for every tiny edit is slow to spread everywhere and dumps extra load back on the server, so it should be used sparingly and precisely, not for every small change.

open as a page

For a live event stream over HLS or MPEG-DASH, how does low-latency chunked delivery cut the delay that whole-segment delivery adds?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Classic live delivery publishes a segment only once it is complete and players hold back several segments, leaving viewers often twenty seconds or more behind; low-latency modes publish sub-second chunks as they are encoded over chunked HTTP, cutting delay to a few seconds.

open as a page

How would you restructure a video-on-demand transcode pipeline where a two-hour upload takes twelve hours to produce six renditions?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Treat transcoding as a queue of small independent tasks: split the source at keyframes into chunks, transcode every chunk of every rendition in parallel on a worker fleet, then assemble, package and publish, with idempotent per-task retries.

open as a page

A CDN with 200 edge PoPs serves a video-on-demand site whose new episode just dropped. At the moment of release, hundreds of PoPs simultaneously experience their first cache miss for the same file within the same second. Without any extra configuration, what happens to the origin, and what's the standard CDN feature that prevents it?

level: seniorimportance: should knowfreq 55%

basics

~20 s

All those PoPs would hit the real server at once and could overwhelm it. The fix is 'origin shielding': one extra layer of caching servers sits between the edge PoPs and the origin, so only that shield fetches from origin once, and every edge PoP gets the file from the shield instead.

open as a page

In progressive-download video playback over HTTP, how does a player seek to minute 40 without downloading everything before it?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

The player uses the file's index to map minute 40 to a byte offset at a nearby keyframe, then sends an HTTP Range request from that offset; a server that supports ranges replies 206 Partial Content with only those bytes.

open as a page

A team wants to run personalized A/B test bucketing logic on every request at the CDN edge (for example, using Cloudflare Workers or a similar edge-compute platform) instead of in their origin application server. What kinds of logic are a good fit for that edge-compute layer, and what would make you push back and say 'this belongs in the origin, not the edge'?

level: principalimportance: nice to knowfreq 35%

basics

~20 s

Edge compute runs small pieces of code on the CDN's servers near the user, so simple, fast decisions like which test group to show, or rewriting a URL, can happen instantly without a trip to the main server. It's a bad fit for anything needing a big database, heavy computation, or long-running work, because edge servers are deliberately limited and don't reliably keep state.

open as a page