For a live event stream over HLS or MPEG-DASH, how does low-latency chunked delivery cut the delay that whole-segment delivery adds?
answer
- live manifests keep growing
- wait, discover, hold back
- deliver the segment while it is written
- parts, preload hints, blocking reloads
- thinner buffer is the price
basics
~20 sClassic live delivery publishes a segment only once it is complete and players hold back several segments, leaving viewers often twenty seconds or more behind; low-latency modes publish sub-second chunks as they are encoded over chunked HTTP, cutting delay to a few seconds.
solid answer
~50 sIn classic live delivery the packager can publish a segment only once it is complete, the player discovers it on its next playlist reload, and players start about three target durations back from the end of the playlist for safety; with 6-second segments that hold-back alone is about 18 seconds, before encode and network time. Low-latency modes attack each term. The encoder emits **CMAF chunks** of a few hundred milliseconds, and the origin serves a still-growing segment with HTTP **chunked transfer encoding**, so frames flow to the player as they are produced. LL-HLS lists these as **partial segments** (`#EXT-X-PART`), announces the next one with `#EXT-X-PRELOAD-HINT`, and supports **blocking playlist reloads** so the player learns of new parts at once; low-latency DASH signals early availability with `availabilityTimeOffset`. Players then keep a buffer of one to a few seconds. The result is typically a few seconds of delay, paid for with a thinner buffer, more requests and a delivery path that must pass partial responses through without holding them back.
go deeper
Recall that a live manifest keeps updating while an on-demand one is complete, and that live viewers sit behind real time because segments must finish before they are published.
Break the delay into encoding, waiting for a segment to close, playlist discovery and player hold-back, and explain how partial segments, preload hints and chunked transfer shrink each part.
Know the operational costs: thin buffers stall more, throughput estimates saturate at the production rate, drift needs playback-rate correction, and every hop must pass partial responses through promptly.
Set the latency target from the product need, weighing a few seconds of HTTP delivery against sub-second real-time protocols on cost, scale, reliability and device reach.
## Live versus on-demand For **on-demand** video every segment exists before anyone presses play; the manifest is complete (in HLS it ends with `#EXT-X-ENDLIST`), and the only latency that matters is **startup time**. For **live** video, segments are created while viewers watch. The manifest is a **sliding window** that keeps growing at the end, players reload it repeatedly, and a new number matters: **latency**, the gap between something happening and a viewer seeing it. ## Where classic live delay comes from With whole-segment delivery, several stages each add delay: - **Capture and encode**: typically around a second or more. - **Waiting for the segment to close**: a segment cannot be published until its last frame is encoded, so the newest published media always trails real time. - **Discovery**: the player sees a new segment only when it next reloads the playlist, which it does roughly once per target duration. - **Hold-back**: HLS players should not start closer than three target durations from the end of the playlist, so they have room to absorb hiccups. With 6-second segments that is about 18 seconds. - **Delivery and player buffer**: network transfer, plus whatever the player keeps buffered. These do not simply add up — hold-back already includes some of the waiting — but in practice classic live streams with 6-second segments commonly sit around twenty to thirty seconds behind real time. ## How low-latency chunked delivery works Low-latency modes keep HTTP and the segment model but change the unit of delivery: 1. The encoder writes each segment as a series of **CMAF chunks** of a few hundred milliseconds, each one immediately usable. 2. The origin serves a segment **while it is still being written**, using HTTP chunked transfer encoding, so a request for the in-progress segment receives chunks as they appear instead of waiting for the whole file. 3. **LL-HLS** lists these pieces as partial segments with `#EXT-X-PART`, announces the next expected part with `#EXT-X-PRELOAD-HINT` so the player can request it before it exists, and, when the server advertises `CAN-BLOCK-RELOAD=YES`, lets the player ask for a playlist that the server holds until a newer update is ready. 4. **Low-latency DASH** uses chunked CMAF too and tells players, through `availabilityTimeOffset`, that segments can be requested before they are complete; the manifest can also declare a target latency. 5. The player holds only a small buffer, often one to a few seconds, and adjusts playback speed slightly to stay near its target distance from the live edge. Not every part has to start with a keyframe: LL-HLS marks parts that do as independent, and players join or switch renditions at those. ## What it costs - **Thin buffers stall more**: a network hiccup longer than the buffer interrupts playback, and each stall pushes the viewer further behind unless the player catches up. - **Harder bandwidth estimation**: a chunk being produced live arrives only as fast as the encoder creates it, so measured throughput saturates near the stream's own bitrate and says little about spare capacity. - **More requests and manifest churn**: sub-second parts mean many more requests and playlist updates per viewer. - **Delivery-path requirements**: every hop must pass chunked responses through promptly instead of buffering the whole object, and manifests need very short freshness. - **Less encoder lookahead**: encoding with minimal delay limits the analysis that improves compression. ## Choosing a latency target | Approach | Typical delay (illustrative) | Unit delivered | Main trade-off | |---|---|---|---| | Classic segmented live | about 20–30 s with 6 s segments | whole segments | stable and cache-friendly, but far behind | | Low-latency chunked HTTP | about 2–5 s | sub-second chunks and parts | thin buffer, more requests, stricter delivery path | | Real-time, conferencing-style protocols | under about 1 s | individual frames over UDP-based transports | different infrastructure, costlier at very large audiences | The right target comes from the product. Broadcast-style events, where the worry is a social feed spoiling a goal, are usually well served by a few seconds. Interactive formats — live auctions, betting, two-way conversation — need sub-second delay, which pushes them beyond HLS and MPEG-DASH altogether.
- Why is bandwidth estimation harder for a low-latency live player?A chunk being produced live arrives only as fast as the encoder creates it, so download time reflects the production rate rather than network capacity. Measured throughput therefore looks capped near the stream's own bitrate, and the thin buffer leaves little room for error. Low-latency players lean more on buffer trends, timing that excludes waits for new data, and conservative switching.
- Why do viewers of the same low-latency stream drift apart, and how do players correct it?Stalls, slow starts and different buffer sizes leave players at different distances from the live edge, and each stall adds delay that stays unless recovered. Players aim at a target latency, often signalled in the manifest, and play slightly faster or slower than normal to close the gap, or jump forward when they fall too far behind.
- When is sub-second latency worth leaving HLS and MPEG-DASH for?When viewers interact in real time, such as live auctions, betting or two-way conversation, a few seconds breaks the experience, so conferencing-style real-time protocols are used despite needing different infrastructure and scaling less cheaply to huge audiences. For broadcast-style events, a few seconds of low-latency HTTP delivery is usually enough and keeps commodity delivery.
saying these in an interview costs you the question
- Live delay comes only from the network, so a faster CDN removes it.
- Shrinking the player buffer to zero is a free way to reach real time.
- Low-latency HLS requires every partial segment to start with a keyframe.
- Live and on-demand streams use the same complete manifest that never changes.
- Low-latency modes replace HTTP with a dedicated streaming protocol.