skip to content

A server streams a product page's HTML, but a recommendations widget in the middle of the page takes 900 ms to load while everything below it is ready immediately. How do teams stop that one slow region from holding up the rest of the streamed document?

level: middleimportance: should knowfreq 45%

answer

  1. document order is not delivery order
  2. hole now, content later
  3. template plus tiny inline script
  4. reserve the size or the page jumps

basics

~20 s

Stream a placeholder element where the slow widget belongs, continue writing the rest of the page, then send the widget's real markup in a later chunk with a small inline script that moves it into the placeholder.

solid answer

~50 s

HTML is a single ordered stream, so if the server writes regions strictly in document order, a slow region in the middle blocks everything after it. The fix is out-of-order streaming. At the widget's position the server writes a placeholder — an empty container with a stable `id`, usually holding a skeleton — and immediately carries on writing the rest of the page. When the recommendations resolve, the server appends their markup near the end of the stream, typically inside a `<template>` so it is inert and invisible, followed by a tiny inline script that clones the template's content into the placeholder. The user sees the full page structure at once, the slow region fills in when it is ready, and nothing after it was delayed. The placeholder must reserve roughly the widget's final size, or the swap causes a layout shift.

code

html · 14 lines
html
<section id="recs" style="min-height:180px" aria-busy="true">
  <p>Loading recommendations…</p>
</section>

<!-- …the rest of the page streams here, unblocked… -->

<template id="recs-data">
  <ul><li>Blue mug</li><li>Red mug</li></ul>
</template>
<script>
  var slot = document.getElementById('recs');
  slot.replaceChildren(document.getElementById('recs-data').content);
  slot.removeAttribute('aria-busy');
</script>

go deeper

for a junior

Know that a streamed page can show an empty slot for a slow section and fill it in later, and that the slot should already be about the right size so nothing jumps when the content arrives.

for a middle

Walk through the three steps concretely: placeholder with a stable id, keep streaming the rest, then a late chunk plus a tiny inline script that moves the content into the slot. Explain why the late markup is inert until swapped.

for a senior

Show judgment about which regions earn a placeholder, keep the count of independent swaps low enough that the page does not feel twitchy, and have an answer for a region that fails after headers are committed.

for a principal

Weigh whether this pattern belongs in the codebase at all versus moving slow work off the request path — caching, precomputation, or a client fetch — and set a convention so teams do not defer primary content chasing a metric.

## Why document order is a constraint A streamed HTML response is one ordered byte sequence. The parser consumes it front to back, so markup written at position N cannot appear before markup at position N−1. If the server renders regions in document order and awaits each one's data, a single slow region stalls every region below it — the footer, the reviews, and any scripts referenced further down all wait behind the recommendations query. That is a scheduling problem, not a bandwidth problem, and the answer is to decouple *where content appears* from *when it is sent*. ## The placeholder-and-swap pattern The technique is old and well proven — Facebook publicised it as BigPipe, and modern streaming frameworks implement the same idea underneath their own APIs. It has three parts: **1. Emit a placeholder at the right position.** When the parser reaches the widget's slot, the server has no data, so it writes a container instead: ```html <section id="recs" class="recs-skeleton" aria-busy="true"><!-- filled later --></section> ``` The container carries a stable identifier and, ideally, a skeleton that occupies the widget's eventual dimensions. **2. Keep streaming everything else.** The server continues writing the rest of the document immediately. The user gets the complete page structure — header, product details, reviews, footer — while the slow query is still running. **3. Send the late chunk and swap it in.** When the data resolves, the server appends the real markup to the still-open response, wrapped in a `<template>` element so the browser parses it but does not render or load anything inside it. Right after it comes a short inline script that moves the content into the placeholder: ```html <template id="recs-data"><ul><li>Blue mug</li></ul></template> <script> var slot = document.getElementById('recs'); slot.replaceChildren(document.getElementById('recs-data').content); slot.removeAttribute('aria-busy'); </script> ``` The inline script runs the instant the parser reaches it — no round trip, no framework required, no waiting for `DOMContentLoaded`. ## Deciding what deserves a placeholder Not every slow region should be deferred this way. The useful test is whether the region is *the reason the user came*. Defer things that are secondary and slow: recommendations, related content, review lists, personalised badges, availability lookups. Do not defer the primary content of the page — a product's own title, price and hero image — because turning the main content into a late swap delays the largest contentful paint and makes the page look empty at the moment the user is judging it. A second consideration is how many placeholders a page has. Each one is a visible transition; a page that resolves eight regions at eight different moments feels twitchy even though every individual chunk arrived promptly. Grouping related slow regions into one placeholder often feels calmer than resolving each independently. ## The failure mode: layout shift The swap replaces a small placeholder with taller real content, and everything below jumps. This is the most common way an out-of-order streaming implementation makes the experience measurably worse while making the timings look better. The placeholder must reserve height — fixed dimensions, an aspect ratio, or a skeleton built from the same layout as the real widget. Where the final size genuinely varies, cap the container and let the content scroll or clamp inside it rather than letting the page reflow. ## Accessibility and correctness details The placeholder should communicate that it is loading, not pretend to be content: `aria-busy="true"` on the container, removed when the swap completes, is a low-cost signal. If the region can fail, the server needs a path that writes an error state into the placeholder rather than leaving a skeleton spinning forever — once headers are committed, a failed region cannot turn the page into an error response. And a client that has JavaScript disabled will never run the swap script, so the placeholder's own contents are what those users see; for content that must exist without scripting, deferral is the wrong tool. ## What good looks like Done well, the network panel shows one long document request whose body arrives in bursts, the first paint shows a complete page skeleton, the deferred region fills in without moving anything around it, and no region below the slow one waited on it. That is the whole goal: preserve the reading order of the page while abandoning the delivery order.

  • Why wrap the late chunk in a <template> instead of writing the markup directly where it lands?
    Content inside a `<template>` is parsed but inert: it is not rendered, and images or iframes inside it are not fetched from that position. That lets the server append the markup anywhere in the stream without it flashing at the bottom of the page before the swap. Cloning `template.content` into the placeholder is what actually activates it.
  • How would you decide whether a slow region deserves a placeholder at all?
    Ask whether it is the content the user came for. Secondary regions — recommendations, related items, availability badges — are good candidates because deferring them costs nothing the user is looking at. The page's primary content is a bad candidate: deferring it delays the largest contentful paint and shows an empty frame at exactly the moment the user is judging the page.
  • What happens to this pattern if the user has JavaScript disabled?
    The swap never runs, so those users see whatever the placeholder itself contains. That is acceptable for enhancement-style regions but unacceptable for content that must be present — for those, render in document order and accept the delay, or move the slow work off the request path entirely.

saying these in an interview costs you the question

  • Says the browser can render later HTML before earlier HTML on its own
  • Leaves the placeholder with no reserved height, causing layout shift
  • Defers the page's main content along with the secondary widgets
  • Thinks the swap needs a second network request
  • Assumes a failed region can still return an error status

context