skip to content

With HTTP/2 server push removed from browsers, how would you decide between HTTP 103 Early Hints, Link: rel=preload on the final response, inlining critical CSS, and doing nothing, for delivering render-blocking resources on a high-traffic site?

level: principalimportance: nice to knowfreq 18%

answer

  1. budget the critical path: setup / think-time / transfer / discovery
  2. Early Hints ceiling = origin think-time
  3. preload = fixes discovery order, nearly free
  4. preconnect = cheapest failure mode
  5. inline only tiny + stable; loses cacheability

basics

~20 s

Decide by where the time actually goes. Long origin think-time favours 103 Early Hints; a fast origin makes preload on the final response the cheap default; a tiny, stable critical set justifies inlining despite losing cacheability; otherwise do nothing and fix caching or the number of critical resources instead.

solid answer

~1 min

Treat it as a measurement question, not a preference. Break the critical path into DNS/TLS, origin think-time, transfer, and parse-then-discover. - **Long origin think-time (hundreds of ms) and a fixed critical set** → **103 Early Hints**: the only option that overlaps fetching with generation. Costs: 1xx must survive every hop, hints must track the build manifest, wrong hints waste bandwidth. - **Fast origin** → `Link: rel=preload` header fields on the final response, or `<link rel=preload>` in the head. No round trip saved over parsing, but it fixes discovery order for late-referenced assets, and it is nearly free. - **Third-party origins on the critical path** → `rel=preconnect`, the highest ratio of benefit to risk, since a wrong guess costs one idle connection. - **Tiny, stable critical CSS (a few KB)** → inline it. You trade cacheability and HTML size for zero round trips; only worth it when the set is small and changes rarely. - **Otherwise do nothing.** If the connection is not the bottleneck, none of this moves the needle; cache the HTML or cut the number of render-blocking resources. Guard the choice with field metrics and a rollback: hints that go stale after a deploy silently regress everyone.

code

http · 8 lines
http
HTTP/2 103 Early Hints
link: <https://fonts.cdn.example>; rel=preconnect
link: </static/critical.8ab3.css>; rel=preload; as=style

HTTP/2 200 OK
content-type: text/html; charset=utf-8
link: </static/critical.8ab3.css>; rel=preload; as=style
link: </static/hero.9c1d.woff2>; rel=preload; as=font; crossorigin

go deeper

for a junior

Know the toolbox and one rule of thumb: preload and preconnect are cheap, Early Hints needs a slow origin to pay off, inlining is for tiny critical CSS.

for a middle

Explain what segment of the critical path each option removes and name the main cost of each.

for a senior

Drive it from measurement, verify 1xx survives the path, tie hint generation to the build, and watch for wasted-fetch and 404 rates after deploys.

for a principal

Decide with a critical-path budget and a failure-cost ordering, own the rollout guardrails and ownership of hint correctness, and be willing to say the right answer is to reduce critical resources or cache the HTML instead.

## Frame the decision as a critical-path budget Every technique here buys back a specific segment of time. Name the segments before choosing: 1. **Connection setup** to your origin and to third parties (DNS + TCP/QUIC + TLS). 2. **Origin think-time**: request received → first byte of HTML produced. 3. **HTML transfer**. 4. **Discovery**: the browser parses HTML, finds `/app.css`, and only then requests it — one extra round trip, plus the queueing behind whatever else is in flight. 5. **Subresource transfer**. Measure them in field data (real-user monitoring) split by device class and network, not in a lab on a warm cache. Then pick per segment. ## Option by option **103 Early Hints.** Buys segment 4 *and* overlaps it with segment 2. This is the only tool that starts subresource fetches before the HTML exists, so its ceiling equals your origin think-time. Requirements: an application framework that can flush an early response, every intermediary forwarding 1xx (buffering proxies and some WAFs drop it), and a hint list generated from the same build artifact as the page so hashed filenames never go stale. Browsers honour it for top-level navigations and for preload/preconnect hints. Wrong hints cost a full wasted fetch each, so keep the list to render-blocking assets. **`Link: rel=preload` on the final response (or `<link rel=preload>` in the head).** Buys part of segment 4 only — the browser still waits for the HTML, but it no longer waits to *parse* to the reference. Valuable when a critical asset is referenced late in the document, imported from inside a JS module graph, or discovered only at runtime (a font referenced by CSS is the classic case: two levels of discovery). Cost is near zero, correctness risk is low, and it works everywhere. This is the sane default. **`rel=preconnect`.** Buys segment 1 for third parties. Best risk profile of anything here: a wrong guess costs one idle socket, a right guess saves a full handshake, often 100–300 ms on mobile. Limit to two or three origins; each one holds resources. **Inlining critical CSS (or a small critical JS shim).** Buys segments 4 and 5 entirely — zero round trips. Costs are real: the bytes are re-sent on every HTML response (uncacheable), the HTML grows so the first congestion window may no longer hold it, and you now maintain a critical-CSS extraction step that silently rots. Justified when the critical set is genuinely small (single-digit KB) and stable, and especially on landing pages where the HTML is already uncacheable. **Do nothing.** The correct answer more often than teams expect. If the origin returns in 20 ms from cache, the connection is fast, and there are two render-blocking assets, all of these techniques compete for single-digit milliseconds. The higher-leverage moves are usually: cache the HTML at the edge, cut the number of render-blocking resources, remove blocking third-party scripts, and stop shipping unused CSS/JS. ## Priorities as the connecting thread Whichever you choose, the connection is shared. Signal urgency with the `Priority` request header (RFC 9218: `u=0..7`, lower is more urgent; `i` for incremental) so the render-blocking assets outrank below-the-fold images. And remember the ceiling: once bytes are committed to a kernel socket buffer or a CDN buffer, no scheduler can reorder them — very large HTTP/2 flow-control windows or aggressive proxy buffering quietly nullify prioritization. Preloading twenty things at once also *creates* the contention that priorities then have to resolve; hinting fewer, better resources beats hinting more with careful priorities. ## Organizational and operational tradeoffs - **Who owns hint correctness?** Early Hints and preload lists are a build-time output. If they live in a proxy configuration edited by a different team from the one that renames the assets, they will break at the next deploy. Generate them from the manifest, and alert on 404s for preloaded URLs. - **Personalization and experiments.** If the critical asset set varies per user segment, a fixed hint list is wrong for a share of traffic. Either derive hints per variant or hint only the invariant subset. - **Blast radius and rollback.** These changes affect every page view. Ship behind a flag, ramp by percentage, and watch Largest Contentful Paint and error rates rather than counting hints emitted. - **Bandwidth cost.** On mobile, wasted preloads are paid by the user; on your side they are egress. A 5% wrong-hint rate on a large site is a real bill. ## The one-paragraph answer Instrument the critical path; pick the technique that attacks the largest segment; prefer the option whose failure mode is cheapest (preconnect, then preload, then Early Hints, then inlining); tie hint generation to the build so it cannot drift; and be willing to conclude that the right move is to reduce the number of critical resources rather than to deliver them cleverly.

  • How would you prove any of this worked?
    Field data, not lab runs: real-user monitoring of Largest Contentful Paint and Time to First Byte, segmented by device and network, compared between a flagged-on and flagged-off population rather than before-and-after in time. Lab tests with warm caches and fast links routinely show nothing, and synthetic runs miss the tail where the benefit actually lives.
  • When is inlining critical CSS the wrong call even though it removes a round trip?
    When the critical set is large or changes often. Inlined bytes are re-sent with every HTML response and never cached, so a 40 KB critical block costs every visitor on every navigation, can push the HTML past the initial congestion window, and duplicates rules the external stylesheet also carries. It also adds an extraction step that drifts out of date silently.

saying these in an interview costs you the question

  • Reaching for Early Hints without measuring whether origin think-time is large enough to overlap
  • Preloading everything, which just moves the contention rather than removing it
  • Inlining large critical CSS and losing cacheability on every navigation
  • Maintaining hint lists in proxy configuration separate from the build manifest, so they rot after deploys
  • Assuming prioritization can fix ordering after bytes are already buffered downstream

context