skip to content

Your site runs over HTTP/2 and a colleague argues that because requests are multiplexed on one connection, you should stop concatenating JavaScript into large bundles and stop worrying about serving assets from a second asset hostname. Which HTTP/1.1-era practices genuinely should go, and where does "requests are free now" break down?

level: seniorimportance: should knowfreq 45%

answer

  1. one connection, many requests
  2. sharding solved a limit that vanished
  3. extra origin means another handshake
  4. each file compresses on its own
  5. deep chains, not wide ones, hurt

basics

~20 s

Domain sharding should go: under HTTP/2 an extra origin adds DNS, TCP and TLS setup and cannot share the connection. Concatenation is looser but not obsolete — bytes still cost parse time, and many tiny files compress worse than one large one.

solid answer

~50 s

Domain sharding should go. It existed to work around the roughly six parallel connections a browser allows per origin under HTTP/1.1; with multiplexing, one connection carries many requests at once, and each extra hostname now costs a fresh DNS lookup, TCP connection and TLS handshake that the existing connection cannot share. Concatenation is a softer call. Requests got much cheaper, so a handful of well-chosen chunks is fine and caches better than one monolith that any change invalidates. But cheap is not free: each file is compressed on its own, so thirty 10 kB chunks compress worse than one 300 kB file; every request still costs server, cache and edge work; and a chunk discovered only after another chunk is fetched and parsed adds a serial round trip. Multiplexing changed nothing about bytes — parse and execute cost is identical either way.

go deeper

for a junior

Know that HTTP/2 carries many requests over a single connection, and that splitting assets across several hostnames was a trick aimed at HTTP/1.1's per-origin connection limit.

for a middle

Explain concretely what an extra origin costs — DNS, TCP and TLS before any byte moves — and why a new stream on a warm connection is far cheaper than a new connection.

for a senior

Demonstrate where the "requests are free" claim fails: per-file compression ratio, chained discovery that serialises fetches, per-request edge work, and cache churn on deploy.

for a principal

Frame it as a delivery standard: how many chunks and how many origins a product may ship, and how you would prove a change helped with field data across real networks rather than a fast local run.

## What HTTP/1.1 forced frontends to do Under HTTP/1.1 a browser opens a small number of parallel connections per origin — around six in practice — and each connection carries one request at a time. That single constraint produced most of the delivery folklore of the era: **concatenate** scripts and stylesheets so there are fewer requests to queue, **sprite** small images into one sheet for the same reason, **inline** critical resources into the HTML so they need no request at all, and **shard** assets across `static1.example.com`, `static2.example.com` and friends so the browser would open several sets of connections and get more parallelism. Every one of those is a workaround for a limit on *concurrent requests*, not a statement about what is inherently good design. Once the limit changes, the workarounds have to be re-argued from scratch. ## What multiplexing changed HTTP/2 carries many concurrent requests over a single connection, so the queueing that motivated concatenation and sprites largely disappears. Two consequences follow directly. First, **domain sharding is now a net cost.** An extra hostname means another DNS resolution, another TCP connection and another TLS handshake before the first byte of the first asset on it can move — round trips that the already-warm main connection has finished paying. It also splits your traffic across connections that cannot be prioritised against each other, so the browser has less ability to decide what matters most. Retiring shards is one of the few unambiguous wins here. Second, **fine-grained splitting became affordable.** Shipping fifteen chunks instead of one is no longer the disaster it was in 2013. That is a real change, and it enables better caching: a monolithic bundle is invalidated in full by a one-line change, whereas separate chunks let unchanged code stay in the user's cache across deploys. ## Where "requests are free" breaks down Four costs survive multiplexing, and a senior answer names them. **Compression ratio.** Each response is compressed independently. Compressors work by finding redundancy within their input, so a single 300 kB file compresses much better than thirty 10 kB files with the same total content — repeated tokens and the compressor's warmed-up state do not carry across file boundaries. Shattering a bundle into very small pieces can measurably increase total bytes on the wire even though it reduced no code. **Per-request overhead that is not connection setup.** Each request still costs header frames, a cache lookup at the edge, possible origin work, and bookkeeping in the browser. Individually small, collectively real once you are into the hundreds. **Discovery chains.** The expensive shape is not many requests — it is *serial* requests. If chunk A must be fetched and parsed before the browser learns chunk B exists, you have added a round trip that no amount of multiplexing removes. A flat set of fifty requests issued at once behaves very differently from five requests issued in a chain of five depths. **Cache churn.** Splitting too finely means a single dependency bump can change many chunk hashes at once, or force the entire chunk graph to be re-fetched — the opposite of the caching benefit that motivated splitting. One more: HTTP/2 puts everything on one TCP connection, so a lossy network degrades all streams together; HTTP/3 improves that situation, but it changes none of the byte, parse or compression arguments above. ## The workable default Drop sharding, drop sprites, keep inlining strictly for genuinely render-blocking critical content (inlining trades cacheability for one fewer round trip, so it should be small and deliberate). Ship a moderate number of chunks — enough that caching works across deploys and that the first screen does not pull in code it never runs, few enough that compression ratio and request bookkeeping stay sane. "A dozen or two" is a very different regime from "three" and from "two hundred". ## Proving it The argument is settled by measurement, not by protocol trivia. Compare total compressed transfer bytes before and after — not raw bytes, since the compression effect is the whole point — and look at whether the request waterfall is wide (parallel, fine) or deep (chained, expensive). Then confirm with field data rather than a single fast-network local run, because the connection-setup savings from removing a shard show up mainly for users on slow, high-latency networks.

  • If requests are cheap, why can a page with thirty chunks still feel slow?
    Usually because the requests are serialised, not parallel. When a chunk is only referenced from inside another chunk, the browser cannot ask for it until the first one has arrived and been parsed, so you pay a round trip per level of depth. Add the lower per-file compression ratio and the per-request edge work, and thirty chunks stops resembling thirty free requests.
  • Does moving to HTTP/3 change this advice?
    Not for bundling. HTTP/3 removes the penalty where one connection's packet loss stalls unrelated streams, so many parallel requests degrade more gracefully on lossy mobile networks. But the decoded bytes, the parse cost and the per-file compression ratio are identical, so the concatenation tradeoff is argued on exactly the same grounds.
  • Is inlining critical CSS or JavaScript into the HTML still worth it under HTTP/2?
    Sometimes, but it is a genuine trade rather than a free win. Inlining removes a round trip for render-blocking content, at the cost of making those bytes uncacheable and re-sent with every HTML response. It pays off for a small, stable, truly render-blocking slice and turns negative as soon as it grows or changes often.

saying these in an interview costs you the question

  • Says HTTP/2 makes request count irrelevant
  • Recommends domain sharding to increase parallelism
  • Claims multiplexing removes the parse cost of the bytes
  • Thinks HTTP/2 compresses multiple files together
  • Argues one giant bundle is always the fastest option

context