Browsers cap parallel HTTP/1.1 connections at roughly six per origin. Why is there a cap at all, and what was domain sharding meant to do about it?
answer
- ~6 per origin = browser convention, not spec
- Cost: sockets, handshakes, slow start, fairness
- Origin = scheme + host + port
- Sharding multiplies the budget
- Anti-pattern under HTTP/2
basics
~20 sMore connections mean more parallel exchanges but also more server sockets and more competing TCP flows, so browsers settled on about six per origin as a compromise. Domain sharding served assets from extra hostnames to multiply that budget.
solid answer
~50 sSince one HTTP/1.1 connection carries one exchange at a time, parallelism means more connections. Browsers cap them per origin — scheme, host and port — at around six because the cost is not borne by the client alone: each connection consumes a server socket and memory, needs its own TCP handshake and TLS negotiation, starts with a fresh congestion window, and competes with the others for the same bottleneck link. Beyond a handful the added parallelism stops paying and starts causing congestion and unfairness. The number is a browser convention, not a protocol rule; the earlier convention was two. **Domain sharding** exploited the fact that the limit is per origin: serve assets from `static1.example.com`, `static2.example.com`, and the browser opens six connections to each. It measurably helped image-heavy HTTP/1.1 pages. It also costs: extra DNS lookups, extra TCP and TLS handshakes, more congestion, and cache and cookie fragmentation. Under HTTP/2 it is actively harmful, because it splits traffic that one multiplexed connection would carry better.
go deeper
Say the browser opens a handful of connections per host so requests can run in parallel, and that sharding used more hostnames to get more of them.
Justify the cap with server sockets, handshake cost and congestion control, define origin precisely, and list sharding's costs.
Quantify the tradeoff, explain slow start and fairness across parallel flows, and know why sharding inverts under HTTP/2 including connection coalescing.
Own the asset-delivery topology: protocol version, number of origins, CDN and certificate strategy, and which historical workarounds should be dismantled during a migration.
## Why more connections at all An HTTP/1.1 connection is serial — one request-response at a time. The only way to fetch assets concurrently is to open more connections. A page with 80 subresources over one connection pays 80 sequential round trips; over six, roughly a sixth of that plus scheduling overhead. Parallel connections were the workaround that made the HTTP/1.1 web usable. ## Why there is a limit The number is chosen by the client, not required by the protocol. RFC 2616 once suggested two per origin; browsers raised it to six (some go to eight for certain conditions) as networks improved. The reasons for any cap: - **Server resources.** Every connection is a socket, a file descriptor, buffers and, for TLS, session state. Multiply by concurrent users and an unbounded client-side appetite becomes a server capacity problem. - **Handshake cost.** Each new connection pays a TCP handshake and, over HTTPS, a TLS handshake — one to three extra round trips before any bytes of the response. Past a point you spend more on setup than you gain in parallelism. - **Congestion control.** Each TCP connection begins in slow start with a small congestion window and probes for capacity independently. Many parallel flows to the same destination over the same bottleneck do not create bandwidth; they split it, cause loss, and behave unfairly toward other traffic on the link — including the user's other tabs and other people on the same network. - **Diminishing returns.** Beyond roughly six, measured page-load improvements flatten out and then reverse on constrained links, especially mobile. The limit is per **origin**: the tuple of scheme, host and port. `https://a.example.com` and `https://b.example.com` have independent budgets — the fact sharding is built on. ## Domain sharding Sharding means publishing static assets across several hostnames so the browser's per-origin budget multiplies: two shards gives about twelve parallel connections, three gives about eighteen. In the HTTP/1.1 era this was standard practice for image-heavy pages and produced genuine wins, typically alongside cookieless asset domains. The costs were always real: - **DNS.** Each new hostname is a fresh resolution, often 20-100 ms, on the critical path. - **Connection setup.** Each shard needs its own TCP and TLS handshakes, so the first request to each shard is expensive; sharding front-loads that cost. - **Congestion.** Twelve or eighteen slow-starting flows through one bottleneck compete with each other, which on mobile links can make the page slower overall. - **Certificates and complexity.** Every shard hostname needs certificate coverage, cache-control consistency and deploy discipline. Getting one shard out of sync produces version skew between assets. - **Cache fragmentation.** The same asset referenced through different shard hostnames is cached separately, wasting storage and misses — and inconsistent shard assignment across pages defeats reuse. The rule of thumb that emerged was: two shards helps, four rarely does, more usually hurts. ## Why it reversed under HTTP/2 HTTP/2 multiplexes many streams over a single connection, so the per-origin connection limit stops mattering — one connection is enough and browsers deliberately use one. Sharding then becomes counterproductive: it splits traffic across several connections, each with its own handshake and its own congestion window, and it defeats HTTP/2's prioritization and header compression, which only work within a connection. Removing sharding is a standard step when migrating a site to HTTP/2 or HTTP/3. There is a subtlety: browsers may coalesce connections for different hostnames that resolve to the same IP and are covered by the same certificate, which softens the penalty — but relying on coalescing is fragile compared with simply not sharding. ## What survives A separate asset origin can still be justified for reasons unrelated to connection counts: keeping cookies off asset requests, isolating untrusted user content on another origin for security, or pointing at a CDN. Those are legitimate; multiplying connection budgets is not, once you are on HTTP/2. ## Answer shape Cap exists because connections cost server resources, handshakes and congestion-control fairness; sharding multiplied a per-origin budget and helped under HTTP/1.1; it is an anti-pattern under HTTP/2 and later.
- Why does opening more parallel TCP connections not simply give you more bandwidth?Each connection starts in slow start with a small congestion window and discovers capacity independently, but they all share the same bottleneck link. Extra flows split the available bandwidth, induce loss as they probe, and pay separate handshakes. Past a small number, added parallelism costs more in setup and congestion than it returns.
- Should a site on HTTP/2 keep its existing shard hostnames?Generally no. HTTP/2 multiplexes everything over one connection, so shards only add handshakes and split prioritization and header compression across connections. Consolidating onto one origin is the usual migration step, unless a separate hostname exists for another reason such as cookieless assets, CDN routing or security isolation.
Opening more checkout lanes speeds up a supermarket — until the lanes outnumber the staff and the doorway everyone must squeeze through.
saying these in an interview costs you the question
- Stating the six-connection limit is mandated by the HTTP specification
- Believing more connections always yield more throughput
- Recommending domain sharding for a site already on HTTP/2 or HTTP/3
- Forgetting that the limit is per origin, so port and scheme matter too
- Ignoring the DNS, TLS and cache-fragmentation costs of extra hostnames