You're asked to estimate outbound network bandwidth for a video-streaming service with 5 million concurrent viewers, each streaming at an average bitrate of 4 Mbps, and to state how confident you are in the resulting number. Walk through the bandwidth calculation, then explain where this kind of back-of-envelope reasoning tends to break down in the real world and what you'd do to compensate.
answer
- bandwidth = concurrency x per-stream bitrate
- watch bit vs byte unit conversion (8x error)
- ABR means bitrate is a distribution, not a constant
- capacity is per-edge/region, not one global pool
- back-of-envelope sets order of magnitude; load testing + telemetry confirm it
basics
~20 sMultiply people watching at once by data used per stream for the total pipe needed. Real traffic isn't evenly spread across servers, so this is a starting point, not the final answer - real measurements and margin are still needed.
solid answer
~50 s5,000,000 concurrent viewers x 4 Mbps = 20,000,000 Mbps = 20 Tbps of instantaneous outbound bandwidth at that peak concurrency. This is the headline number, but I'd immediately flag its precision as low: it assumes every viewer is at exactly the average bitrate simultaneously (real streaming uses adaptive bitrate, so the true distribution spans well below and above 4 Mbps per viewer), it ignores geographic distribution (this isn't one pipe, it's many regional/CDN edges each needing a fraction of this), and it treats 5 million concurrent as a static snapshot rather than a number that itself needs its own peak/headroom estimate. Back-of-envelope math like this is useful for getting the right order of magnitude and for sanity-checking a proposed architecture, but production capacity planning layers real traffic measurement, per-region breakdown, and safety margin on top rather than trusting the single multiplication.
go deeper
Should be able to multiply concurrency by per-unit bitrate to get a total, when the units are given, and get the bits-vs-bytes conversion right with guidance.
Should perform the calculation independently, catch unit-conversion errors unprompted, and recognize in general terms that real traffic isn't perfectly uniform.
Should proactively flag that a single global figure obscures per-region distribution and adaptive-bitrate variance, and connect the estimate to a CDN/multi-region architecture decision.
Should treat the back-of-envelope number explicitly as a bootstrapping tool for architecture shape, articulate precisely where it breaks down (distribution, geography, second-order origin load), and describe the organizational process - load testing, telemetry, pre-scaling - that replaces it for real high-stakes capacity decisions.
## Why bandwidth is the binding constraint here Bandwidth estimation follows the same core discipline as QPS and storage estimation - multiply a per-unit cost by a concurrency or rate figure - but it's worth walking through explicitly because video is the domain where bandwidth, rather than QPS or storage, is usually the binding constraint, and because it's a good vehicle for discussing where back-of-envelope reasoning as a whole starts to mislead rather than inform. ## The calculation The calculation itself is simple: outbound bandwidth = concurrent streams x bitrate per stream. With `5,000,000` concurrent viewers at `4 Mbps` average, that's `20,000,000 Mbps`, or `20 Tbps` (terabits per second), of instantaneous outbound data. Converting units carefully matters here - bandwidth is conventionally expressed in bits per second while storage is in bytes, and mixing them up by a factor of 8 is one of the most common silent errors in this kind of estimate. `20 Tbps` is a genuinely enormous number - for context, it's roughly the scale that only a handful of the largest content delivery networks and cloud backbones operate at in aggregate - which itself should prompt a sanity check: does 5 million truly simultaneous viewers at a shared moment sound right for this product, or is that itself an assumption worth interrogating (a live global sporting final might genuinely produce a number like this; a general-purpose on-demand service spreading views across time zones and content catalog would not concentrate this much simultaneous demand). ## Where the multiplication starts to mislead The reason this simple multiplication both matters and is dangerous is the same reason all back-of-envelope math is valuable and limited: it's excellent at establishing order of magnitude fast, which is exactly what's needed early in a design discussion to rule out obviously wrong architectures (a single server obviously cannot serve 20 Tbps; this must be a globally distributed, CDN-backed design), but it silently launders away every piece of real-world heterogeneity that determines whether an actual deployment survives contact with real traffic. Four specific breakdowns are worth naming explicitly. 1. **First, the average bitrate hides the real distribution**: adaptive bitrate streaming (ABR) means individual viewers sit anywhere from under 1 Mbps (mobile, poor connection, lower-resolution stream) to 15-25 Mbps+ (4K on a strong connection), and the true aggregate load depends on the actual mix of device types, network conditions, and content resolution requested at that moment - not a single averaged figure, which can be wrong in either direction depending on what's popular right now. 2. **Second, the estimate treats bandwidth as one undifferentiated pool**, when in reality traffic must be served from specific points of presence (PoPs) or CDN edge nodes near each viewer; the real engineering question isn't 'how much bandwidth in total' but 'how much bandwidth does each of dozens or hundreds of regional edges need,' and demand is never evenly distributed across those edges - a live event popular in one time zone concentrates load onto specific regional infrastructure far beyond its 'fair share' of the global total. 3. **Third, 5 million concurrent itself is not a fixed input** but a number that has its own peak-vs-average and growth dynamics - concurrency during a live event can spike sharply within a single-digit number of minutes around a kickoff or start time, a pattern flatter time-series-based provisioning can miss entirely. 4. **Fourth, the estimate says nothing about second-order effects like origin load** (even with a CDN caching most traffic, cache misses and live-encoding origin servers still need their own bandwidth and compute capacity, which scales differently from edge delivery). ## Speed against fidelity The trade-off in relying on this kind of estimate is speed versus fidelity: back-of-envelope math takes minutes and is exactly the right tool for an interview or an early design review, where the goal is agreeing on the shape of the solution (needs a CDN, needs multi-region, roughly this order of magnitude of infrastructure) - but it is the wrong tool, used alone, for actually provisioning production infrastructure, where the cost of being wrong (an outage during the exact live event the whole system was built for) is high enough to justify much more rigorous methods: - load testing against realistic traffic shapes - historical telemetry from comparable past events - staged capacity ramps with real-time monitoring during the event itself - pre-negotiated burst capacity with CDN/cloud providers ## The failure mode The failure mode of trusting the single multiplication uncritically is a genuine, recurring production incident pattern: teams that provision based on a single global aggregate number, without breaking it down by region or building in real headroom, get blindsided not by the number being wrong in total, but by its distribution being wrong - the edge serving one popular region saturates long before the 'total capacity is sufficient' spreadsheet would suggest a problem exists. A well-documented real-world case is large live-streaming events - major sports broadcasts on services like ESPN+, Hotstar (which has publicly discussed handling tens of millions of concurrent viewers during cricket finals), or Twitch during major esports events - all of which describe capacity planning as a combination of back-of-envelope sizing to set the initial architecture and infrastructure order-of-magnitude, followed by extensive load testing, regional traffic modeling, and live-event-specific pre-scaling, precisely because a single multiplication of concurrent viewers times average bitrate gets the right ballpark but is never trusted as the final provisioning number for an event where failure is highly visible and costly.
- How would you adjust the 20 Tbps estimate to account for adaptive bitrate streaming instead of assuming every viewer sits at the average bitrate?Rather than one average multiplied uniformly, you'd model the estimate as a weighted sum across bitrate tiers - for example, some percentage of viewers on mobile/low-bandwidth at ~1 Mbps, a majority on standard HD around 3-5 Mbps, and a smaller premium/4K segment at 15-25 Mbps - using real telemetry on device/quality mix if available, since the true aggregate depends heavily on that distribution rather than a single mean.
- Why does breaking the global bandwidth figure down by region or CDN edge matter more than getting the global total right?Infrastructure is provisioned and can fail at the level of individual edges or points of presence, not as one undifferentiated global pool, so a correct global total can still coexist with a specific region's edge being catastrophically under-provisioned if demand concentrates there - which is exactly what happens during a regionally popular live event, making the per-region breakdown the number that actually determines whether an outage occurs.
- At what point in a real capacity-planning process would you move from this kind of back-of-envelope estimate to load testing and real telemetry, and why not skip straight to load testing?The back-of-envelope estimate comes first because it's needed to decide the shape of the architecture (single-region vs multi-region, whether a CDN is required at all, rough infrastructure order of magnitude) before there's anything concrete to load test; once that shape is set, load testing and historical telemetry from comparable events replace the rough estimate with much higher-fidelity numbers for the specific regions and traffic patterns that matter, which back-of-envelope math simply cannot produce on its own.
It's like estimating how much water a city needs by multiplying households by average usage - it gets you the right ballpark for building the main pipeline, but it tells you nothing about which specific neighborhood's pipes will burst during the exact hour everyone showers at once, so you still need real monitoring and local capacity, not just the citywide average.
saying these in an interview costs you the question
- Presents the single multiplication result as a final, precise provisioning number with no caveats
- Confuses bits and bytes when computing bandwidth, silently off by a factor of 8
- Treats bandwidth as one global pool rather than something served per-region/per-edge
- Assumes every stream sits exactly at the average bitrate, ignoring adaptive bitrate's real distribution
- Has no answer for how the estimate would be validated or refined before an actual high-stakes event
- Doesn't sanity-check whether the input concurrency figure itself is plausible for the described product