A photo-sharing service has 2 million new photo uploads per day, each stored as an original file averaging 3 MB plus two generated thumbnails averaging 200 KB combined, and everything is kept for 5 years with 3x replication for durability. How would you estimate the total storage the service needs to provision?
answer
- per-event size = primary object + derived assets
- daily flow x retention days = raw stock
- apply replication/erasure-coding multiplier
- erasure coding ~1.4x vs replication 3x
- growth curve, not flat run-rate
basics
~20 sAdd up what one upload creates (photo plus thumbnails), multiply by uploads per day, then by days you keep them, then multiply again by how many safety copies you store (replication). That total is your storage budget.
solid answer
~30 sPer-upload storage = original (3 MB) + thumbnails (0.2 MB) = 3.2 MB. Daily ingest = 2,000,000 x 3.2 MB = 6.4 TB/day. Over 5 years (~1,825 days) that's roughly 6.4 TB x 1,825 ~= 11.7 PB of raw data. Apply 3x replication for durability: ~35 PB total provisioned storage. I'd also flag that this ignores metadata overhead, filesystem/object-store block padding, and that growth is rarely linear - upload volume itself likely grows year over year, so a static per-day multiplier understates later years and a real plan should model growth rate, not just a flat run-rate.
go deeper
Should be able to multiply per-item size by item count with guidance, and recognize that thumbnails/derived data add to the total when pointed out.
Should independently identify all storage components (original + derived assets), compute the retention-period total, and apply a replication multiplier without prompting.
Should distinguish replication vs erasure coding trade-offs, question unstated retention assumptions, and flag that flat run-rate multiplication understates growth.
Should connect the estimate to cost/durability/availability trade-offs across tiers (hot vs cold storage), propose a growth-curve model instead of a static multiply, and reason about how retention policy is a product/legal decision that should be made explicit, not assumed.
## What makes storage different from QPS Storage capacity estimation follows the same 'multiply the pieces together' discipline as QPS estimation, but the pieces are different: instead of a rate over time you're accumulating a growing total, and instead of one multiplier (peak factor) you typically apply two or three (replication factor, metadata/overhead, and a growth curve). The goal in an interview is not decimal precision - it's demonstrating that you know which multiplicative factors exist and can chain them in the right order without dropping one. ## The chain, one factor at a time 1. **Start with the atomic unit**: what does a single business event (here, one photo upload) actually write to storage? It's rarely just the file the user thinks they're uploading - a photo service typically also generates and stores derived assets, in this case two thumbnails at different sizes, so the real per-event footprint is the sum of every artifact created: `3 MB` original + `0.2 MB` thumbnails = `3.2 MB` per upload. Missing the derived assets is one of the most common estimation errors, because interview prompts often mention only the primary object and expect the candidate to think through what a real system generates around it (thumbnails, video transcodes, search index entries, audit logs). 2. **Next, convert the per-event size into a daily ingest rate** by multiplying by daily volume: `2,000,000` uploads/day x `3.2 MB` = `6,400,000 MB/day`, or `6.4 TB/day`. This is the 'flow' - how fast new data arrives. 3. **Storage estimation then needs to convert flow into a 'stock'** - the total accumulated over the retention window - by multiplying by the number of days data is kept: `6.4 TB/day` x `~1,825` days (5 years) `~= 11.7 PB`. This step is conceptually different from QPS estimation, where you only ever cared about the instantaneous rate; here the rate accumulates because storage, unlike compute, is a durable resource that keeps everything you've ever written until it's explicitly deleted or expires. ## The durability multiplier On top of the raw accumulated size, real systems apply a **durability multiplier**: data is never stored as a single copy, because a single disk or node failure would mean permanent data loss. - **Common replication factors are 3x** (e.g., HDFS's default, or triplicated storage across availability zones). - **Erasure-coding schemes** trade some replication overhead for space efficiency (roughly 1.4-1.5x instead of 3x, at the cost of more complex/slower reconstruction on failure). Applying 3x to `11.7 PB` gives roughly `35 PB` of physical storage that must actually be provisioned - three times the naive raw total, which is exactly the kind of factor that, if forgotten, makes an estimate look deceptively cheap. ## The trade-offs at every layer The trade-offs sit at every layer of this chain. - **Replication factor** trades raw storage cost against durability and read availability (more replicas also means more nodes that can serve reads, which helps QPS, but costs proportionally more disk and network for writes to stay in sync). - **Retention period** trades storage cost against product/legal requirements - a photo service might offer 'keep forever' as a feature, while a logging system might aggressively expire data after 30-90 days specifically to keep storage costs bounded; the retention window is often a product decision disguised as an engineering constraint, and a senior engineer should push back on unstated retention assumptions rather than accept '5 years' passively. - **There's also a hidden linear-growth assumption** in the calculation above: it treats daily ingest as constant for all 5 years, when in reality a growing product sees increasing daily uploads over time, so a flat multiplication understates the true multi-year total; a more rigorous estimate would model a growth curve (e.g., 20% year-over-year) and sum a series rather than multiply a constant. ## The failure mode The failure mode this estimation guards against is capacity planning that only accounts for compute and forgets that storage, unlike a CPU or a cache, never naturally sheds old data - a service that under-provisions storage doesn't degrade gracefully the way an overloaded API might; it hits a hard wall where writes start failing outright once disks fill up, which is a much more catastrophic outage than elevated latency. ## Logical bytes against physical bytes A well-known real-world analog is any large-scale object store like Amazon S3 or a company's internal blob store team, whose capacity planning explicitly separates 'logical bytes stored' from 'physical bytes provisioned' precisely because of the replication/erasure-coding multiplier, and who model growth curves rather than flat run-rates specifically because linear extrapolation reliably under-forecasts a growing product's actual storage bill.
- How would erasure coding change this estimate compared to 3x replication, and why might a system choose one over the other?Erasure coding (e.g., splitting data into k data shards plus m parity shards) typically achieves the same or better durability at roughly 1.4-1.5x overhead instead of 3x, cutting storage cost substantially. The trade-off is that reconstructing lost data from parity shards is more CPU- and network-intensive and slower than just reading a full replica, so systems often use full replication for hot/frequently-accessed data and erasure coding for colder, less frequently accessed data.
- The prompt assumed a flat 2 million uploads/day for all 5 years. How would you adjust the estimate if the service is growing 25% year over year?Instead of multiplying one day's ingest by total days, you'd sum a geometric series across the 5 years - each year's daily rate multiplied by 365 and by the increasing factor. This produces a meaningfully larger total than the flat estimate, especially in the later years, and demonstrates you understand that a static run-rate systematically underestimates a growing product's storage footprint.
- Would you provision for logical bytes or physical bytes when sizing hardware purchases or cloud storage budget?Physical bytes - the number after applying the replication or erasure-coding multiplier - because that is what actually gets written to disks and billed by a cloud provider. Logical bytes (the raw, unreplicated total) is useful for reasoning about the data itself but understates real infrastructure cost if used directly for provisioning.
It's like estimating how big a warehouse you'll need for a subscription box business: figure out what's in one box (including the packing materials, not just the product), multiply by boxes shipped per day, multiply by how many days of inventory you keep on hand, then multiply again because you keep backup stock in two other warehouses in case one floods.
saying these in an interview costs you the question
- Only sizes the primary file and forgets derived assets like thumbnails
- Forgets to apply a replication/durability multiplier entirely
- Treats storage growth as a one-time total rather than an accumulating stock over the retention window
- Assumes a flat daily rate for the entire multi-year period on a clearly growing product without flagging it
- Can't explain the difference between logical and physical (provisioned) storage