skip to content

A team proposes generating an image variant every 50 px of width so that every device downloads an almost pixel-perfect fit. What does that granularity cost, and how many widths would you actually ship?

level: seniorimportance: should knowfreq 40%

answer

  1. granularity buys kilobytes, costs hit rate
  2. each width is its own cache object
  3. diluted request stream, colder objects
  4. cold derivative means the user waits for an encode
  5. warm and slightly large beats cold and perfect

basics

~20 s

Fine granularity buys a few kilobytes per request while splitting traffic across dozens of URLs, so each variant is requested less, cache hit rates fall and cold misses land on real users. Five or six widths per image role captures nearly all the saving.

solid answer

~50 s

The saving is bounded and small: with rungs 50 px apart a device over-downloads by at most a sliver of width, which on a typical photo is a few kilobytes — while a sensible six-rung ladder already caps over-download at roughly 25–30%. The cost is not bounded. Every distinct width is a distinct URL, so requests fan out across dozens of cache objects instead of a handful; each object is requested less often, is more likely to have been evicted, and in an on-the-fly resizing setup a miss means someone waits for the origin to generate the derivative — often on the hero image that defines the page's largest paint. The same fragmentation hits the browser cache, so a returning visitor reuses less across pages. I would ship five or six widths per image role, shared site-wide, and spend the effort on making sure the ceiling and the reported layout size are right, which moves far more bytes than granularity does.

go deeper

for a junior

Know that each image variant is a separate file with its own URL, and that having many of them means more files to build, store and cache rather than a free improvement.

for a middle

Explain that caches key on URL, so more widths split the same traffic across more objects; be able to say why the byte saving from finer rungs is small and bounded.

for a senior

Demonstrate the tradeoff with numbers: bound the transfer saving, then weigh it against hit rate and the cost of a cold derivative landing on the page's largest visible image. Land on five or six shared rungs.

for a principal

Own it as policy and as an argument you can win with data — quantify both sides for the team proposing the change, allow a justified exception on a high-traffic surface, and redirect the effort to the levers that move more bytes.

## The saving is bounded, and small Start by sizing the prize. With a ladder whose rungs grow by about 30% each, the worst case is that a device needs a width just above one rung and takes the next one up, over-downloading by roughly 30% of width — call it 60% more bytes in the worst case, and much less on average. Tighten the rungs to 50 px and the worst case drops to a couple of percent of width. On a 700 px-wide photo weighing 90 KB, that difference is on the order of a few kilobytes per request. It is real, but it is the smallest lever on the page. The ceiling being wrong by a factor of two, or the layout size being described dishonestly, each cost an order of magnitude more. Fine granularity is optimisation effort spent at the flat end of the curve. ## The cost is fragmentation, and it is not bounded Every distinct width is a distinct URL, and caches key on URL. That single fact drives the whole tradeoff. **At the edge.** A CDN node holds each variant as its own object. With six rungs, requests for a popular image concentrate on six objects, all of which stay hot and stay resident. With rungs every 50 px, the same traffic spreads across thirty or more objects; each is requested a fraction as often, each is a weaker candidate against the node's eviction policy, and the tail of them is effectively never warm. Hit rate falls not because the cache got smaller but because the request stream got diluted. **On a miss.** In a pipeline that resizes on demand, a miss is not just a slower fetch from origin — it is an encode. The user waits for the derivative to be produced. And because these are images, the ones most likely to be requested at an unusual width are often the largest visible elements on the page, so the penalty lands precisely on the paint the user is waiting for. Fine granularity converts a rare, cheap event into a common, expensive one. **In the browser.** A returning visitor who saw a 780 px variant on one page and lands on another page where the layout resolves to 830 px gets a fresh download instead of a cache hit. Shared rungs across pages and across image roles are what make repeat visits cheap; per-device-perfect widths destroy that reuse. **In the pipeline.** Generation time, storage, invalidation surface, and the cost of ever changing your mind about encoding settings all scale linearly with the number of variants. Regenerating six rungs across a catalogue is an afternoon; regenerating thirty is a project. ## The shape of the right answer Five or six widths per image role, geometrically spaced, shared across the whole site. That ladder: - caps relative over-download at the growth factor, so no device is badly served; - keeps a small, hot set of objects at the edge; - lets a returning visitor reuse assets across pages; - stays small enough that a human can look at the list and reason about it. The governing principle is that **a slightly-too-large image served from a warm cache usually beats a perfectly-sized one served from a cold one.** Transfer size is one term in the equation; time-to-first-byte on a miss is another, and on a cold derivative it can dwarf the byte saving entirely. ## How to push back well The proposal is not stupid — it optimises the metric it names. The way to argue it down is to name the metric it ignores and put a number on both. Quantify the prize: take a representative image, compute the byte difference between the coarse ladder's worst case and the fine ladder's worst case, and state it in kilobytes. Then quantify the cost: look at the current edge hit rate for images, estimate how request volume splits across the proposed rung count, and look at what a cold-derivative response actually costs in milliseconds. In most real setups the second number is larger than the first, and the conversation ends there without anyone having to appeal to taste. If the team still wants finer fit for a specific high-traffic surface — a product hero that drives conversion, say — that is a defensible exception, because the traffic volume keeps even the extra rungs warm. Granularity is affordable exactly where traffic is dense enough to warm it, which is the opposite of the long tail where it is usually proposed. ## Where to spend the effort instead Ranked by bytes moved: get the ladder's ceiling right so nothing serves absurd widths; make sure the size the browser is told about matches what the layout actually does; encode high-density rungs at a lower quality; adopt a more efficient format for the whole ladder. Every one of those moves more weight than halving the gap between rungs, and none of them costs you cache hit rate.

  • Is there a case where a much finer ladder is defensible?
    Yes — on a single very high-traffic surface, such as the hero of a product page that drives conversion. Volume there is dense enough to keep even unusual rungs warm at the edge, so fragmentation costs little and the fit gain is retained. The rule generalises: granularity is affordable exactly where traffic warms it, which is the opposite of the long-tail catalogue where it is usually proposed.
  • How would you actually measure whether variant count is hurting you?
    Look at cache hit rate for image requests at the edge, broken down by variant, and at the latency difference between a hit and a miss for those URLs. A long tail of rarely-requested widths with a large miss penalty is the signature. Pair it with field data on how long users wait for the largest visible image; if the slow tail correlates with misses rather than with size, granularity is the cause.
  • Does serving a variant that is 30% wider than needed noticeably hurt the user?
    Rarely. On a typical photo it is tens of kilobytes, and on a warm cached object it arrives at full throughput with no origin work. Compare that with a cold miss, where the user waits for the origin to fetch and encode a derivative — often hundreds of milliseconds. The over-download is a small, predictable cost; the miss is a large, unpredictable one.

Stocking every half-size in the back room fits each customer better, but each size now sits on the shelf so long that half the time it has been cleared out and someone has to go make one.

saying these in an interview costs you the question

  • Assumes more variants always means less total transfer
  • Ignores that each width is a separate cache object
  • Treats a cold derivative as merely a slower fetch
  • Forgets repeat visitors reuse assets across pages
  • Optimises fit before fixing the ladder's ceiling

context