skip to content

questions

5

Why would a team move static assets like images, CSS, and JavaScript bundles out of the application servers and into object storage fronted by a CDN, instead of serving them directly from the app tier?

level: juniorimportance: must knowfreq 85%

answer

  1. origin offload
  2. edge PoP cache hit/miss
  3. compute vs storage pricing
  4. geographic latency

basics

~20 s

Static files never change per request, so storing them in cheap storage and letting a CDN cache copies close to users is faster and cheaper than making app servers hand out those same bytes over and over.

solid answer

~40 s

App servers are expensive compute provisioned for running business logic — using them to stream static bytes wastes CPU, memory, and connection slots that could serve dynamic requests, and every request has to travel to wherever the origin lives. Moving assets to object storage (e.g. S3) fronted by a CDN offloads that traffic to infrastructure purpose-built for high-throughput file serving: edge points of presence cache the file near the user so most requests never reach origin at all. This decouples static-asset scaling from app-tier scaling, cuts origin bandwidth costs (CDN egress is typically cheaper per GB than compute egress), and improves latency for end users since bytes travel a shorter network path.

go deeper

for a junior

Should be able to say static files are cached closer to users and that this takes load off app servers, even without naming specific products or headers.

for a middle

Should name concrete technologies (S3/CloudFront-style storage+CDN), describe cache hit/miss at the edge, and mention cost/latency benefits.

for a senior

Should discuss the trade-off of eventual consistency, connect it to real deploy-time staleness issues, and describe cache-busting via versioned/hashed filenames as the standard mitigation.

for a principal

Should reason about this as a capacity-isolation decision (asset traffic isolated from business-logic capacity), discuss cost modeling of egress at scale, and know when the pattern doesn't pay off (e.g. tiny low-traffic internal tools).

## The mechanism The mechanism has two moving parts working together. 1. **First, the static files** — images, compiled CSS, JS bundles, fonts, videos — are uploaded to an **object storage** service (Amazon S3, Google Cloud Storage, Azure Blob Storage) rather than bundled into the application's deployment artifact or served from its local disk. Object storage is a flat key-value store for immutable blobs: durable, horizontally scalable, and priced for storage plus egress rather than compute time. 2. **Second, a CDN** (CloudFront, Cloudflare, Fastly, Akamai) is placed in front of that storage as a distribution. The CDN operates dozens to hundreds of **edge points of presence (PoPs)** around the world. When a client requests a file, DNS or anycast routing sends the request to the nearest PoP. - **Cache hit** — if that PoP already has a cached copy within its TTL, it serves the file directly, with no trip to origin at all. - **Cache miss** — if not, the PoP pulls the file from the object storage origin once, caches it according to the response's `Cache-Control` header, and serves it to the client. Every subsequent request to that PoP is now a hit. ## Why the app tier is the wrong tool This exists because application servers are the wrong tool for this job. - A typical app server process is provisioned around handling business logic — database queries, authentication, computation — and its capacity is measured in requests-per-second of that kind of work. - Streaming a 2MB JS bundle ties up a worker thread, a socket, and outbound bandwidth for the whole transfer, capacity that isn't then available for the requests the app tier actually exists to handle. - Worse, the app tier typically runs in one or a few regions, so every user, regardless of geography, pays the round-trip latency to reach it. Static assets don't need any of the app tier's request-handling machinery; they're just bytes that don't change per request. Splitting them out to storage-plus-CDN lets each system do only what it's optimized for. ## The trade-offs The trade-offs cut in both directions. - **On the cost side**, CDN egress and object storage are usually priced well below compute-tier egress, and offloading the bulk of asset traffic — often the majority of a website's byte volume — away from origin can meaningfully cut both bandwidth and compute-scaling costs. - **Latency** for cache hits is excellent because the edge PoP is topologically close to the user, versus routing every request back to a centralized origin. - **The cost is operational:** you now run two additional pieces of infrastructure — a storage bucket and a CDN distribution — each with its own configuration, access policy, and monitoring. You also trade strong consistency for **eventual consistency**: once a file is cached at the edge, updates to the origin copy won't be visible to users hitting that PoP until the TTL expires or a purge is issued. Deploying a new version of a static asset isn't instantaneous the way updating an app-server response would be. ## Failure modes in production That eventual-consistency trade-off is also the main source of failure modes in production. 1. **Stale-asset bugs after a deploy** — the most common. A team ships a new CSS file under the same filename, but users behind edge caches keep receiving the old bytes for hours until TTL expiry, causing visual breakage or, worse, an HTML page referencing a JS chunk that the old cached bundle doesn't contain (a classic 'chunk load failed' error during rolling frontend deploys). The standard fix is cache-busting — putting a content hash or version number in the filename itself so a new build produces a new URL, and setting a long, `immutable` `Cache-Control` on those hashed files while keeping the entry HTML uncached or short-TTL so it always points at the latest hashed assets. 2. **A cold-cache stampede** — a second failure mode. If a purge invalidates a popular asset across every PoP simultaneously, all of them miss at once and hammer origin concurrently, which can look like a self-inflicted DDoS. 3. **Misconfiguration** — a third. Forgetting CORS headers on cross-origin font or module loads causes silent browser-side blocking, and an overly permissive bucket policy can leak non-public files stored alongside legitimately public assets. ## Where it shows up A concrete, widely used pattern is a modern frontend build pipeline (webpack/Vite output) publishing hashed JS/CSS bundles to an S3 bucket behind CloudFront, with a long, `immutable` `Cache-Control` on the hashed assets and a `no-cache` policy on `index.html`, so browsers and the CDN cache the bundles essentially forever while every deploy is picked up immediately through the always-fresh entry document.

  • What specifically makes serving a video file from an app server worse than serving a JSON API response from it?
    A video file is large and takes a long time to stream, holding a worker thread/connection open for the entire duration and consuming outbound bandwidth proportional to file size and viewer count, whereas a JSON response is typically small and returns quickly. Because app server capacity is usually sized around concurrent connections and CPU for logic, not sustained bulk transfer, a handful of concurrent video downloads can starve the pool available for actual API traffic. Object storage and CDNs are built specifically to handle high-throughput, high-concurrency byte streaming without that contention.
  • If your app is only ever deployed in a single region with all users nearby, does this pattern still help?
    Yes, though the latency win shrinks — the bigger win becomes offloading bandwidth and connection load off the app tier so it can scale independently for business logic, plus cheaper egress pricing through the CDN/storage path versus compute egress. It also still buys resilience: a CDN can keep serving cached assets even if the origin app tier is temporarily down.
  • How does this pattern affect autoscaling of the application tier?
    Because static asset requests — often the majority of a page's total requests/bytes — never reach the app tier, its autoscaling metrics reflect actual business-logic load rather than being skewed by asset-serving spikes, e.g. a viral image doesn't trigger app-server scale-out. This generally lets teams run a smaller, more predictably-loaded app tier.

Like a bakery no longer hand-delivering every loaf from its one central kitchen — it stocks corner stores (edge caches) near neighborhoods so customers grab a loaf locally, and the kitchen only bakes a fresh batch to send out when a store runs out or a new recipe ships.

saying these in an interview costs you the question

  • Says static asset requests should go through the app server for 'consistency'
  • Doesn't mention that separating storage/CDN introduces cache staleness as a trade-off
  • Claims this pattern removes the need for any origin infrastructure
  • Thinks CDN and object storage are the same thing
  • No mention of latency/bandwidth as the core motivation

context

open as a page

When hosting static assets behind a CDN, how do Cache-Control headers and versioned (content-hashed) URLs work together to let you cache aggressively while still rolling out updates safely?

level: middleimportance: must knowfreq 80%

basics

~20 s

Cache-Control tells browsers and the CDN how long to keep a file before checking again. Putting a unique hash in the filename means a changed file gets a brand-new URL, so you can cache old URLs forever without ever serving stale content under a new version.

open as a page

A CDN caching static assets is configured to include a custom request header in its cache key, and shortly after, some users start receiving other users' cached error pages or unexpected content for a shared static URL. What kind of failure is this, and how does it happen mechanically?

level: seniorimportance: should knowfreq 45%

basics

~20 s

This is cache poisoning: the CDN stored a wrong or attacker-influenced response under a cache key that many different users then hit, so instead of one broken response for one weird request, everyone gets served the bad cached copy.

open as a page

A CDN serving static assets from many edge points of presence reports a cache hit ratio of only 40%, meaning most requests still reach the object storage origin. What mechanisms would you look at to raise that ratio, and what's the role of an origin shield in this?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Low hit ratio usually means too many different edge locations each independently asking the origin for the same file, or the cache expiring too fast. An origin shield adds one extra caching layer between the edges and the origin so only one edge has to fetch a file the first time, and every other edge gets it from the shield instead.

open as a page

Under what circumstances would offloading assets to object storage plus a CDN actually be the wrong call, and what would you use instead?

level: principalimportance: nice to knowfreq 40%

basics

~20 s

If content changes per user or per request, is accessed rarely, or must never leave certain servers or regions for legal reasons, caching it broadly on a CDN either doesn't help or actively causes problems, so you'd keep it dynamic or restrict it instead.

open as a page