You own a read-heavy public API fronted by a CDN, and the origin is saturated even though most requests ask for the same handful of things. How would you restructure the API's resources so a much larger share of traffic can be served from the edge?
answer
- measure hit ratio per route, weighted by origin cost
- split public / volatile / per-user resources
- key cardinality: canonicalize, constrain, cap
- byte-stable responses or validators churn
- staleness budget is a product decision
basics
~20 sSplit personalized and volatile data out of the hot public resources, make responses byte-stable, cut cache-key cardinality by canonicalizing and constraining query parameters, and set a per-resource staleness budget. Then measure hit ratio per route and iterate on the worst offenders.
solid answer
~60 sI would work in four passes. **1. Find where the misses are.** Hit ratio and origin requests per route, weighted by cost. Saturation with a small hot set almost always means either key cardinality is too high or the hot resources are not cacheable at all. **2. Decompose the resources.** Any endpoint that mixes a large public payload with a small personalized fragment is uncacheable end to end. Split it: a public resource cached hard at the edge, plus a tiny per-user resource the client composes with it. Same for volatile fields — live counters belong on their own short-lived resource, not embedded in the catalogue. **3. Shrink the key space.** Canonicalize query parameters (order, casing, defaults), strip tracking parameters at the edge, constrain supported sort/filter combinations, and use fixed page sizes. Unbounded combinations mean every key is requested once. **4. Set a staleness budget per resource** with the product owner, and make responses byte-stable so validators do not churn. Then let the edge absorb bursts via revalidate-while-serving and request coalescing. The tradeoffs I would state explicitly: bounded staleness, more client-side composition, and invalidation debt.
go deeper
Recognize that responses mixing per-user data into public payloads cannot be cached at the edge.
Explain resource splitting, canonical query parameters and why volatile fields in the body defeat both caching and revalidation.
Drive it from measurement — hit ratio and cardinality per route — and apply serve-stale-while-revalidating and coalescing to flatten origin load.
Frame it as a portfolio decision: negotiate per-resource staleness budgets, trade client composition for origin offload, and name the invalidation debt and cross-user leak risk you are taking on.
## Diagnose before restructuring "Most requests ask for the same handful of things" but the origin is saturated means the edge is not recognizing that sameness. There are only a few possible reasons, and they need different fixes: - The hot responses are **not storable** (personalized, or never labelled as reusable). - They are storable but the **cache key cardinality** is enormous, so identical content lives under thousands of distinct keys. - They are stored but the **lifetime is so short** that most requests arrive after expiry, so the edge revalidates against the origin constantly. - They are stored and reused, but a **stampede** on expiry sends a burst of identical requests to the origin. The measurement that separates these is hit ratio and origin fetches per route, weighted by origin cost rather than request count — one expensive aggregation endpoint at 0% can outweigh a million cheap hits. ## Decompose resources along their cache characteristics The single highest-leverage change is separating data by how it caches, not by how a screen is laid out. A "view model" endpoint built for one UI page typically bundles: mostly-static reference data, a volatile counter, and a per-user fragment. The union is as uncacheable as its worst component, so the static 95% pays the cost of the personalized 5%. Restructure into: - **Public, slow-changing resources** — large, shared by everyone, long lifetime, excellent edge hit ratio. - **Volatile resources** — small, short lifetime, cheap to regenerate. - **Per-user resources** — private, never in a shared cache, ideally tiny. The client composes them. The cost is more requests per screen — much less painful over HTTP/2 multiplexing than it was in the HTTP/1.1 era — and more client logic. The benefit is that the bulk of the bytes never touch the origin. This is a genuine tradeoff to state out loud: you are trading request count and client complexity for origin offload. ## Attack key cardinality A cache only helps when many requests map to one key. - **Canonicalize the query string** at the edge: sort parameters, normalize case and encoding, drop parameters that do not affect the response (tracking, correlation ids), and drop parameters that merely restate defaults. - **Constrain the combinatorics**: a filter API that accepts arbitrary field combinations, arbitrary sorts and arbitrary page sizes has a key space too large to warm. Supporting a documented set of sorts and a fixed set of page sizes converts a long tail into a warm head. - **Prefer coarse resources for hot paths.** Fetching one page of 50 that everyone requests beats 50 individual resources requested in varying subsets — though the opposite is true when access patterns are truly scattered, because fine-grained resources are reused across many different composites. The judgement is about the observed access distribution, not a rule. - **Watch `Vary`.** Every varying header multiplies entries. Varying on encoding is worth it; varying on a header that has hundreds of values is a hidden cardinality bomb. ## Make the responses stable Edge caching interacts badly with responses that differ byte-for-byte on every generation: server timestamps, request ids, non-deterministic ordering. They defeat revalidation (validators churn, so a revalidation always returns a full body) and they defeat any dedupe the edge might do. Deterministic serialization is a precondition for a high hit rate, and it is worth a test that fetches twice and compares. ## Set the staleness budget deliberately How stale may each resource be? That is a product question, and it is the actual lever on origin load: raising a hot resource from 5 seconds to 60 cuts origin fetches for it by roughly 12x. Ask per resource, get a number, and write it down as part of the contract. Two edge behaviours are worth designing around: serving a slightly stale copy while revalidating in the background (so users never wait on the origin), and coalescing concurrent misses for the same key into one origin request (so expiry does not produce a stampede). Both convert a spiky origin load profile into a flat one. ## Name the costs Be explicit that this is a trade, not a free win: - **Staleness** becomes a user-visible property; someone will report "I updated it and it did not change". - **Invalidation debt**: the more aggressively you cache mutable resources, the more you need a purge or versioning story — a separate design problem, and the one that most often gets deferred until an incident. - **Client complexity** from composition, and more moving parts to reason about when debugging. - **A correctness risk**: any personalization that leaks into a shared-cacheable resource is a cross-user data exposure, so the split between public and private resources must be structural and tested, not conventional. The end state to aim for is simple to describe: a small number of large, public, byte-stable, canonically-keyed resources carrying most of the bytes at a high hit ratio, and a thin private layer that never leaves the origin.
- Coarse-grained or fine-grained resources for a better hit rate?It depends on the access distribution. If most clients request the same composite, a coarse resource gives one hot key with a very high hit rate. If access is scattered, fine-grained resources are reused across many different composites while a coarse one would be requested once per unique combination. Measure the distribution rather than applying a rule.
- What breaks first when you push edge caching aggressively on a public API?Usually the invalidation story: something must change immediately, and there is no purge path or versioned naming to make it happen, so the team ships a lower TTL and loses the offload. The second failure is a personalized field quietly appearing in a shared-cacheable resource, which turns a performance win into a cross-user data exposure.
saying these in an interview costs you the question
- Treating cache hit rate as purely a CDN configuration problem with no API-design input
- Leaving one personalized field in an otherwise public response and expecting edge hits
- Accepting unbounded filter, sort and page-size combinations and then wondering why nothing stays warm
- Assuming a longer lifetime is free without agreeing a staleness budget with the product owner
- Scaling the origin instead of measuring per-route hit ratio and cache-key cardinality