After putting a CDN in front of a site, image and script requests got much faster, but the HTML document's time to first byte barely moved. What explains that, and what would you do about it?
answer
- which requests can the edge answer alone
- the document was never cached
- only the handshakes moved closer
- origin generation time is still in the path
- short edge lifetime plus background refresh
basics
~20 sThe document is almost certainly uncacheable, so every request still travels to the origin and waits for the page to be generated. The edge only removes the connection-setup round trips. Fixing it means making the document cacheable at the edge, moving personalization out of it, or making the origin faster.
solid answer
~50 sStatic assets improved because the edge can answer them itself; the HTML did not because it is personalized or marked uncacheable, so each request is still a full origin fetch. What the CDN did buy on the document is real but small — the handshakes now terminate at a nearby PoP and the edge reaches your origin over an already-warm connection — while the edge-to-origin round trip and the origin's own generation time remain in the user's critical path. From there I would measure the split: if most of the time is the origin generating the page, no delivery change hides it. The options, roughly in order of payoff, are to make the document cacheable at the edge with a short lifetime and serve stale while refreshing in the background, to separate a cacheable shell from the per-user parts, to add a shield tier so all PoPs miss into one cache, and to attack origin render time directly.
go deeper
Understand that a CDN can only answer requests it is allowed to cache, and that a page built individually for each user still has to be produced by your server every time.
Explain what the edge still saves on an uncacheable request — the connection setup and a warm origin connection — and what remains: the origin round trip plus the time the server needs to generate the page.
Demonstrate the diagnosis: split network time from origin time, compare regions, then choose between making the document cacheable, extracting the personalized parts, or fixing origin render time, and say what each costs.
Own the call on cacheability of the main document across the product — which pages may be shared at the edge, what staleness the business will accept, and how personalization is architected so that answer is not permanently "no".
## Why the two behave differently Static assets are cacheable, identical for everyone, and requested repeatedly — the best possible case for an edge cache. The HTML document usually is not: it may be personalized, it may carry a session, and it is very often sent with headers that tell shared caches never to store it. Every request for it is therefore a miss by construction. ## What the edge still gives an uncacheable document It is not nothing, and being precise here is what separates a good answer from a shrug: - The user's DNS, TCP and TLS setup completes against a PoP a few milliseconds away instead of one across the world. That removes a couple of round trips of pure latency. - The edge typically holds warm, already-established connections to your origin, so the fetch itself skips connection setup. - The edge-to-origin leg often rides a better-provisioned path than the user's own network would. What it does **not** remove is the edge-to-origin round trip and the origin's own thinking time. If your server takes 400 ms to build the page, that 400 ms is in every user's TTFB, everywhere, forever. ## Measure the split before choosing a fix Separate the parts before spending effort: ```bash curl -s -o /dev/null \ -w 'tls=%{time_appconnect} ttfb=%{time_starttransfer}\n' \ https://example.com/ ``` A small `tls` with a large gap up to `ttfb` says the time is in the origin fetch, not in reaching the edge. Compare the same request from two regions: if TTFB is nearly identical worldwide, the dominant cost is origin generation, which is location-independent; if it climbs with distance, the miss path is the problem. Server-side timing on the origin closes the loop by telling you how much of that fetch was the application itself. ## The options **Make the document cacheable at the edge.** This is the big one, and it is a product decision more than a technical one. Even a short shared lifetime is transformative for a high-traffic page: a ten-second edge lifetime on a page requested a hundred times a second means the origin renders it once per ten seconds per cache instead of a thousand times. Pair it with serving a stale copy while the edge refreshes in the background, so no user ever waits on the revalidation fetch — the cost you accept is that some users see content up to a short interval old, and the benefit is that the origin is never in the user's path. **Move personalization out of the document.** If the only per-user parts are a name in the header and a cart count, the page does not have to be uncacheable. Serve a cacheable shell to everyone and fetch the small personal payload separately, or compose the personal fragment at the edge. Teams often discover that their HTML is uncacheable because of two lines of output. **Add a shield tier.** With a designated upper-tier cache that all PoPs miss into, one origin fetch fills the shield and every other PoP is served from it. This mostly protects the origin and helps long-tail traffic; it does not help if the document is uncacheable everywhere, since nothing is stored at any tier. **Speed up the origin.** Caching at the layer above is not a substitute for a server that takes half a second to render. And on any page that genuinely cannot be cached — an authenticated dashboard — origin work is the *only* remaining lever. **Reconsider the architecture.** If the page's content changes rarely but is rendered per request, pre-rendering it and serving it as a static object turns the problem into the asset case, which you already know works. ## Why the document matters more than its size suggests The HTML document is the request that gates everything else. The browser cannot discover stylesheets, scripts or the hero image until bytes of the document arrive, so a slow document TTFB delays every subsequent request by the same amount. That is why teams chase document TTFB even though the file is small, and why "we put a CDN in front of it" is only half an answer. ## What a strong answer sounds like Name the cause (an uncacheable document is still an origin fetch), be precise about the residual benefit (setup round trips and a warm origin connection), measure before prescribing (is the time in the network or in the origin?), and then pick a fix whose mechanism matches the measurement — rather than reaching for more edge locations, which changes nothing when the cost is origin render time.
- How would you tell whether the remaining document TTFB is network distance to the origin or the origin's own work?Compare TTFB from several regions and look at server-side timing. If the number is roughly the same worldwide, it is dominated by origin processing, which does not vary with location. If it grows with distance from the origin, the edge-to-origin leg dominates, and caching at the edge or a shield tier is the lever.
- What is the risk of caching the HTML document at the edge for even ten seconds?Serving a response to the wrong user or a stale one. Anything user-specific in the document — a name, a cart, an authenticated fragment — can leak to whoever gets that cached copy, so the page must be genuinely identical for everyone first. Beyond that, freshness slips by the lifetime you chose, which is usually acceptable for content pages and not for a checkout.
- Why is document TTFB worth chasing when the file is only a few kilobytes?Because it gates discovery. The browser cannot request stylesheets, scripts or the main image until document bytes arrive, so every downstream request is delayed by the same amount. Shaving 300 ms off the document typically shifts the whole waterfall left, not just that one response.
saying these in an interview costs you the question
- Assuming a CDN speeds up every response including uncacheable ones
- Adding more PoPs when origin render time is the bottleneck
- Caching a personalized document at a shared cache without checking
- Treating the edge as a substitute for fixing a slow origin
- Ignoring that document TTFB delays every later request