skip to content

Your team is upgrading a large Next.js App Router app to a version where server caching is opt-in rather than applied by default, and routes that used to be static now do work on every request. How do you decide what gets `'use cache'`, and how do you roll that out?

level: principalimportance: should knowfreq 24%

answer

  1. measure before you cache anything
  2. classify data, not routes
  3. shared first, per-user last
  4. tags before generous lifetimes
  5. policy outlives the migration

basics

~20 s

Treat it as a data-classification exercise, not a search-and-replace. Measure which routes actually regressed, cache the expensive user-independent work first with explicit lifetimes and tags, leave request-specific work uncached, and roll out route by route with staleness and hit-rate signals in place.

solid answer

~60 s

The trap is restoring the old numbers by sprinkling the directive everywhere, which converts a performance regression into a correctness risk. I would start by measuring: which routes actually got slower, and what dominates their server time. Then classify the data. Shared and user-independent work — catalogues, published content, navigation, reference data — gets `'use cache'` with an explicit `cacheLife` profile and a `cacheTag` so the write path can invalidate it. Tenant-scoped work is cacheable only with the tenant id as an argument, after authorization has already happened in uncached code. Genuinely per-user work stays uncached; the hit rate would be poor and every entry would be personal data with a retention question. Rollout is route by route, highest traffic first, with hit rate and a staleness signal visible before moving on. I would also treat the implicit defaults we were relying on as the real finding: the old behaviour was invisible in the source, and the migration is the moment to make the freshness policy for each dataset something a reviewer can read.

go deeper

for a junior

Understand that when caching becomes opt-in, previously fast pages can do real work on every request, and that the fix is adding explicit caching to specific data functions rather than a config switch.

for a middle

Be able to walk through classifying data as shared, tenant-scoped or per-user, and explain why the last group is usually left uncached rather than keyed by user id.

for a senior

Show the operational half: measure which routes regressed and what dominates their time, cache the dominant loaders with explicit lifetimes and tags, and verify with hit rate plus a staleness signal.

for a principal

Own the policy and the tradeoff — named profiles by data class, a review rule for what may enter a cache, a stopping point for the long tail, and a clear argument for why legible opt-in caching is worth a one-off migration.

## Reframe the problem The upgrade did not remove a feature; it removed a default. Work that was previously cached because the framework decided to is now uncached because nobody asked for it. The instinct — get the graphs back to where they were, quickly — points straight at the worst available action: adding `'use cache'` broadly and finding out later which entries were user-specific. A performance regression is visible, embarrassing and reversible. A cache serving one customer's data to another is none of those things. So the first move is to slow the loop down enough to classify, and the second is to sequence the work so the highest-value, lowest-risk cases land first. ## Measure before you cache Not every route regressed, and not every regression matters. Worth establishing before touching code: - Which routes changed rendering behaviour, and what their traffic share is. - Where server time actually goes on those routes — one slow upstream call usually dominates, and caching that one function recovers most of the loss. - What was previously being cached that nobody had noticed. Migrations of this kind routinely surface a query that was quietly served from cache for a year and is far more expensive than the team believed. This alone tends to shrink the work from "the whole app" to a handful of loaders. ## Classify the data, not the routes The decision is per dataset: **Shared and user-independent.** Product catalogues, published articles, navigation, pricing, reference data. Cache these first and generously — high hit rate, no confidentiality question, and the freshness requirement is usually measured in minutes. Give each an explicit `cacheLife` profile and a `cacheTag` so publication and edit paths can invalidate immediately rather than waiting on a clock. **Tenant-scoped.** A team's projects, an organization's settings. Cacheable, with the tenant id passed in as an argument so it forms part of the key, and only after authorization has happened in uncached, request-scoped code. The entry now holds one tenant's data, so invalidation on membership and permission changes must exist before this ships. **Per user.** Notification counts, personal drafts, anything derived from the individual. Leave uncached by default. Key cardinality equals the user count so the hit rate is poor, and each entry is personal data with logout, permission-change and deletion semantics attached. If a specific case genuinely needs it, it should be an argued exception with an invalidation plan, not a default. ## Sequence the rollout Route by route, highest traffic first, each change small enough to attribute. For each one: add the directive with an explicit profile and tag, wire the invalidation from the write path, ship, and watch two signals before moving on — the cache hit rate, and something that would reveal staleness (a support-ticket category, a content-team check, a synthetic that publishes and re-reads). A cache with no staleness signal is a cache nobody will trust enough to leave in place. A reasonable stopping rule: once the regression is closed on the routes that carry most traffic, stop. The remaining long tail is usually not worth the invalidation surface it adds. ## Institutionalise the policy The durable output is not the directives; it is the policy. Named `cacheLife` profiles declared once and applied by data class turn freshness into something a reviewer can read — `'catalog'`, `'article'`, `'pricing'` — instead of a number someone guessed at a call site. A short review rule ("nothing request-derived may enter a cached function; tenant-scoped entries need an invalidation path") catches the dangerous cases in code review rather than in production. ## What to say about the tradeoff It is worth being straight that opt-in caching is more work than opt-out, and that the framework made that trade deliberately. Opt-out caching gave good default numbers and an invisible policy — the reason teams shipped stale-data bugs they could not explain was that nothing in the source said what was cached. Opt-in caching costs a migration and ongoing diligence, and buys a codebase where the caching behaviour is legible at the call site. The migration is expensive precisely once; the legibility is permanent. That framing is usually what the question is really testing.

  • What would you say to a teammate who wants to add `'use cache'` at the top of every file in `lib/` to close the regression this week?
    That it converts a visible, reversible latency problem into an invisible, irreversible confidentiality one. File-level directives cache every export, including session-aware helpers added later. I would agree the goal and change the method: measure which loaders dominate, cache those explicitly, and get most of the win in the same week with a fraction of the surface.
  • How do you know afterwards whether the caching you added is actually working?
    Two signals per cached dataset. Hit rate answers whether the key granularity is sane — near-zero means the parameters carry too much cardinality. A staleness signal answers whether the lifetime is honest: a synthetic that publishes and re-reads, or a support-ticket category. Without the second one, nobody will trust the cache enough to keep it.
  • What is the argument that opt-in caching is worth the migration cost at all?
    Opt-out caching gave good default numbers and an unreadable policy — the source did not say what was cached, so stale-data bugs were unexplainable and per-user leaks were possible by accident. Opt-in makes the behaviour legible at the call site. The migration is paid once; the legibility is permanent.
  • Where do you stop, rather than caching the long tail of routes?
    Once the routes carrying most of the traffic have recovered. Each additional cached dataset adds invalidation surface and a staleness risk that someone must own, and the latency saved on a low-traffic admin page rarely pays for it. Leaving work uncached is a legitimate outcome, not an incomplete migration.

saying these in an interview costs you the question

  • Adds the directive broadly and audits afterwards
  • Treats it as a search-and-replace to restore the graphs
  • Caches per-user data to recover a hit rate
  • Sets lifetimes without any invalidation path
  • Ships with no staleness signal and calls it done

context