skip to content

A self-hosted Next.js App Router app runs three container replicas behind a load balancer. An ISR page shows fresh data on some refreshes and stale data on others, and a `revalidateTag` call after a mutation seems to work only sometimes. What is happening, and how do you fix it?

level: seniorimportance: should knowfreq 46%

answer

  1. state lives per container
  2. the load balancer picks the version you see
  3. invalidation reached one instance only
  4. one store all replicas can read
  5. and the same build everywhere

basics

~20 s

Each replica keeps its own cache on its own filesystem, so the three containers hold independent copies of the page and revalidation only invalidates the replica that handled the request. Fix it by pointing all replicas at one shared cache through a custom cache handler.

solid answer

~50 s

Nothing is wrong with the revalidation logic — it is a topology problem. By default a self-hosted Next server stores cached route output and fetch results in its own container filesystem, plus an in-memory layer. Three replicas therefore hold three independent caches, so which version you see depends on which container the load balancer picked. When a Server Action calls `revalidateTag`, that call executes on exactly one replica and invalidates only its cache; the other two keep serving their own copies until their own timers expire. Container restarts and scale-outs make it worse, since a fresh instance starts from the build-time output. The fix is a shared cache: set `cacheHandler` in `next.config` to a module implementing get, set and tag revalidation against Redis, S3 or a database, and set `cacheMaxMemorySize` to zero so no replica answers from a private in-memory copy. All replicas must also run the same build.

go deeper

for a junior

Know that a self-hosted Next server caches pages on the machine it runs on, so several replicas can hold different versions of the same page.

for a middle

Explain why on-demand revalidation only affects the instance that handled the request, and name the shared cache handler as the mechanism that gives every replica one source of cached content.

for a senior

Show the diagnosis path — instance identifiers in responses to prove the split before touching revalidation code — then the fix, including disabling the per-instance memory cache and keeping builds consistent across replicas.

for a principal

Own it as a platform requirement: decide the shared cache store and its failure behaviour, define what happens when it is unreachable, and set the deploy strategy so two builds never serve mismatched assets.

## Where the cache lives when you self-host A Next server caches rendered route output and fetched data so it does not redo work on every request. Self-hosted, that cache is local: files under the app's cache directory in the container's own filesystem, fronted by an in-memory layer for speed. On a single instance this is invisible and works fine. The moment you scale to more than one instance it becomes the dominant source of "why is the data wrong sometimes". ## Why replicas diverge The load balancer sends each request to whichever replica it likes. Replica A regenerated the page thirty seconds ago; replica B has not regenerated since the deploy; replica C restarted five minutes ago and is back to the build-time output. All three are behaving exactly as designed, and the user sees whichever one answered. Refreshing gives a different answer because it hits a different container. On-demand revalidation makes the split visible in a way people notice. A Server Action that calls `revalidateTag` or `revalidatePath` runs on a single replica — the one that received the POST. It invalidates that replica's entries. The other two were never told anything happened. So the user who submitted the mutation often sees the correct fresh page (their request stuck to the same instance) while everyone else sees stale content, which reads as an intermittent bug rather than a systemic one. Two more effects compound it: an ephemeral container filesystem loses the cache on restart, and horizontal scale-out adds a cold instance that must warm up from scratch — so the number of divergent versions grows with your autoscaling. ## Confirming it before fixing it Do not start by rewriting revalidation code. Prove the topology first. Set a per-instance identifier as an environment variable and emit it, plus a render timestamp, in a response header or a debug element. Then request the page repeatedly and correlate: if the stale and fresh responses cleanly partition by instance id, you have per-instance caches, not a bad `revalidate` value or a misapplied tag. This distinction matters because the two causes have completely different fixes and the symptoms look identical from a browser. ## The fix: one cache for all replicas Next supports replacing the default cache implementation. Point `cacheHandler` at a module that exports a class implementing the cache interface — `get`, `set` and tag revalidation — backed by a store every replica can reach, typically Redis. Also set `cacheMaxMemorySize` to zero, otherwise each replica may still answer from its own in-memory copy and miss an invalidation that already landed in the shared store. ```js // next.config.js module.exports = { cacheHandler: require.resolve('./cache-handler.js'), cacheMaxMemorySize: 0, } ``` This is the piece a managed platform was providing for you: a cache shared across all the machines serving your app, with invalidation that propagates. The reason the problem appears on day one of self-hosting is that it was never your code's responsibility before. A version note: the option is named `cacheHandler` in current Next.js; in versions before 14.1 the equivalent was `incrementalCacheHandlerPath`. ## The other half: build identity A shared cache assumes the replicas are the same application. If a rolling deploy leaves old and new containers running side by side, they have different build ids and different asset filenames, so a document served by a new replica can request assets from an old one that no longer has them, and cache entries written by one build may not be meaningful to the other. Treat this as part of the same problem: deploy atomically where you can, keep the previous version's static assets available during the rollover, and serve those assets from a CDN or shared object storage rather than from whichever container happens to answer. ## What to say in an interview Lead with the diagnosis — per-instance caches and a revalidation that only reached one instance — because it shows you understand where state lives. Then name the fix as a shared cache handler plus disabling the per-instance memory cache, and close with the deploy-identity caveat. Candidates who jump straight to "increase the revalidate time" or "add sticky sessions" are patching the symptom; sticky sessions in particular make the bug less visible while leaving every non-sticky client, including crawlers, seeing the same divergence.

  • Why not just enable sticky sessions on the load balancer?
    It hides the symptom for one user and leaves it for everyone else. A user's own requests then land on the replica that handled their mutation, so they see fresh data, while every other visitor and every crawler still gets whichever replica answered them. It also couples correctness to session affinity, which breaks on scale-out, restarts and cache-warming.
  • What does setting `cacheMaxMemorySize` to zero buy you once a shared cache handler is in place?
    It stops each replica from answering out of its own in-memory copy. Without it, an entry invalidated in the shared store can still be served locally by an instance that already has it resident, so invalidation looks unreliable again. Zero forces every read through the shared handler, which is the behaviour you actually wanted.
  • What breaks if two different builds of the app are serving simultaneously during a rolling deploy?
    They have different build ids and different hashed asset names, so a document rendered by a new replica can request assets an old replica cannot serve, producing 404s and failed hydration for the users caught mid-rollover. Serving static assets from a CDN or shared storage that holds both versions during the overlap removes most of it.
  • Where does this problem go on a managed platform?
    The platform provides a single cache shared by every machine serving the app, with invalidation propagating across them, so a revalidation triggered anywhere is visible everywhere. Nothing about your code changes — which is exactly why the failure surfaces on the first self-hosted deploy of an app that worked fine before.

saying these in an interview costs you the question

  • Blames the revalidate value instead of per-instance caches
  • Proposes sticky sessions as the fix
  • Assumes revalidateTag broadcasts to every running instance
  • Thinks the ISR cache lives in the database by default
  • Expects a container restart to preserve regenerated pages

context