After a content edit, the instance that handled the write serves the new page while other instances and the CDN still serve the old one. How do you diagnose and fix it?
answer
- the instruction landed where it arrived
- enumerate every place a copy exists
- ask each instance, then the public URL
- two layers, two purges
- application store first, cache in front second
basics
~20 sThe purge ran in one process and cleared only what that process holds. Locate each stale copy by comparing instances against the public URL, then share the cache store and purge the cache in front as a second explicit step.
solid answer
~50 sTwo independent facts produce this. First, if cached pages live in each instance's memory, a purge is executed by whichever instance received the write and clears nothing anywhere else; the peers keep their own copies until their own lifetimes run out. Second, a CDN holds a separate copy keyed by URL that no code inside the application can touch. Diagnose by separating the layers: request each instance directly and compare, then request the public URL and look at how old the response says it is. Divergence between instances points at per-instance memory; agreement between instances with a stale public URL points at the CDN. The fix is structural, not a bigger purge: one shared store so a purge is a single write the fleet observes, plus an explicit CDN purge on the publishing path, ordered so the app store is clean before the CDN refetches.
go deeper
Take away that a purge only clears the cache it can reach. There is more than one place a page is stored, so seeing the new page yourself does not mean everyone does.
Be able to list the layers holding a copy and explain why per-instance memory makes a purge local. Describe how you would tell an instance-level problem from one in the cache in front.
Drive the diagnosis by isolating layers, then propose the structural fix: a shared store so a purge is fleet-wide, an explicit purge of the cache in front, correctly ordered, on a single publishing path.
Define the publishing contract: how quickly a correction must be visible, which lifetimes that permits, who owns the purge path, and how it is verified from outside rather than trusted.
## Read the symptom literally The report contains its own diagnosis: *one* instance is right and everything else is wrong. That is the signature of an instruction that was executed **where it arrived** rather than somewhere every reader can see. A write request lands on exactly one process, and if that process's caching side effects are local to it, so is the result. Layer that on top of a second, entirely separate fact — that a cache in front of the application holds its own copy of the response — and you get the classic publishing failure: the edit is live, the author can see it because their session happened to reach the right instance, and most of the world cannot. ## Where the copies can be Before touching anything, enumerate the places a copy of that page can exist: 1. **Each instance's process memory**, if the framework's cache lives in process. 2. **A store shared by the instances**, if one has been configured. 3. **The CDN or other shared cache in front**, one stored response per URL, ageing on its own lifetime. 4. **The visitor's own browser**, per whatever the response's headers permitted. A purge triggered inside the application can reach 1 (only where it ran) and 2. It cannot reach 3 or 4 unless someone explicitly makes a call to the CDN or the response's lifetime expires. ## How to prove which layer it is The method is to take the layers apart and ask them separately: - **Ask each instance directly**, bypassing the load balancer, for the same path. If they disagree with each other, the application's cache is per-instance and the case is closed for layer 1. - **Ask the origin once more after the purge.** If the origin is now uniformly fresh but the public URL is stale, the copy is in front of the application. - **Look at how old the answer claims to be.** A response served from a shared cache reports its `Age`, and an age that keeps growing across requests is a stored copy being replayed rather than the application answering. - **Change nothing but the key.** Requesting the same page with an extra query parameter usually forces a cache in front to treat it as a new entry and fetch from the origin; if that version is correct, everything downstream is stale by definition. Remember this is a diagnostic, not a fix, because it also populates a new entry. Do not stop at the first confirmed layer. Both problems commonly exist at once, and fixing only the visible one produces a second incident an hour later. ## The fix, in the order it matters 1. **Make the application's cache one copy.** Move entries to a store every instance reads and writes. A purge then becomes a single write that the whole fleet observes, which removes the entire class of per-instance drift, not just this symptom. 2. **Put the CDN purge on the publishing path.** The path from a write to a correct page has two stops, and the code that performs them belongs in one place that runs after the write has committed, not scattered across handlers. 3. **Order them: application store first, CDN second.** They are not atomic. If you clear the CDN while the application still holds the old entry, the very next request repopulates the CDN with stale content and you have purged nothing. Clearing the application first means any refetch during the window fills the CDN from a correct origin. 4. **Bound what you cannot purge.** Browsers already hold copies and no purge reaches them; shared lifetimes you cannot remove on demand must be short enough that the worst case is acceptable. 5. **Verify from outside.** Publishing is not done when the write commits; confirm the public URL, not a direct origin request, and keep that check in the flow that matters. ## What separates a senior answer Junior answers reach for a longer purge or a shorter lifetime. The senior answer names **two independent layers with two owners**, insists that the application's cache stop being process-local so a purge means something fleet-wide, and treats the ordering of the two purges as a correctness property rather than a detail. It also acknowledges the residue: copies already in browsers, and the small window between the two steps in which someone can still be served the old page.
- Why purge the application's store before the cache in front, and not the other way around?Because the two steps are not atomic. Clearing the cache in front first means the next request refetches from an application that still holds the old entry, so the stale content is immediately restored and the purge accomplished nothing. Clearing the application first guarantees any refetch during the window pulls a correct response.
- What residue is left even after both layers are purged correctly?Copies already held by browsers, which no server-side purge can reach, and anything served during the gap between the two steps. Both are bounded by the lifetimes previously sent, which is why a response with a long private lifetime is a commitment. If a page must be correctable within minutes, its headers have to say so.
- The team suggests sticky routing so the author always hits the instance that wrote the change. Good idea?No. It hides the symptom for the author and leaves every other visitor sampling divergent caches, while adding a routing constraint that makes deploys and instance loss messier. It treats a consistency problem as a load-balancing problem. Sharing the cache fixes the cause; pinning traffic only moves who notices.
- How would you catch this class of failure before an author reports it?Check from outside, the way a visitor does. After a publish, request the public URL and confirm the change is present and the response is not being replayed from an old stored copy. An automated check on the public URL after each publish turns a support report into an alert, and its failures point straight at whichever layer was missed.
Correcting the notice pinned in your own office changes nothing about the copies already pinned up in every other building, or the one in the lobby everybody actually reads.
saying these in an interview costs you the question
- Assuming one purge call clears every instance
- Believing an application-side purge reaches the cache in front
- Purging the cache in front before the application's own store
- Suggesting sticky routing so the author sees their own write
- Declaring it fixed after verifying against a single instance
- Treating a shorter lifetime as a substitute for a purge path