A news website publishes a breaking-news article, then five minutes later the editor fixes a factual error in the headline. Readers in different countries keep seeing the old, wrong headline for varying amounts of time after the fix. What CDN mechanisms would you use to get the correction out quickly and reliably, and what's the trade-off of using them aggressively for every edit?
answer
- purge forces a cache-miss on next request
- surrogate key / cache tag enables bulk purge
- propagation across PoPs takes seconds to tens of seconds, not zero
- purge storm = correlated miss burst back to origin
- prefer scoped purge over wildcard
basics
~20 sTell the CDN to throw away its cached copy of that page (purge/invalidate) so the next request fetches the fixed version from the real server. Doing this for every tiny edit is slow to spread everywhere and dumps extra load back on the server, so it should be used sparingly and precisely, not for every small change.
solid answer
~50 sUse the CDN's purge/invalidation API, targeted narrowly at the specific URL — or better, a cache tag (surrogate key) like article:12345 attached to every cached representation of that article, including AMP or mobile variants — rather than a broad wildcard purge. A targeted purge on a typical commercial CDN propagates to most PoPs within seconds to low tens of seconds, after which the next request per PoP is a cache miss that pulls the corrected copy from origin. The trade-off: purges aren't instant or free — propagation lag means some regions briefly serve stale content, purge requests are usually rate-limited, and firing one sends a correlated burst of cache-miss traffic back to origin across every PoP at once. Purging on every minor edit, or purging overly broad wildcards, is therefore expensive and can degrade origin health, which is why moderate TTLs plus purge-on-demand for urgent corrections is the usual balance.
go deeper
Should know that purging exists as a way to manually force-refresh cached content ahead of its TTL.
Should know the difference between scoped and wildcard purges and that purges take real, non-zero time to propagate.
Should design a purge strategy per content type — tags for CMS content, versioned URLs for build assets — and understand purge-storm risk on the origin.
Should own the invalidation strategy for a whole platform, balancing purge rate limits, origin protection during purge storms, and legal/compliance requirements for fast global takedown.
## What purging is **Purging** (also called **invalidation**) is the CDN mechanism for forcing cached content to be treated as expired before its TTL would naturally run out, so the next request becomes a cache miss and pulls a fresh copy from origin. CDNs typically expose a few flavors: - **Single-URL purge** — invalidate exactly one cached object. - **Wildcard or prefix purge** — invalidate everything under a path pattern, e.g. `/articles/*`. - **Tag-based purging** — using what Fastly popularized as **surrogate keys**: an identifier the origin attaches to a cached response (via a response header) that groups together every cached variant representing the same underlying content. An article page might have separate cached copies for desktop HTML, AMP, and a JSON API representation; tagging all three with `article:12345` means one purge-by-tag call invalidates all of them across the entire PoP fleet in a single operation, instead of the caller having to know and enumerate every URL variant. ## Why it exists This exists because TTL expiry alone is too blunt an instrument for content that changes unpredictably. These all need a way to force freshness on demand rather than waiting out a TTL that was chosen for the common case, not the emergency case: - **Editorial corrections.** - **Legal takedown requests.** - **Security-incident response** — pulling a compromised or leaked asset immediately. - **Configuration fixes.** Purge propagation is not instantaneous: a purge request goes to the CDN's control plane, which then has to fan the instruction out to every PoP holding a copy — for large fleets this can take anywhere from a few seconds to roughly a minute depending on the vendor and purge type (single-URL purges tend to be faster than broad wildcard ones). During that propagation window it is entirely normal — and often surprising to newcomers — for a reader in one region to already see the correction while a reader in another region still sees the old headline. ## The trade-off The central trade-off is how aggressively to rely on purging versus other invalidation strategies. Purging on every edit gives strong control over freshness but costs real infrastructure pressure: purge APIs are usually rate-limited precisely because a purge forces a wave of near-simultaneous cache misses back to origin — a correlated **mini stampede** that is worse the broader the purge's scope. - **An alternative** for content that changes on a predictable schedule is simply running a shorter TTL so staleness self-heals without any manual action, but that trades away cache-hit ratio and increases baseline origin traffic even when nothing changed. - **For build artifacts and static assets**, the versioned-URL approach (giving every changed file a new, content-hashed name) sidesteps purging almost entirely, but it doesn't work for content that must live at a stable, human-facing canonical URL, like a news article's permalink — you can't rename a news URL every time a typo is fixed. In practice, most platforms combine all three: long TTLs plus immutable, versioned URLs for build assets; moderate TTLs plus purge-by-tag for CMS-driven content; and short TTLs with no purging at all for content too cheap or too fast-changing to bother invalidating manually. ## Failure modes - **The most visible failure mode** is exactly the one in the scenario — inconsistent, geography-dependent staleness right after an edit, which is confusing to users and hard to reproduce for support teams unless they know to check which PoP served a given response. - **A more operationally dangerous failure mode** is a broad or wildcard purge triggering a **purge storm**: if a large fraction of a site's cached objects are invalidated simultaneously, every PoP's next request for each of those objects becomes a correlated cache miss, and if the origin isn't sized (or shielded) to absorb that burst, response times and error rates spike right when the team least expects it — right after what should have been a routine content fix. ## Where it shows up A concrete real-world pattern is a CMS publish/update webhook that automatically calls the CDN's purge-by-tag endpoint with the content's ID whenever an editor saves changes, so corrections propagate within seconds without anyone manually issuing purges through a dashboard. Security teams use the same mechanism defensively — instantly purging a leaked or compromised asset from every PoP worldwide as part of incident response, rather than waiting for a TTL to expire.
- What is a surrogate key (cache tag) and why is it useful for purging?It's an identifier the origin attaches to a cached response — for example, tagging every cached variant of article 12345 with that same tag — so a single purge-by-tag call invalidates every representation of that content across every PoP, instead of the caller having to enumerate and purge each URL individually.
- What operational risk does a broad wildcard purge introduce that a scoped, single-URL purge doesn't?It forces a large fraction of cached objects to miss simultaneously across the whole PoP fleet, sending a correlated burst of traffic back to origin (a purge storm) that can overload it, whereas a scoped purge only affects a small, known set of objects and produces a much smaller miss burst.
- If purges take up to roughly a minute to fully propagate, what read-your-write guarantee does the correction offer immediately afterward?Essentially none globally — it's eventually consistent; a user hitting a PoP that hasn't yet received the purge instruction can still see stale content for up to that propagation window, so genuinely time-critical corrections, like legal takedowns, often need a secondary safeguard beyond the purge alone.
Purging a CDN is like recalling a magazine issue from newsstands across a country: head office can radio every stand to pull the issue, but it takes real time for that call to reach every stand, and if head office recalls every issue for every typo, the delivery trucks never get a break.
saying these in an interview costs you the question
- Believes purge is instantaneous across every edge location worldwide
- Defaults to wildcard/full-site purge for every small content change
- Doesn't distinguish purging (manual force-refresh) from TTL-based expiry (automatic)
- Unaware that a purge causes a correlated cache-miss burst back to origin
- Confuses purging with cache eviction due to storage pressure at a PoP