skip to content

Explain surrogate-key (cache-tag) based purging in a CDN such as Fastly or Varnish, and how it differs from purging by URL.

level: seniorimportance: should knowfreq 42%

answer

  1. Origin tags responses with entity keys
  2. Cache keeps key → objects reverse index
  3. One purge evicts all variants, any URL
  4. Entity keys + collection keys + structural keys
  5. Soft purge avoids origin stampede; keep a TTL backstop

basics

~20 s

The origin tags each response with surrogate keys naming the data it contains. The CDN indexes cached objects by tag, so one purge of a key evicts every representation containing that data — regardless of URL, query string or variant — instead of you enumerating URLs.

solid answer

~60 s

With **URL purging** you must know every URL that embodies a changed entity. That set is usually unknowable: the same product appears in `/products/42`, in `/products?category=shoes&page=3`, in `/search?q=running`, in a homepage payload, and in every query-string and content-negotiated variant of those. **Surrogate keys** invert the problem. The origin emits a header — `Surrogate-Key: product-42 category-shoes catalog` in Fastly, `xkey`/tag headers in Varnish — listing the entities a response depends on. The CDN keeps a reverse index from key to cached objects. When product 42 changes, the origin issues one purge for `product-42` and the CDN evicts every object carrying that tag, across all URLs and variants, typically in well under a second globally. The design work is choosing the key vocabulary: an entity key per record, plus coarse keys for collections and cross-cutting concerns. Keys must be emitted by whatever code builds the response, so they stay accurate as endpoints evolve. Watch key cardinality per response and make purges idempotent and retryable, because a lost purge is unbounded staleness.

code

http · 4 lines
http
HTTP/1.1 200 OK
Content-Type: application/json
Cache-Control: public, s-maxage=3600, stale-while-revalidate=60
Surrogate-Key: product-42 product-88 category-shoes products

go deeper

for a junior

Know that a CDN can tag cached responses and purge everything with a given tag, rather than purging one URL at a time.

for a middle

Explain the tag header, the reverse index, and why URL purging fails once query strings, pagination and variants multiply.

for a senior

Design the key vocabulary, discuss cardinality, soft versus hard purge, purge reliability via retries or an outbox, and keep a finite TTL as a backstop.

for a principal

Position it as the thing that decouples freshness from TTL — long edge retention with sub-second write visibility — and own the invariant that read-path tagging and write-path purging come from one shared mapping.

## The problem URL purging cannot solve Invalidation by URL requires the writer to enumerate every cached URL affected by a change. In a real API that set is combinatorial and often unknown at write time. Change one product and you have potentially invalidated: its own detail representation, every paginated collection page that contained it, every filtered and sorted variant of those collections, search results, aggregate or summary endpoints, and — multiplying all of the above — every content-negotiated or parameterised variant the cache stores separately. Nobody can maintain that list correctly as the API grows, and the failure mode is silent: forgotten URLs keep serving stale data until their TTL expires. ## How surrogate keys work The mechanism is a reverse index maintained by the cache. 1. **Tagging.** When the origin builds a response, it attaches a header listing the logical entities the response depends on. In Fastly this is `Surrogate-Key: <space-separated keys>`; Varnish achieves the equivalent with the `xkey` module or a custom tag header plus VCL; other CDNs use "cache tags" with the same shape. 2. **Indexing.** The cache stores the object and records, for each key, that this object carries it. The key header itself is stripped before the response is sent to the client — it is metadata between origin and cache, not part of the public contract. 3. **Purging.** On a write, the origin issues a purge for the affected keys. The cache looks up the index and evicts (or marks stale) every matching object, in every point of presence, without needing any URL. Because the index is per-object, one purge can evict thousands of variants, and it does so regardless of how the URL was formed. ## Designing the key vocabulary This is where the engineering judgement lies. - **Entity keys** — one per record touched: `product-42`, `user-9f2`, `order-771`. Emitted by whatever code loaded that record into the response. If the response embeds ten products, it carries ten entity keys. - **Collection keys** — a coarse key such as `products` or `category-shoes` attached to any response whose membership could change: listings, search results, counts. A create or delete purges the collection key even though no single entity key covers "the set changed". - **Structural keys** — for changes that affect rendering rather than data: a template or schema version, a feature-flag epoch. Purging one key can flush an entire class of responses after a deploy. Two rules keep this honest. First, **the tag must be emitted by the code that reads the data**, ideally automatically in the data-access or serialization layer, so a new endpoint that loads a product inherits the right key without anyone remembering. Manually curated key lists in controllers rot within months. Second, **write paths must purge every key their change implies**, which means the mapping from "entity changed" to "keys to purge" belongs in one shared place, not scattered across handlers. ## Operational concerns **Cardinality.** Keys per response cost index memory, and enormous key sets on a single object are a known way to hurt a cache. A collection page listing 100 items with 100 entity keys is usually fine; a response carrying thousands is a smell suggesting you should tag it with a collection key instead. **Purge reliability.** A purge is a network call that can fail while the database write has already committed. Treat purging as an at-least-once side effect: retry it, or drive it from a durable outbox or change-data-capture stream rather than inline in the request. Purges are naturally idempotent — evicting an already-evicted object is harmless — which makes retries safe. **Soft purge.** Many CDNs distinguish hard purge (evict immediately; next request goes to origin) from soft purge (mark stale; keep serving under `stale-while-revalidate` or `stale-if-error` while refreshing). Soft purge is usually the better default for high-traffic content because it avoids a synchronous origin stampede at the moment of a write — precisely when the origin is already doing work. **Ordering.** A purge issued before the write is visible to readers can re-cache the old value. Purge after commit, and be aware that with read replicas the origin may still serve pre-write data for a moment; a small delay or a read-your-writes strategy on the refresh path avoids re-poisoning the cache. **Still keep a TTL.** Surrogate keys reduce but do not eliminate the need for expiry. Keep a finite `s-maxage` so a lost purge is a bounded incident rather than permanent staleness. ## Why it changes the caching economics Once invalidation is precise, you no longer have to buy freshness with a short TTL. You can hold content at the edge for hours and still reflect writes in under a second, which is the combination expiry-only caching cannot deliver. That is the actual argument for the extra machinery.

  • What is the difference between a hard purge and a soft purge, and when do you prefer each?
    A hard purge evicts the object outright, so the next request is a miss and goes to the origin. A soft purge marks it stale, so the cache can keep serving the old copy under stale-while-revalidate or stale-if-error while it refreshes in the background. Soft purge is the better default for high-traffic content because it avoids a synchronous origin stampede at the exact moment of a write; hard purge is right when serving even briefly stale data is unacceptable, such as after removing content for legal or privacy reasons.
  • How do you make sure surrogate keys stay accurate as the API grows?
    Emit them automatically from the layer that loads data, rather than curating lists in controllers. If the repository or serializer records every entity it touched during a request and the framework assembles those into the header, any new endpoint inherits correct tagging for free. The complementary half is a single shared mapping from a domain change to the keys it purges, so write paths cannot drift from read paths.
  • A purge call fails after the database write has already committed. What should the system do?
    Treat purging as an at-least-once side effect rather than a best-effort inline call. Retry it, and for durability drive purges from an outbox or change-data-capture stream so a crash between commit and purge does not lose the invalidation. Purges are idempotent — evicting an already-evicted object is harmless — so retries are safe, and a finite s-maxage remains the backstop that bounds staleness if every attempt fails.

URL purging is deleting library books by shelf position; surrogate keys are a catalogue card per title that lists every shelf it sits on, so one lookup finds them all.

saying these in an interview costs you the question

  • Believing you can enumerate all affected URLs after a write in a real API.
  • Tagging only the entity and forgetting collection keys, so listings stay stale after a create or delete.
  • Curating key lists by hand in controllers instead of emitting them from the data layer.
  • Treating a purge as guaranteed and dropping the TTL backstop entirely.
  • Issuing the purge before the write is committed and visible, which lets the old value be re-cached.

context