Your CI pipeline creates a CloudFront invalidation for /* after every deploy, and the team deploys dozens of times a day. What problems does that create, and what would you do instead?
answer
- wildcard is cheap in dollars, costly in misses
- hit rate resets to zero every deploy
- new bytes should get a new URL
- only unversioned entry paths need purging
- deploy assets before the document
basics
~20 sInvalidating /* dumps the entire edge cache on every deploy, so the origin absorbs a burst of misses and users see uncached latency; frequent wildcard invalidations also queue against concurrency quotas. Ship content-hashed filenames and invalidate only the unversioned entry documents.
solid answer
~50 sA `/*` invalidation is cheap on the bill — a wildcard counts as one path — but expensive in behaviour. It discards every cached object, so the next wave of requests all miss and hit the origin at once: latency spikes, origin cost rises, and an origin that was comfortably shielded by a 95% hit rate suddenly sees full traffic. Do it dozens of times a day and the cache never warms up, plus in-progress wildcard invalidations run into a concurrency quota, so deploys start queueing behind each other. It is also unreliable as a correctness mechanism: propagation takes minutes and does not touch browser caches, so users can be served a new HTML document that references assets they still have old copies of. The durable fix is **versioning instead of purging** — content-hashed asset filenames with a long TTL, a short TTL or no-store on the HTML entry document, and an invalidation limited to that handful of stable paths.
code
bash · 15 linesset -euo pipefail
DIST=E1EXAMPLE12345
BUCKET=my-site-bucket
# 1. hashed assets first, long-lived
aws s3 sync ./dist/assets "s3://$BUCKET/assets" \
--cache-control 'public, max-age=31536000, immutable'
# 2. then the unversioned entry document, short-lived
aws s3 cp ./dist/index.html "s3://$BUCKET/index.html" \
--cache-control 'public, max-age=60'
# 3. purge only what could not change its name
aws cloudfront create-invalidation \
--distribution-id "$DIST" --paths '/index.html' '/sw.js'go deeper
Know that hashed asset filenames are why most deploys need no purge at all, and that the HTML file is the one that usually does.
Explain the mechanics you are trading: a wildcard purge empties the edge caches, so every subsequent request is a miss against the origin until the cache refills.
Demonstrate the production judgment — deploy order, which few paths still need purging, what the hit-rate and origin-request graphs look like across a deploy, and when a wildcard purge is genuinely warranted.
Set the release contract for the org: fingerprinted assets, an origin-governed entry document, and purging reserved for incidents — so cache behaviour is a property of the build system rather than a step someone can forget.
## Why `/*` on every deploy is the wrong default It looks like the safe choice — purge everything, nothing can be stale. In production it trades a rare correctness problem for a constant performance and reliability one. **The origin takes the whole load, repeatedly.** A CDN's value is the hit rate. Dumping the cache resets it to zero at each edge, so every popular object is refetched. With dozens of deploys a day the cache spends most of its life cold, and the origin is sized for CDN-shielded traffic it no longer receives. This is a self-inflicted thundering herd. **Users feel it.** The first request for each object after the purge pays full origin latency, including cross-region round trips. Your p99 gets a sawtooth that lines up exactly with deploy times — a signature worth recognising in a dashboard. **It queues.** CloudFront limits how many invalidation requests can be in progress at once, with a tighter limit on wildcard requests. A pipeline that fires `/*` on every merge eventually has invalidations waiting on invalidations, and the deploy step either blocks or reports success before the purge has propagated. **It does not even guarantee freshness.** Invalidation is asynchronous and reaches CloudFront's caches only. During propagation different edges can be in different states, and a browser holding a long-lived copy is untouched. So the failure it is supposed to prevent — a user seeing mixed old and new assets — can still happen. ## The pattern that replaces it The idea is that **new content gets a new URL**, so there is nothing to purge. 1. **Content-hash every static asset at build time.** `app.js` becomes `app.8f3a1c.js`. A new build produces a filename CloudFront has never cached, so its first request is a miss by construction, and the old filename keeps serving anyone mid-session. Give these a long TTL — the URL is a promise that the bytes never change. 2. **Keep the entry document unversioned and short-lived.** `/index.html` must keep its name, so it is the one thing that has to change in place. Serve it with a short lifetime, or with no-store if you want every load to consult the origin, and let the cache policy's Minimum TTL be 0 so the origin's intent is honoured. 3. **Deploy in the right order: assets first, then the document.** Upload the hashed files, confirm they are readable at the origin, and only then publish the HTML that references them. The reverse order guarantees a window where the document points at assets that 404. 4. **Invalidate the short list only.** `/index.html`, a service-worker script, maybe a manifest — the handful of stable paths. That is a few paths per deploy, well inside the free monthly allowance, and it purges almost nothing, so the hit rate survives. ```bash aws s3 sync ./dist/assets s3://my-bucket/assets --cache-control 'public, max-age=31536000' aws s3 cp ./dist/index.html s3://my-bucket/index.html --cache-control 'public, max-age=60' aws cloudfront create-invalidation --distribution-id E1EXAMPLE12345 --paths '/index.html' ``` ## Where `/*` is still legitimate As an **incident tool**, not a deploy step: a bad build that must stop being served now, content pulled for legal reasons, a misconfiguration that cached responses it should not have. Being deliberate about that keeps the operational cost where it belongs — occasional, not per-commit. ## Answering the interviewer's follow-through Expect "what if the app is server-rendered and there are no hashed filenames?" Then the honest answer is that most paths should not be long-cached at the edge at all: give dynamic documents a short TTL so staleness self-heals in seconds, and reserve edge caching with long lifetimes for the assets you *can* fingerprint. Blanket purging is a symptom of caching things whose freshness you cannot express. Expect also "how do you know it's working?" — watch the CloudFront cache hit-rate metric across a deploy. Under the versioned scheme the hit rate should dip slightly (new asset URLs) and recover in minutes; under `/*` it falls off a cliff every time. ## The service-worker caveat If the site registers a service worker, the worker script itself must never be long-cached, or clients keep running an old worker that serves old assets from its own storage. That script belongs on the same short-TTL, always-invalidated list as the HTML — an easy one to forget and a painful one to debug, because no amount of CDN purging fixes a stale worker on the client.
- Which files still need invalidating once every asset has a content-hashed filename?Only the unversioned, stable entry points: the HTML document, a service-worker script, and anything like a manifest whose URL cannot change. Everything hashed is new-by-construction — CloudFront has never cached that filename, so there is nothing to purge.
- Why must hashed assets be uploaded before the HTML that references them?Because the HTML is what points users at the new filenames. Publish it first and there is a window where viewers request assets the origin does not have yet, producing 404s that CloudFront may then cache as negative responses. Assets first, document last, makes the switch effectively atomic.
- How would you tell from metrics that the /* invalidations were hurting you?Look at the CloudFront cache hit-rate metric and origin request volume around deploy times: a wildcard purge produces a cliff in hit rate and a matching spike in origin requests and latency, repeating on the deploy cadence. Under versioned filenames you see a shallow dip that recovers within minutes.
saying these in an interview costs you the question
- Says /* is fine because wildcards are billed as one path
- Thinks invalidation makes a deploy atomic for all users
- Publishes the HTML before the assets it references
- Long-caches the service-worker script alongside other assets
- Treats purging as the primary freshness mechanism rather than TTLs