skip to content

In a runtime-integrated microfrontend setup where the host loads a remote's JavaScript bundle by a fixed URL, such as a stable remoteEntry.js path, what cache-invalidation problem can arise, and what are the common strategies to avoid serving stale code?

level: seniorimportance: should knowfreq 55%

answer

  1. stable URL + long cache = stale bundle risk
  2. content-hash chunks, short-TTL manifest
  3. two-tier cache policy
  4. CDN propagation lag causes transient 404s

basics

~20 s

If the browser or a CDN caches the old file at that same web address, users can keep getting yesterday's broken code even after the team ships a fix, because nothing tells the cache the file changed.

solid answer

~40 s

Because runtime integration means the host fetches a remote's bundle by URL rather than embedding it at compile time, that URL becomes a caching hazard: if the manifest file and its chunks are served with long cache lifetimes at a stable path, browsers and CDNs can keep serving a stale, already-superseded version after a new deploy, so a fix that shipped isn't actually visible to users. The standard fixes mirror normal static-asset cache-busting: content-hash the filenames of the actual JS chunks so a new deploy is a new URL, cached forever safely, and keep only the small manifest/entry file on a short TTL or no-cache policy so the host always discovers the latest hashed chunk names. Some setups add an explicit version or deploy-timestamp query parameter to the entry URL itself for extra safety.

go deeper

for a junior

Can recognize that a cached old file is a possible reason a fix isn't showing up for some users.

for a middle

Knows the general fix is short-cache the manifest and long-cache the hashed chunks, even if fuzzy on why.

for a senior

Can explain precisely why content-hashing makes long caching safe and articulate the two-tier policy, plus name the 'works for me' support-ticket symptom.

for a principal

Can design the end-to-end deploy/cache/rollback policy for a multi-team runtime-integrated system, including CDN purge strategy and propagation-lag handling.

## Where the problem comes from The problem starts from how runtime integration necessarily works: the host doesn't have the remote's code baked into its own build, so at some point in the browser it issues an HTTP request for the remote's code, by URL. Everything downstream of that URL is subject to ordinary web-caching machinery — browser HTTP cache, any CDN edge cache in front of the origin, and sometimes intermediate proxies — none of which know anything about your deploy pipeline. If that URL is stable, the same manifest path on every deploy, which is the natural way to configure a host's remotes map so you don't have to redeploy the host every time, and it's served with a long cache header, common for JS assets because it's normally safe and great for performance, then after the checkout team deploys a new version, browsers and CDN edges that already have the old file cached will keep serving it, potentially for up to a year or until the cache entry happens to be evicted, even though the new code is live at the origin. Users experience this as 'we shipped the fix but it's not actually fixed for some users,' one of the most confusing classes of bug in a runtime-integrated system because the origin server is correct and the deploy pipeline reports success. ## Why the tension is genuine This is a genuine tension, not a bug in the tooling: long cache lifetimes on static JS assets are a core web-performance best practice, avoiding re-downloading unchanged code on every page view, but they only work correctly when combined with content-addressed URLs — the URL itself changes whenever the content changes, typically via a content hash in the filename. - **When that discipline holds**, you can cache the hashed chunk forever, safely, because a cache hit for that exact URL is guaranteed to be the exact content you want; a new deploy simply produces new hashed filenames that nobody has cached yet. - **The part of the runtime-integration setup that breaks this discipline** is the manifest or entry point itself, because by design its whole job is to be discovered at a stable, predictable URL so the host doesn't need redeploying to find it. That stable file is exactly the one piece that must not be cached long-term, or must be cache-busted on every deploy, because it's the pointer that tells the host which hashed chunk is current. ## The two-tier mitigation The standard mitigation, then, is a two-tier caching policy that matches the two-tier fetch a manifest-based setup already does: | Tier | Policy | Why | |---|---|---| | **Content-hash every actual code chunk** | serve those with aggressive, effectively-immutable caching | safe because the URL changes whenever the content does | | **The small entry or manifest file** | serve with a no-cache or very short max-age policy | so that every host page load re-checks with the origin or CDN for the latest manifest, which in turn points at the latest hashed chunks | Some teams go further and add an explicit deploy-timestamp or version query parameter to the entry URL itself so even a misconfigured CDN can't serve a genuinely stale entry file, since the URL literally changed. ## Failure modes Failure modes show up in a few characteristic ways. - **'Works for me, broken for some users'** is the most common — a support ticket says a bug is still happening after a confirmed deploy, and it turns out to be a subset of users whose browser or a specific CDN edge node still has the old manifest cached; this is hard to reproduce because it depends on cache state, not code state, and clearing your own browser cache fixes it for you but not for the affected user. - **Partial-deploy inconsistency** is another failure mode: because chunk files are content-hashed and immutable, if the manifest updates but an edge node hasn't yet propagated a newly-referenced chunk, a user can get a manifest pointing at a chunk URL that 404s for a few seconds — teams handle this with retry logic around federated dynamic imports. ## Where it shows up A well-known real pattern here: large single-page-app operators publicly describe the same content-hash-plus-short-TTL-manifest pattern for client bundle rollout, and it's the same principle runtime-integrated microfrontend setups borrow directly from general static-asset deployment practice — the microfrontend wrinkle is just that the manifest is now per-independently-deployed-team, so a stale-manifest bug can be scoped to one team's remote while the rest of the host is fine.

  • Why is it safe to cache a content-hashed JS chunk forever, but unsafe to do the same for the manifest file?
    A content-hashed chunk's URL is a function of its content, so the same URL can never legitimately point at different content — a cache hit is always correct, making infinite caching safe and free performance. The manifest's whole job, by contrast, is to live at a stable, predictable URL so hosts don't need redeploying to find the remote, which means its content does change over time at that same URL, so long caching there directly causes staleness.
  • A user reports a bug that was supposedly fixed yesterday, but only some users see the fix. What would you check first?
    First check the cache-control headers on the manifest/entry file and whether the fix actually shipped new content-hashed chunk URLs — if the manifest is cached with a long TTL, some browsers or CDN edges are likely still serving yesterday's manifest pointing at yesterday's chunks. I'd also check CDN cache-purge status for that specific file rather than assuming the deploy itself failed.

It's like a restaurant menu board that never gets refreshed even though the kitchen changed the recipe — customers keep ordering off the old menu and get confused when what arrives doesn't match what they expected.

saying these in an interview costs you the question

  • thinks a successful deploy means all users immediately see it
  • doesn't distinguish caching policy for the manifest vs the actual chunks
  • suggests fixing this by disabling all caching, killing performance rather than fixing the actual issue
  • can't explain why content-hashed filenames make long caching safe

context