A team that owns a checkout microfrontend wants to add a required new field to the payload of an event it emits (an event other teams' fragments already listen for). What does a safe versioning policy for rolling out this kind of change look like, from proposal to old-version retirement?
answer
- semver mental model for runtime contracts
- dual-emit old+new shape
- announce + deadline + telemetry
- only retire on zero usage, not calendar alone
- same discipline as backend API versioning
basics
~20 sDon't flip the switch all at once. Send both the old and new versions of the data for a while, tell other teams the old one is going away on a date, watch to see who's still using it, and only remove the old version once nobody is.
solid answer
~40 sTreat it like semantic versioning applied to a runtime contract: an additive-only change (new optional field) is a minor/patch bump and can ship immediately; adding a *required* field or changing an existing one is a major/breaking change and needs a deprecation cycle — emit both old and new shapes concurrently, announce the change and a removal date to known consumers (ideally discoverable via a contract registry, not tribal knowledge), instrument usage of the old shape so you can measure who's still on it, and only remove the old shape once telemetry confirms zero live consumers or the announced deadline passes with sign-off from remaining consumers. The policy exists because there's no shared deploy gate to force synchronized cutover — the old and new shapes must coexist safely for as long as any consumer needs.
go deeper
Can articulate the basic idea of 'warn people before you change something they depend on' without necessarily naming semver or dual-emit mechanics.
Describes the semver-style additive-vs-breaking distinction and the dual-emit/announce/deadline pattern for a breaking change.
Adds instrumentation/telemetry as the actual gate for retirement (not the calendar), and can trace a concrete field-level example end to end.
Designs the org-wide mechanism (contract registry, ownership metadata, telemetry tooling) that makes this policy enforceable across many teams rather than dependent on individual discipline.
## The one question a policy answers A versioning policy for a microfrontend contract answers one question precisely: for any given change, is it safe for the producing team to ship unilaterally, or does it require a negotiated transition with consumers who can't be forced to update on the same schedule? ## The semver mental model The cleanest mental model borrows directly from semantic versioning (semver), applied not to an npm package but to the runtime shape of props, events, and shared tokens: - **patch/minor changes** are additive and non-breaking (a new optional prop, a new event field nobody was relying on, a new event type); - **major changes** are anything that removes, renames, retypes, or changes the meaning of something a consumer might already depend on. Adding a *required* field to an existing event — the scenario here — falls squarely into "major," because any consumer still expecting the old shape either ignores the new field harmlessly (fine) or, worse, has downstream logic that now receives a payload it wasn't built to validate against (e.g. a schema check that now fails, or logic that silently misbehaves because the new field was expected to always be present). ## The rollout mechanics The rollout mechanics for a breaking change follow a predictable shape, and skipping any step is where teams get burned. 1. **First, propose and announce:** the field addition is documented — ideally in a shared contract registry or schema package, not a Slack message — with a clear rationale and a target removal date for the old shape. 2. **Second, dual-emit:** the producing team ships a version of the event that includes *both* the new required field and continues to populate whatever the old shape needed, so that consumers who haven't updated keep working exactly as before, while consumers who have updated can start relying on the new field. This is the step that actually buys safety — it decouples the producer's deploy from every consumer's deploy, letting each consumer migrate on their own schedule. 3. **Third, instrument and monitor:** log or emit a metric whenever the old shape's compatibility path is actually exercised (i.e. a consumer is still reading pre-migration semantics), so the team can see the graph of stragglers trend to zero rather than guessing. 4. **Fourth, retire:** only once usage of the old shape has genuinely dropped to zero (or the announced deadline has passed and remaining consumers have explicitly signed off, e.g. because they've deprecated their own fragment), the compatibility shim is removed and the dual-emit code deleted. ## The trade-off The trade-off this policy makes explicit is velocity versus safety, and it's a real cost on both sides. - The dual-emit period means carrying two code paths, more complex tests, and a real deferred cleanup task that's easy to let rot (the classic "temporary" migration shim that's still there two years later because nobody owns closing the loop) — that's the cost of safety. - The alternative, a hard cutover with an announced date and no compatibility shim, is cheaper to write but pushes all the migration risk onto consumers, forcing every dependent team to redeploy in a synchronized window — which is precisely the coordination tax the whole microfrontend architecture exists to avoid. - Skipping the policy altogether (just ship the breaking change) is the failure mode covered in the CI/CD independence discussion: it produces runtime breakage in someone else's fragment, discovered by users rather than tests. ## The failure mode the policy has of its own A failure mode specific to versioning policy itself, distinct from just "shipping a breaking change," is *silent* policy violation: a team correctly dual-emits and announces a deadline, but the removal date arrives and the shim gets deleted anyway because "the calendar says it's time," without actually checking the usage telemetry — and it turns out one consumer team was on vacation, missed the announcement channel, or the announcement never reached a newly-formed team that started depending on the old shape after the original announcement went out. This is why the instrumentation step matters more than the calendar date: the date is a target, the telemetry is the ground truth that gates the actual removal. ## A worked scenario A concrete, realistic worked scenario: a checkout fragment's `checkout:completed` event currently carries `{ orderId, total }`. The team needs to add a required `currency` field because the company is launching multi-currency support and `total` alone is now ambiguous. Rather than shipping `{ orderId, total, currency }` and assuming `currency` will just be ignored safely by old listeners (risky if any listener does strict schema validation), the safer sequence is: - ship `{ orderId, total, currency }` where `currency` defaults to `'USD'` if genuinely unset (additive, non-breaking for now), - announce that `currency` will become a required, always-populated field with no default in six weeks, - instrument which consumers are reading `currency` versus falling back, - and only remove the default once every known consumer confirms they read and handle the field. This mirrors exactly how backend teams version REST or event-driven APIs — the browser context doesn't change the discipline, just where the breakage shows up if you skip it.
- How do you handle a consumer team that misses the deprecation announcement entirely and only notices when the old shape is removed?This is exactly why usage telemetry should gate the removal, not just the calendar date — if the old shape's usage metric never dropped to zero, that's the signal to chase down the remaining consumer before removing anything, not after. A contract registry with a required 'owning team' field per consumer also helps, since it turns 'who's still using this' from a mystery into a queryable list.
- Is a deprecation window always necessary, or are there cases where a hard, unannounced breaking change is acceptable?It can be acceptable when the producing team can prove, via the contract registry or code search, that there are zero live consumers of the shape being changed — e.g. a fragment that's being sunset entirely, or an event type that was only ever used internally within one team's own components. The moment any other team's fragment is a known or even suspected consumer, skipping the window trades a small time saving for a real chance of an untraceable production incident.
It's like a utility company switching a whole city from one voltage standard to another: you don't flip the grid overnight — you run both voltages in parallel, tell every building owner the old one is being phased out by a date, monitor which buildings still draw the old voltage, and only decommission it once the meters confirm nobody's plugged in anymore.
saying these in an interview costs you the question
- Proposes just shipping the breaking change and telling other teams afterward
- No mention of a deprecation/dual-support window for a breaking change
- Treats a calendar deadline as sufficient without checking actual usage telemetry before removing the old shape
- Doesn't distinguish additive changes (safe immediately) from breaking changes (need a policy)
- Assumes all consumer teams will see and act on an announcement made once