When is publishing a v2 data mart alongside v1 better than migrating consumers in place?
answer
- can the serving layer hide it?
- who pays the migration cost, one team or thirty?
- the second version rarely dies on its own
- what makes divergence impossible rather than monitored?
- name the date before you publish
basics
~20 sVersion in parallel when the change is structural or semantic, the consumer set is large or partly unknown, and each team must migrate on its own schedule. Only do it with a named owner, a sunset date, and v1 derived from v2 so the two cannot diverge.
solid answer
~50 sParallel versions buy consumer-paced migration and a side-by-side reconciliation you can show stakeholders; they cost double compute and storage, double the surface for bugs, and — the real risk — they tend to become permanent. Publish a v2 when the change is one that cannot be absorbed behind a view: a different grain, a different key strategy, a materially different set of measures. Absorb it in the view or semantic layer instead when the rows are the same and only names or shapes are moving, which is far cheaper. If you do version, set the conditions up front: a named owner, a dated sunset agreed before launch, a cap of two live versions, usage telemetry so you can see who is left, and wherever possible v1 rebuilt as a derivation of v2 so there is one pipeline rather than two that drift. Versioning by default is a smell; versioning as the exception, with an expiry, is engineering.
code
sql · 8 lines-- one pipeline, two contracts: v1 becomes a compatibility surface
CREATE OR REPLACE VIEW fct_orders_v1 AS
SELECT order_number,
MIN(order_date) AS order_date,
MAX(customer_sk) AS customer_sk,
SUM(extended_price + allocated_shipping) AS order_total
FROM fct_order_line_v2
GROUP BY order_number;go deeper
Know that a mart people depend on cannot simply be replaced, and that one option is to publish a new version next to the old one so teams can move when they are ready.
Explain the tradeoff: parallel versions cost double compute and double maintenance and can drift, while an in-place change forces every consumer to move at once.
Show which changes genuinely need a parallel version — grain, key strategy, materially redefined measures — versus what a serving view can absorb, and how a reconciliation harness gets cautious consumers to accept the new one.
Own the policy: a cap on live versions, mandatory sunset date and owner before publication, telemetry-driven retirement, and deriving the old version from the new so the platform never carries two divergent pipelines for the same facts.
## The decision is about blast radius, not about elegance Every change to a published mart is on a spectrum from "absorbable" to "a different thing wearing the same name". The organisational question is who absorbs the cost: the producer, once, behind an indirection layer; or every consumer, on their own schedule, with a parallel version to migrate onto. ## Absorb it in the view or semantic layer when you can If the underlying rows and their meaning are unchanged and only the presentation moves — a rename, a reordering, a column split that is derivable, a table reshaped underneath — you do not need a v2 anything. Change the base model, keep the view's output identical, and consumers never learn about it. This is by far the cheapest option and it should be the default attempt. Two live physical versions is what you fall back to when the view cannot preserve the contract. ## Version in parallel when the change is genuinely a different object The cases that justify it: - **The grain changes.** No view over an order-line table reproduces order-grain semantics for every possible query, only for the aggregates you thought of. - **The key strategy changes.** Surrogate keys are reissued, or a dimension becomes history-bearing, so persisted keys and saved filters held outside the warehouse stop pointing at what they used to. - **The measure set changes materially.** Several measures are redefined at once and consumers need time to validate each against their own reconciliations. - **The consumer set is large, cross-organisational, or partly unknown**, and a flag-day cutover would mean coordinating dozens of teams into one window. - **Reconciliation is a requirement.** Regulated or finance-facing consumers often must run old and new side by side for a period and prove agreement before they will accept the new one. Parallel publication is the only way to give them that. ## What parallel versions actually cost - **Compute and storage, roughly doubled** for the overlap period. - **Two places for a bug fix.** Every correction must be applied to both, or the versions diverge and the divergence itself becomes an incident. - **Cognitive load.** New consumers pick the wrong version; documentation must say which one is current in every place it appears. - **Permanence.** This is the failure mode that matters. Without a forcing function, v1 lives forever, a third version eventually appears, and the platform accumulates a museum. The half-migrated estate is worse than either endpoint because every question now needs the qualifier "in which version?". ## The preconditions that make it work 1. **A sunset date agreed before v2 launches**, not discovered afterwards. If you cannot name the date, you are not versioning, you are forking. 2. **A named owner** accountable for retiring v1, with the migration on their roadmap rather than in their backlog. 3. **Usage telemetry** from query history, so "who is still on v1" is a query rather than a survey — and so the sunset can be evidence-based. 4. **A cap of two live versions.** A v3 while v1 is still up is a policy failure; retire before you add. 5. **v1 derived from v2 wherever the semantics allow.** Rebuilding v1 as an aggregate or a view over v2 collapses two pipelines into one, makes divergence structurally impossible, and turns v1 into a compatibility surface rather than a second product. 6. **A reconciliation harness** that proves the two agree where they should, run on a schedule and visible to consumers. This is what converts "trust us" into evidence and is usually what unblocks the last reluctant team. ## Where to put the version number Versioning the physical tables (`fct_orders_v2`) is blunt but honest: it is impossible to use the wrong one by accident. Versioning at the view or semantic layer is subtler and cheaper — the same underlying tables serve both contracts — and is the right answer whenever the difference is expressible as a projection. What you should avoid is versioning implicitly: two tables with different names and no stated relationship, where only tribal knowledge says which is current. ## Migration is a programme, not an announcement The producing team's job does not end at publishing v2. Provide the mapping — old column to new, old formula to new — as something a consumer can apply mechanically. Offer to migrate the highest-traffic dashboards yourself; the top handful of consumers is usually most of the traffic, and moving them yourself both removes the bulk of the risk and demonstrates the migration is safe. Report progress against the sunset date rather than waiting for it. ## The interview-shaped answer Say: default to absorbing the change behind the serving layer; version in parallel only for grain, key or semantic changes that a view cannot hide, or where consumers must reconcile; and treat sunset date, named owner, usage telemetry, a two-version cap and v1-derived-from-v2 as the entry price, not as follow-up work.
- What most reliably prevents a v2 mart from becoming permanent?A sunset date and a named owner agreed before v2 is published, plus usage telemetry from query history so remaining consumers are a list rather than a guess. Add a hard cap of two live versions, so publishing a third requires retiring one first. Without a forcing function the old version simply persists.
- When should the version live in the view layer rather than in separate physical tables?When the difference is expressible as a projection over the same rows — renames, reshapes, derivable splits, a narrower column set. One set of tables then serves both contracts with no duplicated storage or compute and no possibility of divergence. Physical duplication is for changes a view cannot reproduce, such as a grain change.
- Why derive v1 from v2 instead of running both pipelines from source?Two independent pipelines encoding the same business logic drift, and every fix has to land twice. Deriving the old version from the new one leaves a single place where logic lives, makes agreement structural rather than something you monitor, and reduces v1 to a compatibility surface that is cheap to drop.
saying these in an interview costs you the question
- Versions every change by default instead of using the serving layer
- Publishes v2 with no sunset date or owner
- Runs v1 and v2 as independent pipelines from source
- Lets a third version appear before the first is retired
- Treats publishing v2 as the end of the migration work