How do you decide whether a warehouse needs a Data Vault layer at all?
answer
- Which layer does it replace? Trick question
- Count the layers you end up operating
- What problem do eight source systems create?
- Replay without hubs and links is possible
- Retrofitting history you never kept
basics
~20 sAdopt it when many sources must be integrated on business keys, schemas churn, and history must be provable — those are the problems it solves. Skip it with few stable sources: it adds a whole layer and marts are still required on top.
solid answer
~50 sData Vault is an **integration and retention layer, not a consumption layer**, so adopting it means at least three layers — raw landing, vault, marts — and roughly three to five times the table count of an equivalent dimensional model. That cost buys specific things, so check whether you have the matching problems. It pays when: many source systems assert overlapping business keys and must be integrated; regulation or restatement requires proving what a source said on a date; upstream schemas change often, so additive evolution beats refactoring a published star; multiple teams need to load independently and in parallel. It does not pay when: there are one or two stable sources, one team, and BI is the only consumer. Then you carry the modelling and query cost with no integration problem to solve. The middle path is an immutable raw landing zone plus dimensional marts rebuilt from it, which buys replay and audit without hubs, links and satellites.
code
text · 14 linesWITHOUT VAULT WITH VAULT
sources sources
| |
raw landing (immutable) raw landing (immutable)
| |
marts (star schemas) raw vault -> business vault
| |
BI marts (star schemas)
|
BI
Buys: integration on business keys, per-source provable history,
additive evolution, parallel loads
Costs: one more layer, 3-5x tables, longer path to first reportgo deeper
Not an interview question at this stage. Just know Data Vault is a layer some enterprises put between raw data and their reporting models, chiefly for integration and audit.
Be able to name the pattern's benefits and costs concretely: additive schema evolution and parallel loading against a much larger table count and a mandatory reporting layer on top.
Expect to be asked whether a described environment justifies it. Match the specific pains — many sources, audit obligations, churning schemas — to the specific properties, and say plainly when they are absent.
Own the full lifecycle cost: layers to operate, engineers to train, latency added, time to first value delayed, and the asymmetry that missing history cannot be retrofitted while an abandoned vault loses nothing.
## Frame the decision correctly The common mistake is comparing Data Vault against a dimensional model as if they compete for the same slot. They do not. A vault is an **integration and retention** layer; it is not queried by analysts and it is not the thing BI tools read. Adopting it therefore never removes the star schemas — it adds a layer beneath them. The real comparison is: - staging → **vault** → marts, versus - staging or immutable raw landing → marts. So the question is never "vault or star"; it is "is the middle layer worth building and operating". ## What the layer actually buys **Integration on business keys across many sources.** A hub is the one place a business key exists, no matter how many systems supply it. When eight systems each have their own notion of a customer, that single point of integration is genuinely valuable and genuinely hard to retrofit later. **Provable history.** Insert-only satellites, split per source, with load dates and record sources, mean you can reconstruct what each system asserted on any past date, and show the evidence. Regulated environments and any business that has to restate published figures care about this a great deal. **Additive schema evolution.** New source becomes a new satellite. New relationship becomes a new link. New attribute becomes a column on a new satellite. None of these require rewriting or backfilling existing tables. Against a landscape of upstream systems that change quarterly, this is the pattern's strongest practical argument, because the alternative is repeatedly refactoring a published dimensional model that consumers depend on. **Parallel, restartable loading.** With hash keys computed from business keys, every table loads independently, so many teams and many schedules coexist without a global ordering. ## What it costs **Table count and cognitive load.** What might be six dimensional tables becomes dozens of hubs, links and satellites. Every developer must learn the pattern, and half of them will get satellite splitting and link cardinality wrong for the first six months. **A mandatory layer on top.** Nobody queries a raw vault. You will build marts anyway, so you are building and testing two models, plus point-in-time and bridge tables to make the vault queryable by the mart-building jobs. **Latency and cost.** Three layers means more transformation, more storage of the same facts in several shapes, and a longer path from source to dashboard. **Time to first value.** A dimensional mart can answer a business question in weeks. A vault programme spends its early months building infrastructure that no consumer can see, which is an organizational risk as much as a technical one. ## The signals that say yes - Many source systems, with overlapping or conflicting definitions of the same business entities. - A regulatory, audit or restatement requirement to prove what data looked like historically. - Upstream systems whose schemas change frequently, or a roadmap of acquisitions and system replacements. - Multiple engineering teams loading concurrently, with no appetite for a globally ordered pipeline. - Long retention horizons where the sources cannot be trusted to keep their own history. ## The signals that say no - One or two stable sources — you would be integrating nothing. - A small team, where the pattern's overhead consumes the capacity that would have delivered reports. - BI as the only consumer, with no audit obligation. - Strong time-to-value pressure and an unproven analytics function. ## The middle path Most of what teams actually want from a vault is **replay**: the ability to rebuild marts from an untouched record of what arrived, without re-ingesting. You can have that without hubs, links and satellites — keep an immutable, append-only landing zone of raw extracts, partitioned by load date and never mutated, and make every downstream model a deterministic, rebuildable transformation of it. That gives audit of what arrived and full reprocessing, and it costs one storage layer rather than a modelling discipline. What it does **not** give you is integration on business keys, per-source history in a queryable shape, or additive evolution. If the pain you actually have is eight systems disagreeing about who a customer is, raw files will not solve it and a vault will. ## Half measures and reversibility Adopting it need not be all-or-nothing. It is legitimate to vault the genuinely contested core — customer, product, account, the entities several systems assert — and leave peripheral sources loading straight into marts from staging. The core is where integration and audit value concentrates, and it is also where retrofitting later is most painful. That asymmetry is the real reason to decide deliberately. Adding a vault later means rebuilding history you no longer have; removing one later means throwing away work but losing nothing you cannot regenerate. When the sources are many and the obligations are real, the cost of being wrong in one direction is much higher than in the other.
- Does adopting Data Vault remove the need for dimensional marts?No. The vault is an integration and retention layer that nobody queries directly — its joins and insert-only history make it hostile to ad-hoc analysis. Star schemas are still built on top for consumption, which is why adopting the pattern adds a layer rather than replacing one.
- What is the cheapest way to get replay and audit without a Data Vault?An immutable, append-only raw landing zone: keep every extract partitioned by load date, never mutate it, and make every downstream model a deterministic rebuild from it. That buys reprocessing and evidence of what arrived. It does not buy integration on business keys or per-source history in queryable form.
- Can you adopt it for part of the warehouse only?Yes, and it is often the right call. Vault the contested core entities that several systems assert — customer, product, account — where integration and audit value concentrates, and let peripheral single-source data go straight from staging into marts. That targets the overhead at the problem it actually solves.
- Why is the decision harder to reverse in one direction?Because history cannot be retrofitted. Adding a vault later means you never captured what the sources said in the intervening years, and those systems may have overwritten it. Abandoning a vault later wastes effort but loses no information, since everything downstream can be rebuilt.
saying these in an interview costs you the question
- Says Data Vault replaces the need for star schemas
- Recommends it for a single-source, small-team warehouse
- Ignores the table-count and query cost entirely
- Treats it as a performance optimization for reporting
- Claims it can be retrofitted later at no cost