Why does a Kimball dimensional mart reach business users sooner than an Inmon warehouse?
answer
- Compare the unit of delivery, not the table shape
- How much must you understand before shipping?
- Speed comes from deferring something — what?
- The bill arrives at the third mart
- Planning shared dimensions is cheap; retrofitting is not
basics
~20 sBecause the bottom-up unit of delivery is one business process, not the enterprise. You model only the sources that process needs and publish a usable star schema; the top-down approach must integrate the enterprise model before any report exists.
solid answer
~50 sIt is a scope difference, not a speed trick. Kimball's unit of work is **one business process** — choose it, declare the grain, pick dimensions, pick facts, ship. You only have to understand the handful of sources that process touches, so a first mart can be live in weeks and every subsequent process is another increment. Inmon's unit of work is the **enterprise model**. Reports wait until the integration layer covers the entities they need, and modelling those entities means cross-departmental agreement on what a customer or an order is — slow, political work that pays back later. The cost is symmetrical, which is what the interviewer is really testing. Bottom-up front-loads value and defers integration risk: the third and fourth marts are the ones that expose whether the shared dimensions were planned properly, and retrofitting a definition across several published fact tables is expensive. Top-down front-loads the integration and defers the value.
go deeper
Recall that the bottom-up unit of delivery is a single business process while the top-down unit is the whole enterprise model, and that scope difference is why one produces a usable report sooner.
Explain what the early delivery defers — integration across subject areas — and name the moment it comes due, typically when the third or fourth mart needs a shared entity defined differently.
Show the two-sided refactor calculus: cheap changes inside one subject area versus expensive cross-cutting definition changes applied to already-published fact tables, and what planning prevents the expensive case.
Own the delivery strategy — a thin vertical slice that ships value while the shared-dimension roadmap is agreed — and be able to defend it to a sponsor who wants either everything now or perfect integration first.
## The mechanism: scope, not speed Nothing about a star schema is intrinsically faster to build than a normalized table. The difference is **how much you must understand before you can ship anything**. In the bottom-up approach the deliverable is one business process. The design sequence is fixed and small: choose the business process, declare the grain of the fact table, identify the dimensions, identify the facts. To do that for *orders* you need to understand the order sources and the entities orders touch. You do not need to have resolved how the support system identifies a customer, or what the supply-chain team means by a shipment. Those are next quarter's problems, and they are somebody's problem only when a mart needs them. In the top-down approach the deliverable that unblocks everything else is the integrated enterprise model. Its scope is the business, not a process. Until it covers the entities a report needs, that report does not exist — and covering an entity means every source that touches it has been mapped and reconciled. So the honest phrasing is: bottom-up ships value earlier because it *scopes smaller*, and it scopes smaller because it defers integration. ## What gets deferred, and when the bill arrives Deferred is not free. The bottom-up bet is that shared dimensions can be planned up front and reused as each process is added. When that planning happens, the second and third marts are cheap: they attach to dimension tables that already exist. When it does not, the bill arrives around the third or fourth mart. Each was built with a locally convenient customer table, and now finance and support disagree. Retrofitting one agreed definition means changing dimension keys and reloading the fact tables that reference them — across models people already have dashboards and saved queries against. That is materially harder than the same change made once in an upstream layer before anything was published. The symmetrical risk on the top-down side is that value never arrives. Six months of enterprise modelling with no report to show erodes sponsorship, and departments respond by building their own extracts — which recreates the fragmentation the architecture existed to prevent, but now outside it. ## Refactor cost, both directions Interviewers like this comparison because the costs mirror each other: - **Change inside one subject area.** Bottom-up wins. Add a fact, change a mart's grain, touch one model and its consumers. - **Change to a cross-cutting definition.** Top-down wins. One change in the integration layer propagates; bottom-up must apply it to every fact table that references the dimension, and reconcile published history. - **Adding an eleventh source.** Top-down wins. Map it into existing entities once; bottom-up maps it into each mart that cares. - **Adding a new consumer for existing data.** Roughly even, if the shared dimensions are in good shape. ## What actually shortens time to value In practice, teams get the early delivery without abandoning integration by doing a **thin slice of both**: land raw data, build only the integration a first mart needs, publish the star, then widen the integration layer as each subsequent process arrives. The bus-planning discipline — deciding which shared dimensions exist across the whole roadmap *before* building the first mart, even though you only build the ones the first mart needs — is what keeps the second and third increments cheap. Planning shared dimensions is cheap; retrofitting them is not. Cheap storage and ELT help too, but be precise about why: they make **re-processing** cheap, so an early modelling mistake can be corrected by rebuilding from retained raw data rather than by re-extracting from source systems that may no longer hold the history. They lower the cost of being wrong; they do not remove the need to agree on definitions. ## Answering in an interview Name the scope difference first, then the deferred integration cost, then say what you would actually do: deliver a first process quickly *and* plan the shared dimensions across the roadmap so the later increments do not collide. Saying only "Kimball is faster" is the answer of someone repeating a slide; naming when the deferred bill comes due is the answer of someone who has shipped a second mart.
- Does cheap cloud storage and ELT remove the time-to-value advantage of building bottom-up?It narrows it rather than removing it. Landing and retaining raw data makes reprocessing cheap, so an early modelling mistake can be rebuilt rather than re-extracted, which lowers the risk of publishing early. But the slow part of the top-down approach was never the compute — it was reaching cross-departmental agreement on entity definitions, and no storage price changes that.
- How do you keep the second and third marts from becoming expensive after a fast first delivery?Plan the shared dimensions across the whole roadmap before building the first mart, even though you only build what the first mart needs. Decide up front which entities are shared, what their keys and grain are, and who owns them. Then each new process attaches to existing dimension tables instead of introducing a competing one.
- When is the slower, integrate-first path clearly the right call?When the sources genuinely conflict and reconciliation is the hard part — regulated reporting, post-merger source consolidation, many systems describing the same entities. There, publishing a fast mart on unreconciled data produces confident wrong numbers, which costs more than waiting. Existing central governance also makes the upfront modelling cheaper than it looks.
saying these in an interview costs you the question
- Says star schemas are simply faster to build than normalized tables
- Presents the early delivery as free with no deferred cost
- Claims cheap storage makes upfront integration unnecessary
- Thinks the difference is query performance rather than build scope
- Ignores that retrofitting a shared dimension touches published fact tables