How do you decide how many medallion layers a lakehouse needs and who owns each?
answer
- Where does responsibility change hands?
- Name the guarantee, name the owner
- Every extra hop costs freshness
- Who is allowed to publish a consumable table?
- A second copy is a second set of numbers
basics
~20 sPut a layer where responsibility changes hands and a guarantee is added, not where a diagram says three. Count the real handoffs — ingestion, conformance, business definition — give each an accountable owner with a tested contract, and delete any layer whose promise nobody can name.
solid answer
~50 sLayers are **contract boundaries between owners**, so the count follows the number of genuine handoffs rather than doctrine. The common three map to three accountabilities: ingestion owns landing fidelity and delivery SLAs; a platform or analytics-engineering group owns conformance — grain, identity resolution, shared reference data; domain teams own business definitions in the published layer. Where those roles collapse into one team and one clean source, two layers are enough. Where a domain merges six overlapping systems, an extra conformance stage may earn its place. The test for each layer is a stated guarantee, an owner who is paged, a tested contract and a change policy. Against that, weigh cost: every copy adds storage, latency and another place for logic to hide. Governance matters more than the count — who may publish a consumable table, whether metric definitions live in one place, and how a breaking change to a published model is announced.
code
yaml · 12 lineslayer: gold
table: fct_daily_sales
owner: finance-analytics
grain: one row per order_date, product_key, store_key
upstream: [silver_orders, silver_product]
guarantees:
- grain uniqueness enforced on every build
- not-null on all key columns
- available by 07:00 UTC on business days
change_policy:
breaking_change_notice_days: 30
consumers: [exec_dashboard, finance_close, partner_extract]go deeper
Know that three layers is a convention rather than a rule, and that a layer only earns its place when it adds a guarantee someone is responsible for.
Be able to state the four admission questions for a proposed layer — added guarantee, owner, tested contract, change policy — and name the cost of each extra copy in storage and freshness.
Show judgment about collapsing or splitting layers for a specific source landscape, and about which tables are publishable interfaces versus internal ones that may be refactored freely.
Own the governance: who may publish, where a metric is defined once, how breaking changes to published models are announced, and the change-speed problem that drives teams to fork their own copies.
## Start from handoffs, not from a diagram The useful question is not "do we do medallion?" but "where in this pipeline does responsibility change hands?" Each such point wants a contract: an interface where the upstream party promises something specific and the downstream party is entitled to rely on it without reading upstream code. Layers are where those contracts live. Count the handoffs honestly and the layer count falls out — often three, sometimes two, occasionally four in a large domain. ## The three accountabilities behind the classic three layers **Landing** is owned by whoever runs ingestion. Their promise is fidelity and delivery: this is what the source sent, it arrives by this time, schema drift is detected and surfaced rather than silently dropped. Notably they promise nothing about meaning. **Conformance** is usually owned by a platform or analytics-engineering group. Their promise is a declared and tested grain, resolved identity across systems, standardised code lists, units and currencies, and shared reference data. This layer exists because identity resolution is expensive and must be done once, not per consumer. **Business definition** is owned by domain teams — finance, growth, supply chain. Their promise is semantics: what counts as revenue, how returns are treated, which grain a fact table sits at. They are the right owners because they are the ones who can adjudicate the definition, and the wrong ones to be maintaining deduplication logic. When those three accountabilities genuinely live in one small team and the source is clean, forcing three layers just adds copies. When a single domain draws on six overlapping systems, splitting conformance into two stages — per-source cleaning, then cross-source integration — can be the honest answer. ## The admission test for a proposed layer Four questions, all of which must have answers: 1. **What guarantee does it add** that no existing layer provides? Name it in one sentence a consumer would recognise. 2. **Who owns it** — a named team that gets paged when it breaks and reviews changes to it. 3. **What is the contract** — declared grain, enforced schema, freshness target, tests that fail the build. 4. **What is the change policy** — how a breaking change is announced and how long consumers have. If any answer is missing, the proposal is a copy with a nicer name, and copies are how platforms silt up. ## The costs you are buying against Every layer costs storage, compute to build, and pipeline latency — a three-hop chain cannot be fresher than the sum of its hops, which matters when the business wants hourly numbers. It also costs cognitive load: more tables to catalogue, longer lineage to reason about during an incident, more places a filter can hide. And it costs organisational friction, because each boundary is a queue a change must pass through. Layering is not free virtue; it is bought coupling-reduction. ## Serving: does the published layer replace a separate mart copy? A live decision on lakehouse platforms is whether business-ready tables are consumed directly by BI or copied again into a separate serving store. The honest criteria are concurrency and latency of the consuming workload, cost model of each option, freshness requirements, and whether governance and access control can be expressed once rather than twice. The strong argument against the extra copy is definitional: a second copy is a second place for the numbers to diverge, and reconciling two stores becomes somebody's permanent job. The argument for it is when the consuming workload genuinely has requirements the primary platform cannot meet. Decide that on measurements, and if you copy, make the copy strictly derived and non-authoritative. ## Governance is the part that actually decays Three policies keep layering alive past its first year. **Publication control**: only a reviewed, owned set of tables is consumable; everything else is internal and may be refactored without notice. **Definition uniqueness**: a metric is defined once — in a shared model or a metrics layer — rather than recomputed in each dashboard, so "active customer" cannot mean three things. **Change policy for published models**: a notice period and a communication path, because consumers build on the published layer precisely because they were told to. Also decide who may *skip* a layer, and make it an explicit, time-bounded exception rather than a quiet habit. Every degraded platform started with reasonable one-off exceptions nobody revisited. ## How to argue this in an interview Strong answers refuse the doctrinal framing. They say layers are contract boundaries, count the handoffs in the specific organisation, admit the cost of each copy in latency and cognition, name owners rather than schemas, and identify the organisational failure mode — teams forking their own copies when the shared model changes too slowly — as the thing that actually determines whether the architecture survives contact with reality. Weak answers recite bronze, silver and gold and stop.
- When are two layers enough instead of three?When the accountabilities collapse. One team owning ingestion and modelling, a single well-typed source with no cross-system identity problem, and a small consumer set means the conformance stage adds no guarantee anyone else relies on. Land raw immutably, publish business-ready tables, and skip the middle copy until a second source makes conformance real.
- How do you keep teams from forking their own copies of a shared model?Make the shared model faster to change than to fork: a clear owner, a short review path, tests that make changes safe, and the ability to add a column without a quarter of negotiation. Then enforce publication control so forks cannot become consumable. Forking is usually a rational response to a slow process, not indiscipline.
- What changes about layer ownership when domain teams own their own pipelines end to end?The horizontal platform layer thins out and the contracts become inter-domain rather than inter-stage. You still need shared conformance for cross-domain entities — customer, product, calendar — otherwise every domain resolves identity differently and cross-domain analysis stops working. Centralise the shared vocabulary; decentralise everything specific to a domain.
- How would you decide whether business-ready tables are served directly to BI or copied into a separate store?Measure the consuming workload — concurrency, latency targets, freshness — against what the primary platform delivers, and price both options. Copying buys performance and costs a second place for numbers to diverge plus permanent reconciliation work. If you copy, make it strictly derived, non-authoritative, and rebuilt from the published tables rather than independently transformed.
Layers are like handoffs in a supply chain: you put a checkpoint where custody and accountability transfer, not every few metres because the manual has three columns.
saying these in an interview costs you the question
- Prescribes exactly three layers regardless of the organisation
- Assigns layers to schemas instead of to accountable owners
- Ignores the freshness and cost penalty of each extra hop
- Lets any team publish consumable tables without review
- Treats a second serving copy as authoritative alongside the original