What does the lakehouse architecture add on top of raw files sitting in a data lake?
answer
- a directory of files is not a table
- something must define the current version
- think atomicity, schema, deletes, statistics
- the files stay open and shared
basics
~20 sA lakehouse adds an open table-format metadata layer over lake files, so a directory of files behaves like a table: atomic commits, consistent snapshots for readers, schema enforcement and evolution, row-level updates and deletes, and statistics for pruning — while staying readable by many engines.
solid answer
~50 sA raw lake is a directory of files, and a directory is not a table: readers can see a half-rewritten partition, concurrent writers clobber each other, there is no reliable row-level delete, and nothing stops a producer changing a column's type. The lakehouse keeps the files and the open formats but puts a **table-format metadata layer** over them that records which files constitute the table at each version. From that you get atomic commits and snapshot-consistent reads, schema enforcement and controlled evolution, `UPDATE`/`DELETE`/`MERGE` semantics, file-level statistics the engine can use to prune scans, and time travel to an earlier version. Crucially the bytes stay in open formats in your own object storage, so several engines can read the same tables without copying. The bet is that you get most of a warehouse's semantics without the closed format or the per-engine copy.
go deeper
Know that a lakehouse means lake storage in open formats plus a metadata layer that makes those files behave like real tables. Being able to name atomic commits and schema enforcement is enough here.
Explain why a bare directory of files fails — partial reads, no row-level delete, silent schema drift — and how a versioned file-list gives each guarantee back. This is the layer interviewers probe most.
Talk about running it: compaction of small files, snapshot expiry, orphan cleanup, and making one set of permissions hold across several engines. Say what you would still keep in warehouse-managed storage and why.
Frame it as an architectural bet on open formats and multi-engine reads versus a single vendor's integrated stack, and be explicit about the operational headcount and governance work that bet buys you.
## The problem the lakehouse is answering A data lake gives you cheap, open, elastic storage and lets any engine read your files. What it does not give you is a **table**. A directory of Parquet files has no notion of a version, so: - A reader scanning while a job rewrites the same prefix can see a mixture of old and new files, or files that vanish mid-scan. - Two writers touching the same partition have no arbitration; the last one to finish wins, silently. - There is no row-level `UPDATE` or `DELETE`. Correcting one bad row means rewriting whichever files contain it and hoping nobody read in between — which is a real problem the first time somebody exercises a deletion right. - Nothing enforces the schema. A producer adds a field or widens a type and consumers discover it as a cast failure days later. - The engine has no statistics beyond what it can infer from directory names and file footers, so pruning is coarse. The traditional answer was to copy curated data out of the lake into a warehouse, where all of that exists. The lakehouse answer is to add the missing semantics **in place**. ## What the metadata layer provides A lakehouse table is data files in an open columnar format plus a metadata layer that records, for each version of the table, exactly which files belong to it, what the schema is, and summary statistics about the files. How the format encodes that metadata is the table format's own business; architecturally, what matters is that this pointer exists and can be advanced atomically. From that single idea the guarantees follow: **Atomic commits and snapshot isolation.** A writer stages new files, then swaps the table's current version pointer in one atomic step. Readers resolve the version once at query start and read exactly the file set it names, so they never observe a partial write. Writers that conflict can be detected and one retried, rather than both succeeding into an inconsistent state. **Row-level mutations.** Because the table is defined by a file list, an `UPDATE`, `DELETE` or `MERGE` can rewrite affected files (or record deletions separately) and publish the result as a new version. Late-arriving corrections, GDPR-style erasures and change-data-capture merges become ordinary operations rather than bespoke rewrite scripts. **Schema enforcement and evolution.** The metadata carries the authoritative schema, so a write that does not match is rejected instead of silently corrupting the dataset, and adding or renaming a column is a recorded, versioned change rather than an accident of what the last producer happened to write. **Statistics and pruning.** Per-file value ranges in the metadata let the engine eliminate files before opening them, giving lake tables the kind of data skipping that warehouses have always had. This is often the biggest practical performance win over listing a prefix and reading footers. **Time travel and rollback.** Old versions remain resolvable until expired, so you can query yesterday's state, diff two versions, or roll back a bad load without restoring from backup. ## What stays open — and why that is the point The data files remain a public format in **your** object storage, and the table metadata is a public specification too. That means several engines can read — and increasingly write — the same table without anyone copying it: a batch engine, a streaming engine, a Python/dataframe tool and one or more SQL warehouses. Compare that with the warehouse-only model where each system that needs the data gets its own copy, its own refresh pipeline, and its own opportunity to disagree with the others. Cutting copies cuts storage, cuts pipeline maintenance and cuts the number of places a number can be wrong. ## What it is not A lakehouse is not automatically as fast or as easy to operate as a mature managed warehouse. You still own compaction of small files, expiry of old versions and orphaned files, and cluster or engine sizing. Catalog and permission integration across multiple engines is a real project, not a checkbox — the metadata layer defines the table, but who is allowed to read which column is still enforced by whatever catalog and storage permissions you wire up. Query performance on well-tuned warehouse-managed storage is often still ahead, particularly for high-concurrency BI. And multi-engine *write* support is less uniform than multi-engine read support, so "any engine can write it" deserves verification per engine rather than assumption. ## How to say it in an interview One sentence: the lakehouse keeps lake economics and open formats, and buys back the table semantics — atomicity, schema, mutations, statistics — that a bare directory of files cannot provide. Then name the operational cost you take on in exchange: compaction, metadata maintenance and cross-engine governance are now yours.
- Why can a plain data lake not give you a reliable DELETE of one row?Because the table is only a set of files, deleting a row means rewriting every file that contains it. There is no atomic way to publish that rewrite, so concurrent readers can see the row present, absent, or duplicated across old and new files. A version pointer over the file set is what makes the swap atomic.
- If lakehouse tables are readable by many engines, why do teams still load some data into a warehouse?Latency-sensitive, high-concurrency BI usually still runs better on warehouse-managed storage, where the engine controls layout, caching and clustering end to end. Lakehouse tables also push compaction, version expiry and cross-engine permissions onto your team. Many platforms curate in the lakehouse and serve the hottest marts from the warehouse.
- What new operational chores does adopting an open table format introduce?Compacting small files produced by frequent commits, expiring old snapshots and cleaning orphaned files so storage and metadata do not grow without bound, and keeping catalog registration and permissions consistent across every engine that reads the tables. None of these are automatic the way a managed warehouse's housekeeping is.
saying these in an interview costs you the question
- Calling a lakehouse just a lake with a SQL engine on top
- Claiming open table formats make queries faster than any warehouse
- Assuming the file format itself provides transactions
- Thinking time travel means unlimited free history retention
- Believing adoption removes all operational maintenance