Why is the bronze layer of a lakehouse kept immutable and append-only?
answer
- What can you not get back later?
- Sources have retention windows too
- Downstream tables should be a function of it
- Was it them or us — how do you prove it?
- Duplicates on arrival are expected, not a defect
basics
~20 sBecause it is the only copy of what the source actually sent. Keeping it append-only lets you recompute every downstream table after a logic bug or schema change, prove whether an error came from the source or from your code, and survive sources you cannot re-read.
solid answer
~50 sBronze is the platform's evidence locker. If it is append-only, every downstream table is a **function** of it, so any cleaning bug, wrong business rule or newly discovered field can be fixed by recomputing rather than by re-requesting data. That matters because sources are frequently not re-readable: API windows expire, event streams have retention limits, CDC feeds only carry changes once, files get rotated. Immutability also settles the incident question "did they send it wrong or did we transform it wrong?" — you can show the original bytes with their arrival time. In practice you append records with ingestion metadata (load timestamp, batch or file id) and never update or delete in place; corrections are expressed as new rows plus downstream logic. The realistic exceptions are legally mandated deletion and retention expiry, both handled as deliberate, audited operations rather than routine edits.
code
sql · 8 linesCREATE TABLE bronze_orders_raw (
ingest_batch_id VARCHAR(64) NOT NULL,
source_file VARCHAR(512) NOT NULL,
ingested_at TIMESTAMP NOT NULL,
source_row_number BIGINT NOT NULL,
payload VARCHAR(65535) NOT NULL -- exactly as received
);
-- loads only ever INSERT; nothing here is updated or repaired in placego deeper
Remember the one-line reason: bronze is the only copy of what the source actually sent, so it is never edited in place. Cleaning and deduplication happen in the next layer.
Explain the mechanics — append with ingestion metadata, corrections arrive as new rows, duplicates are expected and resolved downstream — and give the re-readability argument about API windows, stream retention and CDC feeds.
Show you have used bronze in an incident: proving whether the source or the pipeline was at fault, and reprocessing a bounded window after a logic fix. Be ready to discuss retention limits as a deliberate loss of reprocessing range.
Own the policy: per-dataset retention tiers, cost of fidelity versus optionality, privacy-deletion design in an append-only store, and whether bronze counts as a system of record for audit purposes.
## What "immutable bronze" means in practice An immutable, append-only bronze layer means loads only ever **add** rows. You never update a bronze row to fix a value, never delete rows that look wrong, and never overwrite yesterday's landing table with a corrected extract. Each row carries provenance columns — an ingestion timestamp, a batch or file identifier, sometimes a source offset or row number — so it is always answerable when a record arrived and in which delivery. Corrections from the source arrive as *new* rows, and it is downstream logic, not an in-place edit, that decides which version wins. ## Reason one: downstream tables become recomputable The layering only pays off if silver and gold can be thrown away and rebuilt. That property holds exactly when the raw input is still there and unchanged. A currency conversion applied with the wrong rule, a deduplication that picked the wrong tie-break, a field you parsed as local time when it was UTC — all of these are recoverable by rerunning the transform over retained bronze. If bronze was cleaned on ingest or overwritten each night, the bug is baked into the only copy you have, and "fixing" it means going back to the source system's owner with an apology and a data request. ## Reason two: sources are usually not re-readable This is the reason candidates forget, and it is the strongest one. API endpoints expose rolling windows and then stop serving old data. Event streams have retention measured in days. Change-data-capture feeds emit each change once. Vendors rotate or purge extract files. Operational databases are mutable by definition — reading them again next month gives you *today's* state, not the state you loaded from. The moment you drop raw, you have made the past unrecoverable, and no amount of downstream design compensates. ## Reason three: blame isolation and audit When a finance number is wrong, the first question is whether the source sent bad data or the pipeline mangled good data. With an unmodified bronze row plus its arrival timestamp you answer that in a query rather than an argument. The same property serves audit and regulatory review: you can demonstrate what was received and when, independently of what your transformations concluded. Some teams treat bronze as the system of record for exactly this reason. ## Reason four: schema drift and "we did not need that column yet" Sources add fields. If bronze stores only the columns your current model uses, a field that becomes important next quarter has no history — you can model it going forward but not backwards. Landing the full payload, often as a single semi-structured column plus a few extracted keys, means new requirements can be served retroactively. This is the cheapest form of optionality a platform has. ## What you *do* add on the way in Immutability is not the same as adding nothing. Ingestion metadata is the accepted addition: ```sql CREATE TABLE bronze_orders_raw ( ingest_batch_id VARCHAR(64) NOT NULL, source_file VARCHAR(512) NOT NULL, ingested_at TIMESTAMP NOT NULL, payload VARCHAR(65535) NOT NULL ); ``` That is enough to reconstruct which load produced which downstream row, to reprocess a single batch, and to detect a re-delivery of the same file. ## Duplicates are expected, not a defect Because loaders retry and sources re-send, bronze legitimately contains duplicates. Deduplicating on the way in destroys the record of what actually arrived and hides delivery problems. The duplicates are resolved downstream, where the grain is declared and the tie-break rule is written once and tested. ## The exceptions Two things do modify bronze, and both are deliberate. **Deletion requests** under privacy regulation must be honoured, which means bronze has to be designed so a subject's rows can be located and purged — hard deletes, audited, not routine edits. **Retention expiry** removes data past a defined age, which is a cost and compliance decision made with eyes open: the day you set a 90-day bronze retention, you have decided that reprocessing older than 90 days is impossible. Say that out loud when you set it. ## Cost, and the honest trade-off Keeping everything forever is not free. The usual answers are cheap object storage, retention tiers by dataset value, and compressing or archiving cold partitions — plus being selective about landing genuinely enormous, low-value telemetry. What you should not do is trade away immutability for a modest storage saving; the failure it protects against — an undetectable bug in the only copy of the truth — is far more expensive than the bytes.
- If bronze is append-only, how do you handle a source that re-sends a corrected record?Append it as a new row with its own ingestion timestamp, then resolve it downstream: silver picks the latest version per business key using the arrival metadata. Bronze keeps both the wrong and the corrected delivery, which is what lets you explain a number that changed and reprocess either version deliberately.
- How do privacy deletion requests coexist with an immutable bronze layer?They are the deliberate exception. Bronze is designed so a subject's rows are locatable — a stable subject identifier, or storage partitioned so purges are targeted — and deletions are executed as audited, recorded operations rather than casual edits. The rule is that bronze is immutable to *pipelines*, not exempt from law.
- What is the cost of keeping bronze forever, and how do teams bound it?Storage plus the compute to scan it during reprocessing. Teams bound it with per-dataset retention tiers, compression and archiving of cold partitions, and by declining to land genuinely enormous low-value telemetry at full fidelity. Setting a retention window is an explicit decision that reprocessing beyond it becomes impossible — document it as such.
It is the sealed evidence bag: you can photograph it, transcribe it and reason about it, but the moment someone edits the contents, nothing derived from it can be defended.
saying these in an interview costs you the question
- Says validation and cleansing should happen on ingest into bronze
- Assumes the source can always be re-extracted later
- Deduplicates on the way into bronze to save storage
- Overwrites the landing table on each load to keep it tidy
- Thinks immutability means privacy deletions can be refused