Engineers want to keep raw event data with identifiers forever so pipelines can always be rebuilt; how do you decide what the platform actually keeps?
answer
- the benefit of forever is rebuilds
- the cost is risk, erasure, storage
- how often do you really rebuild
- keep less identifying forms longer
- legal floors and ceilings decide the edges
basics
~20 sWeigh the real value of full rebuilds against the risk, erasure cost and legal limits of keeping identifiable raw data. Usually keep identifiable raw data for a bounded rebuild window, then keep pseudonymised or aggregated forms, with legal minimums as explicit exceptions.
solid answer
~50 s"Forever" optimises one benefit — the ability to rebuild any pipeline from scratch — against several costs: breach impact, the erasure burden of every retained copy, storage, and privacy law's expectation that identifiable data is kept **no longer than necessary**. I would first measure the benefit: how far back rebuilds and backfills actually go, and how often. Usually that gives a **bounded rebuild window** — say 13 months of identifiable raw data. Beyond it, keep **reduced forms**: pseudonymised events with identifiers replaced by keyed hashes, or aggregates, which still support most long-range analysis. **Legal minimums** for specific records are explicit exceptions with their own class. The decision is **documented** per data class with legal and the data owners, reviewed yearly, and enforced by the retention platform, so "forever" becomes a deliberate exception, not the default.
go deeper
Know that keeping identifiable data forever carries privacy risk and erasure cost.
Explain how pseudonymised or aggregated history can replace identifiable raw data for most long-range uses.
Measure how far back rebuilds really go and propose retention tiers with automated enforcement.
Own the trade-off with legal and data owners, document it per data class, and defend it against both over-retention and over-deletion.
## The two sides **For keeping everything raw:** - any pipeline can be rebuilt after a logic bug; - new metrics can be computed over the full history; - machine-learning teams want long histories. **Against:** - **breach impact** grows with every year of identifiable data held; - **erasure** gets harder: every retained copy must be found and rewritten; - **storage limitation**: privacy regimes such as the GDPR expect identifiable data to be kept only as long as necessary for its purpose; - **cost**: storage, scans and compaction over data rarely read. There is no universal answer; the principal's job is to make the trade-off **explicit, measured and enforced**. ## A decision method 1. **Measure actual need.** From job history, how far back have rebuilds and backfills gone in the last two years? How often did anyone need identifiable raw data older than a year? 2. **Separate what needs identity.** Many long-range uses — trends, seasonality, model features over time — work on **pseudonymised** events (identifiers replaced by keyed hashes) or on **aggregates**. 3. **Identify legal floors.** Some records must be kept for a minimum period (financial records, for example); those become explicit retention classes owned by legal. 4. **Define tiers.** | Age of data | Form kept | Supports | |---|---|---| | 0 to 13 months | identifiable raw | rebuilds, backfills, erasure handled per request | | 13 months to 5 years | pseudonymised events, identifying payload dropped | long-range analysis, model training | | beyond 5 years | aggregates only | trend reporting | | legal-minimum records | as the duty requires | the specific obligation, then deletion | 5. **Document and review.** Record the reasoning per data class, with the owners and legal, and review yearly. 6. **Enforce automatically** through the retention platform, so the tiers are real. ## Arguments to expect - *"We cannot predict what we will need."* True, but storing identifiable data indefinitely for unknown future uses is exactly what minimisation principles argue against; pseudonymised history covers most unknown analytical needs. - *"Pseudonymised data is still personal data."* Correct — it lowers risk rather than removing obligations, which is why it has its own retention class too. - *"Rebuilding needs exact raw inputs."* Then define the rebuild window from real incidents, and snapshot outputs so older periods need not be rebuilt. ## Why interviewers ask it It is a principal-level judgement with legal, cost and engineering dimensions. A strong answer **measures the claimed benefit**, offers **reduced forms** instead of all-or-nothing, puts **legal floors** in their place, and turns the decision into **enforced, reviewed policy**.
- What evidence would convince you to lengthen the identifiable raw window?Repeated incidents where correct results required rebuilding from identifiable raw data older than the window, with no workable alternative from pseudonymised data or snapshotted outputs. Each extension should be recorded with its reason and reviewed.
- Does pseudonymising old events remove the need for erasure handling?No. Pseudonymised data remains personal data while someone can link it back, so erasure and retention still apply, but the linking key can be controlled and destroyed, which makes both cheaper.
saying these in an interview costs you the question
- Keeping identifiable raw data forever by default without a documented reason
- Treating pseudonymised history as outside privacy obligations
- Deciding retention without legal and data-owner input
- Setting a policy that no automated job enforces