A legal hold lands on a log platform that deletes on a schedule. How do you design for that?
answer
- Two constraints point in opposite directions
- Classify streams before they reach storage
- The deletion job must consult the hold
- Every copy has to obey the policy
- Prove enforcement, not just intent
basics
~20 sA hold must suspend deletion where deletion is enforced — scoped to a stream and a time range, dated and releasable — not recorded in a ticket. Classify streams by obligation and make every copy obey the policy.
solid answer
~50 sRetention has a floor and a ceiling that pull in opposite directions: obligations that require keeping certain records for a stated period, and budget plus minimisation duties that require not keeping them longer than needed. Both attach to what a record contains, so classification has to happen at or before routing. If everything lands in one undifferentiated store, the only defensible policy is the longest floor applied to everything, which is simultaneously the most expensive option and a minimisation breach. A hold is then a suspension the deletion job consults before it acts: scoped to a stream and time range, attributed, dated, explicitly releasable, and verified against the next deletion cycle. Defensibility means being able to show the policy, evidence that it was enforced, and that every copy obeys it — backups, the archive tier, downstream exports, dead-letter queues and whatever a vendor retains for you.
code
json · 16 lines{
"class": "catalogue-access",
"obligation": "contractual",
"floor_days": 395,
"delete_after_days": 395,
"owner": "legal",
"holds": [
{
"id": "H-2291",
"range": "2026-02-01/2026-04-30",
"placed_by": "counsel",
"placed_at": "2026-05-14",
"released_at": null
}
]
}go deeper
Know that some logs must be kept for a fixed period by obligation, and that a legal hold pauses deletion for a subset of them. Never change a stream's retention setting without checking who owns that decision.
Explain how a hold is enforced: the deletion job checks a scoped, dated registry entry before removing anything. Be ready to say why per-stream classes are required for anything other than one blanket retention period.
Show that you hunt the copies — snapshots, archive tiers, exports, replay queues, vendor-side retention — and that you verify a hold took effect on the next deletion cycle rather than trusting a configuration setting.
Own the tension between a preservation duty and a deletion duty, and decide which classes get immutable storage. Be ready to say how a mid-contract platform migration keeps a multi-month retention claim true, and what you accept losing.
Retention policy is where a logging platform stops being an engineering optimisation and becomes an obligation with an owner outside engineering. The design problem is that the constraints point in opposite directions and only one of them is visible on the invoice. ## Four constraints, not one | Constraint | What it says | Cost of getting it wrong | |---|---|---| | Retention floor | keep this class of record for at least a stated period | you cannot answer a question you were required to be able to answer | | Budget ceiling | keep no more than the platform can afford | the platform gets cut under pressure, usually indiscriminately | | Minimisation / deletion duty | do not keep personal data longer than it is needed | over-retention is itself the liability | | Legal hold | stop deleting this subset now, until released | deletion destroys material you were told to preserve | The asymmetry to notice is that **"keep everything forever" is not the safe default it feels like.** Some obligations require deletion, and indefinite retention makes every future request, hold and export larger, slower and more expensive. ## Classification is the actual design decision A retention floor attaches to what a record *contains*, not to which team emitted it. That means a routing key carrying the class has to exist in the pipeline before the records reach storage — tagged at the source, or derived at the collection tier — so each class can be routed to a stream or store that carries its own deletion date. If you cannot separate classes, the only defensible policy is the longest floor applied to everything. That is simultaneously the most expensive option available and a breach of minimisation for every record that should have been deleted earlier. Write the classes down as data: the class, the obligation it derives from, its floor, its deletion date, an owner and a review date. Settings spread across per-store configuration are not a policy — nobody can read them and nobody owns them. ## What a hold has to be, mechanically 1. **A record the deletion job reads before it acts** — not a note in a ticket, an email to the platform team, or a temporarily changed setting somebody must remember to change back. 2. **Scoped** — a stream, a time range, and where the store supports it a subject or tenant. A hold that cannot be scoped degrades into holding everything, which is precisely the bill you pay for coarse classification. 3. **Attributed, dated and releasable** — who placed it, under what matter, when, and who may lift it. A hold nobody is empowered to release becomes permanent by accident — the second most common failure after one that never took effect. 4. **Verified** — after placing it, confirm that the next deletion cycle actually skipped the held range. A hold you did not test is a belief, not a control. 5. **Predictable against tiering** — a held range may still move to cheaper storage, because a hold forbids deletion rather than relocation, provided the data stays retrievable within whatever time the obligation implies. ## Defensibility is a property of the pipeline Three things have to be demonstrable, and only the first is usually written down: - **The policy** — which classes, which floors, which deletion dates, agreed with whoever owns the obligation. - **Enforcement** — the deletion actually ran on schedule and produced an auditable record of what it removed and what it skipped because of a hold. - **Completeness** — every copy obeys the same rules. The copies are where this fails, and there are always more of them than the architecture diagram shows: snapshots and backups of the live store, the archive tier, a downstream analytics export, a dead-letter or replay queue holding records that failed processing, materialised summaries and result caches, whatever the platform vendor retains on its own schedule, and an engineer's ad-hoc export sitting in a bucket. A claim that "these logs are kept for thirteen months and then deleted" is only as true as the least-controlled of those copies. ## The migration case, where all of it bites at once Moving off a hosted logging vendor mid-quarter exercises every part of this. The outgoing vendor still holds the historical records under a contract with its own deletion behaviour, so a hold has to cover both platforms, and the contract has to say what happens to the data on termination and how quickly. Export rarely round-trips: what comes out is usually raw records, and re-ingesting several terabytes into the new platform pays the write-time processing charge again. A "thirteen months, always searchable" claim across a migration therefore resolves to one of three honest answers — run both platforms until the old window ages out, accept a searchability gap where the old data exists only as a restorable archive, or pay to re-ingest. Choose deliberately and record the choice; the failure mode is discovering, while a hold is live, that the old vendor deleted the range some fixed period after termination. ## Immutability, and who owns the floor Write-once storage makes tampering hard and deletion hard in the same stroke: use it where preservation is the obligation, and avoid it for classes carrying personal data with a deletion duty. The standing organisational failure is engineering guessing the floor and encoding the guess. Ask for the obligation in writing, per class, review it on a schedule, and make the platform's default behaviour deletion — with retention as the explicit, owned exception.
- Why is keeping everything forever not the safe default it appears to be?Because some obligations require deletion rather than retention: personal data kept beyond its purpose is itself a breach. Indefinite retention also makes every future hold, export and subject request larger, slower and more expensive, and it removes the ability to say confidently what the platform does and does not hold.
- How would you prove that a range under hold was genuinely not deleted?The deletion job emits an auditable record per cycle naming what it removed and what it skipped and why, which you read alongside the dated registry entry. Then verify empirically by querying the held range after a cycle has run. A hold that was configured but never observed to take effect is a belief, not a control.
- Where do retention claims most often turn out to be false?In the copies. Snapshots and backups of the live store, the archive tier, a downstream analytics export, a dead-letter queue of records that failed processing, materialised summaries, an engineer's ad-hoc export, and whatever the platform vendor retains on its own schedule. The claim is only as strong as the least-controlled copy.
A hold is a stop-work order handed to the shredder operator, not a memo filed with management.
saying these in an interview costs you the question
- Treats a legal hold as a ticket for the platform team, not a check the deletion job performs
- Assumes keeping everything forever is always the safe option
- Applies one retention period to every stream because nothing is classified
- Forgets backups, archives and downstream exports when claiming data was deleted
- Never verifies that a placed hold actually stopped the next deletion cycle
- Lets engineering guess the retention obligation instead of getting it in writing