When does crypto-shredding, encrypting each person's data with their own key and destroying the key on erasure, beat physically deleting rows in a lakehouse, and what does it cost?
answer
- destroy the key, not the rows
- reaches copies you cannot rewrite
- a key per person at scale
- decrypt costs at query time
- only the encrypted fields go dark
basics
~20 sCrypto-shredding encrypts each person's identifying fields with a per-person key; erasure destroys the key, making every copy unreadable, including backups and old snapshots. It wins where rewriting copies is impractical, but costs key management at scale and query-time decryption.
solid answer
~50 sWith **crypto-shredding**, identifying fields are encrypted with a **key per person**, held in a key store outside the lake. Erasure **destroys that key**, and every ciphertext copy — current files, old snapshots, backups, extracts that kept the encrypted form — becomes unreadable at once, without rewriting any file. It beats physical deletion where copies are **many, immutable or hard to reach**: long snapshot histories, backups you cannot rewrite, append-only archives. It costs: a key store holding **millions of keys** that must itself be highly available and backed up without resurrecting destroyed keys; **decryption at read time**, which slows queries and prevents filtering or joining on the encrypted values unless they are decrypted; and it only protects **what was encrypted** — anything left in clear, or quasi-identifiers that still single the person out, remains. Whether key destruction satisfies a particular legal duty is for the privacy team to confirm.
go deeper
Know that destroying the only key to encrypted data makes every copy of it unreadable.
Explain which copies crypto-shredding reaches that physical deletion struggles with, such as snapshots and backups.
Weigh key-store scale, query cost and partial coverage against rewrite cost, and design a hybrid with keyed-hash join keys.
Decide whether the platform adopts per-person encryption, who owns the key store, and how legal acceptance is confirmed and documented.
## The idea Physical deletion has to find and rewrite every copy. **Crypto-shredding** inverts it: instead of removing the data, make it permanently unreadable by destroying the only key that decrypts it. 1. Each person gets a **data key**, stored in a key store outside the lake and indexed by their identifier. 2. Their **identifying fields** (name, email, address, free-text notes) are encrypted with that key before landing. 3. Authorised reads look up the key and decrypt. 4. On erasure, the key is **destroyed**. Every copy of the ciphertext — in current tables, old snapshots, backups, exports that carried it — is now unreadable. ## When it beats physical deletion | Situation | Why crypto-shredding helps | |---|---| | Long time-travel histories that must be kept for audit | old snapshots need not be rewritten | | Backups and archives that cannot be edited | the ciphertext in them goes dark | | Append-only raw zones replayed often | replays read ciphertext they can no longer decrypt | | High request volume with huge tables | no file rewrite per request | ## What it costs - **Key management at scale.** Millions of per-person keys, with lookups on the read and write paths. The store must be highly available, since losing it makes all data unreadable. - **Backups of the key store.** They must not bring destroyed keys back on restore — a subtle design constraint, often handled by wrapping per-person keys under keys that rotate and by excluding destroyed keys from restores. - **Query cost and capability.** Decrypting at read time slows scans, and encrypted values cannot be filtered, grouped or joined on without decrypting first. Many designs keep a **keyed-hash pseudonym** in clear for joins and encrypt only the identifying payload. - **Partial coverage.** Only encrypted fields go dark. Non-encrypted columns — behaviour, timestamps, rare attributes — may still single the person out, so the remaining data is not automatically anonymous. - **Legal acceptance.** Whether rendering data unreadable counts as erasure under a given regime is a question for the privacy team; many organisations accept it with documented key destruction, but it should be confirmed rather than assumed. ## A hybrid in practice - Encrypt identifying payloads per person; keep a keyed-hash join key in clear. - Physically delete from the **current** tables on a batch schedule. - Rely on crypto-shredding for **snapshots, backups and archives** that are expensive to rewrite. ## Why interviewers ask it It tests whether a senior candidate can reason about erasure beyond "run a DELETE". A strong answer explains **why destroying a key reaches unreachable copies**, is honest about **key-store scale, query cost and partial coverage**, and leaves the **legal judgement** where it belongs.
- Why keep a keyed-hash pseudonym in clear alongside the encrypted fields?Analysts and pipelines need to join and count by person without decrypting every row. The keyed hash gives a stable join key while the identifying payload stays encrypted; after erasure, the remaining pseudonymous rows can be deleted or left if they no longer identify anyone.
- What happens if the key store is restored from last month's backup?Keys destroyed since then could come back, re-enabling decryption of erased people. The design must prevent that, for example by recording destructions and re-applying them after any restore, or by excluding destroyed keys from backups.
It is like many identical locked boxes scattered across warehouses that all open with one key: destroy that key and every box stays where it is, but none can be opened again.
saying these in an interview costs you the question
- Assuming crypto-shredding makes all of a person's remaining data anonymous
- Storing per-person keys in the same lake as the ciphertext
- Restoring a key-store backup without re-applying key destructions
- Encrypting join columns and expecting joins to keep working