Leadership asks you to encrypt all sensitive data with symmetric encryption. Since the application must hold the key in order to use the data, explain what threats this actually removes, what it does not, and how you would decide where the encryption boundaries go.
answer
- encryption relocates data confidentiality into key custody
- removes: media, backups, replicas, low-privilege readers
- removes nothing from the process that decrypts
- name the excluded adversary, put the key outside its domain
- deterministic ciphertext = searchable = leaks equality/frequency
basics
~20 sIt converts a data problem into a key-custody problem. It removes readers who get bytes without keys - backups, replicas, stolen media, low-privilege accounts, dumps. It removes nothing from an attacker inside the process that legitimately decrypts.
solid answer
~60 sEncryption relocates the problem: whoever can reach the key can read the data, so the only question that matters is **which compromise domains sit on the wrong side of the key**. It genuinely removes: stolen or decommissioned media, backups and snapshots, a read replica or export shipped to analytics, a database account or operator with storage access but no key access, and accidental disclosure through dumps and logs. It does not remove: compromise of the application process, which will call decrypt for the attacker; a flaw in the application's own read path; an attacker who holds the app's credentials to the key service; and inference from deterministic ciphertext, which leaks equality and frequency whenever you make values searchable. So the method is: for each dataset, name the domain that must not be able to read it and put the key outside that domain. If the answer is 'the application itself', encryption buys nothing and the right control is a different one - tokenisation with surrogates, or client-side encryption where the server never holds the key.
go deeper
Say that encryption only helps against someone who gets the data without the key, and that where the key lives is the real question.
Enumerate concretely what it removes (backups, replicas, stolen media, lower-privileged readers) and what it does not (compromised app, authorisation bugs), and note the searchability trade.
Drive the decision method: name the excluded adversary per dataset, place the key outside that domain, and add rate limits and audit on the decrypt path as the residual control when you cannot.
Own the programme-level trade-offs: key scoping and crypto-shred capability, availability coupling to the key service, the cost of losing query capability, and honest reporting of which mandates buy detection rather than prevention.
## The reframe 'Encrypt everything sensitive' sounds like a security control. It is really a *relocation*: after encryption, reading the data requires the key, so the confidentiality of the data is exactly the confidentiality of the key plus the trustworthiness of everything that can invoke decryption. If the key is reachable from the same place the data is reachable, nothing changed except CPU cost and operational complexity. So the first question in any such conversation is not 'which cipher' but 'which principal must be unable to read this, and is the key outside that principal's reach?'. ## What it genuinely removes There is a real and valuable class of attacker who obtains bytes without obtaining the key: - **Storage-layer exposure**: stolen or improperly decommissioned drives, a misconfigured object store, a cloud snapshot shared too widely. - **Backups and copies**: backup files travel further and live longer than anyone intends, and they are frequently restored into less protected environments. - **Replicas and exports**: the analytics copy, the data-warehouse feed, the anonymised set that was not anonymised. These are typically read by systems that were never in scope for the original authorisation model. - **Lower-privileged principals inside the system**: a database account, a DBA, a support tool or an operator that can read rows but was never granted key use. This is where field-level encryption earns its keep, because it splits 'can read the store' from 'can read the value'. - **Accidental disclosure**: crash dumps, debug logs, an error page. If the field was encrypted at the boundary, the accident leaks ciphertext. Note the shape: every item is an attacker who is *outside* the decrypting process. ## What it does not remove - **Compromise of the decrypting process.** If an attacker runs code in the application, it has the application's identity and can call the same decrypt path. The key being in a hardware module does not help - the module is happy to serve the authorised caller. - **Flaws in the application's own read path.** An authorisation bug or an injection that reaches the app's legitimate query surface produces decrypted data, because the app decrypts on the way out. Encryption is orthogonal to authorisation. - **Credential theft against the key service.** If the app's credentials or role can be assumed, so can its decryption ability. The key service boundary is only as strong as the identity that crosses it. - **Inference from ciphertext.** The moment you require equality search or an index over an encrypted column, you need deterministic encryption for that column, and deterministic ciphertext leaks equality and frequency: which rows share a value, and how the values are distributed. On low-cardinality fields that is often enough to recover the plaintext by frequency alone. This is a real trade, not a bug, and it must be made explicitly. ## The decision method For each dataset, in order: 1. **Name the adversary you are excluding.** Concretely: 'someone with the backup', 'the analytics team', 'the DBA', 'a compromised app instance'. Vague answers produce vague designs. 2. **Place the key outside that adversary's domain.** If the adversary is the storage layer, keys held by the application suffice. If the adversary is the application operator, the key must belong to the customer or to a separate service the operator cannot invoke. If the adversary is the application process itself, encryption is the wrong control and you need the data not to be there: tokenisation, where the app holds only a surrogate and a separate vault holds the mapping, or client-side encryption where the key never reaches your infrastructure at all. 3. **Rank the work** by sensitivity multiplied by reachability multiplied by whether custody separation is achievable. A high-sensitivity field whose key must live in the app is a poor investment compared with a moderate field where custody can genuinely be split. 4. **Choose the key scope.** One key per tenant, per dataset or per record changes what a single leak costs and whether you can perform a per-tenant crypto-shred. One key for everything is the version of this project that gets shipped by default and is worth almost nothing. ## The control ordering The usual ranking applies, weakening from guarantee to heuristic. **Structural separation** is strongest: the key never exists in the compromised domain - client-side keys, a separate service with its own identity boundary, hardware that never exports material. Next, **transformation**: wrapping, tokenisation and surrogates, which reduce what the compromised domain holds but do not eliminate its access to the mapping. Next, **validation**: policy on who may call which key operation on which data, which is enforcement by rules that can be misconfigured or bypassed by an authorised-but-compromised caller. Last, **detection**: audit trails, anomaly detection and rate limits on the decrypt path, which do not prevent anything but make sustained bulk decryption visible. When custody cannot be separated, the detection rung is what you are actually buying, and it should be described that way rather than dressed up as prevention. ## The costs you must state Encryption is not free, and a serious answer names the bill: keys have a lifecycle that someone must own; the key service becomes an availability dependency, so its outage becomes your outage and you need caching and failure policy; searching, sorting and indexing on encrypted fields is lost or degraded, and restoring it via deterministic encryption reintroduces leakage; and broad mandates typically produce one key used everywhere, which delivers the operational cost of the programme with none of the separation that justified it.
- An attacker achieves remote code execution in the application. Which of your encryption work still pays off?Almost none of the field encryption whose key the application can use, because the attacker inherits the application's identity and can call the same decrypt path. What still pays is anything whose key lives outside that process - customer-held client-side keys, data belonging to another service's key scope - plus the detection and rate limits on the decrypt path, which turn a silent full dump into a visible, slowed one.
- Product wants to search on an encrypted column. What are you agreeing to?Exact-match search requires the same plaintext to produce the same ciphertext, which means deterministic encryption for that column. That leaks which rows share a value and the frequency distribution of values, and on low-cardinality data frequency analysis can recover the plaintext without any key. It is a deliberate trade: state it, restrict it to columns where equality leakage is acceptable, and do not let it spread to the sensitive high-cardinality fields.
- How do you decide the key scope - one key overall, per tenant, or per record?By what a single key compromise should cost and what operations you need. One key overall makes rotation a full re-encryption and any leak total. Per-tenant keys bound the blast radius to one customer and enable crypto-shredding on deletion by destroying that key. Per-record keys give the tightest bound and the most metadata and unwrap traffic, so they suit small volumes of very sensitive records rather than everything.
Putting documents in a safe helps against a burglar. It helps not at all against the clerk who opens the safe forty times a day - for that you either stop giving the clerk the documents, or you count the openings.
saying these in an interview costs you the question
- Claiming that encrypting a datastore protects against application-level compromise or injection reaching the app's read path.
- Presenting 'encrypted at rest' as an answer to a threat model without saying who holds the key.
- Encrypting a column and then adding an unencrypted index or search field over the same values.
- Using one key for everything, so nothing is separated and rotation is impossible.
- Ignoring that the key service is now on the critical path for availability.