skip to content

Take a piece of sensitive data that enters at the edge of a distributed system and ends up encrypted in storage. Explain how you would reason about where it exists in cleartext along the way, and what design moves reduce the number of components that ever see the plaintext.

level: principalimportance: should knowfreq 25%

answer

  1. plaintext locality = who can read it, for how long
  2. every termination/parse/queue/retry is a cleartext point
  3. queues are durable plaintext
  4. reference early, resolve in one custodian
  5. state the invariant + what enforces it

basics

~20 s

Draw the plaintext's path and count the components that can see it. Every hop that terminates a connection, parses, queues, caches, retries or logs holds a cleartext copy. Reduce the count: substitute a reference for the value early, and decrypt only inside the one component that needs it.

solid answer

~60 s

**Reason by locality.** Draw the value's route from entry to storage and mark every place it exists in the clear: each connection termination point, each parsed request object, each queue payload, each retry buffer, each cache, each diagnostic capture. That set — not the storage encryption — is the true exposure surface, and it is usually far larger than the team expects because intermediaries that merely route the data still hold it in memory and often in logs. **Then shrink the set.** - **Substitute a reference early.** Replace the sensitive value with an opaque token at the first component that can, so downstream services carry a reference and only a single custodian can resolve it. - **Encrypt at the field level as early as possible**, so intermediate hops move ciphertext they cannot read, with only the intended consumer holding the key. - **Keep the plaintext off durable intermediates**: queues, retry stores, dead-letter topics, caches, audit records. - **Collapse hops** that exist only to forward. The design goal is a stated invariant: plaintext exists in exactly these N components, and here is what enforces it.

go deeper

for a junior

Understand the idea that data is readable in the clear inside every service that handles it, even when every connection is encrypted, and that a queue message is a stored copy.

for a middle

Be able to enumerate the cleartext points along a path and suggest the basic reduction: carry a reference instead of the value, and keep the plaintext out of queues, caches and logs.

for a senior

Reason about field-level encryption for a specific consumer, moving the operation to the data, and treating retention policies of logs and queues as security parameters; be explicit about which hops can be collapsed.

for a principal

Deliver a stated invariant — plaintext exists in exactly these components — with a real enforcement mechanism (boundary schema, resolution permission, reviewed changes) and an honest account of the concentration, latency, availability and reporting costs you are accepting.

## The question behind the question Teams routinely say a value is "encrypted end to end" when what they mean is that each hop was encrypted and the database column is encrypted. Between those, at every hop, the value existed as plaintext in a process, and often in that process's buffers, queue payloads, retry stores and diagnostics. The design discipline is to make that set explicit and then make it small. The unit of reasoning is **plaintext locality**: for a given sensitive value, the set of components that can observe it in the clear, and for how long. Exposure grows with the size of that set because each member contributes its own memory, its own logs, its own crash artefacts, its own operators and its own dependencies. ## Step 1 — draw the path and mark the cleartext points Walk the value from where it is created to where it comes to rest, and mark every place it is readable: - **Connection termination points.** Wherever a transport-encrypted connection ends and is re-established, the payload exists in the clear in that component's memory. A gateway, a proxy or a service mesh sidecar that inspects or routes traffic sees it. - **Parsing and object mapping.** The parsed request object holds the value in whatever containers the framework chose, with lifetimes it controls. - **Asynchronous hand-offs.** A message placed on a queue is *durable* plaintext: it persists, it is replicated, it goes to a dead-letter store on failure, and it is readable by anyone with queue-inspection access. This is one of the most commonly missed cleartext locations because the code looks like a function call. - **Retry and outbox stores.** A retry buffer or transactional outbox holding the request payload keeps plaintext on disk with a lifetime governed by failure handling. - **Caches**, whose eviction policies were chosen for hit rate. - **Diagnostics.** Every component on the path can serialise the value into a log, an error report or a trace attribute — so each hop multiplies not just memory exposure but log exposure. - **Backups and replicas** of any of the durable items above. Counting these honestly is the deliverable; the count is usually two to five times what people guess. ## Step 2 — the moves that shrink the set **Substitute a reference for the value, as early as possible.** The highest-leverage move: at the first component capable of doing so, exchange the sensitive value for an opaque reference and let everything downstream carry the reference. A single custodian component holds the mapping and the plaintext; every other service handles a value that is worthless if disclosed. This is structural separation applied at architecture scale — downstream services do not *fail* to leak the secret, they have nothing to leak. The costs are honest ones: the custodian becomes a high-value target and an availability dependency, and any function that genuinely needs the real value must call it. **Encrypt at the field level for the intended consumer.** Where a reference will not work, encrypt the field so that only the consumer that must read it holds the key. Intermediate hops then transport ciphertext. This preserves the routing topology while removing every intermediary from the plaintext set — but only if the key is genuinely not available to them, which is the part that gets fudged when everyone shares one key. **Move the operation to the data, not the data to the operation.** If a component needs a decision derived from the value rather than the value itself — is this card valid, does this password match, is this person over eighteen — send the question to the custodian and get the answer back. The plaintext never leaves. **Keep plaintext off durable intermediates.** If a value must cross an asynchronous boundary, send a reference and let the consumer resolve it, so the queue, its replicas and its dead-letter store hold nothing sensitive. Where that is impossible, encrypt the payload and treat the queue's retention as a security parameter. **Collapse pass-through hops.** A service whose only role is forwarding still contributes memory, logs, crash artefacts and operators to the exposure set. Removing it is a security improvement as well as a latency one. **Minimise at the edge.** The cheapest reduction is not to collect the field, to collect a coarser version of it, or to discard it immediately after the single operation that needed it. Every downstream question about protecting it disappears. **Bound retention everywhere it does land.** Cleartext in a log, a queue or a cache is defined by its retention policy, and those policies are frequently defaults nobody chose. ## Step 3 — state and enforce the invariant A principal-level answer ends with an invariant a team can hold itself to, in this shape: > Plaintext for this value exists only in the edge component that receives it, in the custodian, and in the consumer that acts on it. Every other component receives a reference. This is enforced by the schema at the service boundary and by the fact that no other service holds the key or the resolution permission. Then name what would violate it and how you would notice: a new service added to the path that resolves the reference "for convenience", a debugging change that logs the resolved value, a queue payload that starts carrying the real field for performance, a cache introduced in front of the custodian. Enforcement lives in the contract (the boundary schema does not have a field for the plaintext), in the permission model (only the custodian may resolve), and in review of any change that adds a resolution call — with scanning of logs and queues as a backstop rather than a plan. ## The trade-offs to name out loud Centralising plaintext into one custodian concentrates risk and creates an availability dependency and a latency cost; that is a deliberate trade of *many weak boundaries for one strong one*, and it is only correct if the custodian is genuinely hardened, small, and monitored. Field-level encryption complicates search, indexing, analytics and support workflows. Reference substitution breaks naive joins and reporting. A principal answer states these costs and the conditions under which each is worth paying, rather than presenting the pattern as free. ## Answering the question Name plaintext locality as the unit of analysis, enumerate the cleartext points including the ones people miss (queues, retry stores, dead-letter topics, diagnostics at every hop), give the shrinking moves in order — do not collect it, substitute a reference early, encrypt for the intended consumer, move the operation to the data, collapse pass-through hops — and finish with a stated invariant, its enforcement mechanism, and the costs you accepted.

  • Centralising plaintext in one custodian service concentrates the risk. Why is that an improvement rather than a single point of failure?
    Because the alternative is not zero risk but many diluted ones: a dozen services each holding plaintext in memory, logs, queues and crash artefacts, each with its own dependencies and operators. Concentration lets you spend real hardening effort — a small codebase, a tight attack surface, strong authorisation, full audit, dedicated monitoring — in one place instead of thin effort in twelve. You do take on an availability dependency and a latency cost, which must be designed for explicitly with caching of references, not of plaintext.
  • Which cleartext location do teams most often miss when they draw this diagram?
    Asynchronous ones: message queues, retry buffers, transactional outboxes and dead-letter topics. The code looks like a hand-off, but the payload is durable, replicated, retained by a policy nobody reviewed, and readable by anyone with queue inspection rights — often a broader group than has database access. The second most-missed is diagnostics at intermediate hops, because each pass-through service can log the payload it is merely forwarding.
  • How do you keep the invariant from eroding over time?
    Put it in the contract rather than in a document: the boundary schema for downstream services has no field for the plaintext, so carrying it requires a visible schema change; only the custodian holds the resolution permission, so a new resolver requires an authorisation change someone must approve; and any change adding a resolve call is a reviewed event. Log and queue scanning is a backstop that tells you the invariant already broke, not a control that keeps it.

Encrypting each leg of a journey is like sealing an envelope for each courier while every relay office opens it to read the address. What you want is a package the relays cannot open and a claim ticket instead of the contents.

saying these in an interview costs you the question

  • Calling a system "end to end encrypted" when each hop terminates the connection and handles plaintext.
  • Forgetting that queues, retry stores and dead-letter topics hold durable plaintext.
  • Sharing one field-encryption key across all services, so encrypting for the consumer protects nobody.
  • Assuming storage-level encryption reduces the in-flight exposure surface.
  • Presenting tokenisation or a custodian service as free, without naming the latency, availability and reporting costs.

context