skip to content

Encryption at rest is off on a data store created six hours ago that already holds data — what now?

level: seniorimportance: should knowfreq 42%

answer

  1. exposure before fix
  2. what was written, who could reach it
  3. creation-time property, no toggle
  4. rebuild, migrate, cut over, destroy
  5. snapshots and rotation are part of done

basics

~20 s

Establish the exposure first: what was written, who could reach it, for how long. On most managed data stores encryption at rest is fixed at creation, so remediation is a migration, not a setting change.

solid answer

~50 s

Start with exposure, not the fix. Six hours of writes are already durable, so establish what landed there, whether the store was reachable beyond its intended callers, and what the provider's access logs show. Contain next: tighten network and identity access, and stop new writes if the service can tolerate it. Only then remediate — and on most managed data stores encryption at rest is a create-time property, so there is no toggle. Remediation means standing up an encrypted store, migrating, cutting over, destroying the original, and rotating any credentials or secrets that were written into the unencrypted one. That is an availability-affecting change owned by the service team, not something on-call does unilaterally at 3am. Record the window — created at, detected at, contained at, remediated at — because that window is what any later breach or audit assessment turns on.

go deeper

for a junior

Know that encryption at rest is usually chosen when a store is created and cannot simply be switched on later, and that the first move is to find out what data is in there rather than to start changing settings.

for a middle

Be ready to lay out the order — exposure, containment, remediation, record — and to explain why remediation means migrate and rebuild, including the snapshots and the credential rotation that people forget.

for a senior

Demonstrate that you can run this without owning the service: bound the window from provider logs, contain, assess exposure against what the store holds, and hand a clearly framed decision to the owning team rather than taking an outage alone.

for a principal

Own the distinction between mitigated and never-happened. Be able to state what a realised exposure obliges the organisation to do, and to argue where the remediation cost of a create-time property justifies closing the ungated creation path.

## Why this is a triage problem, not a fix problem The finding arrived from a control that looks at what already exists, so by the time you are reading it the non-conforming state has been real for six hours and has been *used*. The instinct to reach for the setting is the wrong first move, and on a managed data store it usually is not even available. Work in the order exposure → containment → remediation → record. ## Step 1: establish the exposure The question you must answer is not "is it encrypted" but "what is now at risk, and to whom". - **What landed there.** Ask the owning team what the store holds. Customer records, tokens, secrets and anything covered by a regulatory regime change the severity by an order of magnitude compared with derived cache data. - **Who could reach it.** Encryption at rest defends against a specific threat: someone obtaining the underlying storage without going through the service's access controls. It does nothing about an over-permissive access policy. Check whether the store was additionally reachable from outside its intended callers, because if it was, you have two incidents, and the second one is worse. - **What the logs show.** The provider's activity and data-access logs establish who created it, with which identity, and who has read from it. That evidence is also what tells you whether this is a genuine incident or a control gap with no realised exposure. - **How long.** Creation time and detection time bound the window. Write both down before anyone starts changing things, because they get harder to reconstruct once remediation is under way. ## Step 2: contain Containment reduces further accumulation while the real fix is planned. Tighten identity and network access to the minimum the service needs. Consider whether new writes can be paused or diverted — often they cannot without an outage, and that is a decision for the service owner, not on-call. Containment is not remediation and should never be reported as such. ## Step 3: remediate, knowing what remediation actually is On many managed data stores, encryption at rest is fixed when the store is created. There is no property to flip afterwards. Remediation is therefore a project: 1. create a new store with encryption enabled 2. migrate the data 3. cut clients over 4. verify 5. destroy the original, including any backups and snapshots taken during the window 6. rotate anything sensitive that was written into the unencrypted store — its confidentiality cannot be restored, only invalidated Every one of those steps can affect availability, so this is a planned change owned by the service team with a window, not something executed unilaterally overnight. The honest on-call output is often a contained system, a written exposure assessment, and a decision handed to the owner: take the migration now, or accept documented risk until a window. **Step 5 is the one people forget.** Deleting the store but leaving an unencrypted snapshot behind means the exposure survives the remediation, and the sweep may well report the store as fixed. ## Step 4: record the window, and resist calling it undone Capture created-at, detected-at, contained-at, remediated-at, what was written, and who had access. This matters for three separate audiences: the breach-assessment question of whether notification obligations were triggered, the audit question of what the control actually did, and the engineering question of how long your detective path takes end to end. And hold the line on language. "Remediated" means the state is now conforming. It does not mean the six hours did not happen. Data was written unencrypted, and if that data was sensitive, the correct statement is that the risk was realised and then mitigated — not that the control worked. ## Step 5: the finding is also evidence of a coverage gap Separately from the incident, this store came into existence without a pre-change control ever seeing it. That is the durable lesson, and it belongs in the follow-up, not in the middle of triage: which creation path produced it, and is it worth closing. The provider's activity log names the identity that made the create call, which usually answers that question immediately. ## What a weak answer looks like Jumping straight to "enable encryption" — which frequently is not possible — declaring the incident closed once the new store is live, forgetting snapshots and backups, and treating the fix as though it retroactively protected the six hours of writes. The interviewer is testing whether you understand that a detective control buys you a shorter exposure, never no exposure.

  • The owning team asks to leave it and enable encryption at the next rebuild. What do you say?
    That it is their risk to accept but it must be an explicit, time-bounded, written acceptance naming what the store holds and who can reach it — not a silence. If the data is sensitive or regulated, escalate rather than accept, and keep the containment measures in place regardless of when the migration lands.
  • What do you record so the finding is still useful three months later?
    Created-at, detected-at, contained-at and remediated-at; what the store held; which identity created it and by which path; who read from it during the window; and whether snapshots existed. That set answers the breach-notification question, the audit question and the how-slow-is-our-detection question at once.
  • How does encryption at rest actually reduce risk here?
    It defends against access to the underlying storage that bypasses the service's own controls — recovered media, a copied volume, a mishandled backup. It does nothing against an over-broad access policy or stolen credentials, so do not report the migration as having closed exposures it never addressed.

saying these in an interview costs you the question

  • Starts by trying to enable encryption on the existing store
  • Declares the incident closed once the new store is live
  • Leaves unencrypted snapshots and backups from the window
  • Skips rotating secrets that were written unencrypted
  • Reports the fast fix as proof the control worked

context