You need to be able to state that every EBS volume, S3 object and RDS database created across a 200-account AWS organization is encrypted at rest. How do you get from "mostly encrypted" to a defensible guarantee, and what does each control you add actually buy?
answer
- prevent, detect, and backfill are different jobs
- defaults cover the common case only
- SCPs bind the account admin too
- condition-key coverage varies per service
- report coverage, never blanket compliance
basics
~20 sStack three layers: service defaults so the encrypted path is the easy one, preventive organization policies that deny creation when an encryption condition key is false, and detective scanning for the gaps and the backlog. Only the preventive layer, which binds account admins too, supports a guarantee.
solid answer
~50 s"Mostly encrypted" comes from relying on defaults and good intentions, and it fails the moment one team creates a resource by hand. I would build three layers. First, make the encrypted path the default everywhere — S3 encrypts new objects by default, and EBS has an account-and-Region setting that encrypts every new volume — so the common case needs no discipline. Second, add preventive service control policies that deny the creating API call when the relevant condition key says unencrypted, for example denying `ec2:CreateVolume` when `ec2:Encrypted` is false. That is the layer that produces a guarantee, because an SCP binds every principal in the account including its administrators. Third, run detection to find what predates the controls and the services that expose no usable condition key. Then I would report coverage per service rather than claiming a blanket compliance statement, and treat the pre-existing unencrypted estate as a separate migration.
go deeper
Know the difference between a control that stops something happening and one that reports it afterwards, and that turning on a default does not touch resources that already exist.
Describe the layers concretely: service defaults, a deny on the creating action conditioned on an encryption condition key, and continuous scanning for the rest. Be able to explain why each layer alone is insufficient.
Show the rollout judgment — scan first to learn what will break, stage by organizational unit, and treat the pre-existing estate as a separate migration with a per-service cost, including the snapshot-copy-and-restore path for databases.
Own the claim you make to an auditor and the exception process behind it: what is genuinely guaranteed versus reported, how exceptions are isolated rather than granted in place, and how you sequence 200 accounts so the programme survives its first broken deployment.
## Why "mostly" is the honest starting point Encryption coverage decays for structural reasons, not lazy ones. Defaults changed at different times across services, so old resources predate them. Some resources are created outside the paved path — a console click during an incident, a one-off script, a vendor's role provisioning something in your account. And services differ in whether encryption can be added later at all: an object can be rewritten, but an unencrypted RDS instance cannot be encrypted in place — you take a snapshot, copy the snapshot with encryption enabled, and restore from the copy, which is a cutover with downtime. Any plan that treats this as a single sweep will be wrong. ## Layer one: make the encrypted path the default The cheapest coverage comes from removing the decision. S3 applies encryption to new objects without anyone asking. EC2 exposes an account-level, Region-level setting that causes every newly created EBS volume to be encrypted; because it is per Region, turning it on means turning it on in every Region you use, in every account — which is itself a fleet-wide task, and a good early test of whether you can actually change 200 accounts. What defaults buy you is the *common case* for free. What they do not buy you is a guarantee: a default can be overridden by an explicit parameter in the API call, and a default set today says nothing about what existed yesterday. ## Layer two: preventive policy, which is where the guarantee comes from A guarantee needs a control that cannot be bypassed by the person creating the resource. That is a service control policy attached to an organizational unit: it constrains every principal in every account beneath it, including those accounts' own administrators and root users, and it applies regardless of which tool made the call. The shape is a `Deny` on the creating action with a condition on the service's encryption condition key — for example, denying `ec2:CreateVolume` when `ec2:Encrypted` is `false`. The important properties are that it fires at creation time, before any unencrypted data exists, and that it produces a hard, visible failure rather than a report someone may or may not read. Three things to plan for: - **It breaks things.** Some automation, some vendor integration, some CI job will be creating a resource without the parameter, and it will start failing. Roll out to a non-production OU first, watch what breaks, fix the callers, then widen. A preventive control deployed org-wide on day one is how these programmes get rolled back and never retried. - **Coverage is per service.** The condition key has to exist and has to be evaluated at the right call for this to work. Where it does not, this layer simply has no reach, and pretending otherwise is how the guarantee becomes false. - **It says nothing about the past.** SCPs are evaluated on requests, so every resource created before the policy is untouched. ## Layer three: detection, for the gaps the preventive layer cannot cover Detection is what fills the two holes above: the pre-existing estate, and the services with no usable preventive hook. Continuous configuration scanning with aggregated findings across the organization gives you an inventory of unencrypted resources with an owner attached, which is the input to the migration and the evidence for the audit. Its weakness is inherent — it finds things after they exist, so it is a reporting control, not a guarantee. Run it *before* the preventive layer as well as after. The findings tell you what the SCP is about to break. ## The backlog is a separate project Separate "nothing new is unencrypted" from "nothing old is unencrypted". The first is achievable in weeks and is what stops the problem growing. The second is a migration with a per-service cost model: rewriting S3 objects, snapshot-copy-and-restore for EBS volumes and RDS instances with a maintenance window each, and per-team coordination. Sequencing them the other way round — trying to clean up while the estate keeps growing — is the classic failure. ## What you actually report The defensible statement is not "everything is encrypted". It is: here are the services in scope; for each, here is the preventive control and the date it went live; here is current coverage from continuous detection; here is the remaining backlog with owners and dates; and here are the services where no preventive control exists, with what compensates. That version survives an auditor asking how you know, because each claim names the mechanism that makes it true. A blanket claim backed by a dashboard nobody can explain does not. ## The judgment being tested What this question really probes is whether you distinguish a control that *prevents* from a control that *reports*, whether you know that the preventive layer binds the account administrator, and whether you can say out loud which parts of the estate your guarantee does not cover. Candidates who claim total coverage are marked down; candidates who name their gaps and how they are shrinking are not.
- Why does a preventive organization policy give a stronger claim than a scanning rule that flags unencrypted resources?Because it changes what can exist rather than what you know about. The API call fails before unencrypted data is created, and it fails for every principal in the account including its administrator. A scanner reports after the fact, depends on someone acting on the finding, and leaves a window in which the data is real and unprotected.
- A team says the encryption SCP is blocking a legitimate deployment. How do you handle exceptions without hollowing out the guarantee?Not by carving a hole in the policy for a principal — that is the escape hatch everyone eventually uses. Move the workload to an OU with a different, documented policy set, time-box it, and record why. The guarantee then reads "encrypted everywhere except this named OU", which is still a true and auditable statement.
- What is the honest answer when an auditor asks whether every service in the account enforces encryption at rest?That enforcement is per service and you report it that way: the services with a preventive control and its live date, the services covered only by detection, and what compensates on the remainder. Claiming uniform coverage you cannot demonstrate is worse than naming a gap with a plan, because the first failure found destroys the credibility of the rest.
saying these in an interview costs you the question
- Claims a scanning rule proves nothing unencrypted can be created
- Assumes enabling a default retroactively encrypts existing resources
- Rolls a preventive policy org-wide with no staged rollout
- Believes an account administrator can be excluded from an SCP
- States blanket compliance without naming per-service coverage