You are switching on AWS detection and compliance services across a 200-account AWS Organization. How do you decide what to enable where, and what keeps the bill and the finding volume from becoming unmanageable?
answer
- start at the org, not the account
- new accounts must inherit it
- each service bills on a different unit
- one mandatory standard, not all of them
- the silent failure is no findings
basics
~20 sEnable through Organizations with a delegated security account and auto-enable for new accounts, cover every allowed Region, and treat AWS Config's recorder scope and Macie's scan targets as the main cost dials. Suppress known exceptions at ingestion, not by ignoring dashboards.
solid answer
~60 sStart from the org, not the account: nominate a delegated administrator in a dedicated security account for GuardDuty, Security Hub, Config, Inspector and Macie, and turn on **auto-enable for new accounts**, otherwise coverage decays the moment someone creates account 201. Cover every Region you actually permit, and use an SCP to deny the ones you do not, so "enable everywhere" stays a finite list. Then tune the cost dials, which are not the same per service: **Config** bills per configuration item and rule evaluation, so scope the recorder's resource types and consider daily recording for eligible ones; **GuardDuty** bills per GB of telemetry analysed, and its optional protection plans are separately priced; **Macie** bills per GB inspected, so run automated discovery on sampled buckets rather than full jobs across a data lake; **Security Hub** bills per finding ingested and per control check. For volume, enable a small mandatory standard rather than all of them, rely on consolidated control findings, and use Security Hub automation rules to suppress or downgrade recorded exceptions at ingestion. Coverage itself should be monitored — a service silently off in one account is the failure that matters most.
go deeper
Know that these services are enabled per account and per Region and that AWS Organizations provides a central way to switch them on rather than doing it account by account.
Explain the delegated administrator and auto-enable model, and name what each service actually bills for — configuration items, gigabytes analysed, resources scanned, findings ingested.
Reason about the specific cost dials, especially Config recorder scope, and about routing only actionable findings into a response path while suppressing recorded exceptions at ingestion.
Own the tradeoff explicitly: define a centrally-funded mandatory floor, an opt-in tier above it, an exception process with owners and expiry, and coverage metrics that catch the silent gaps a green dashboard hides.
## Frame it as three separate problems A 200-account rollout is not one decision. It is (1) how enablement happens and stays true, (2) what it costs and which dials control that, and (3) whether anyone can act on the output. Candidates who answer only the first have described a project, not a programme. ## 1. Enablement that does not decay The org-native pattern is the same for every service in this family: - **Delegated administrator** in a dedicated security-tooling account — not the Organizations management account — so day-to-day security operations never need credentials in the most privileged account. - **Auto-enable for new accounts**, so account creation cannot outrun the security baseline. Manual enablement at 200 accounts is guaranteed drift. - **Region strategy**: every one of these services is regional and blind outside where it is enabled. Decide the list of permitted Regions, deny the rest with a Service Control Policy, and enable across the permitted list. This turns "enable everywhere" from an unbounded cost into a bounded one, and closes the classic gap where activity happens in a Region nobody watches. - **A Config aggregator and a Security Hub aggregation Region** so the result is one queue, not 200. The deployment mechanism itself (templates, pipelines) belongs to your infrastructure-as-code and delivery practice; what you own here is the requirement that enablement is declared centrally rather than performed by hand. ## 2. The cost dials, per service They do not share a pricing shape, and the naive "turn everything on everywhere" produces a bill dominated by one or two of them: - **AWS Config** is usually the surprise. It charges per configuration item recorded and per rule evaluation. An account with rapidly-scaling Auto Scaling groups generates enormous CI volume from ENIs and instances. Dials: scope the recorder to the resource types you actually have rules for, exclude the noisiest types, and use daily rather than continuous recording where eligible. Note the tension: Security Hub controls depend on Config, so cutting the recorder too aggressively quietly disables controls. - **GuardDuty** charges per GB of CloudTrail events, flow logs and DNS analysed. Foundational coverage is generally worth it unconditionally; the **optional protection plans** (EKS, RDS, Lambda, malware scanning, runtime monitoring agents) are the per-workload judgment call, and are reasonable candidates for enabling only where the workload exists. - **Macie** charges per GB inspected. Automated sensitive-data discovery with sampling across buckets is the sane default; full classification jobs across a multi-petabyte data lake are a deliberate, scoped exercise, not a standing configuration. - **Amazon Inspector** charges per scanned resource, so its cost tracks fleet size predictably, and per-scan-type enablement lets you skip targets you do not run. - **Security Hub** charges per finding ingested and per control check, which is why enabling every available standard multiplies both. A useful principle: the services with **flat, predictable, per-resource** pricing (GuardDuty foundational, Inspector) are good candidates for a non-negotiable org-wide baseline. The services with **volume-driven** pricing (Config, Macie) need per-account scoping and an owner watching the spend. ## 3. Output nobody can act on is not security At 200 accounts, aggregate volume becomes the binding constraint before coverage does. The levers: - Enable a **small mandatory standard** — one foundational best-practices set — before adding regulator-driven ones on the accounts that need them, rather than every standard everywhere. - Use **consolidated control findings** so a control shared by several standards produces one finding. - Use **Security Hub automation rules** to suppress or re-severity findings at ingestion against recorded exceptions, and GuardDuty suppression rules for known-benign detections. An exception should be a written, expiring decision, not an unread row. - Route only what has an owner: high-severity findings to a real response path via EventBridge; the rest to a periodic review. - Distinguish **findings** from **posture**: a slowly-improving control score is a programme metric; an active GuardDuty finding is an event. ## 4. The metric that actually matters Monitor **coverage**, not only findings. Every service here fails silently: switched off in one account, or in one Region, or with a Config recorder scoped past the resource type in question — and the dashboard stays green. Compare enabled-account counts against the Organizations account list, and Inspector or Config coverage against actual resource inventory, and alarm on divergence. The genuinely dangerous state in a large org is not too many findings; it is a corner of the estate that produces none because nobody is looking at it. ## The tradeoff to state out loud There is no configuration that is simultaneously complete, cheap and quiet. Complete coverage costs money and produces noise; cutting cost creates blind spots; cutting noise risks suppressing the one finding that mattered. The defensible position is an explicit floor that is mandatory and centrally funded, a tier above it that teams opt into for their workload profile, and every exception recorded with an owner and an expiry.
- Teams complain that AWS Config is the single largest line item in their security spend. What do you change?Look at configuration-item volume first — usually a small number of high-churn resource types in scaling fleets. Scope the recorder to types you have rules for, exclude the noisiest, and use daily recording where eligible. Then check which rules the spend is buying: rules nobody acts on should be removed. Keep in mind that trimming too far disables the Security Hub controls that depend on Config.
- How do you stop the org-wide baseline from being quietly disabled by an account owner?Deny the disabling actions with a Service Control Policy so member accounts cannot turn off the detectors or delete the recorder, leaving only the delegated administrator able to change them. Pair it with alarming on coverage divergence, because SCPs cover the API path while monitoring covers the cases an SCP does not reach, such as a Region nobody enabled in the first place.
- Which parts of this baseline would you make mandatory versus optional per team?Mandatory: GuardDuty foundational coverage, Security Hub with one foundational standard, Config recording of the resource types those controls need, and Inspector for compute — all centrally funded so cost is never a reason to opt out. Optional per workload: GuardDuty's protection plans for EKS, RDS or runtime monitoring, Macie scanning scope, and additional compliance standards driven by a specific regulator.
saying these in an interview costs you the question
- Enables everything everywhere with no cost model
- Leaves enablement manual so new accounts drift uncovered
- Uses the Organizations management account as the security account
- Ignores that all of these services are per-Region
- Measures success by number of findings rather than coverage