Your team enabled AWS CloudTrail on day one. After a suspected data exfiltration you can see who changed an S3 bucket policy, but you cannot find which objects were downloaded. Why, and what would have made those reads visible?
answer
- control plane versus data plane
- one category is opt-in
- volume and per-event billing
- object reads were never recorded
- scope selectors by ARN prefix
basics
~20 sCloudTrail records management events (control-plane calls such as PutBucketPolicy) by default, but data events — S3 object-level GetObject and PutObject, Lambda Invoke, DynamoDB item operations — are off unless you explicitly enable them, and they are billed per event.
solid answer
~40 sCloudTrail splits events into two categories. Management events are control-plane calls — creating a bucket, changing its policy, attaching an IAM policy — and a trail logs them by default. Data events are the data plane: `s3:GetObject` and `s3:PutObject` on objects, `lambda:InvokeFunction`, DynamoDB item-level calls. They are **off by default** and charged per event, which is why almost nobody has them on before their first incident. So the bucket-policy change is in the trail and the object reads are simply not recorded anywhere. The fix is to add a data-event selector on the trail — using advanced event selectors you can scope it to `AWS::S3::Object` under a specific bucket or key prefix, and to write-only or read-only, so you pay for the sensitive data rather than for every object in the estate.
code
bash · 3 linesaws cloudtrail put-event-selectors \
--trail-name org-audit-trail \
--advanced-event-selectors '[{"Name":"Regulated bucket object reads","FieldSelectors":[{"Field":"eventCategory","Equals":["Data"]},{"Field":"resources.type","Equals":["AWS::S3::Object"]},{"Field":"resources.ARN","StartsWith":["arn:aws:s3:::regulated-bucket/"]},{"Field":"readOnly","Equals":["true"]}]}]'go deeper
Be able to say that CloudTrail logs management (control-plane) events by default and that S3 object reads are data events you must switch on, so "CloudTrail is enabled" does not mean everything is recorded.
Explain the volume and per-event-cost reason behind the default, name the data-event resource types (S3 objects, Lambda invocations, DynamoDB items), and show how advanced event selectors scope logging to a bucket or prefix.
Demonstrate the incident-response consequence: the gap cannot be filled after the fact, so data events are a threat-model decision made per bucket in advance. Include the readOnly filter choice and the fact that PutEventSelectors replaces the whole set.
Own the policy question — which data classes justify object-level audit across the whole estate, how that cost is budgeted and charged back, and what standing control keeps a newly created regulated bucket from being born without a selector.
## Two categories, two defaults CloudTrail classifies everything it can record into event categories, and the two that matter here behave completely differently. **Management events** describe operations on the resources themselves — the control plane. `CreateBucket`, `PutBucketPolicy`, `PutBucketAcl`, `AttachRolePolicy`, `RunInstances`, `AssumeRole`. When you create a trail, management events are included by default, and CloudTrail also keeps the last 90 days of management events per region in the console's **Event history** view even with no trail at all. That default is why the bucket-policy change was there waiting for you. **Data events** describe operations on the data *inside* a resource — the data plane. The common ones are S3 object-level calls (`GetObject`, `PutObject`, `DeleteObject`), `lambda:InvokeFunction`, and DynamoDB item-level operations such as `GetItem` and `PutItem`. These are **not logged unless you configure a selector**, they are not in Event history, and they are billed per event delivered. The asymmetry is deliberate and is about volume. A busy account might generate a few thousand management events a day. The same account's S3 fleet might serve millions of object reads a minute. Charging and logging those by default would make CloudTrail unusable, so AWS made the high-volume half opt-in. ## Why this is the classic incident-response surprise The failure mode is not "CloudTrail was off". It is "CloudTrail was on and everyone assumed that meant everything was recorded". During the incident you can reconstruct the *setup* perfectly — the moment the bucket policy was loosened, by which principal, from which IP, with which user agent — and then hit a wall at the only question that matters to legal and to your customers: which objects actually left. That gap cannot be filled retroactively. There is no archive to enable after the fact; the events were never generated. This is the reason data events belong in a threat model, not a budget review. The decision is "for which buckets is object-level read history worth paying for", and the honest answer for the two or three buckets holding regulated data is always yes. ## Scoping data events so the bill stays sane Modern trails use **advanced event selectors**, which let you filter on fields rather than toggling a whole category. The fields you will actually use are `eventCategory` (`Management` or `Data`), `resources.type` (`AWS::S3::Object`, `AWS::Lambda::Function`, `AWS::DynamoDB::Table`), `resources.ARN` with operators such as `StartsWith`, and `readOnly`. ```bash aws cloudtrail put-event-selectors \ --trail-name org-audit-trail \ --advanced-event-selectors '[{"Name":"Reads in the regulated bucket","FieldSelectors":[{"Field":"eventCategory","Equals":["Data"]},{"Field":"resources.type","Equals":["AWS::S3::Object"]},{"Field":"resources.ARN","StartsWith":["arn:aws:s3:::regulated-bucket/"]}]}]' ``` That records object-level activity for one bucket and nothing else. Two refinements are worth knowing. Filtering on `readOnly` lets you keep writes and drop reads, or the reverse — for an exfiltration threat model, reads are the ones you want, which is the opposite of the instinct to log only mutations. And scoping by key prefix means a bucket that mixes a sensitive prefix with a high-volume public one can be logged partially. Note that a trail's event selectors are replaced wholesale by `PutEventSelectors`, so a change must resubmit the full set, not just the addition. ## What is still missing Even with data events on, CloudTrail records the API call, not the payload. You learn that a principal called `GetObject` on `s3://regulated-bucket/customers/2026-08.csv` at a time and from an IP; you do not learn what was in the file. For an exfiltration report that is usually enough, because the object key plus your own data catalogue tells you what was in it. Delivery is also not instantaneous — events reach the S3 bucket in minutes, so a trail is forensic evidence rather than a real-time trip wire. If you need alerting, the trail feeds a detection product; the trail itself is the record. Finally, do not confuse S3 server access logging with CloudTrail data events. Server access logs are a separate, best-effort S3 feature with no delivery guarantee and a different record format. When an auditor or an incident asks who read an object, the answer that holds up is the CloudTrail data event. ## The interview shape The expected answer is three beats: name the split (management on by default, data events opt-in and metered), explain the volume reason behind the default, and then show that you would scope selectors by resource ARN prefix and read/write rather than turning everything on. Candidates who stop at "you should have enabled data events" miss the part interviewers actually probe — the cost decision that made everyone skip it in the first place.
- For an exfiltration threat model, would you log read-only or write-only data events, and why?Reads. The instinct is to log mutations because those change state, but exfiltration is a read: the attacker's whole goal is to copy data out without altering it. Filtering on the `readOnly` field keeps GetObject history while dropping the write traffic, which in a mostly-read bucket is the more expensive half but the one that answers the question.
- Someone proposes using S3 server access logging instead of CloudTrail data events to save money. What do you say?They are not equivalent. S3 server access logging is best-effort with no delivery guarantee and a different record format; CloudTrail data events are the auditable record, integrate with the trail's integrity validation and organization-wide delivery, and carry the calling principal in a form an auditor accepts. Save money by scoping selectors to sensitive prefixes, not by downgrading the evidence.
- Your trail already has selectors and you add a new one with PutEventSelectors. What is the trap?PutEventSelectors replaces the trail's selector set wholesale rather than appending. Submitting only the new selector silently drops the existing ones, so a change must resubmit the complete set. It is a quiet way to stop logging the bucket you were already watching.
saying these in an interview costs you the question
- Assumes enabling CloudTrail records every API call
- Thinks Event history includes object-level reads
- Cannot name the management vs data event split
- Proposes enabling data events on all buckets estate-wide
- Believes data events can be recovered retroactively