You need to answer ad-hoc questions about AWS API activity spanning the past year. Compare querying with Amazon Athena over your CloudTrail trail's S3 objects against using CloudTrail Lake, and say what decides it.
answer
- the raw objects are already in S3
- bytes scanned is the meter
- partitioning is the whole game
- managed store shifts cost to ingest
- query frequency decides the crossover
basics
~20 sAthena over the trail's S3 objects is cheap to store and pay-per-query on bytes scanned, but you own the table and its partitioning. CloudTrail Lake is a managed event data store queried with SQL, priced on ingestion and retention, and needs no plumbing. Query volume and setup effort decide it.
solid answer
~50 sWith a trail you already have the raw material: gzipped JSON objects in S3, partitioned by account, region and date. **Athena** reads them in place — you define a table, and the crucial detail is partitioning, because Athena bills on bytes scanned and an unpartitioned table makes every query read a year of objects. Set up partition projection and a query for one region and one week touches a sliver of the data. Storage stays at S3 prices, which is as cheap as this gets. **CloudTrail Lake** flips the cost model: you create an event data store, pay on data ingested and retained, and get SQL over it with no table, no partitions and no Athena setup — including the option to import events a trail already delivered to S3. Rule of thumb: high volume with occasional queries favours Athena; frequent investigation where analyst time is the scarce resource favours Lake.
code
sql · 9 linesSELECT eventtime,
eventname,
useridentity.arn AS principal,
sourceipaddress
FROM cloudtrail_logs
WHERE eventname = 'DeleteBucket'
AND awsregion = 'eu-west-1'
AND timestamp BETWEEN '2025/03/01' AND '2025/03/31'
ORDER BY eventtime;go deeper
Know that trail events land as compressed JSON objects in S3 and that querying them means pointing a query engine such as Athena at the bucket rather than browsing the console.
Explain how the S3 path layout becomes partitions, why bytes scanned is the meter, and what CloudTrail Lake removes from the setup in exchange for ingestion-based pricing.
Show the crossover reasoning with numbers in mind — volume ingested versus queries run per month — and name the third surface, CloudWatch Logs, as the short alerting window rather than the archive.
Own the layering decision for the estate: one durable system of record, one analysis surface chosen on query economics, and a short hot window, with a clear rule on who maintains each.
## What you already have A trail writes CloudTrail events into S3 as gzipped JSON, one object per batch, laid out under a path like `AWSLogs/<account-id>/CloudTrail/<region>/YYYY/MM/DD/`. That directory structure is the whole story for query cost, because it is what a query engine can use to avoid reading data. ## Athena over the trail bucket Athena runs SQL directly over those objects. The CloudTrail console can generate the table definition for you, mapping the JSON into columns — `eventtime`, `eventname`, `eventsource`, `awsregion`, `sourceipaddress`, `useragent`, `errorcode`, and the nested `useridentity` struct. The economics are dominated by one thing: **Athena bills on bytes scanned**. A table with no partitions means every query reads every object in the bucket, so "which principal called `DeleteBucket` last March?" scans a year of data to return four rows. With partition projection configured, Athena computes the partition values from the path convention instead of maintaining a catalogue of them, and a predicate that pins the region and date range restricts the scan to the matching prefixes. ```sql SELECT eventtime, eventname, useridentity.arn, sourceipaddress FROM cloudtrail_logs WHERE eventname = 'DeleteBucket' AND timestamp BETWEEN '2025/03/01' AND '2025/03/31'; ``` What you own in return: the table definition, the projection configuration, keeping it correct when accounts or regions are added, and the fact that columns like `requestparameters` arrive as JSON strings you have to pick apart. In exchange, the data at rest costs S3 prices and nothing else. ## CloudTrail Lake CloudTrail Lake replaces the plumbing with a managed **event data store**. You choose what it ingests, choose a retention period measured in years, and query it with SQL from the CloudTrail console or API. There are no partitions to design, no catalogue to maintain, and no separate query service to configure. It can also ingest events beyond a single trail — activity across an organization, and events copied in from data a trail already delivered to S3 — which matters when your history predates the decision to use Lake. The cost model moves to the front of the pipeline: you pay on the volume of data **ingested** into the store and on retaining it, with queries charged separately on the data scanned. That is a very different curve from S3-plus-Athena. If you ingest a firehose of data events and query twice a quarter, you have paid for a lot of ingestion you never used. If a security or platform team is running investigations weekly, the time saved is worth real money and the answer arrives in minutes rather than after a table-definition project. ## The third option, and its limits A trail can also deliver into CloudWatch Logs. That is the right surface for the *recent* window — it is where metric filters and alarms on API activity live, and where an operator looking at the last few hours will naturally be. It is the wrong surface for a year of history: you pay ingestion on everything you put in, and keeping a year of high-volume API activity there is usually the most expensive of the three. A sane layered answer to "where do CloudTrail events live" is therefore: S3 as the durable system of record, a short CloudWatch Logs window for alerting, and either Athena or Lake as the analysis surface over the long tail. ## What actually decides it - **Query frequency versus data volume.** Rare queries over huge volume → Athena. Frequent queries over moderate volume → Lake. - **Who is asking.** If the people running the queries are not the people who would maintain a partitioned Athena table, Lake removes a dependency. - **Existing investment.** A team already running a data-lake pattern with Glue and Athena has near-zero marginal cost to add one more table. - **How far back, and where the history is.** Lake retention is configured on the event data store and reaches years; if your existing history is in S3, factor in the ingestion cost of copying it in. - **Whether the data is already there.** You cannot query what you never delivered. Neither option retroactively creates events for a period when no trail existed. The answer an interviewer wants is not a winner but the axis: Athena keeps storage cheap and moves cost and effort to query time; Lake pays up front for ingestion and buys back setup and analyst time.
- Why is an unpartitioned Athena table over CloudTrail logs a cost problem rather than just a slow one?Because Athena charges on the bytes it scans, not on the rows it returns. Without partitions there is no way to skip objects, so a query for one day of one region reads every object in the bucket, and the bill scales with total history rather than with the question. Partition projection derives the partitions from the S3 path convention so predicates on region and date prune the scan.
- When would you keep CloudTrail events in CloudWatch Logs rather than only in S3?When you need to alarm on API activity or investigate the last few hours interactively. Metric filters over a log group turn matching events into a metric an alarm can watch, which S3 delivery cannot do. Keep the retention on that log group short, because you pay ingestion on everything you send and the durable copy already lives in the trail's S3 bucket.
- Your organisation adopts CloudTrail Lake today but needs to investigate something from eight months ago. Is that possible?Only if the events were delivered somewhere at the time. Lake can import events a trail already wrote to S3, so if that trail existed you can copy the relevant range into an event data store and query it — paying ingestion on what you import. If no trail existed and the window is outside the 90-day Event history, the record does not exist anywhere and no product recovers it.
saying these in an interview costs you the question
- Thinks Athena charges per query, not per byte scanned
- Skips partitioning because S3 storage looks cheap
- Assumes CloudTrail Lake is free for existing trail data
- Believes Lake can recover activity never delivered anywhere
- Keeps a year of API activity in CloudWatch Logs