skip to content

Observability & Cost

The two things that decide whether a system survives in production: knowing what it is doing and knowing what it costs. You cover CloudWatch metrics, logs, alarms and dashboards, X-Ray tracing, and the cost tooling, purchasing models, and Well-Architected reviews behind the bill.

part ofAWSoverview, primer and where to startread it →
on this pageshow

explore

questions

page 1 of 2

AWS CloudTrail shows recent API activity in the console under Event history without any setup. What does Event history actually cover, and what does creating a trail add?

level: juniorimportance: must knowfreq 72%

answer

  1. free, on by default, nothing to enable
  2. a rolling ninety-day window
  3. control-plane calls only
  4. a trail is delivery, not recording
  5. opt-in data events, S3 lifecycle retention

basics

~20 s

CloudTrail Event history is free, always on, and covers only the last 90 days of management events in the current region. A trail is the delivery configuration that writes events to an S3 bucket you own, giving you retention you control and the option to log data events.

solid answer

~40 s

CloudTrail always records management-plane API calls, and the console's **Event history** exposes the last 90 days of them for the region you are looking at, free and with nothing to enable. Three things it will not do: go back further than 90 days, show data events such as S3 object-level `GetObject` or Lambda invocations, or feed anything downstream. A **trail** fixes all three — it continuously delivers events as gzipped JSON into an S3 bucket you own (optionally also into CloudWatch Logs), so retention becomes an S3 lifecycle decision rather than a fixed 90 days, and you can opt into data events and Insights events. What a trail does not buy you is speed: delivery is on the order of minutes, so anything that must react quickly should hang off EventBridge instead.

code

bash · 5 lines
bash
# Query the 90-day Event history directly - no trail required
aws cloudtrail lookup-events \
  --lookup-attributes AttributeKey=EventName,AttributeValue=ConsoleLogin \
  --max-results 10 \
  --query 'Events[].{Time:EventTime,User:Username,Name:EventName}'

go deeper

for a junior

Be ready to say plainly that management events are recorded automatically and visible for 90 days, and that a trail is what you create to keep them longer or to send them to S3.

for a middle

Explain the mechanics: what a trail delivers, where the objects land in the bucket, why data events are a separate opt-in, and why delivery takes minutes rather than seconds.

for a senior

Show that you would have created the multi-region trail before the incident, and that you know which questions Event history can still answer during triage when no trail exists.

for a principal

Own the standard: a management-event trail in every account as a baseline because the first copy is effectively free, with data events treated as a per-workload cost decision rather than a default.

## What is already on CloudTrail records management-plane API activity in every AWS account with nothing enabled. For each call it captures who made it (the IAM principal and, for assumed roles, the session), what was called (`eventSource` and `eventName`), when, from which source IP and user agent, the request parameters, and whether it failed and why. The console surfaces this as **Event history**, and the same data is reachable programmatically through the `LookupEvents` API. ```bash aws cloudtrail lookup-events \ --lookup-attributes AttributeKey=EventName,AttributeValue=TerminateInstances \ --max-results 5 ``` This costs nothing, and you cannot turn it off. ## The three limits of Event history **Ninety days, rolling.** As of 2025 the window is 90 days. Activity older than that is simply not there — no setting extends it, and if you had no trail configured at the time, the record is gone for good. This is the single fact that catches teams out: the first day anyone needs six-month-old activity is the day they discover nobody created a trail. **Management events only.** Management events are control-plane calls — `RunInstances`, `CreateBucket`, `AttachRolePolicy`, `AssumeRole`. Data events, the data-plane calls such as S3 object-level `GetObject`/`PutObject`, `lambda:Invoke`, and DynamoDB item-level operations, are off by default and never appear in Event history at all. Neither do Insights events. **One region at a time, and global services are special.** The console view is scoped to the region you are in, so a call made in eu-west-1 will not show up while you are looking at us-east-1. Globally-scoped services such as IAM and CloudFront record their events in us-east-1 regardless of where you were sitting when you made the call. Event history also only lets you look events up by a single attribute — an event name, a user name, a resource — not run an arbitrary aggregation across them. ## What a trail is A trail is a *delivery* configuration, not a switch that starts the recording. You point it at an S3 bucket you own and CloudTrail writes batches of events there as gzipped JSON objects, under a path shaped like: ``` AWSLogs/<account-id>/CloudTrail/<region>/YYYY/MM/DD/ ``` From that one change, several things follow: - **Retention becomes yours.** How long the record survives is now an S3 lifecycle decision, not a fixed 90 days. - **The data becomes queryable in bulk.** Objects in a bucket can be read by Athena, copied into CloudTrail Lake, or processed by anything else you like. - **Data events become available**, if you explicitly opt in and scope them — they are billed per event, so this is a deliberate choice rather than a default. - **CloudWatch Logs delivery becomes possible**, which is what you configure when you want metric filters and alarms over API activity. - **A multi-region trail** captures activity from every region into that single bucket, which is what makes "look everywhere" practical. ## What a trail does not add: latency A common misconception is that a trail makes CloudTrail real time. It does not. Events reach S3 in minutes rather than seconds, because CloudTrail batches them into objects. That is fine for investigation and for reporting, and unfit for anything that must respond immediately. When you need a fast reaction to an API call, the mechanism is Amazon EventBridge: management events are published to the account's default event bus with a detail type of `AWS API Call via CloudTrail`, and a rule can match on the event source and name and invoke a target directly. For object-level S3 activity, either enable S3's own EventBridge notifications or have a trail logging the relevant data events. ## The cost shape Event history is free. For a trail, the first copy of management events delivered per trail is free; a second trail recording the same management events is billed, as are data events (per event, with no free allowance) and the S3 storage and any CloudWatch Logs ingestion you add on top. That asymmetry is why the standard advice is to keep management events on everywhere and to treat data events as something you turn on deliberately, per resource. ## When Event history is enough For "which principal terminated that instance yesterday afternoon?", Event history answers in under a minute and needs nothing. As soon as the question spans more than 90 days, needs data-plane calls, or has to be answered by a machine rather than a human clicking, you need a trail — and you needed it *before* the event you are investigating.

  • How quickly does a trail deliver an event, and what would you use if you needed to react to an API call faster than that?
    Trail delivery to S3 is measured in minutes, because CloudTrail batches events into objects. For a fast reaction, match the call in EventBridge instead: management events arrive on the account's default bus with the detail type `AWS API Call via CloudTrail`, and a rule can invoke a Lambda function or an SNS topic directly. Use the trail as the durable record and EventBridge as the trigger.
  • If someone deletes a trail, what happens to the log files it already delivered?
    Nothing — the objects already written stay in the S3 bucket, because they are ordinary S3 objects governed by your bucket's lifecycle and retention settings, not by CloudTrail. Deleting the trail only stops future delivery. Event history is likewise unaffected and keeps showing the last 90 days of management events.
  • You have a multi-region trail. Where do IAM API calls appear?
    Globally-scoped services such as IAM and CloudFront record their events in us-east-1, whatever region you were calling from. A multi-region trail captures those global service events and delivers them to the same bucket, which is one of the practical reasons to create trails as multi-region rather than one per region.

saying these in an interview costs you the question

  • Thinks CloudTrail keeps all history forever by default
  • Expects S3 object GET calls in Event history
  • Believes nothing is recorded until you create a trail
  • Treats trail delivery as real-time, seconds not minutes
  • Assumes one region's console view shows every region

context

open as a page

What is an Amazon CloudWatch dashboard actually made of, what kinds of widget can it hold, and why does putting a metric on a CloudWatch dashboard not mean anyone will be told when it moves?

level: juniorimportance: must knowfreq 50%

basics

~20 s

A CloudWatch dashboard is a JSON document of widgets — metric graphs, log tables, alarm-status tiles, text — and each widget carries its own region, statistic and period. Dashboards only draw data; notification comes from CloudWatch alarms, never from a dashboard.

open as a page

An application on an EC2 instance writes to a log file on disk and you want those lines in Amazon CloudWatch Logs. What are the ways log data gets into CloudWatch Logs, and what does each path require you to install or grant?

level: juniorimportance: must knowfreq 65%

basics

~20 s

Three paths: the CloudWatch agent tailing files on a host, AWS services delivering logs natively (Lambda, the ECS awslogs driver, VPC Flow Logs), or your code calling the PutLogEvents API. All three need IAM permission to create streams and put events.

open as a page

Walk through the anatomy of a CloudWatch Logs Insights query — the pipeline of commands and the @-prefixed fields that are always available — and show how you would pull the 20 most recent lines containing ERROR out of one log group.

level: juniorimportance: must knowfreq 60%

basics

~20 s

A Logs Insights query is a pipeline of commands joined by the pipe character: fields or display picks columns, filter narrows rows, sort orders them, limit caps the output. Every event exposes @timestamp, @message, @logStream and @ingestionTime.

open as a page

Your AWS bill is up sharply month over month and nobody knows why. Walk through how you would use AWS Cost Explorer — granularity, Group by and filters — to narrow the increase down to a specific cause.

level: juniorimportance: must knowfreq 70%

basics

~20 s

Set AWS Cost Explorer to Daily granularity to find the day spend jumped, Group by Service to name the culprit, then filter to that service and regroup by Usage Type, Region or Linked Account until the charge line is identified.

open as a page

AWS bills EC2 compute at On-Demand rates by default. What are the main ways to pay less for the same compute, and what does each one ask of you in return?

level: juniorimportance: must knowfreq 76%

basics

~20 s

Three purchase axes exist. On-Demand pays the list rate for total flexibility. Commitments — Savings Plans or Reserved Instances — trade a one- or three-year pledge for a large discount. Spot buys spare capacity cheaply, but AWS can reclaim it.

open as a page

The AWS Well-Architected Framework organises its guidance into pillars. Name the pillars and say what each one asks about a workload.

level: juniorimportance: must knowfreq 62%

basics

~10 s

The AWS Well-Architected Framework has six pillars: Operational Excellence, Security, Reliability, Performance Efficiency, Cost Optimization and Sustainability. Each pillar groups design questions and best practices used to assess a workload from one angle.

open as a page

AWS CloudTrail data events are off by default. Explain how they differ from management events, why switching them on account-wide can dominate the CloudTrail bill, and how you would scope them.

level: middleimportance: must knowfreq 62%

basics

~20 s

Management events are control-plane calls and their first copy per trail is free; data events are data-plane calls such as S3 object reads and Lambda invocations, billed per event with no free allowance. Volume is why they dominate the bill, so scope them with event selectors.

open as a page

You already alarm on your load balancer's 5xx count and target response time. What does an Amazon CloudWatch Synthetics canary add on top of that, and what does running one actually cost you in moving parts?

level: middleimportance: must knowfreq 58%

basics

~20 s

A CloudWatch Synthetics canary is a scheduled script that exercises your endpoint the way a user would, so you detect outages with no traffic to measure and catch failures outside your servers — DNS, certificates, CDN, third parties. It is a managed Lambda, with its own role, logs and S3 artifacts.

open as a page

In Amazon CloudWatch Logs, what is the relationship between a log group and a log stream, and at which of the two do you configure retention? What happens to data already stored when you change that setting?

level: middleimportance: must knowfreq 72%

basics

~20 s

A log group is the container and the unit of configuration; a log stream is one ordered sequence of events from one source inside it. Retention is set per log group, defaults to never expire, and applies retroactively — older events are deleted.

open as a page

You need a business metric — orders placed per minute — visible in CloudWatch from a Lambda function. Compare publishing it with the CloudWatch PutMetricData API against the CloudWatch Embedded Metric Format, and explain what adding a dimension does to the metric and to the bill.

level: middleimportance: must knowfreq 58%

basics

~20 s

PutMetricData is a synchronous API call you wait on and pay per request; the Embedded Metric Format lets you print one structured JSON log line that CloudWatch converts into a metric asynchronously. Every distinct combination of dimension values becomes its own billable metric.

open as a page

A CloudWatch alarm on your service's error-count metric stayed in INSUFFICIENT_DATA throughout a real outage instead of firing. Explain how a CloudWatch alarm evaluates a metric — period, evaluation periods, datapoints to alarm — and how the missing-data treatment decides what happens when the metric stops arriving.

level: middleimportance: must knowfreq 74%

basics

~20 s

A CloudWatch alarm counts how many of the last N periods breached the threshold and fires when M of them did. During the outage the metric stopped being published, so there was nothing to compare, and the default missing-data treatment leaves the alarm unevaluated rather than breaching.

open as a page

Your EC2 instances and S3 buckets already carry a `CostCenter` tag, but AWS billing reports still show that spend as untagged. Explain what a cost allocation tag is, how user-defined tags differ from AWS-generated ones, and what has to happen before a tag can group costs.

level: middleimportance: must knowfreq 60%

basics

~20 s

Tagging a resource does not change billing by itself. A tag key only groups cost after it is activated as a cost allocation tag in the Billing console of the management (payer) account, which takes up to 24 hours and applies to usage from then on.

open as a page

A team asks you to "cap" their AWS spend at a fixed amount each month. Explain what AWS Budgets can and cannot do for them, including actual versus forecasted alert thresholds and what budget actions add.

level: middleimportance: must knowfreq 62%

basics

~20 s

AWS Budgets notifies, it does not cap. A cost budget alerts on actual or forecasted thresholds; only an attached budget action — applying a restrictive IAM policy or SCP, or stopping EC2 and RDS instances — changes anything on its own.

open as a page

On an AWS bill, which network traffic is free and which shows up as a data-transfer charge? Cover traffic between two instances in one Availability Zone, traffic between Availability Zones in the same Region, and traffic out to the internet.

level: middleimportance: must knowfreq 65%

basics

~20 s

Inbound internet traffic is free; outbound is charged per gigabyte. Traffic crossing Availability Zones inside a Region is billed in both directions. Same-AZ traffic is free over private IPv4 but charged when it uses public or Elastic IP addresses.

open as a page

Compare an AWS Compute Savings Plan with an EC2 Instance Savings Plan: what does each one commit you to, and what flexibility do you give up for the deeper discount?

level: middleimportance: must knowfreq 64%

basics

~20 s

Both commit a fixed dollar-per-hour spend for one or three years. A Compute Savings Plan applies to EC2 in any family or region and also to Fargate and Lambda. An EC2 Instance Savings Plan is locked to one instance family in one region and discounts more deeply.

open as a page

When instrumenting a service with the AWS X-Ray SDK, when should a key/value go on a segment as an annotation versus as metadata, and what does that choice change about finding the trace later?

level: middleimportance: must knowfreq 58%

basics

~20 s

Annotations are indexed by X-Ray and searchable in filter expressions; metadata is stored with the segment but not indexed. Put the few fields you will search traces by in annotations, and everything else in metadata.

open as a page

A marketing launch will multiply your AWS workload's traffic next month. Using Service Quotas and Trusted Advisor, how do you make sure an AWS account limit is not the thing that fails first?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Enumerate the quotas each service in the path consumes, read the applied values per account and Region in Service Quotas, and compare them against projected peak. Request increases on adjustable quotas weeks ahead, design around non-adjustable ones, and alarm on quota utilisation before launch day.

open as a page

What changes when you switch an AWS Lambda function's tracing mode from PassThrough to Active, what segments does the resulting trace contain, and why might that function still appear as a trace of its own instead of joining its caller's?

level: seniorimportance: must knowfreq 62%

basics

~20 s

Active tracing makes the Lambda service record the invocation itself, producing a service segment plus a function segment with Initialization, Invocation and Overhead subsegments. The function joins its caller's trace only if the caller propagated the trace header and the execution role can write to X-Ray.

open as a page

An EC2 instance shows CPU and network metrics in CloudWatch, but there is no memory-used or disk-space metric for it anywhere. Why not, and what do you have to run to get them?

level: juniorimportance: should knowfreq 62%

basics

~20 s

CloudWatch's built-in EC2 metrics are collected outside the instance, at the virtualization layer, so they cover CPU, network and EBS activity but never anything inside the guest. Memory and filesystem usage require installing the CloudWatch agent.

open as a page

In AWS X-Ray, what is the difference between a segment and a subsegment, and how does X-Ray turn the segments it receives into the service map shown in the console?

level: juniorimportance: should knowfreq 60%

basics

~20 s

A segment is what one service records for one request; subsegments break that segment into finer steps such as downstream calls. X-Ray links all segments sharing a trace id and aggregates them into the service map graph.

open as a page

An EC2 security group's rules are not what your team expected. What do AWS CloudTrail, AWS Config and Amazon CloudWatch Logs each record, and which question is each one the right tool for?

level: middleimportance: should knowfreq 54%

basics

~20 s

CloudTrail records the API calls made against AWS itself. AWS Config records the resulting configuration state of a resource over time. CloudWatch Logs stores log lines emitted by applications and services. Use CloudTrail for the call, Config for the state, CloudWatch Logs for the effect on your workload.

open as a page

A front-end team wants to know what real browsers experience on your site, not just what the servers logged. What does Amazon CloudWatch RUM collect, how does a browser get permission to send that data, and which levers control its volume?

level: middleimportance: should knowfreq 32%

basics

~20 s

CloudWatch RUM embeds a JavaScript client in your pages that reports page-load performance, Core Web Vitals, JavaScript errors, failed HTTP calls and session/browser context to an app monitor. The browser authorizes with temporary guest credentials from an Amazon Cognito identity pool, and a session sample rate controls volume.

open as a page

During an incident you need to search a dozen CloudWatch log groups belonging to different services, some of them in a second AWS account, from a single Logs Insights query. How do you do that, and what has to be in place first?

level: middleimportance: should knowfreq 40%

basics

~20 s

One Logs Insights query can span many log groups: select them individually or by name prefix, and use the @log field to tell events apart. Reaching another account first requires CloudWatch cross-account observability linking that account to your monitoring account.

open as a page

An application writes plain-text lines such as `2026-03-01T10:00:00Z level=ERROR path=/checkout latency=812ms` into a CloudWatch log group. Using CloudWatch Logs Insights, how would you turn those unstructured lines into a 95th-percentile latency per path in five-minute buckets?

level: middleimportance: should knowfreq 55%

basics

~10 s

Use parse to lift path and latency out of @message into named fields, then aggregate with stats: parse @message "path=* latency=*ms" as path, latency | stats pct(latency, 95) as p95 by path, bin(5m).

open as a page

What is a metric filter in Amazon CloudWatch Logs, and why does a metric filter you just created often report nothing even though matching lines are visibly present in the log group?

level: middleimportance: should knowfreq 48%

basics

~20 s

A metric filter watches a log group for a pattern and publishes a CloudWatch metric when lines match. It reports nothing at first because filters apply only to events ingested after creation — they never backfill — and because non-matching periods publish no data point unless defaultValue is set.

open as a page

AWS Cost Explorer lets you chart cost as Unblended, Blended or Amortized. What does each metric mean, and which one do you use to explain a month that contained a large upfront Savings Plan or Reserved Instance payment?

level: middleimportance: should knowfreq 52%

basics

~20 s

Unblended is the rate an account was actually charged as usage occurred; blended averages rates across a consolidated billing family; amortized spreads upfront Reserved Instance and Savings Plan fees over the commitment term. Use amortized for the upfront month.

open as a page

An AWS Lambda function is billed in GB-seconds. A colleague proposes cutting cost by lowering its memory setting from 1024 MB to 512 MB. Explain when that actually lowers the bill and when it raises it.

level: middleimportance: should knowfreq 48%

basics

~20 s

Lambda bills memory multiplied by duration, and CPU is allocated in proportion to memory. Halving memory halves the per-millisecond price but can more than double the runtime of CPU-bound work, raising total cost. It only saves money for functions that mostly wait on I/O.

open as a page

You inherit an AWS account whose monthly bill keeps climbing although no new workload has shipped for a year. Which classes of resource keep charging after whatever needed them is gone, and how would you find them?

level: middleimportance: should knowfreq 45%

basics

~20 s

Storage and reserved capacity outlive the compute that created them: unattached EBS volumes, accumulating snapshots and AMIs, public IPv4 addresses now billed hourly whether used or not, idle load balancers, and non-production environments running around the clock.

open as a page

Savings Plans have largely replaced Reserved Instances for new AWS EC2 commitments. What can a Reserved Instance still do that a Savings Plan cannot, and when would you buy a standard rather than a convertible RI?

level: middleimportance: should knowfreq 52%

basics

~20 s

Reserved Instances can reserve capacity when scoped to an Availability Zone, can be exchanged if convertible, and can be sold on the Reserved Instance Marketplace if standard. Savings Plans do none of these — and services such as RDS and ElastiCache still only offer reserved instances.

open as a page

showing 1–30 of 60