Observability & Cost
The two things that decide whether a system survives in production: knowing what it is doing and knowing what it costs. You cover CloudWatch metrics, logs, alarms and dashboards, X-Ray tracing, and the cost tooling, purchasing models, and Well-Architected reviews behind the bill.
part ofAWSoverview, primer and where to startread it →on this pageshowhide
explore
- X-Ray Distributed Tracing6 questions
- CloudTrail Audit Events5 questions
- Cost Management & Optimization21 questions
- Cost Allocation & Chargeback5 questions
- Cost Explorer, Budgets & Anomalies5 questions
- Pricing Models & Commitments6 questions
- Waste & Data-Transfer Levers5 questions
- Well-Architected Reviews & Trusted Advisor5 questions
- CloudWatch23 questions
- Metrics & Alarms6 questions
- CloudWatch Logs6 questions
- Logs Insights Querying5 questions
- Dashboards, Synthetics & RUM6 questions
questions
page 2 of 2Your team is asked to run a Well-Architected Review of a production workload using the AWS Well-Architected Tool. What does the review actually consist of, what is a milestone, and what do you have at the end?
basics
~20 sA review defines a workload in the AWS Well-Architected Tool, applies one or more lenses, and answers each lens design question by selecting the best practices you follow. Unselected practices become rated risks in an improvement plan; a milestone freezes that state as an immutable snapshot.
What does AWS Trusted Advisor check, how does your AWS Support plan change what you see, and why is it not a substitute for a Well-Architected Review?
basics
~20 sAWS Trusted Advisor runs curated checks over an account's configuration and usage across categories including cost optimization, performance, security, fault tolerance, service limits and operational excellence. Basic and Developer Support see only a subset; the full catalogue needs Business, Enterprise On-Ramp or Enterprise Support.
Explain the X-Amzn-Trace-Id header that AWS X-Ray uses to correlate one request across services: which fields it carries, which components set it first, and what the Sampled flag values mean to a service that receives it.
basics
~20 sX-Amzn-Trace-Id carries Root (the X-Ray trace id), Parent (the caller's segment or subsegment id) and Sampled (1, 0 or ?). AWS entry points such as an Application Load Balancer or a REST API Gateway stage add it; every hop must forward it or the trace breaks.
What does Amazon CloudWatch Application Signals give you that building your own service dashboards from custom metrics does not, and what does an Application Signals SLO actually track?
basics
~20 sCloudWatch Application Signals auto-instruments supported runtimes to emit standardised latency, error and fault metrics per service and operation, plus a service map and dependency view. Its SLO object tracks a goal over an interval, reporting attainment and remaining error budget rather than just a threshold breach.
Your workloads are split across a dozen AWS accounts and an on-call engineer has to assume a role into each one to read its CloudWatch data. How does AWS cross-account observability change that, and what does it not give you?
basics
~20 sCloudWatch cross-account observability designates a monitoring account that creates an Observability Access Manager sink; each source account creates a link to that sink naming the resource types it shares. The monitoring account then reads metrics, log groups and traces in place — one console, no role hopping.
An AWS bill is dominated by Amazon CloudWatch Logs charges. Explain how CloudWatch Logs charges for data, and which levers the service itself gives you to bring the number down.
basics
~20 sCloudWatch Logs charges mainly for ingestion per GB, with much cheaper per-GB-month storage on compressed data. Levers: ingest less, route high-volume vended logs to S3 instead, use the Infrequent Access log class, and shorten retention — which only touches storage.
A team says their CloudWatch Logs Insights queries take minutes to return during incidents, and the bill now shows a growing Logs Insights charge. What actually drives the cost and latency of a Logs Insights query, and how would you make the same investigation cheaper and faster?
basics
~20 sCloudWatch charges Logs Insights per gigabyte of log data scanned, and scan volume is set by the log groups and time range selected — not by how selective the filter is. Narrow the selection, not the query.
You need every line landing in an Amazon CloudWatch Logs log group delivered to another system in near real time. Explain what a subscription filter is, which destinations it can send to, and how you would choose between them.
basics
~20 sA subscription filter is a standing rule on a log group that pushes matching events, as they arrive, to Lambda, Kinesis Data Streams, Amazon Data Firehose, or a cross-account destination. Firehose suits bulk delivery to storage; Lambda suits per-event transformation.
CloudWatch offers static-threshold alarms, anomaly-detection alarms and composite alarms. What does each one actually evaluate, and what would make you pick one over the others?
basics
~20 sA static alarm compares a metric statistic to a fixed number. An anomaly-detection alarm compares it to a band that CloudWatch learns from the metric's own history. A composite alarm evaluates no metric at all — it is a boolean rule over the states of other alarms.
A CloudWatch alarm on the raw count of 5xx responses from your load balancer pages at every traffic peak and stays silent during a quiet-hour outage. How would you use CloudWatch metric math to alarm on an error rate instead, and what must you handle for the alarm to behave when traffic drops to zero?
basics
~20 sAlarm on a ratio rather than a count: use CloudWatch metric math to divide the 5xx count by the request count and threshold the percentage. Then guard the expression so that zero traffic produces a defined value instead of a gap that leaves the alarm unevaluated.
Finance wants hourly cost per team tag, broken down by usage type, for the last twelve months — more detail than the Billing console will show. Explain what the AWS Cost and Usage Report gives you that the console does not, and how you would query it.
basics
~20 sThe Cost and Usage Report delivers the raw billing line items to an S3 bucket you own, at hourly granularity with activated tags and optionally resource IDs as columns. You query it with Athena over the S3 data, which gives detail and joins the console cannot express.
Teams in your AWS Organization keep launching resources with no `CostCenter` tag, or spelling it `costcenter`. Explain what an AWS Organizations tag policy actually enforces, and how you would prevent untagged resources from being created at all.
basics
~20 sAn Organizations tag policy standardises tag keys and allowed values and reports non-compliance; with enforcement enabled for named resource types it blocks non-compliant tagging operations. It never blocks creating an untagged resource — that needs a deny policy on the aws:RequestTag condition key.
A monthly AWS cost budget only tells you once spend is already high. How does AWS Cost Anomaly Detection differ — what a monitor watches, how alert thresholds are configured, and what it still cannot catch?
basics
~20 sAWS Cost Anomaly Detection learns each monitored dimension's normal spend pattern and alerts on unexpected deviation, catching a spike days before a monthly threshold. It cannot see spend that has always been high, or a slow creep that resembles growth.
The largest line on a VPC's monthly bill is 'NAT Gateway data processing'. The workload is a fleet of containers in private subnets that pull images from Amazon ECR and read and write objects in Amazon S3 all day. Explain why that charge is so large and what you would change to reduce it.
basics
~20 sA NAT Gateway bills an hourly rate plus a per-GB data-processing fee on everything it forwards — including in-Region AWS traffic. Route S3 and DynamoDB through free gateway VPC endpoints and other services through interface endpoints so that bulk traffic never touches the NAT.
Finance reports that your AWS Savings Plans utilization is 100% but coverage is around 40%. What do those two numbers measure, and what does that combination tell you to do next?
basics
~20 sUtilization is the share of your committed hourly spend that eligible usage actually consumed; coverage is the share of eligible compute usage that a commitment discounted. Full utilization with low coverage means the commitment is sound but too small — most compute is still billing at On-Demand rates.
An ECS service on Fargate is instrumented with the AWS X-Ray SDK, but no traces appear in the X-Ray console. Walk through how segment data actually reaches X-Ray from a container, and where you would look for the break.
basics
~20 sThe SDK does not call X-Ray directly: it sends segments over UDP to a local collector — the X-Ray daemon or an ADOT collector sidecar on port 2000 — which batches them and calls PutTraceSegments. Check that the sidecar exists, that the app points at it, and that the task role grants X-Ray write.
Across a large AWS estate the CloudTrail-related spend has become a material line on the bill, while compliance wants API activity kept for years. How do you decide what to log, where it lives, and for how long?
basics
~20 sKeep management events everywhere because the first trail copy is free and the record is irreplaceable. Treat data events as a per-workload decision scoped by resource and write-only. Hold one durable copy in S3 with a lifecycle policy, and size any hot query or alerting surface separately.
You own FinOps for an AWS Organization of around 120 accounts and leadership wants each product team charged for what it uses. How would you design the allocation model, and what would you do about spend that no tag can attribute?
basics
~20 sMake the account boundary carry most of the allocation, since all spend in an account belongs to it without tagging discipline, then use a small mandatory tag set inside accounts. Split genuinely shared cost by rule, and start with showback before charging anyone.
Your company wants a multi-year AWS compute commitment, but the fleet is mid-migration to containers and part of it is moving to a different instance architecture. How do you decide what to commit to, and for how long?
basics
~20 sCommit only to the part of the baseline that survives every plausible roadmap outcome, prefer Compute Savings Plans because they follow workloads across families and into Fargate and Lambda, ladder shorter terms over the uncertain portion, and buy in tranches rather than one irreversible purchase.
Your team keeps one near-identical Amazon CloudWatch dashboard per environment and per region, and every widget change has to be made four times. How do CloudWatch dashboard variables let you collapse those copies into one dashboard?
basics
~20 sCloudWatch dashboard variables add a control at the top of a dashboard that rewrites its widgets on selection. A property variable swaps a value such as a dimension, region or account across all widgets; a pattern variable substitutes a matched string anywhere in the dashboard JSON.
How do you encrypt an Amazon CloudWatch Logs log group with a customer managed KMS key, and what must that key's policy contain for it to work?
basics
~20 sAssociate a symmetric KMS key with the log group at creation or with AssociateKmsKey. The key policy must let the logs.<region>.amazonaws.com service principal use the key, normally scoped by a condition on the kms:EncryptionContext:aws:logs:arn context key.
What are AWS Cost Categories in the Billing and Cost Management console, how do they differ from cost allocation tags, and what problem do their split charge rules solve?
basics
~20 sA Cost Category is a billing dimension you define with rules over accounts, tags, services, regions and charge types, rather than one that resources carry. It groups spend that tags cannot reach, and its split charge rules distribute shared costs across the teams that caused them.
You need to answer ad-hoc questions about AWS API activity spanning the past year. Compare querying with Amazon Athena over your CloudTrail trail's S3 objects against using CloudTrail Lake, and say what decides it.
basics
~20 sAthena over the trail's S3 objects is cheap to store and pay-per-query on bytes scanned, but you own the table and its partitioning. CloudTrail Lake is a managed event data store queried with SQL, priced on ingestion and retention, and needs no plumbing. Query volume and setup effort decide it.
CloudWatch Contributor Insights lets you define a rule over log groups instead of running an ad-hoc Logs Insights query. What does such a rule actually produce, and when would you build one rather than just querying?
basics
~20 sA Contributor Insights rule evaluates log events as they arrive and produces a continuously updated ranking of top contributors by a key you choose, plus graphable metrics you can alarm on — unlike a query, which reads history once.
AWS Cost Explorer provides both a utilization report and a coverage report for Savings Plans and Reserved Instances. What is the difference between the two, and what does a low value on each one tell you?
basics
~20 sUtilization is the share of a purchased commitment that was actually used; coverage is the share of eligible on-demand usage that a commitment discounted. Low utilization means you are paying for unused commitment; low coverage means discountable usage is still billed at On-Demand rates.
A manager wants to buy a three-year AWS Savings Plan immediately to cut the EC2 bill. Why should rightsizing and a Graviton evaluation happen first, and what does AWS Compute Optimizer contribute?
basics
~20 sA commitment discounts the fleet you have, so buying first locks in three years of existing waste. Rightsizing and moving eligible workloads to cheaper Graviton instance types lower the baseline first; AWS Compute Optimizer supplies the per-resource evidence for those changes.
How does an AWS X-Ray sampling rule decide whether a given request is traced? Explain the reservoir and fixed-rate fields, how rule priority works, and how one reservoir is shared across many instances of the same service.
basics
~20 sAn X-Ray sampling rule matches requests by service, host, method and URL path, then traces a per-second reservoir of them and a fixed percentage of the rest. Rules are evaluated by priority, lowest number first, and the reservoir is handed out to instances as quotas by the X-Ray service.
Your organisation wants every CloudWatch metric from dozens of AWS accounts to land in a third-party observability platform. Compare CloudWatch Metric Streams with polling the CloudWatch GetMetricData API, and explain what should drive the decision.
basics
~20 sMetric Streams push metric updates continuously through Amazon Data Firehose with near-real-time latency and a per-update charge. Polling GetMetricData pulls on your schedule, costs per API request and throttles as the account count grows. Volume, latency and cost decide it.
Cross-AZ data transfer between chatty internal services is the largest line on your AWS bill. An engineer proposes making every service call only same-Availability-Zone instances of its dependencies. How would you evaluate that proposal?
basics
~20 sZone-local routing removes a real per-gigabyte charge but trades away the load spreading and headroom that multi-AZ deployment buys. Quantify the saving first, then require per-zone capacity headroom and an automatic fallback to remote zones before accepting it.
You own architecture governance for dozens of teams on AWS. How would you run Well-Architected reviews so they change what gets built, instead of becoming an annual paperwork exercise?
basics
~20 sTie reviews to lifecycle events rather than a calendar, encode organisation-specific standards as custom lenses, save milestones so progress is measurable, and require that high risk issues become funded, owned backlog items. A review nobody is resourced to act on produces documents, not change.
showing 31–60 of 60