skip to content

CloudWatch

CloudWatch is the default observability plane for AWS resources: metrics, log groups, alarms, dashboards, and events. Interviewers ask what you would alarm on, so knowing the metric surface matters far more than knowing the console.

part ofAWSoverview, primer and where to startread it →
on this pageshow

explore

questions

23

What is an Amazon CloudWatch dashboard actually made of, what kinds of widget can it hold, and why does putting a metric on a CloudWatch dashboard not mean anyone will be told when it moves?

level: juniorimportance: must knowfreq 50%

answer

  1. saved drawing, not a watcher
  2. JSON body of widgets
  3. region and stat live per widget
  4. nobody is paged by a graph

basics

~20 s

A CloudWatch dashboard is a JSON document of widgets — metric graphs, log tables, alarm-status tiles, text — and each widget carries its own region, statistic and period. Dashboards only draw data; notification comes from CloudWatch alarms, never from a dashboard.

solid answer

~50 s

A CloudWatch dashboard is stored as a JSON body — a `widgets` array written by the `PutDashboard` API, which is what the console editor produces when you drag things around. Each widget has a `type` (`metric`, `log`, `alarm`, `text`, `custom`), a size and position on a 24-column grid, and a `properties` object holding the metrics it graphs, the `view` (time series, single value, gauge, bar, pie, table), the `stat`, the `period` and — importantly — a `region`. Because region is per widget, one dashboard can show several regions side by side. The thing candidates get wrong is treating a dashboard as monitoring: a dashboard is pull-based and only renders while someone is looking at it. Nothing on it evaluates a threshold or sends a notification. That is a CloudWatch alarm's job; an `alarm` widget on the dashboard merely displays the state an alarm already computed.

go deeper

for a junior

Be able to say that a dashboard is a saved set of widgets over existing CloudWatch data, name a few widget types, and state plainly that alarms notify while dashboards only display.

for a middle

Explain the dashboard body as JSON written by PutDashboard, what lives in a widget's properties — metrics, view, stat, period, region — and why region being per widget makes multi-region panes possible.

for a senior

Show the operational instinct: alarms decide when a human is woken, dashboards are what that human opens next. Call out log widgets as the usual cause of a slow, expensive dashboard.

for a principal

Own the estate-level view — dashboards as versioned code, a consistent layout teams can navigate under pressure, and a clear split between what pages someone and what merely informs.

## What a dashboard actually is A CloudWatch dashboard is not a stateful monitoring object. It is a **saved JSON document** describing how to draw data that CloudWatch already holds. The console's drag-and-drop editor is a front end over one API call, `PutDashboard`, which takes a dashboard name and a `DashboardBody` string. `GetDashboard` gives you the same JSON back, which is why dashboards are easy to keep in source control and to template across environments. The body's top-level key is `widgets`, an array. Every widget has: - `type` — the widget kind. - `x`, `y`, `width`, `height` — placement on a grid that is **24 columns wide**; height is in grid rows. - `properties` — everything type-specific. ```json { "widgets": [ { "type": "metric", "x": 0, "y": 0, "width": 12, "height": 6, "properties": { "region": "eu-west-1", "view": "timeSeries", "stat": "p99", "period": 60, "metrics": [ ["AWS/ApplicationELB", "TargetResponseTime", "LoadBalancer", "app/prod-alb/abc123"] ] } } ] } ``` ## The widget types - **`metric`** — the workhorse. Its `metrics` array lists `[namespace, metricName, dimName, dimValue, {options}]` tuples, and entries may also be **metric math** or **search expressions** rather than literal metrics. `view` selects the rendering: `timeSeries`, `singleValue`, `gauge`, `bar`, `pie`, `table`. - **`log`** — runs a CloudWatch Logs Insights query and renders the result set. It executes when the dashboard loads, so it costs money per view and is slower than a metric widget. - **`alarm`** — shows the current state of one or more alarms, or an alarm-status grid. It reflects state; it does not create it. - **`text`** — Markdown. Underrated: this is where the runbook link and "what to do when this graph is red" belong. - **`custom`** — invokes a Lambda function that returns HTML or JSON, for data CloudWatch does not natively hold. ## Region and account live on the widget Each `metric` widget's `properties.region` decides which region's metrics are fetched, independently of where the dashboard resource itself was created. That single field is what makes a genuinely multi-region dashboard possible without duplicating anything. Widgets can likewise be pointed at another AWS account once cross-account observability is set up, so one pane can span an estate. ## Statistic and period are display choices `stat` and `period` on the widget are how the data is *rolled up for this graph* — `Sum` at a 60-second period tells a different story from `Average` at 300 seconds over exactly the same underlying metric. Two widgets can disagree visually while both being correct. Always read the statistic and period before concluding anything from a CloudWatch graph. ## Why a dashboard is not monitoring This is the point interviewers are probing. A dashboard: - is **pull-based** — the queries run when a browser opens it, and never otherwise; - **evaluates nothing** — no threshold, no state machine, no notification target; - **wakes nobody** — there is no SNS topic anywhere in its definition. Detection and notification come from CloudWatch alarms, which evaluate on their own schedule whether anyone is watching. The healthy pattern is: alarms decide when a human is needed, dashboards are what that human opens once they have been paged. A team that only builds dashboards has built a wall of graphs that reliably shows an outage to nobody at 3 a.m. ## Practical notes - Dashboards are cheap and code-friendly; treat the JSON as the source of truth and let the console be an editor, not a database. - `GetMetricWidgetImage` renders a single metric widget to a PNG server-side, which is how graphs get attached to incident tickets or chat messages without a login. - A `log` widget with a wide time range is the usual reason a dashboard is slow and unexpectedly expensive to open, because Logs Insights bills by data scanned.

  • A colleague wants a dashboard that turns red and emails the team when latency is high. What do you tell them?
    The dashboard cannot do that. Create a CloudWatch alarm on the latency metric with an SNS action for the notification, then add an `alarm` widget to the dashboard so the same state is visible when someone opens it. The alarm does the evaluating and paging; the widget just mirrors it.
  • Two widgets graph the same metric but look completely different. What do you check first?
    The `stat` and `period` on each widget. `Sum` over 60 seconds and `Average` over 5 minutes render the same underlying data as very different shapes, and a longer period also hides short spikes through aggregation. Check the region property too — the widgets may not be looking at the same account or region at all.
  • Why would you keep a dashboard's JSON in a repository rather than editing it in the console?
    Because the JSON *is* the dashboard — `PutDashboard` takes the whole body, so a file in git is a faithful, reviewable, diffable definition. It lets you template the same layout across environments, restore a dashboard someone deleted, and see who changed a threshold or a metric and when.

saying these in an interview costs you the question

  • Thinking a dashboard notifies you when a graph spikes
  • Believing a dashboard is limited to a single region
  • Assuming a dashboard runs continuously in the background
  • Confusing an alarm widget with the alarm itself
  • Reading a CloudWatch graph without checking its statistic and period

context

open as a page

An application on an EC2 instance writes to a log file on disk and you want those lines in Amazon CloudWatch Logs. What are the ways log data gets into CloudWatch Logs, and what does each path require you to install or grant?

level: juniorimportance: must knowfreq 65%

basics

~20 s

Three paths: the CloudWatch agent tailing files on a host, AWS services delivering logs natively (Lambda, the ECS awslogs driver, VPC Flow Logs), or your code calling the PutLogEvents API. All three need IAM permission to create streams and put events.

open as a page

Walk through the anatomy of a CloudWatch Logs Insights query — the pipeline of commands and the @-prefixed fields that are always available — and show how you would pull the 20 most recent lines containing ERROR out of one log group.

level: juniorimportance: must knowfreq 60%

basics

~20 s

A Logs Insights query is a pipeline of commands joined by the pipe character: fields or display picks columns, filter narrows rows, sort orders them, limit caps the output. Every event exposes @timestamp, @message, @logStream and @ingestionTime.

open as a page

You already alarm on your load balancer's 5xx count and target response time. What does an Amazon CloudWatch Synthetics canary add on top of that, and what does running one actually cost you in moving parts?

level: middleimportance: must knowfreq 58%

basics

~20 s

A CloudWatch Synthetics canary is a scheduled script that exercises your endpoint the way a user would, so you detect outages with no traffic to measure and catch failures outside your servers — DNS, certificates, CDN, third parties. It is a managed Lambda, with its own role, logs and S3 artifacts.

open as a page

In Amazon CloudWatch Logs, what is the relationship between a log group and a log stream, and at which of the two do you configure retention? What happens to data already stored when you change that setting?

level: middleimportance: must knowfreq 72%

basics

~20 s

A log group is the container and the unit of configuration; a log stream is one ordered sequence of events from one source inside it. Retention is set per log group, defaults to never expire, and applies retroactively — older events are deleted.

open as a page

You need a business metric — orders placed per minute — visible in CloudWatch from a Lambda function. Compare publishing it with the CloudWatch PutMetricData API against the CloudWatch Embedded Metric Format, and explain what adding a dimension does to the metric and to the bill.

level: middleimportance: must knowfreq 58%

basics

~20 s

PutMetricData is a synchronous API call you wait on and pay per request; the Embedded Metric Format lets you print one structured JSON log line that CloudWatch converts into a metric asynchronously. Every distinct combination of dimension values becomes its own billable metric.

open as a page

A CloudWatch alarm on your service's error-count metric stayed in INSUFFICIENT_DATA throughout a real outage instead of firing. Explain how a CloudWatch alarm evaluates a metric — period, evaluation periods, datapoints to alarm — and how the missing-data treatment decides what happens when the metric stops arriving.

level: middleimportance: must knowfreq 74%

basics

~20 s

A CloudWatch alarm counts how many of the last N periods breached the threshold and fires when M of them did. During the outage the metric stopped being published, so there was nothing to compare, and the default missing-data treatment leaves the alarm unevaluated rather than breaching.

open as a page

An EC2 instance shows CPU and network metrics in CloudWatch, but there is no memory-used or disk-space metric for it anywhere. Why not, and what do you have to run to get them?

level: juniorimportance: should knowfreq 62%

basics

~20 s

CloudWatch's built-in EC2 metrics are collected outside the instance, at the virtualization layer, so they cover CPU, network and EBS activity but never anything inside the guest. Memory and filesystem usage require installing the CloudWatch agent.

open as a page

A front-end team wants to know what real browsers experience on your site, not just what the servers logged. What does Amazon CloudWatch RUM collect, how does a browser get permission to send that data, and which levers control its volume?

level: middleimportance: should knowfreq 32%

basics

~20 s

CloudWatch RUM embeds a JavaScript client in your pages that reports page-load performance, Core Web Vitals, JavaScript errors, failed HTTP calls and session/browser context to an app monitor. The browser authorizes with temporary guest credentials from an Amazon Cognito identity pool, and a session sample rate controls volume.

open as a page

During an incident you need to search a dozen CloudWatch log groups belonging to different services, some of them in a second AWS account, from a single Logs Insights query. How do you do that, and what has to be in place first?

level: middleimportance: should knowfreq 40%

basics

~20 s

One Logs Insights query can span many log groups: select them individually or by name prefix, and use the @log field to tell events apart. Reaching another account first requires CloudWatch cross-account observability linking that account to your monitoring account.

open as a page

An application writes plain-text lines such as `2026-03-01T10:00:00Z level=ERROR path=/checkout latency=812ms` into a CloudWatch log group. Using CloudWatch Logs Insights, how would you turn those unstructured lines into a 95th-percentile latency per path in five-minute buckets?

level: middleimportance: should knowfreq 55%

basics

~10 s

Use parse to lift path and latency out of @message into named fields, then aggregate with stats: parse @message "path=* latency=*ms" as path, latency | stats pct(latency, 95) as p95 by path, bin(5m).

open as a page

What is a metric filter in Amazon CloudWatch Logs, and why does a metric filter you just created often report nothing even though matching lines are visibly present in the log group?

level: middleimportance: should knowfreq 48%

basics

~20 s

A metric filter watches a log group for a pattern and publishes a CloudWatch metric when lines match. It reports nothing at first because filters apply only to events ingested after creation — they never backfill — and because non-matching periods publish no data point unless defaultValue is set.

open as a page

What does Amazon CloudWatch Application Signals give you that building your own service dashboards from custom metrics does not, and what does an Application Signals SLO actually track?

level: seniorimportance: should knowfreq 28%

basics

~20 s

CloudWatch Application Signals auto-instruments supported runtimes to emit standardised latency, error and fault metrics per service and operation, plus a service map and dependency view. Its SLO object tracks a goal over an interval, reporting attainment and remaining error budget rather than just a threshold breach.

open as a page

Your workloads are split across a dozen AWS accounts and an on-call engineer has to assume a role into each one to read its CloudWatch data. How does AWS cross-account observability change that, and what does it not give you?

level: seniorimportance: should knowfreq 42%

basics

~20 s

CloudWatch cross-account observability designates a monitoring account that creates an Observability Access Manager sink; each source account creates a link to that sink naming the resource types it shares. The monitoring account then reads metrics, log groups and traces in place — one console, no role hopping.

open as a page

An AWS bill is dominated by Amazon CloudWatch Logs charges. Explain how CloudWatch Logs charges for data, and which levers the service itself gives you to bring the number down.

level: seniorimportance: should knowfreq 52%

basics

~20 s

CloudWatch Logs charges mainly for ingestion per GB, with much cheaper per-GB-month storage on compressed data. Levers: ingest less, route high-volume vended logs to S3 instead, use the Infrequent Access log class, and shorten retention — which only touches storage.

open as a page

A team says their CloudWatch Logs Insights queries take minutes to return during incidents, and the bill now shows a growing Logs Insights charge. What actually drives the cost and latency of a Logs Insights query, and how would you make the same investigation cheaper and faster?

level: seniorimportance: should knowfreq 45%

basics

~20 s

CloudWatch charges Logs Insights per gigabyte of log data scanned, and scan volume is set by the log groups and time range selected — not by how selective the filter is. Narrow the selection, not the query.

open as a page

You need every line landing in an Amazon CloudWatch Logs log group delivered to another system in near real time. Explain what a subscription filter is, which destinations it can send to, and how you would choose between them.

level: seniorimportance: should knowfreq 44%

basics

~20 s

A subscription filter is a standing rule on a log group that pushes matching events, as they arrive, to Lambda, Kinesis Data Streams, Amazon Data Firehose, or a cross-account destination. Firehose suits bulk delivery to storage; Lambda suits per-event transformation.

open as a page

CloudWatch offers static-threshold alarms, anomaly-detection alarms and composite alarms. What does each one actually evaluate, and what would make you pick one over the others?

level: seniorimportance: should knowfreq 42%

basics

~20 s

A static alarm compares a metric statistic to a fixed number. An anomaly-detection alarm compares it to a band that CloudWatch learns from the metric's own history. A composite alarm evaluates no metric at all — it is a boolean rule over the states of other alarms.

open as a page

A CloudWatch alarm on the raw count of 5xx responses from your load balancer pages at every traffic peak and stays silent during a quiet-hour outage. How would you use CloudWatch metric math to alarm on an error rate instead, and what must you handle for the alarm to behave when traffic drops to zero?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Alarm on a ratio rather than a count: use CloudWatch metric math to divide the 5xx count by the request count and threshold the percentage. Then guard the expression so that zero traffic produces a defined value instead of a gap that leaves the alarm unevaluated.

open as a page

Your team keeps one near-identical Amazon CloudWatch dashboard per environment and per region, and every widget change has to be made four times. How do CloudWatch dashboard variables let you collapse those copies into one dashboard?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

CloudWatch dashboard variables add a control at the top of a dashboard that rewrites its widgets on selection. A property variable swaps a value such as a dimension, region or account across all widgets; a pattern variable substitutes a matched string anywhere in the dashboard JSON.

open as a page

How do you encrypt an Amazon CloudWatch Logs log group with a customer managed KMS key, and what must that key's policy contain for it to work?

level: middleimportance: nice to knowfreq 26%

basics

~20 s

Associate a symmetric KMS key with the log group at creation or with AssociateKmsKey. The key policy must let the logs.<region>.amazonaws.com service principal use the key, normally scoped by a condition on the kms:EncryptionContext:aws:logs:arn context key.

open as a page

CloudWatch Contributor Insights lets you define a rule over log groups instead of running an ad-hoc Logs Insights query. What does such a rule actually produce, and when would you build one rather than just querying?

level: seniorimportance: nice to knowfreq 25%

basics

~20 s

A Contributor Insights rule evaluates log events as they arrive and produces a continuously updated ranking of top contributors by a key you choose, plus graphable metrics you can alarm on — unlike a query, which reads history once.

open as a page

Your organisation wants every CloudWatch metric from dozens of AWS accounts to land in a third-party observability platform. Compare CloudWatch Metric Streams with polling the CloudWatch GetMetricData API, and explain what should drive the decision.

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Metric Streams push metric updates continuously through Amazon Data Firehose with near-real-time latency and a per-update charge. Polling GetMetricData pulls on your schedule, costs per API request and throttles as the account count grows. Volume, latency and cost decide it.

open as a page