skip to content

You already alarm on your load balancer's 5xx count and target response time. What does an Amazon CloudWatch Synthetics canary add on top of that, and what does running one actually cost you in moving parts?

level: middleimportance: must knowfreq 58%

answer

  1. you generate the traffic yourself
  2. no users at 4 a.m., still a signal
  3. covers DNS, TLS and the CDN
  4. a Lambda with a role, logs and a bucket
  5. VPC canary needs a NAT route

basics

~20 s

A CloudWatch Synthetics canary is a scheduled script that exercises your endpoint the way a user would, so you detect outages with no traffic to measure and catch failures outside your servers — DNS, certificates, CDN, third parties. It is a managed Lambda, with its own role, logs and S3 artifacts.

solid answer

~60 s

Server-side metrics only tell you about requests that reached your servers. A Synthetics canary is **active probing**: AWS runs a script you supply — a Node.js Puppeteer or Python Selenium runtime — on a schedule from outside the application, and publishes `SuccessPercent` and `Duration` into the `CloudWatchSynthetics` namespace so you can alarm on them. That gives you three things load-balancer metrics cannot: a signal at 4 a.m. when there is no real traffic, coverage of everything in front of your servers (DNS, TLS certificate, CloudFront, WAF, a third-party login provider), and a check of a whole user *journey* rather than an individual request. The cost is that a canary is real infrastructure: it runs as a managed Lambda in your account with an execution role, writes a log group per canary and writes screenshots and HAR files to an S3 bucket that grows forever without a lifecycle rule. Put one in a VPC to reach a private endpoint and it needs a NAT route like any other VPC Lambda, or it cannot reach the internet.

code

javascript · 26 lines
javascript
const synthetics = require('Synthetics');
const log = require('SyntheticsLogger');

const apiCanary = async function () {
  await synthetics.executeHttpStep('health', {
    hostname: 'api.example.com',
    method: 'GET',
    path: '/health',
    port: 443,
    protocol: 'https:'
  });

  await synthetics.executeHttpStep('search', {
    hostname: 'api.example.com',
    method: 'GET',
    path: '/v1/search?q=shoes',
    port: 443,
    protocol: 'https:'
  });

  log.info('canary steps complete');
};

exports.handler = async () => {
  return await apiCanary();
};

go deeper

for a junior

Know that a Synthetics canary is a scheduled script AWS runs against your endpoint from outside, so you find out it is down even when nobody is using it.

for a middle

Explain the runtimes and blueprints, that results arrive as SuccessPercent and Duration in the CloudWatchSynthetics namespace, and that the canary is a Lambda with a role, a log group and an S3 artifact bucket.

for a senior

Demonstrate operating them: alarm on consecutive failures, lifecycle the artifacts, expect the VPC NAT trap, and design canaries that do not pollute production data or page people over their own flakiness.

for a principal

Own the policy — which journeys are worth probing, the detection latency you are buying against the run cost, who maintains scripts as the UI changes, and how synthetic traffic is kept out of business metrics.

## Why active probing exists Every metric your application emits shares one blind spot: it only describes requests that **arrived**. If DNS is broken, if the certificate expired overnight, if CloudFront is serving a stale error, if the identity provider your login page depends on is down — your load balancer's 5xx count is flat, your latency looks great, and you are down. The second blind spot is low traffic. Alarms built on ratios and counts go quiet when the denominator goes to zero. At 4 a.m. an outage and a quiet night look identical. CloudWatch Synthetics answers both by **generating the traffic itself**. ## What a canary is A canary is a script plus a schedule. AWS runs it as a managed Lambda function in your account. Two runtime families: - **Node.js with Puppeteer** — drives headless Chrome, so it can click through a real page. - **Python with Selenium** — the same idea in Python. Runtime versions are named like `syn-nodejs-puppeteer-<version>` and `syn-python-selenium-<version>`, and you pin one; AWS deprecates old ones, so runtime upgrades are ongoing maintenance, not a one-off. AWS ships **blueprints** for the common shapes so you rarely start from an empty file: heartbeat monitoring (load a URL, screenshot it), API canary (a sequence of HTTP calls), broken-link checker, GUI workflow, canary recorder, and visual monitoring (compare a screenshot against a baseline and fail on visual drift). The library gives you step helpers rather than raw HTTP: ```javascript const synthetics = require('Synthetics'); await synthetics.executeHttpStep('login', { hostname: 'api.example.com', method: 'POST', path: '/session', port: 443, protocol: 'https:' }); ``` Each named step is timed and reported separately, so a failure says *which* step broke, not just that the run failed. ## What it emits Canary results land in the **`CloudWatchSynthetics`** namespace, dimensioned by canary name — most importantly `SuccessPercent` and `Duration`, plus per-step timings. You alarm on those exactly as you would on any other metric. `SuccessPercent` is the classic: a single failed run out of a handful in the evaluation window is noise; several consecutive failures are an outage. ## The moving parts you now own This is the half candidates forget, and it is what separates "I have read about canaries" from "I run them": - **An execution role.** The canary needs to write its artifacts (`s3:PutObject`), write logs (`logs:CreateLogStream`, `logs:PutLogEvents`) and publish metrics (`cloudwatch:PutMetricData`, conventionally restricted by a condition on the `CloudWatchSynthetics` namespace). - **An S3 artifact bucket.** Every run writes screenshots, HAR files and logs. This is genuinely useful during an incident — you get to *see* the page as it failed — and it grows without bound. Put a lifecycle rule on that prefix on day one. - **A log group per canary**, holding the script's own output. Set retention; the default is to keep it forever. - **Cost per run.** You pay per canary run, plus the Lambda, the storage and the metrics. A one-minute schedule against twenty canaries is a real line item — pick the interval from how fast you need detection, not by reflex. ## The VPC trap A canary can be attached to a VPC to reach a private endpoint. The moment you do that, it is a VPC Lambda and inherits that world's rules: it has no public internet access unless its subnets route through a NAT gateway (or the relevant VPC endpoints exist), and it needs a security group whose egress reaches the target. A canary that worked publicly and starts timing out the day it is moved into a VPC is almost always this and not the script. ## Where canaries mislead - **They are synthetic.** One scripted journey from a handful of AWS regions is not what your users experience across their networks and devices. Canaries prove your path is *reachable*; they do not prove it is *fast for humans*. Real-user telemetry answers that. - **They rot.** A canary that logs in with a test account breaks when the account expires, the login flow gets a new field, or a CAPTCHA is added. A flapping canary that everyone ignores is worse than none, because it teaches the team to ignore an alarm. - **They mutate state.** A canary that places an order places a real order. Design for it: dedicated test accounts, idempotent or clearly-tagged data, and a way for downstream teams to filter synthetic traffic out of their analytics. ## How to choose the coverage Canary the journeys whose failure means the business has stopped — sign-in, search, checkout, the public API's health path — not every endpoint. A small set of well-maintained canaries on critical paths, alarmed on consecutive failures, is worth more than fifty flaky ones.

  • Your canary alarms on a single failed run and the team has started ignoring it. How do you fix that?
    Alarm on consecutive failures rather than one run — several evaluation periods breaching before the alarm fires — so a transient network blip does not page anyone. Then fix the underlying flakiness: pin the test account, avoid brittle selectors, and split long journeys into named steps so you can see which one is actually unstable.
  • Where do you look first when a canary fails and the application seems fine?
    The canary's S3 artifacts and its log group. The screenshot and HAR from the failing run usually show it immediately — an expired certificate, a redirect to a maintenance page, a third-party script that hung, or a login form that changed shape. That evidence is the main reason to keep the artifact bucket, and the main reason to lifecycle it rather than delete it.
  • A canary is moved into a VPC to test a private endpoint and every run now times out. What is the likely cause?
    It is a VPC Lambda now, so it has no route to the public internet unless its subnets go through a NAT gateway or the relevant VPC endpoints exist, and its security group must permit egress to the target. Anything in the script that reached out publicly — a CDN asset, a third-party call — dies first, usually as a timeout rather than a clean error.
  • Why keep canary coverage small rather than probing every endpoint?
    Because each canary is real cost and real maintenance — a run charge, a Lambda, artifacts, a log group, and a script that breaks when the UI changes. Canary the journeys whose failure means the business has stopped, and let ordinary metrics and alarms cover the rest.

saying these in an interview costs you the question

  • Believing server-side 5xx metrics would catch a DNS or certificate failure
  • Alarming on a single failed canary run
  • Leaving the artifact bucket and log group with no retention
  • Forgetting a VPC canary needs a NAT route for public calls
  • Assuming a canary tells you what real users experience
  • Scripting a checkout canary that writes real production orders

context