skip to content

Your organisation wants every CloudWatch metric from dozens of AWS accounts to land in a third-party observability platform. Compare CloudWatch Metric Streams with polling the CloudWatch GetMetricData API, and explain what should drive the decision.

level: principalimportance: nice to knowfreq 26%

answer

  1. push versus pull
  2. per-request cost versus per-update cost
  3. throttling is the pull ceiling
  4. streams carry no history
  5. filter by namespace or overpay

basics

~20 s

Metric Streams push metric updates continuously through Amazon Data Firehose with near-real-time latency and a per-update charge. Polling GetMetricData pulls on your schedule, costs per API request and throttles as the account count grows. Volume, latency and cost decide it.

solid answer

~50 s

Polling is the legacy integration model: you enumerate metrics with `ListMetrics`, then call `GetMetricData` on a schedule, up to 500 metric queries per request. It is pull-based, so you control what you fetch and can backfill history, but it scales badly — every account and region multiplies the call volume, you pay per request, you hit throttling as the fleet grows, and newly created metrics stay invisible until the next enumeration. Metric Streams invert it: you create a stream, filter it by namespace with include or exclude rules, and CloudWatch pushes metric updates through Amazon Data Firehose to an HTTP endpoint or S3 in JSON or OpenTelemetry format, typically within a couple of minutes. You stop paying per request and start paying per metric update plus Firehose delivery, which means an unfiltered stream over noisy namespaces can cost more than the polling it replaced. Streams carry no history, so they solve ingest, not backfill.

go deeper

for a junior

Know that CloudWatch metrics can leave the account two ways — something calls the API to pull them, or CloudWatch pushes them out through a stream — and that the push route goes via Amazon Data Firehose.

for a middle

Explain the mechanics on both sides: ListMetrics plus GetMetricData on a schedule versus CreateMetricStream with namespace filters, an output format and a Firehose destination.

for a senior

Show the operational picture — polling latency and throttling at fleet scale, the per-update plus Firehose cost of streaming, the absence of history, and the need to alarm on the delivery path itself.

for a principal

Own the decision and its economics: which namespaces justify continuous streaming, where a hybrid keeps backfill possible, how ninety streams stay in step across the organisation, and whether workloads should bypass CloudWatch as a source entirely.

## The two integration shapes Every third-party observability platform faces the same problem: CloudWatch holds the AWS-vended metrics and they need to be somewhere else. There are exactly two ways out. **Pull.** The vendor assumes a role in each of your accounts and calls the CloudWatch API on a schedule. `ListMetrics` discovers what exists; `GetMetricData` fetches datapoints, accepting up to 500 `MetricDataQuery` structures per request and supporting metric math on the way out. **Push.** You create a CloudWatch metric stream in each account and region. CloudWatch continuously writes metric updates into an Amazon Data Firehose delivery stream, which delivers them to an HTTP endpoint the vendor operates, or to S3. `CreateMetricStream` takes an output format — JSON, or OpenTelemetry — and include or exclude filters that select namespaces, so you can stream `AWS/Lambda` and `AWS/RDS` and leave the rest behind. A `statistics_configurations` block asks for additional statistics beyond the default set of sample count, average, sum, minimum and maximum. ## Where pull hurts at scale **Latency compounds.** AWS publishes the metric, then the poller waits for its next tick, then the vendor ingests it. A five-minute poll interval means detection is minutes behind reality before anyone's alerting logic even runs. **Cost is per call, and calls multiply.** Requests scale with accounts × regions × namespaces × poll frequency. Teams end up choosing between coverage and frequency, and usually degrade one silently. **Throttling is the real ceiling.** CloudWatch API rate limits are per account and region; a wide fleet, a vendor integration, an internal exporter and someone's dashboard all draw on the same budget. When it saturates you get `ThrottlingException`, gaps in the vendor's graphs, and an argument about whose poller caused it. **Discovery lags.** A metric that did not exist at the last `ListMetrics` call is not fetched. A new Lambda function or a new queue is invisible until the next enumeration cycle. ## Where push hurts **You pay for what you stream, not what you use.** Charging is per metric update delivered, plus Firehose ingestion and whatever the destination costs. An unfiltered stream across every namespace in a busy account can quite easily cost more than the polling it replaced — the filters are not an optimisation, they are the design. **No history.** A stream carries data from the moment it is created. If you need last quarter for a comparison, that is still a `GetMetricData` job, which is why mature setups keep a small polling path alongside the stream. **Delivery is a pipeline you now own.** Firehose buffers before delivering, adding delay; when the destination rejects a batch it lands in the configured S3 backup; failures show up as Firehose metrics that somebody has to watch. You have traded an API you called for infrastructure you operate. **No server-side computation.** Streams carry raw metric updates. Metric math, `SEARCH()` and alarm evaluation stay behind in CloudWatch — everything derived has to be recomputed in the destination. **It is per account and per region.** Thirty accounts across three regions is ninety streams plus ninety Firehose deliveries, deployed and kept in step by whatever account-baseline mechanism you already use. ## What should actually drive the decision First, **what the destination is for.** If it is the primary place engineers look and alert from, streaming latency is worth paying for. If it is a quarterly capacity or cost review, polling at low frequency is perfectly adequate and much cheaper. Second, **how much of CloudWatch you genuinely need.** Most estates need a handful of namespaces continuously and everything else rarely. That points to a hybrid: stream the namespaces that drive detection, poll the long tail on a slow cycle. Third, **whether you are throttling today.** Existing `ThrottlingException` rates and gaps in vendor dashboards are a hard argument for push; without them, polling may be fine for years. Fourth, **what the vendor supports.** OpenTelemetry-format streams are widely consumed, but formats and required Firehose destination shapes vary, and the answer is a fact about their integration, not a preference of yours. Fifth — and it is the question people skip — **whether CloudWatch should be the source at all.** Application telemetry can be emitted directly to the destination from the workload, leaving CloudWatch as the source only for AWS-vended metrics nobody else can produce. That usually shrinks the streaming bill more than any filter. ## The shape most organisations land on Stream the namespaces behind detection and dashboards, with explicit include filters, deployed by the same mechanism that bootstraps every account. Keep a low-frequency polling path for backfill, for one-off analysis and for namespaces not worth streaming. Watch the Firehose delivery metrics, because a silently failing stream looks exactly like a healthy quiet system.

  • Why do teams keep a polling path even after moving to Metric Streams?
    Because a stream only carries data from the moment it exists. Backfilling a new dashboard, investigating something that happened before the stream was created, or pulling a quarter of history for a capacity review all still need `GetMetricData`. Polling also covers namespaces that are not worth the per-update cost of streaming, and gives you metric math server-side, which the stream does not.
  • A metric stream is delivering nothing and nobody noticed for a week. What would have caught it?
    Firehose's own delivery metrics — records delivered, delivery failures, and objects landing in the configured S3 backup bucket — with an alarm on them. A stopped stream looks identical to a quiet system in the destination, so the monitoring has to be on the pipeline, not on the data. This is why the delivery path counts as infrastructure you operate.
  • How would you cut the cost of an existing metric stream without losing detection coverage?
    Tighten the include filters to the namespaces that actually drive alerts and dashboards, and drop the noisy long tail — per-resource namespaces in busy accounts usually dominate. Then check `statistics_configurations`: extra statistics multiply the updates delivered. Finally ask whether application metrics should bypass CloudWatch entirely and be emitted straight to the destination, which removes them from both bills.

saying these in an interview costs you the question

  • Assuming Metric Streams are cheaper than polling regardless of volume
  • Expecting a stream to backfill historical datapoints
  • Forgetting that streams are per account and per region
  • Streaming every namespace without include or exclude filters
  • Not monitoring the Firehose delivery path itself

context