skip to content

You need a business metric — orders placed per minute — visible in CloudWatch from a Lambda function. Compare publishing it with the CloudWatch PutMetricData API against the CloudWatch Embedded Metric Format, and explain what adding a dimension does to the metric and to the bill.

level: middleimportance: must knowfreq 58%

answer

  1. namespace + name + dimensions = identity
  2. API call versus structured log line
  3. dimensions do not roll up
  4. unbounded dimension values explode cost
  5. context belongs in the log, not the dimension

basics

~20 s

PutMetricData is a synchronous API call you wait on and pay per request; the Embedded Metric Format lets you print one structured JSON log line that CloudWatch converts into a metric asynchronously. Every distinct combination of dimension values becomes its own billable metric.

solid answer

~50 s

`PutMetricData` sends a metric datum straight to CloudWatch. In a Lambda handler that means an outbound HTTPS call on the request path — added latency, a failure mode, billed duration, and a per-request API charge — and if you fire and forget it before the freeze, it may never leave. The Embedded Metric Format avoids all of that: you write a JSON log line containing an `_aws` node that declares the namespace, dimension sets and metric names, CloudWatch Logs ingests it as normal stdout, and the metric is extracted asynchronously on the service side. You get the metric *and* a searchable log line carrying the request id and any other context, at no latency cost. Dimensions are part of the metric's identity — `OrdersPlaced` with `Service=checkout` and `OrdersPlaced` with `Service=cart` are two distinct metrics, each billed monthly — so a dimension whose values are unbounded, such as a customer id, is a cost incident waiting to happen.

code

json · 16 lines
json
{
  "_aws": {
    "Timestamp": 1739452800000,
    "CloudWatchMetrics": [
      {
        "Namespace": "Orders",
        "Dimensions": [["Service"], ["Service", "PaymentProvider"]],
        "Metrics": [{ "Name": "OrdersPlaced", "Unit": "Count" }]
      }
    ]
  },
  "Service": "checkout",
  "PaymentProvider": "stripe",
  "OrdersPlaced": 1,
  "orderId": "ord-8812"
}

go deeper

for a junior

Know that a custom metric is created just by publishing it, that PutMetricData is the direct API, and that a dimension is part of the metric's name rather than a filter applied afterwards.

for a middle

Explain the EMF _aws node — namespace, dimension sets, metric names — and why printing a log line beats an API call inside a Lambda handler on latency, failure modes and cost.

for a senior

Demonstrate cardinality judgment: which fields earn dimension status versus staying as log context, why no rollup exists across a dropped dimension, and how a VPC-attached function without egress turns PutMetricData into a timeout.

for a principal

Own the standard — a naming and dimension convention every service follows, a budget for custom metrics, and the decision about whether business metrics belong in CloudWatch at all or in the platform your dashboards already live on.

## What a custom metric actually is A CloudWatch metric is identified by three things: the **namespace**, the **metric name**, and the **exact set of dimension name/value pairs**. Change any one of them and you have a different metric with its own history. That single rule explains most custom-metric behaviour, including the billing. Crucially, dimensions do not roll up. If you publish `OrdersPlaced` with dimensions `{Service=checkout, PaymentProvider=stripe}`, there is *no* `OrdersPlaced` with only `{Service=checkout}` and no un-dimensioned `OrdersPlaced` — you cannot query a subset of the dimensions and get a total. If you want the aggregate you must publish it as a separate series too. ## Route 1: PutMetricData The direct API call: ```bash aws cloudwatch put-metric-data \ --namespace Orders \ --metric-name OrdersPlaced \ --unit Count --value 1 \ --dimensions Service=checkout,PaymentProvider=stripe ``` Each `MetricDatum` carries a metric name, optional dimensions, a timestamp, a `Unit`, and either a single `Value`, a `StatisticValues` summary (sample count, sum, min, max), or matched `Values` and `Counts` arrays for compressing many observations. Setting `StorageResolution` to 1 instead of the default 60 makes it a high-resolution metric with sub-minute granularity. What it costs you operationally: a network round trip to the CloudWatch endpoint. In a VPC-attached Lambda with no NAT route and no interface endpoint, that call simply hangs until the function times out — a classic way to turn adding a metric into an outage. It is billed per API request as well as per metric, it consumes function duration, and it is subject to `PutMetricData` throttling under high call rates. Batching several data items into one call helps, but only if you have several to send. Use it when there is no log pipeline to lean on: a long-running process that batches many observations, a script, or a component whose logs do not go to CloudWatch Logs. ## Route 2: Embedded Metric Format EMF turns a structured log line into metrics. You print JSON to stdout containing an `_aws` node; CloudWatch Logs ingests it and the service extracts the declared metrics for you: ```json { "_aws": { "Timestamp": 1739452800000, "CloudWatchMetrics": [ { "Namespace": "Orders", "Dimensions": [["Service"], ["Service", "PaymentProvider"]], "Metrics": [{ "Name": "OrdersPlaced", "Unit": "Count" }] } ] }, "Service": "checkout", "PaymentProvider": "stripe", "OrdersPlaced": 1, "orderId": "ord-8812", "requestId": "3f0c..." } ``` Three things are worth reading closely. `Dimensions` is an array *of arrays*: each inner array is one dimension set, so the example publishes both a per-service total and a per-service-per-provider breakdown from one line. Fields referenced by a dimension set must exist as top-level properties. And any *other* top-level field — `orderId`, `requestId` — is not a dimension: it stays in the log, queryable later, without creating a metric. That is the property that makes EMF so useful: high-cardinality context lives in the log, low-cardinality aggregation lives in the metric. The official `aws-embedded-metrics` libraries exist for several runtimes and build this JSON for you, but the format is plain JSON and writing it yourself is entirely legitimate. The trade: metrics appear after the log line is ingested and processed, so it is not instantaneous, and you pay CloudWatch Logs ingestion and storage on top of the custom-metric charge. ## Resolution and retention The default storage resolution is 60 seconds. High-resolution metrics (`StorageResolution: 1`, also settable per-metric in EMF) accept sub-minute data and allow alarms with 10- or 30-second periods, at a higher price. As of 2025 CloudWatch keeps sub-minute data for 3 hours, one-minute data for 15 days, five-minute data for 63 days and one-hour data for 455 days, rolling older data up rather than deleting it — so a high-resolution metric is not a long-term high-resolution archive. ## The cardinality trap Every unique dimension-value combination is a metric, and every metric has a monthly charge (roughly $0.30 per custom metric per month in common regions for the first tier, as of 2025). Add `CustomerId` as a dimension on a metric emitted for ten thousand customers and you have created ten thousand metrics. The damage is not only financial: the console becomes unusable, and no single alarm covers the fleet because each customer is a separate series. The discipline is simple. Dimensions answer "which bucket do I want to alarm on and compare" — service, environment, operation, status class. Anything unbounded belongs in the log line beside the metric, which is exactly what EMF gives you for free.

  • You published OrdersPlaced with dimensions Service and PaymentProvider. Why can you not graph the total across all providers?
    Because the dimension set is part of the metric's identity, not a tag you can drop. `{Service=checkout, PaymentProvider=stripe}` and `{Service=checkout, PaymentProvider=paypal}` are two independent metrics and no `{Service=checkout}` series exists. Either publish the extra dimension set as well — EMF's array-of-arrays does this in one line — or sum the series with a metric math expression at query time.
  • What breaks if a Lambda function calls PutMetricData and does not await the response?
    The execution environment is frozen as soon as the handler returns, so an in-flight request can be suspended and may never complete, or may complete unpredictably on the next invocation. The metric silently goes missing under load, which is worse than not having it. Awaiting it costs latency and billed duration — which is precisely why EMF, where a printed line is enough, suits Lambda better.
  • When is a high-resolution metric actually worth the extra cost?
    When you need to detect something faster than a minute — a sharp latency or saturation spike driving an automated response — and you will act on it within the three hours that sub-minute data is retained. It also unlocks 10- and 30-second alarm periods. For dashboards, capacity trends or anything you look at the next day, the resolution is aggregated away anyway and you have paid for nothing.

saying these in an interview costs you the question

  • Treating dimensions as free labels you can filter or drop later
  • Adding a request id or customer id as a dimension
  • Calling PutMetricData without awaiting it in a Lambda handler
  • Believing EMF metrics are free because they come from a log line
  • Expecting an un-dimensioned rollup to exist automatically

context