skip to content

What is a metric filter in Amazon CloudWatch Logs, and why does a metric filter you just created often report nothing even though matching lines are visibly present in the log group?

level: middleimportance: should knowfreq 48%

answer

  1. text in, time series out
  2. evaluated as events arrive
  3. it never looks backwards
  4. no match means no data point, not zero
  5. dimensions multiply into custom metrics

basics

~20 s

A metric filter watches a log group for a pattern and publishes a CloudWatch metric when lines match. It reports nothing at first because filters apply only to events ingested after creation — they never backfill — and because non-matching periods publish no data point unless defaultValue is set.

solid answer

~50 s

A metric filter is attached to a log group and turns matching log lines into a numeric CloudWatch metric you can graph or alarm on. You give it a filter pattern — space-delimited terms for plain text, or a JSON selector such as `{ $.level = "ERROR" }` for structured logs — plus a metric transformation naming the namespace, metric name and `metricValue`, which can be a literal `1` for counting or an extracted field like `$.durationMs` for measuring. Two behaviours catch people out. First, a filter evaluates only events **ingested after it was created**; it will never process history, so it looks dead until new matching lines arrive. Second, when nothing matches in a period, no data point is published at all — the metric is sparse. Setting `defaultValue: 0` makes it emit zero instead, which is usually what you want before building anything on top of it.

code

bash · 15 lines
bash
# Count ERROR lines from JSON-formatted application logs
aws logs put-metric-filter \
  --log-group-name /myapp/api \
  --filter-name app-errors \
  --filter-pattern '{ $.level = "ERROR" }' \
  --metric-transformations \
      metricName=AppErrors,metricNamespace=MyApp,metricValue=1,defaultValue=0

# Extract a numeric field instead of counting
aws logs put-metric-filter \
  --log-group-name /myapp/api \
  --filter-name slow-requests \
  --filter-pattern '{ $.durationMs > 1000 }' \
  --metric-transformations \
      metricName=SlowRequestMillis,metricNamespace=MyApp,metricValue='$.durationMs'

go deeper

for a junior

Know what a metric filter is for — turning matching log lines into a CloudWatch metric — and that it is attached to a log group with a filter pattern and a metric name.

for a middle

Explain the mechanics: the two filter-pattern dialects, metricValue as a literal or an extracted field, and the two gotchas — no backfill, and a sparse metric unless defaultValue is set.

for a senior

Show the operational judgment: prefer direct metric emission where you own the code, watch dimension cardinality because each combination is a billed custom metric, and recognise that a silent filter can mean stalled ingestion rather than zero errors.

for a principal

Own the boundary between logs-as-metrics and real instrumentation across teams — where paying to parse text back into numbers is acceptable, what the cardinality guardrails are, and how the cost of derived custom metrics is attributed.

## What it is A metric filter is the bridge from unstructured text to a numeric time series. It lives on a log group, it is evaluated at ingestion time, and it emits a CloudWatch metric. The classic use is counting: how many `ERROR` lines per minute, how many `OutOfMemoryError`s, how many 5xx entries in an access log — signals your application never bothered to emit as a metric. You create one with `PutMetricFilter`, supplying: - `filterPattern` — what to match. - `metricTransformations[]` — `metricNamespace`, `metricName`, `metricValue`, optionally `defaultValue`, `unit` and `dimensions`. ```bash aws logs put-metric-filter \ --log-group-name /myapp/api \ --filter-name app-errors \ --filter-pattern '{ $.level = "ERROR" }' \ --metric-transformations \ metricName=AppErrors,metricNamespace=MyApp,metricValue=1,defaultValue=0 ``` ## Filter patterns Two dialects, and mixing them up is a common error: - **Unstructured** — bare terms are matched against the message, space-delimited and case-sensitive: `ERROR` matches any line containing that token; `"Access Denied"` matches the quoted phrase; `?ERROR ?WARN` is an OR. You can also use the space-delimited form `[ip, user, ..., status=5*]` to bind positional fields. - **JSON** — when the message is a JSON document, `{ $.level = "ERROR" }` selects on a field, and you can combine with `&&`/`||` and use numeric comparisons: `{ $.durationMs > 1000 }`. The JSON dialect is why structured logging pays off here: extracting a field is trivial, while parsing it out of prose is not. ## The two behaviours that surprise everyone **No backfill.** A metric filter is applied as events are ingested. Create it at 14:00 and it sees nothing that arrived at 13:59, ever. If you need the historical number, you run a query over the stored data instead — that is a different tool. So the correct interpretation of "my new filter shows nothing" is usually "no matching line has been ingested since I created it", and the correct test is to generate one and watch. **Sparse metrics.** If no line matches during a period, CloudWatch Logs publishes *nothing* — not zero. The metric simply has gaps. Graphs look broken, and anything built on the metric has to decide what a gap means. `defaultValue` fixes this: set it to `0` and the transformation publishes a zero for periods with no matches, giving you a continuous series. As a rule, set `defaultValue: 0` on any counting filter you intend to build on; leave it unset only when a gap is genuinely more meaningful than a zero. ## Dimensions and the cardinality trap A transformation can attach dimensions whose values are pulled from extracted fields — for example dimensioning an error count by `$.service`. This is powerful and dangerous: **every distinct combination of dimension values is a separate custom metric**, and custom metrics are billed per metric per month. Dimension by something unbounded — a request id, a user id, a URL path with ids in it — and you generate an unbounded metric count and a bill to match. Dimension only on values you can enumerate: environment, service name, error class. The filter itself costs nothing; the metrics it creates are ordinary custom metrics and are charged as such. ## Where it sits against the alternatives - If you want to *ask a question* of stored logs interactively, a metric filter is the wrong tool — it is not a query engine, it is a standing rule evaluated forward. - If your application can emit the number itself, emitting a metric directly is cheaper and cleaner than writing a line and paying to parse it back out. - If you need the raw events elsewhere in real time, that is a subscription filter, not a metric filter — same filter-pattern syntax, completely different output. The honest positioning: metric filters are for signals you can only get from text you do not control — a library's stack trace, a managed service's log line, a legacy application nobody will re-instrument. ## Operational notes A log group can carry multiple metric filters, up to a service quota, and each is independent. Editing a pattern does not re-evaluate history either. And because evaluation happens at ingestion, a filter on a log group whose ingestion stalls produces silence that is indistinguishable from "nothing matched" — which is exactly why `defaultValue: 0` plus an independent signal that ingestion is alive is the mature setup.

  • When should the application emit a metric directly instead of you adding a metric filter?
    Whenever you control the code. Emitting the number directly avoids paying to ingest a log line and then paying again for the custom metric, and it gives you dimensions chosen deliberately rather than scraped from text. Metric filters earn their place on output you cannot change — third-party libraries, managed-service logs, legacy applications nobody will re-instrument.
  • What is the risk of adding a dimension to a metric filter transformation?
    Cardinality. Each distinct dimension-value combination becomes its own custom metric, billed monthly, so dimensioning on a request id, user id or raw URL path can generate thousands of metrics from one filter. Restrict dimensions to enumerable values such as environment or service name, and treat any field derived from user input as unsafe.
  • You need the count of errors from last week, but the filter was only created today. What do you do?
    Query the stored events instead — the data is still in the log group even though the filter never saw it. Metric filters are forward-only standing rules and cannot be replayed, so historical counts come from querying the retained logs, and the filter covers you from now on.

saying these in an interview costs you the question

  • Expecting a new metric filter to backfill historical log data
  • Assuming a period with no matches publishes zero
  • Dimensioning a filter on a high-cardinality field like request id
  • Confusing metric filters with subscription filters
  • Thinking the metric filter itself is what costs money

context