skip to content

What does Amazon CloudWatch Application Signals give you that building your own service dashboards from custom metrics does not, and what does an Application Signals SLO actually track?

level: seniorimportance: should knowfreq 28%

answer

  1. standard golden signals, no code change
  2. ADOT auto-instrumentation via the CloudWatch agent
  3. service, operation, environment, remote service
  4. budget spent, not threshold crossed
  5. rolling versus calendar interval

basics

~20 s

CloudWatch Application Signals auto-instruments supported runtimes to emit standardised latency, error and fault metrics per service and operation, plus a service map and dependency view. Its SLO object tracks a goal over an interval, reporting attainment and remaining error budget rather than just a threshold breach.

solid answer

~60 s

The difference is **standardisation without per-team effort**. Application Signals uses AWS Distro for OpenTelemetry auto-instrumentation, enabled through the CloudWatch agent on EKS, ECS, EC2 or Lambda, so supported runtimes emit the same metrics — latency, errors and faults, dimensioned by service, operation, environment and the remote service being called — without anyone hand-writing a `PutMetricData` call. That buys a consistent service map and a dependency view across the estate, which hand-rolled custom metrics never achieve because every team names things differently. On top of those metrics it adds a first-class **SLO** object: you pick an indicator (a service operation's latency or availability, or any CloudWatch metric), a goal such as 99.9%, and an interval that is rolling or calendar-aligned. Application Signals then tracks **attainment and remaining error budget over that interval** — a level of service consumed to date — rather than the instantaneous "is it above the line right now" that a plain alarm answers. It is glue, not magic: it still only covers supported runtimes and platforms.

go deeper

for a junior

Know that Application Signals adds standard latency and error metrics plus a service map for supported AWS workloads without you writing instrumentation code.

for a middle

Explain that it rides on ADOT auto-instrumentation via the CloudWatch agent, emits consistently dimensioned service and operation metrics, and that a service map is only possible because the naming is standardised.

for a senior

Show the SLO mechanics: an indicator, a goal, a rolling or calendar interval, and attainment plus remaining error budget as alarmable metrics — and be honest about coverage gaps and ingestion cost.

for a principal

Own the estate decision — where a standardised AWS-native baseline beats a bespoke OpenTelemetry pipeline, what service naming governance it requires, and what the organisation actually does when a budget runs out.

## The problem with hand-rolled service metrics Any team can emit their own latency and error metrics. The trouble appears at estate scale: one team calls it `LatencyMs`, another `duration`, one dimensions on `endpoint`, another on `route`, one counts a 4xx as an error, another does not. Nothing can be compared, no cross-service view can be built, and the person on call has to learn each service's conventions. **Application Signals** attacks that by making the metrics a platform concern rather than an application concern. ## How it is enabled It uses **AWS Distro for OpenTelemetry (ADOT) auto-instrumentation**, switched on through the CloudWatch agent, for supported runtimes across EKS, ECS, EC2 and Lambda. On Kubernetes that means the CloudWatch agent's operator injecting the instrumentation into pods; elsewhere it is agent configuration plus the auto-instrumentation attached to the process. The application code does not change. Because it rides on OpenTelemetry, hand-written spans you have already added continue to work alongside it. ## What it produces Standardised metrics in the **`ApplicationSignals`** namespace — latency, errors and faults — dimensioned consistently by: - **`Service`** — the logical service name; - **`Operation`** — the endpoint or handler within it; - **`Environment`** — which deployment this is; - **`RemoteService`** / **`RemoteOperation`** — the dependency being called. That last pair is what makes the **service map** possible: because every service reports its callees using the same names, AWS can assemble the topology without anyone drawing it. The map is the successor to the older **ServiceLens** view, which stitched X-Ray traces to CloudWatch metrics and logs for the same purpose. The distinction between **error** and **fault** is worth stating: faults are server-side failures you own; errors include client-caused failures. Conflating them is the classic reason an availability number looks bad because clients are sending bad requests. ## The SLO object An alarm answers *"is this metric over the line right now?"* An SLO answers *"how much of the allowed failure have we spent this period?"* — a different question with different operational consequences. An Application Signals SLO is configured with: - **An indicator.** Either a service operation's latency or availability, discovered automatically from the standardised metrics, or an arbitrary CloudWatch metric or metric-math expression when what you care about is not request-shaped. - **A goal** — the target level, for instance 99.9%. - **An interval** — either **rolling** (a trailing window of N days) or **calendar-aligned** (this month, this quarter). Rolling never resets; calendar gives you a clean budget on the first of the month. Which you choose changes team behaviour more than anything else in the configuration. As of 2025 the indicator may be **period-based** — each evaluation period is judged good or bad and attainment is the fraction of good periods — or **request-based**, where attainment is computed from good requests over total requests. The two give visibly different numbers on bursty traffic, because a period-based SLI weights a quiet minute the same as a busy one. Application Signals then reports **attainment** and **remaining error budget** for the interval, and those are themselves metrics, so you can alarm on the budget rather than on the raw signal. ## Where the boundary is Application Signals is a **measurement mechanism**. Choosing what to measure, what the target should be, and what the team does when the budget is nearly gone is service-level management practice and sits outside the tool. AWS gives you the meter; the policy is yours. ## Honest limitations - **Coverage is conditional.** Only supported runtimes and platforms are auto-instrumented. A legacy service, an unsupported language, or something running outside AWS produces nothing, and a service map with holes in it is trusted less than one that is admittedly partial. - **It is not free.** Application Signals bills for what it ingests, and auto-instrumenting an entire estate at full fidelity is a real number. Estimate before enabling fleet-wide. - **Naming still needs governance.** Auto-instrumentation infers a service name from the environment. If two deployments of the same service report the same name across environments, the map merges them. Set the service name and environment deliberately. - **It is not a full APM.** It gives standardised golden signals, a topology and SLOs. Deep code-level profiling and continuous profiling are separate concerns. ## When to reach for it Use it when you want a consistent, low-effort baseline across many services on AWS-managed compute, particularly if teams have not instrumented themselves. If you already run a mature OpenTelemetry pipeline into a backend you like, the value is narrower — you are mostly buying the SLO object and the AWS-native service map.

  • How is an SLO's error budget operationally different from an alarm on the same metric?
    An alarm is instantaneous — it fires while the metric is over the line and clears afterwards, so a series of short breaches leaves no trace. An error budget accumulates: each bad period or bad request consumes budget for the interval, so ten brief incidents in a month are visible as one nearly-exhausted budget even though every alarm cleared.
  • What changes if you pick a calendar-aligned interval instead of a rolling one?
    A rolling interval never resets, so a bad week keeps weighing on attainment until it ages out — steadier, but there is no clean slate. A calendar interval resets on the first of the period, which is easier to report against and to reason about, but creates a perverse incentive near the end of a month when the budget is already spent.
  • Application Signals shows a service map with obvious gaps. What are the likely reasons?
    Services on unsupported runtimes or platforms, workloads running outside AWS-managed compute, or instrumentation not enabled on that deployment. Dependencies are only drawn from callers that report a remote service, so an uninstrumented caller makes a real edge invisible even when both ends exist.
  • Why does Application Signals distinguish errors from faults?
    Because they mean different things for who is at fault. Faults are server-side failures the service owner must fix; errors include client-caused failures such as malformed requests. Rolling them together makes availability look bad when the cause is a misbehaving client, and hides genuine regressions in the noise.

saying these in an interview costs you the question

  • Thinking Application Signals instruments every workload regardless of runtime
  • Treating an SLO as just an alarm with a nicer chart
  • Ignoring the error versus fault distinction in availability numbers
  • Assuming a service map is complete when callers are uninstrumented
  • Enabling it fleet-wide without estimating ingestion cost

context