What does Amazon CloudWatch Application Signals give you that building your own service dashboards from custom metrics does not, and what does an Application Signals SLO actually track?
answer
- standard golden signals, no code change
- ADOT auto-instrumentation via the CloudWatch agent
- service, operation, environment, remote service
- budget spent, not threshold crossed
- rolling versus calendar interval
basics
~20 sCloudWatch Application Signals auto-instruments supported runtimes to emit standardised latency, error and fault metrics per service and operation, plus a service map and dependency view. Its SLO object tracks a goal over an interval, reporting attainment and remaining error budget rather than just a threshold breach.
solid answer
~60 sThe difference is **standardisation without per-team effort**. Application Signals uses AWS Distro for OpenTelemetry auto-instrumentation, enabled through the CloudWatch agent on EKS, ECS, EC2 or Lambda, so supported runtimes emit the same metrics — latency, errors and faults, dimensioned by service, operation, environment and the remote service being called — without anyone hand-writing a `PutMetricData` call. That buys a consistent service map and a dependency view across the estate, which hand-rolled custom metrics never achieve because every team names things differently. On top of those metrics it adds a first-class **SLO** object: you pick an indicator (a service operation's latency or availability, or any CloudWatch metric), a goal such as 99.9%, and an interval that is rolling or calendar-aligned. Application Signals then tracks **attainment and remaining error budget over that interval** — a level of service consumed to date — rather than the instantaneous "is it above the line right now" that a plain alarm answers. It is glue, not magic: it still only covers supported runtimes and platforms.
go deeper
Know that Application Signals adds standard latency and error metrics plus a service map for supported AWS workloads without you writing instrumentation code.
Explain that it rides on ADOT auto-instrumentation via the CloudWatch agent, emits consistently dimensioned service and operation metrics, and that a service map is only possible because the naming is standardised.
Show the SLO mechanics: an indicator, a goal, a rolling or calendar interval, and attainment plus remaining error budget as alarmable metrics — and be honest about coverage gaps and ingestion cost.
Own the estate decision — where a standardised AWS-native baseline beats a bespoke OpenTelemetry pipeline, what service naming governance it requires, and what the organisation actually does when a budget runs out.
## The problem with hand-rolled service metrics Any team can emit their own latency and error metrics. The trouble appears at estate scale: one team calls it `LatencyMs`, another `duration`, one dimensions on `endpoint`, another on `route`, one counts a 4xx as an error, another does not. Nothing can be compared, no cross-service view can be built, and the person on call has to learn each service's conventions. **Application Signals** attacks that by making the metrics a platform concern rather than an application concern. ## How it is enabled It uses **AWS Distro for OpenTelemetry (ADOT) auto-instrumentation**, switched on through the CloudWatch agent, for supported runtimes across EKS, ECS, EC2 and Lambda. On Kubernetes that means the CloudWatch agent's operator injecting the instrumentation into pods; elsewhere it is agent configuration plus the auto-instrumentation attached to the process. The application code does not change. Because it rides on OpenTelemetry, hand-written spans you have already added continue to work alongside it. ## What it produces Standardised metrics in the **`ApplicationSignals`** namespace — latency, errors and faults — dimensioned consistently by: - **`Service`** — the logical service name; - **`Operation`** — the endpoint or handler within it; - **`Environment`** — which deployment this is; - **`RemoteService`** / **`RemoteOperation`** — the dependency being called. That last pair is what makes the **service map** possible: because every service reports its callees using the same names, AWS can assemble the topology without anyone drawing it. The map is the successor to the older **ServiceLens** view, which stitched X-Ray traces to CloudWatch metrics and logs for the same purpose. The distinction between **error** and **fault** is worth stating: faults are server-side failures you own; errors include client-caused failures. Conflating them is the classic reason an availability number looks bad because clients are sending bad requests. ## The SLO object An alarm answers *"is this metric over the line right now?"* An SLO answers *"how much of the allowed failure have we spent this period?"* — a different question with different operational consequences. An Application Signals SLO is configured with: - **An indicator.** Either a service operation's latency or availability, discovered automatically from the standardised metrics, or an arbitrary CloudWatch metric or metric-math expression when what you care about is not request-shaped. - **A goal** — the target level, for instance 99.9%. - **An interval** — either **rolling** (a trailing window of N days) or **calendar-aligned** (this month, this quarter). Rolling never resets; calendar gives you a clean budget on the first of the month. Which you choose changes team behaviour more than anything else in the configuration. As of 2025 the indicator may be **period-based** — each evaluation period is judged good or bad and attainment is the fraction of good periods — or **request-based**, where attainment is computed from good requests over total requests. The two give visibly different numbers on bursty traffic, because a period-based SLI weights a quiet minute the same as a busy one. Application Signals then reports **attainment** and **remaining error budget** for the interval, and those are themselves metrics, so you can alarm on the budget rather than on the raw signal. ## Where the boundary is Application Signals is a **measurement mechanism**. Choosing what to measure, what the target should be, and what the team does when the budget is nearly gone is service-level management practice and sits outside the tool. AWS gives you the meter; the policy is yours. ## Honest limitations - **Coverage is conditional.** Only supported runtimes and platforms are auto-instrumented. A legacy service, an unsupported language, or something running outside AWS produces nothing, and a service map with holes in it is trusted less than one that is admittedly partial. - **It is not free.** Application Signals bills for what it ingests, and auto-instrumenting an entire estate at full fidelity is a real number. Estimate before enabling fleet-wide. - **Naming still needs governance.** Auto-instrumentation infers a service name from the environment. If two deployments of the same service report the same name across environments, the map merges them. Set the service name and environment deliberately. - **It is not a full APM.** It gives standardised golden signals, a topology and SLOs. Deep code-level profiling and continuous profiling are separate concerns. ## When to reach for it Use it when you want a consistent, low-effort baseline across many services on AWS-managed compute, particularly if teams have not instrumented themselves. If you already run a mature OpenTelemetry pipeline into a backend you like, the value is narrower — you are mostly buying the SLO object and the AWS-native service map.
- How is an SLO's error budget operationally different from an alarm on the same metric?An alarm is instantaneous — it fires while the metric is over the line and clears afterwards, so a series of short breaches leaves no trace. An error budget accumulates: each bad period or bad request consumes budget for the interval, so ten brief incidents in a month are visible as one nearly-exhausted budget even though every alarm cleared.
- What changes if you pick a calendar-aligned interval instead of a rolling one?A rolling interval never resets, so a bad week keeps weighing on attainment until it ages out — steadier, but there is no clean slate. A calendar interval resets on the first of the period, which is easier to report against and to reason about, but creates a perverse incentive near the end of a month when the budget is already spent.
- Application Signals shows a service map with obvious gaps. What are the likely reasons?Services on unsupported runtimes or platforms, workloads running outside AWS-managed compute, or instrumentation not enabled on that deployment. Dependencies are only drawn from callers that report a remote service, so an uninstrumented caller makes a real edge invisible even when both ends exist.
- Why does Application Signals distinguish errors from faults?Because they mean different things for who is at fault. Faults are server-side failures the service owner must fix; errors include client-caused failures such as malformed requests. Rolling them together makes availability look bad when the cause is a misbehaving client, and hides genuine regressions in the noise.
saying these in an interview costs you the question
- Thinking Application Signals instruments every workload regardless of runtime
- Treating an SLO as just an alarm with a nicer chart
- Ignoring the error versus fault distinction in availability numbers
- Assuming a service map is complete when callers are uninstrumented
- Enabling it fleet-wide without estimating ingestion cost