skip to content

An EC2 instance shows CPU and network metrics in CloudWatch, but there is no memory-used or disk-space metric for it anywhere. Why not, and what do you have to run to get them?

level: juniorimportance: should knowfreq 62%

answer

  1. hypervisor sees outside, not inside
  2. memory is not externally observable
  3. agent required for guest metrics
  4. CWAgent namespace, mem_used_percent
  5. detailed monitoring changes period only

basics

~20 s

CloudWatch's built-in EC2 metrics are collected outside the instance, at the virtualization layer, so they cover CPU, network and EBS activity but never anything inside the guest. Memory and filesystem usage require installing the CloudWatch agent.

solid answer

~50 s

The `AWS/EC2` namespace is filled in by the platform, not by your instance. AWS can see how much CPU time the hypervisor handed you, how many bytes crossed the network interface and how the EBS volume behaved, but it has no view inside your operating system — so there is no memory metric, no free-disk metric and no per-process metric. To get those you install the CloudWatch agent, attach an instance profile carrying the `CloudWatchAgentServerPolicy` managed policy, and give the agent a JSON configuration (commonly stored in SSM Parameter Store). By default it publishes into the `CWAgent` namespace with names such as `mem_used_percent` and `disk_used_percent`, and its `procstat` plugin can add per-process metrics. Two consequences follow: every unique agent metric is a billable custom metric, and EC2's own metrics still arrive every five minutes unless you enable detailed monitoring for one-minute granularity.

code

json · 15 lines
json
{
  "agent": { "metrics_collection_interval": 60 },
  "metrics": {
    "namespace": "CWAgent",
    "append_dimensions": {
      "InstanceId": "${aws:InstanceId}",
      "AutoScalingGroupName": "${aws:AutoScalingGroupName}"
    },
    "metrics_collected": {
      "mem": { "measurement": ["mem_used_percent"] },
      "disk": { "measurement": ["disk_used_percent"], "resources": ["/"] },
      "swap": { "measurement": ["swap_used_percent"] }
    }
  }
}

go deeper

for a junior

Say plainly that AWS measures EC2 from the outside, so memory and free disk are not there by default, and that the CloudWatch agent is what you install to get them.

for a middle

Explain the mechanics: the agent publishes custom metrics via PutMetricData into the CWAgent namespace, needs an instance profile with CloudWatchAgentServerPolicy, and takes a JSON config typically delivered from SSM Parameter Store.

for a senior

Show you have rolled this out: the private-subnet egress requirement, diagnosing a silent agent from its own log, and the fact that per-instance dimensions turn an agent rollout into hundreds of billable custom metrics.

for a principal

Own the fleet-wide policy — which guest measurements are worth collecting everywhere versus on demand, how the collection interval and dimension design set the monthly bill, and whether the CloudWatch agent or a collector you already run should own host telemetry.

## Two different collectors, two different vantage points CloudWatch metrics for EC2 come from two entirely separate sources, and most confusion about "missing" metrics is really confusion about which source you are looking at. The first source is AWS itself. The virtualization layer that runs your instance measures what it can see from the outside and publishes it into the `AWS/EC2` namespace with an `InstanceId` dimension: `CPUUtilization`, `NetworkIn` / `NetworkOut`, `NetworkPacketsIn` / `NetworkPacketsOut`, `StatusCheckFailed` and its system/instance variants, and (for instance-store-backed devices) disk read/write counters. EBS volume behaviour appears separately under `AWS/EBS` keyed by `VolumeId`. You do not install anything, you cannot turn it off, and you cannot extend it. What that vantage point fundamentally cannot see is anything the guest operating system knows: how much RAM the kernel has handed out, how full a filesystem is, how many file descriptors a process holds, whether swap is being used. Memory in particular is not derivable from outside — the hypervisor allocated the instance its RAM at launch and has no idea what Linux or Windows did with it afterwards. ## What the CloudWatch agent adds The second source is the CloudWatch agent (`amazon-cloudwatch-agent`), a process you install and run inside the instance. It reads the same places you would read by hand — `/proc/meminfo`, `statfs`, `/proc/<pid>/stat` — and calls the CloudWatch `PutMetricData` API to publish them as custom metrics. Its default namespace is `CWAgent`, and typical metric names are `mem_used_percent`, `swap_used_percent`, `disk_used_percent`, `disk_inodes_free` and `diskio_io_time`. The `procstat` plugin adds per-process CPU and memory. The same agent binary also ships logs, which is why teams often already have it running and simply have not enabled the `metrics` section of its config. A minimal configuration looks like this: ```json { "agent": { "metrics_collection_interval": 60 }, "metrics": { "namespace": "CWAgent", "append_dimensions": { "InstanceId": "${aws:InstanceId}", "AutoScalingGroupName": "${aws:AutoScalingGroupName}" }, "metrics_collected": { "mem": { "measurement": ["mem_used_percent"] }, "disk": { "measurement": ["disk_used_percent"], "resources": ["/"] } } } } ``` Store that in SSM Parameter Store and point the agent at it: ```bash sudo /opt/aws/amazon-cloudwatch-agent/bin/amazon-cloudwatch-agent-ctl \ -a fetch-config -m ec2 -c ssm:AmazonCloudWatch-linux -s ``` ## Permissions and network path The agent needs credentials, and on EC2 that means an instance profile. The `CloudWatchAgentServerPolicy` managed policy grants what it needs: `cloudwatch:PutMetricData`, log ingestion permissions, and `ssm:GetParameter` for pulling the config. If the instance sits in a private subnet with no route to the internet, the agent also needs a path to the CloudWatch endpoint — either a NAT gateway or an interface VPC endpoint for `com.amazonaws.<region>.monitoring`. A silently non-reporting agent is far more often a permissions or routing problem than a configuration one, and the agent's own log at `/opt/aws/amazon-cloudwatch-agent/logs/amazon-cloudwatch-agent.log` will say which. ## Granularity is a separate axis Basic monitoring publishes `AWS/EC2` metrics at five-minute periods. Detailed monitoring, enabled per instance or in a launch template, publishes them at one-minute periods for an additional charge. It changes only the period — it does not add memory or disk-space metrics, and no amount of detailed monitoring will make them appear. Agent metrics have their own cadence set by `metrics_collection_interval`. ## What it costs Every unique combination of namespace, metric name and dimension values is one billable custom metric per month. A fleet of 200 instances each publishing four agent metrics with an `InstanceId` dimension is 800 custom metrics, and adding a second dimension such as device or mount point multiplies it again. This is the standard way an agent rollout turns into a surprising CloudWatch bill: aggregate where you can (the agent's `aggregation_dimensions` lets you also publish rolled-up series), collect only the measurements you will actually alarm on, and resist adding a dimension with unbounded values. ## Traps worth naming `CPUUtilization` is hypervisor-measured CPU time, not the guest's load average, and on burstable T-family instances you also want `CPUCreditBalance` to explain a slow instance. A stopped agent leaves no gap warning — the metric simply stops, which is exactly the case a naively configured alarm will fail to catch.

  • The agent is installed and the instance has a role, but nothing appears under CWAgent. Where do you look?
    Start with the agent's own log at `/opt/aws/amazon-cloudwatch-agent/logs/amazon-cloudwatch-agent.log`. The usual causes are: the config was never fetched (`amazon-cloudwatch-agent-ctl -a status`), the instance profile lacks `cloudwatch:PutMetricData` or `ssm:GetParameter`, or the instance is in a private subnet with no NAT route and no interface endpoint for `com.amazonaws.<region>.monitoring`, so every publish call times out.
  • Does enabling detailed monitoring on an EC2 instance give you any metrics you did not have before?
    No. It changes the publication period for the existing `AWS/EC2` metrics from five minutes to one minute, for a per-instance charge. The metric set is identical. Its practical value is alarm responsiveness — a one-minute period lets an alarm react in a couple of minutes rather than fifteen — and finer-grained scaling signals.
  • You want one alarm for 'any instance in this Auto Scaling group is low on disk'. What makes that awkward with per-instance agent metrics?
    Each instance publishes its own series keyed by `InstanceId`, so a single alarm on one series only watches one host, and instances come and go. Options are publishing an aggregated series via the agent's `aggregation_dimensions`, or building the alarm on a metric-math expression that reduces the group's series to one — a `SEARCH()` expression can display them all, but you cannot create an alarm directly on a `SEARCH()` result.

saying these in an interview costs you the question

  • Claiming CloudWatch reports EC2 memory usage out of the box
  • Thinking detailed monitoring adds memory or disk-space metrics
  • Believing the CloudWatch agent only ships logs, not metrics
  • Assuming agent metrics are free like the built-in EC2 ones
  • Forgetting the instance profile and the network path to CloudWatch

context