skip to content

An application on an EC2 instance writes to a log file on disk and you want those lines in Amazon CloudWatch Logs. What are the ways log data gets into CloudWatch Logs, and what does each path require you to install or grant?

level: juniorimportance: must knowfreq 65%

answer

  1. nothing reads the file for you
  2. agent, native delivery, or direct API
  3. the instance role has to allow it
  4. CreateLogStream plus PutLogEvents

basics

~20 s

Three paths: the CloudWatch agent tailing files on a host, AWS services delivering logs natively (Lambda, the ECS awslogs driver, VPC Flow Logs), or your code calling the PutLogEvents API. All three need IAM permission to create streams and put events.

solid answer

~40 s

There are three. First, an agent on the host: you install the unified CloudWatch agent, point its config at the file, and give the instance role `logs:CreateLogStream` and `logs:PutLogEvents` — the managed policy `CloudWatchAgentServerPolicy` covers it. Second, native service delivery: Lambda writes its own logs given `AWSLambdaBasicExecutionRole`, ECS ships container stdout through the `awslogs` log driver in the task definition, and AWS-vended logs such as VPC Flow Logs or API Gateway access logs are delivered by the service once you name a destination log group. Third, direct API — your process calls `PutLogEvents` itself, which couples the app to AWS and means you own batching and retries. For a file on disk, the agent is the normal answer; nothing reads the file for you.

code

json · 16 lines
json
{
  "logs": {
    "logs_collected": {
      "files": {
        "collect_list": [
          {
            "file_path": "/var/log/myapp/app.log",
            "log_group_name": "/myapp/app",
            "log_stream_name": "{instance_id}",
            "timestamp_format": "%Y-%m-%d %H:%M:%S"
          }
        ]
      }
    }
  }
}

go deeper

for a junior

Be able to name the three routes in one breath — agent on a host, the AWS service delivering its own logs, or a direct PutLogEvents call — and say that each needs IAM permission to write.

for a middle

Explain where each path's permissions actually live: the instance profile for the agent, the execution role for Lambda, the task execution role for the ECS awslogs driver, and a service-linked role or log-group resource policy for vended logs.

for a senior

Show the triage habit: empty log group means IAM or network. Talk through checking the role, then the private-subnet route to the CloudWatch Logs endpoint, and know when a Fluent Bit sidecar beats the stock driver.

for a principal

Own the fleet-wide decision of which telemetry lands in CloudWatch Logs at all versus S3, and the standards that make it consistent — naming conventions for log groups, a single agent configuration channel, and permissions attached at the role template rather than per workload.

## The model you are writing into CloudWatch Logs stores events in a **log group** (the logical container, e.g. `/myapp/app`) which holds **log streams** (one sequence of events from one source). Nothing arrives by magic: something must call the service's ingestion API. The three paths below are just three different things doing that call. ## Path 1 — an agent on the host For an application writing to a file on EC2 or on-premises, you install the **unified CloudWatch agent** (package `amazon-cloudwatch-agent`). Its configuration file has a `logs.logs_collected.files.collect_list` section naming the file path, the target log group, and the stream name — the placeholder `{instance_id}` is commonly used so each instance gets its own stream. The agent authenticates with the instance's IAM role (via the instance profile). It needs at least `logs:CreateLogGroup`, `logs:CreateLogStream`, `logs:PutLogEvents` and `logs:DescribeLogStreams`; AWS ships `CloudWatchAgentServerPolicy` for exactly this. The single most common failure is an instance role that has no logs permissions — the agent runs, its own log fills with access-denied errors, and the log group never appears. An alternative on containers is a log-forwarding sidecar (ECS `awsfirelens` with Fluent Bit), which is the same idea with more routing flexibility. ## Path 2 — the service delivers its own logs Many AWS services write to CloudWatch Logs for you: - **Lambda** writes each invocation's stdout/stderr to `/aws/lambda/<function-name>`, provided the execution role allows the log actions — the managed `AWSLambdaBasicExecutionRole` grants them. Attach nothing and the function still runs, but you get no logs. - **ECS/Fargate** uses the `awslogs` log driver, configured in the task definition's `logConfiguration` with options `awslogs-group`, `awslogs-region` and `awslogs-stream-prefix`; the **task execution role** needs the log permissions. - **Vended logs** — VPC Flow Logs, Route 53 Resolver query logs, API Gateway access and execution logs, RDS engine logs — are produced by the AWS service itself and delivered to a log group you nominate. AWS bills these on a separate, cheaper *vended logs* pricing tier, and for high-volume ones such as Flow Logs you can usually choose S3 as the destination instead. A useful mental note: with vended logs you are not granting your own principal anything, you are configuring the *service* to deliver, which usually means a service-linked role or a resource policy on the log group. ## Path 3 — call the API directly Your process can call `PutLogEvents` with a batch of `{timestamp, message}` entries against a named group and stream. Logging libraries with CloudWatch appenders do this. It works, but you inherit batching, retry, ordering and back-off yourself, and the application now needs AWS credentials and network reachability to the CloudWatch Logs endpoint. Older SDK code threaded a `sequenceToken` returned by the previous call back into the next one; AWS has since removed that requirement, so modern code should not be building fragile token-chaining logic. Direct API calls are the right choice mainly when there is no file and no host you control — for example inside a process that already batches structured events. ## The private-subnet trap All three paths reach a public AWS endpoint. Something in a private subnet with no NAT route and no `com.amazonaws.<region>.logs` interface VPC endpoint cannot ship logs at all — the symptom is a workload that is clearly running and a log group that stays empty. This bites Lambda-in-a-VPC and locked-down EC2 fleets constantly, and it is a network problem, not a logging one. ## What an interviewer is checking That you know logs do not collect themselves, that you can name the agent for host files and the `awslogs` driver for containers, and that you go straight to IAM when nothing shows up. Saying "CloudWatch just picks up the file" is the answer that ends the topic badly.

  • A Lambda function is clearly running but has no log group at all. What do you check?
    Its execution role. Lambda creates `/aws/lambda/<name>` and writes to it using the function's own role, so a role without `logs:CreateLogGroup`, `logs:CreateLogStream` and `logs:PutLogEvents` — typically a hand-rolled role missing `AWSLambdaBasicExecutionRole` — produces a silently log-free function. If the role is fine, check whether logging was disabled or redirected in the function's logging configuration.
  • Why does a workload in a private subnet sometimes fail to ship logs even with correct IAM permissions?
    Because the CloudWatch Logs API is a public endpoint. Without a NAT gateway route or a `com.amazonaws.<region>.logs` interface VPC endpoint, the agent or SDK cannot reach it, so calls time out and events are dropped. IAM is fine; routing is not. The give-away is a healthy workload with an empty or missing log group.
  • When would you prefer sending VPC Flow Logs to S3 rather than to CloudWatch Logs?
    When the volume is high and you query it rarely. Flow Logs are billed on the vended-logs ingestion tier per GB in CloudWatch Logs, whereas S3 storage is far cheaper and Athena can query the files on demand. Choose CloudWatch Logs when you want the records alongside application logs for immediate, interactive investigation.

saying these in an interview costs you the question

  • Assuming CloudWatch automatically picks up files on an instance
  • Thinking the CloudWatch agent needs no IAM permissions
  • Confusing the ECS awslogs driver with installing an agent in the image
  • Believing a private subnet can reach CloudWatch Logs without NAT or an endpoint
  • Saying Lambda logs appear regardless of the execution role

context