skip to content

An Amazon ECS task definition can reference both a task role and a task execution role. What is each one for, who uses it, and how do the symptoms differ when one of them is missing a permission?

level: middleimportance: must knowfreq 68%

answer

  1. two actors, two roles
  2. one runs before your code exists
  3. secrets injection is not the app's role
  4. failure to start vs AccessDenied at runtime
  5. per-task-definition least privilege

basics

~20 s

The task execution role belongs to the ECS agent and Fargate infrastructure: pulling the image, writing logs, injecting secrets before the container starts. The task role is the application's own identity for AWS API calls at runtime. Missing execution-role permissions fail the start; missing task-role permissions cause AccessDenied inside the app.

solid answer

~50 s

They serve two different consumers. The **task execution role** (`executionRoleArn`) is assumed by the ECS agent or the Fargate infrastructure *on your behalf, before your code runs* — it pulls the image from ECR, creates the CloudWatch log stream, and resolves any Secrets Manager or SSM parameters referenced in the task definition's `secrets` block. The AWS managed policy `AmazonECSTaskExecutionRolePolicy` is the usual starting point. The **task role** (`taskRoleArn`) is the identity the containers themselves get: ECS injects a container credentials environment variable, the SDK fetches temporary credentials from the local container credentials endpoint, and calls to S3, DynamoDB and so on are signed as that role. The symptoms separate cleanly: an execution-role gap shows up as a task that never reaches RUNNING — `CannotPullContainerError`, a log-configuration failure, or a secret that cannot be fetched — while a task-role gap shows up as AccessDenied in your application logs from a task that started fine.

code

json · 10 lines
json
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": "secretsmanager:GetSecretValue",
      "Resource": "arn:aws:secretsmanager:eu-west-1:111122223333:secret:orders/db-AbCdEf"
    }
  ]
}

go deeper

for a junior

Know that ECS tasks get their AWS permissions from a role rather than from keys in the image, and that the task definition names two separate roles for two separate purposes.

for a middle

Explain what each role does — image pull, log stream and secret resolution before startup versus the application's own API calls — and map a given error message to the role that owns it.

for a senior

Demonstrate the diagnosis and the scoping judgment: one task role per service, secrets on the execution role, and awareness that on the EC2 launch type the instance profile can shadow the task role entirely.

for a principal

Own the platform rule that every workload gets its own narrowly scoped task role, that shared cluster-wide roles are treated as a defect, and that guardrails prevent teams from widening the execution role into a general-purpose grant.

## Two roles because there are two actors A container running on ECS involves two parties that call AWS. Before your process exists, something must fetch the image, wire up logging, and resolve secrets into the container's environment. After it starts, your code calls AWS services. AWS refuses to blur those: the first actor is the ECS/Fargate infrastructure and uses the **task execution role**; the second is your application and uses the **task role**. Both are ordinary IAM roles whose trust policy allows `ecs-tasks.amazonaws.com` to assume them; only their consumers and their timing differ. ## The task execution role Set as `executionRoleArn` in the task definition, this role is used by the agent/infrastructure at task-launch time for: - **Pulling the image** from Amazon ECR — `ecr:GetAuthorizationToken`, `ecr:BatchGetImage`, `ecr:GetDownloadUrlForLayer`. Public images from other registries need no permission here. - **Log delivery** for the `awslogs` driver — `logs:CreateLogStream` and `logs:PutLogEvents` (plus `logs:CreateLogGroup` if you let ECS create the group). - **Secret injection.** If the container definition has a `secrets` block pointing at Secrets Manager ARNs or SSM parameters, the *execution* role — not the task role — must be allowed `secretsmanager:GetSecretValue` or `ssm:GetParameters`, and `kms:Decrypt` on the key if the secret is encrypted with a customer managed key. This is the single most common ECS IAM mistake: the permission is added to the task role, where it does nothing. The managed policy `AmazonECSTaskExecutionRolePolicy` covers the ECR and CloudWatch Logs basics; secret access is yours to add. ## The task role Set as `taskRoleArn`, this is the workload's identity. ECS injects an environment variable into every container in the task pointing at a credentials path on the local agent endpoint; the AWS SDKs know to call it as part of their default credential lookup and receive temporary STS credentials scoped to that role, refreshed automatically. Your code configures nothing. ```json { "family": "orders", "executionRoleArn": "arn:aws:iam::111122223333:role/ecsTaskExecutionRole", "taskRoleArn": "arn:aws:iam::111122223333:role/orders-task", "containerDefinitions": [{ "name": "api", "image": "111122223333.dkr.ecr.eu-west-1.amazonaws.com/orders:1.4.2", "secrets": [{ "name": "DB_PASSWORD", "valueFrom": "arn:aws:secretsmanager:eu-west-1:111122223333:secret:orders/db-AbCdEf" }] }] } ``` The task role is per task definition, so it is the natural least-privilege boundary: the orders service gets its own role with its own bucket and table, and a compromise there does not reach the payments service's data. Sharing one fat role across every task in a cluster throws that away. ## Reading the failure The two roles fail at different moments, which makes diagnosis fast: - Task stops before RUNNING with `CannotPullContainerError`, or with a message about the log configuration, or `ResourceInitializationError` mentioning a secret → **execution role** (or, on a private subnet, no route to ECR/Secrets Manager, which needs a NAT path or interface endpoints). - Task is healthy, your logs show `AccessDenied` calling S3/DynamoDB/SQS → **task role**. Confirm with `aws sts get-caller-identity` from inside the container (`ecs execute-command`) — it should print the task role's assumed-role ARN. ## The EC2 launch-type trap On the Fargate launch type a task is isolated and only ever sees its task role. On the EC2 launch type the container shares a host that has its own instance profile, and unless the agent is configured to block it, a container in bridge mode can reach the instance metadata endpoint and pick up the *instance* role. That means a task might quietly work while its task role grants nothing — and it means the node role's permissions become every container's permissions. Keep the instance profile minimal (what the ECS agent itself needs) and block container access to the metadata endpoint; the credential shadowing is otherwise invisible until an audit. ## Rule of thumb If a permission is needed to *start* the container, it belongs on the execution role. If it is needed by code *inside* the container, it belongs on the task role. Every task definition should name both, and neither should be shared more widely than the workload it serves.

  • A container fails to start with ResourceInitializationError while fetching a secret, and the team has already granted secretsmanager:GetSecretValue to the task role. What do you tell them?
    The permission is on the wrong role. Secrets referenced in the task definition's `secrets` block are resolved by the ECS infrastructure before the container process exists, so the grant must be on the task execution role — and if the secret uses a customer managed KMS key, that role also needs `kms:Decrypt` on it. The task role is irrelevant until the container is already running.
  • How does the application inside an ECS task actually receive the task role's credentials?
    ECS injects a container credentials environment variable into every container in the task, pointing at a path on the agent's local credentials endpoint. The AWS SDKs consult that source as part of their default credential lookup, retrieve temporary STS credentials for the task role, and refresh them before expiry. Nothing is configured in application code.
  • Why is it risky to run ECS tasks on the EC2 launch type with a permissive instance profile?
    Containers on that host can reach the instance metadata endpoint unless it is blocked, so they can obtain the node's instance-profile credentials in addition to — or instead of — their task role. Every container effectively inherits the node's permissions, and the task role stops being a real boundary. Keep the instance profile to what the ECS agent needs and block container access to metadata.

saying these in an interview costs you the question

  • Calling them the same role with two names
  • Putting secret-fetch permissions on the task role
  • Assuming the task role pulls the container image
  • One shared role for every task in the cluster
  • Ignoring that EC2-launch-type containers can reach the instance role

context