skip to content

A new Fargate service deployed into private VPC subnets never reaches RUNNING: every task stops with a CannotPullContainerError naming the ECR registry. What are the likely causes, and how would you work through them?

level: seniorimportance: should knowfreq 50%

answer

  1. the platform pulls, not your code
  2. private subnets need a deliberate path out
  3. two ECR endpoints are not enough
  4. public IP is useless without a gateway route
  5. reproduce from the same subnet

basics

~20 s

The task cannot reach the registry. A Fargate task in a private subnet needs a NAT route or ECR, S3 and CloudWatch Logs VPC endpoints, outbound 443 allowed in its security group, and an execution role with ECR permissions.

solid answer

~50 s

The pull happens over the task's own ENI, inside your subnet, using the **execution role** — so the failure is almost always network reachability or execution-role permissions, not the image. Work through four things. First, egress: a private subnet needs a route to a NAT gateway, or interface endpoints for `com.amazonaws.<region>.ecr.api` and `com.amazonaws.<region>.ecr.dkr` **plus** a gateway endpoint for `com.amazonaws.<region>.s3`, because ECR serves layer blobs from S3. Second, `assignPublicIp` — `ENABLED` only helps in a subnet with an internet gateway route; in a private subnet it is useless, and `DISABLED` in a public subnet with no NAT fails the same way. Third, the security group must allow outbound 443, and interface endpoints' own security groups must allow 443 in from the task. Fourth, the execution role needs the ECR pull actions and private DNS must be enabled on the interface endpoints so the registry hostname resolves to them. Add the `com.amazonaws.<region>.logs` endpoint too, or the `awslogs` driver fails the same way.

go deeper

for a junior

Know that a Fargate task pulls its image over the network from your subnet, so a private subnet needs either NAT or VPC endpoints for the pull to work at all.

for a middle

Name the specific endpoints — ecr.api, ecr.dkr, the S3 gateway endpoint and logs — and explain why the execution role rather than the task role is the identity doing the pull.

for a senior

Show a disciplined bisection: read the stopped reason, check route tables and both sides of the security group pair, reproduce from an instance in the same subnet, and confirm with CloudTrail.

for a principal

Own the pattern so it stops recurring: bake endpoints or NAT into the VPC module, decide the endpoint-versus-NAT economics for the whole estate, and set endpoint policies as a guardrail.

## What the error actually tells you `CannotPullContainerError` in the task's `stoppedReason` means the platform tried to fetch the image and could not. Crucially, the pull is done by the ECS/Fargate infrastructure using the **task execution role**, over the **task's own ENI in your subnet**. That single sentence eliminates most wrong guesses: it is not your application, not the task role, and not (usually) a broken image. It is reachability or authorization for the registry. A related sibling error, `ResourceInitializationError: unable to pull secrets or registry auth`, points at the same class of problem for the `secrets` block or for authentication itself. ## Cause one: no path out of the subnet A private subnet by definition has no internet gateway route. There are exactly two supported ways for the task to reach ECR: 1. **NAT.** A route for `0.0.0.0/0` to a NAT gateway in a public subnet of the same AZ. Simple, works for everything, and bills per GB processed — which for image pulls at scale is a real line item. 2. **VPC endpoints.** Interface (PrivateLink) endpoints for `com.amazonaws.<region>.ecr.api` and `com.amazonaws.<region>.ecr.dkr`, **and** a gateway endpoint for `com.amazonaws.<region>.s3` on the subnet's route table. The S3 one catches people out constantly: ECR stores layer blobs in S3, so the two ECR endpoints alone will authenticate and then fail mid-download. If the task also uses the `awslogs` driver — nearly all do — add the `com.amazonaws.<region>.logs` interface endpoint, otherwise the task may pull fine and then fail during log-stream setup instead. Secrets injection needs `secretsmanager` or `ssm` endpoints on the same logic. ## Cause two: assignPublicIp is wrong for the subnet `awsvpcConfiguration.assignPublicIp` takes `ENABLED` or `DISABLED`. A public IP only functions if the subnet's route table sends `0.0.0.0/0` to an internet gateway. So: - Private subnet + `ENABLED` → still no path out. The address is assigned and unusable. - Public subnet + `DISABLED` → no path out either, and this is the more common accident, because "it's a public subnet" feels like it should be enough. ## Cause three: security groups The task's security group must permit **outbound** TCP 443. Default groups allow all egress, but hardened environments frequently do not. When you use interface endpoints, there is a second hop: the endpoint's own security group must allow **inbound** 443 from the task's security group or CIDR. A one-way rule here produces exactly the same symptom as having no endpoint at all. ## Cause four: identity and DNS Check that `executionRoleArn` is set and that the role grants `ecr:GetAuthorizationToken`, `ecr:BatchCheckLayerAvailability`, `ecr:GetDownloadUrlForLayer` and `ecr:BatchGetImage` — the managed `AmazonECSTaskExecutionRolePolicy` covers these. If the repository lives in another account, its **repository policy** must also allow the pulling account; an allow in the role alone is not enough for a cross-account resource. For endpoints, private DNS must be enabled, otherwise the ECR hostname still resolves to public addresses that the subnet cannot reach. And if you attached an endpoint policy, confirm it does not exclude the ECR or S3 actions you need. ## A working order of investigation 1. `aws ecs describe-tasks` and read `stoppedReason` verbatim — the sibling errors point at different subsystems. 2. Check the subnet's route table for a NAT route, and list the VPC endpoints attached to that VPC. 3. Check the task security group's egress and the endpoint security groups' ingress. 4. Confirm `executionRoleArn` is present and its policy is attached; check CloudTrail for a denied ECR call from that role. 5. Reproduce from an EC2 instance in the *same subnet with the same security group* — if `aws ecr get-login-password` and a pull work there, the network is fine and the problem is the role or the task definition. ## Design guidance Endpoints versus NAT is a genuine tradeoff, not just a fix. Interface endpoints cost an hourly charge per endpoint per AZ plus data processing, but they keep traffic off the internet, remove a NAT bottleneck, and let you scope access with endpoint policies. A NAT gateway is one thing to run but charges every gigabyte, and image pulls are gigabytes. For a cluster of any size, the endpoints usually win on both cost and security posture. Whichever you pick, encode it in the VPC template so a new private subnet is never created without it — this failure is nearly always a subnet that missed the pattern.

  • Why is an S3 gateway endpoint required when using ECR interface endpoints?
    Because ECR only serves the API and authentication over its own endpoints; the actual image layer blobs are downloaded from Amazon S3. Without a route to S3, the task authenticates successfully and then stalls fetching layers, which reads as a pull failure. The S3 gateway endpoint is added to the subnet's route table and is free of hourly charges.
  • The same image pulls fine from a public subnet but fails in the private one. Does that prove the image is fine?
    Yes — it proves the image, the repository policy and the execution role's ECR permissions all work, because the public-subnet task used the identical execution role. That narrows the problem to the private subnet's egress path: route table, NAT, endpoints, or security group rules. It is a fast, high-value bisection when you are unsure which layer is at fault.
  • How would you decide between a NAT gateway and interface endpoints for this traffic?
    On volume and posture. NAT bills per gigabyte processed, and image pulls plus log shipping are steady gigabytes, so a busy cluster often pays more for NAT than for the hourly endpoint charges. Endpoints also keep the traffic off the public internet and can be constrained with endpoint policies. NAT stays necessary for genuinely arbitrary internet egress.
  • The task now pulls the image but stops before any application log appears. Where do you look next?
    Log delivery. With the awslogs driver the platform needs to reach CloudWatch Logs, so either a `logs` interface endpoint or a NAT route, plus `logs:CreateLogStream` and `logs:PutLogEvents` on the execution role. The stopped reason typically names a resource-initialization failure rather than an image pull, which is the tell that you have moved one step further along startup.

saying these in an interview costs you the question

  • Blames the task role instead of the execution role
  • Adds ECR endpoints but forgets the S3 gateway endpoint
  • Thinks assignPublicIp ENABLED works in a private subnet
  • Assumes the image tag or Dockerfile is at fault
  • Opens the task security group without checking endpoint ingress

context