AWS SDKs resolve credentials through a default provider chain. What sources does it consult, in roughly what order, and why does that ordering explain a container that ignores its attached role and calls AWS as some other principal?
answer
- explicit first, ambient last
- environment beats the attached role
- a baked profile is still a credential
- who am I, really
- one command names the principal
basics
~20 sThe chain walks from most explicit to most ambient: credentials passed in code, then environment variables, then the shared config and credentials files, then container or instance metadata last. Anything left in the environment or a baked profile therefore silently outranks the workload's attached role.
solid answer
~50 sEvery AWS SDK, and the CLI, resolve credentials by walking a chain from the most explicit source to the most ambient one and stopping at the first that yields credentials. Roughly: credentials passed explicitly in code, then process/system properties in some SDKs, then the environment variables `AWS_ACCESS_KEY_ID`/`AWS_SECRET_ACCESS_KEY`/`AWS_SESSION_TOKEN`, then a web identity token file (`AWS_WEB_IDENTITY_TOKEN_FILE` with `AWS_ROLE_ARN`), then the shared config and credentials files under `~/.aws` selected by `AWS_PROFILE`, then the container credentials endpoint that ECS and EKS Pod Identity inject, and finally the EC2 instance metadata service. The ordering is deliberate — an explicit override should beat the ambient identity — but it means a stale `AWS_ACCESS_KEY_ID` baked into an image, or an `AWS_PROFILE` copied from a developer laptop, silently wins over the task or instance role. The symptom is AccessDenied under an unexpected principal; `aws sts get-caller-identity` names it immediately.
code
bash · 10 lines# Inside a container that should be using its task role
aws sts get-caller-identity
# {"Arn": "arn:aws:iam::111122223333:user/legacy-deploy", ...} <- not the role
aws configure list
# Name Value Type Location
# access_key ****************ABCD env AWS_ACCESS_KEY_ID
# region eu-west-1 config-file
env | grep -i '^AWS_' # find what is shadowing the rolego deeper
Know that the SDK finds credentials for you and that leftover AWS_ environment variables can override the role you attached — and that aws sts get-caller-identity tells you who you actually are.
Lay out the chain from explicit to ambient and explain why a role attached by the platform sits at the bottom, then trace a shadowing bug from symptom to cause.
Show the operational discipline: treat stray credentials in a production container as an incident, verify identity from inside the workload, and recognise resolution failures that masquerade as permission failures.
Own the guardrails that make shadowing impossible — image scanning for baked credentials, no long-lived keys issued at all, and a platform contract that workloads receive identity from the runtime rather than from configuration.
## What the chain is for The same application binary should run on a laptop, in CI, and on ECS without code changes, and pick up whatever credentials that context provides. The AWS SDKs achieve this with a **default credential provider chain**: an ordered list of providers, each asked in turn, the first one that returns credentials wins, and later providers are never consulted. Nothing about this is magic — it is a plain for-loop, and knowing its order is what makes the surprising cases obvious. ## Roughly the order Exact ordering differs slightly between SDK languages and versions, so state it as a shape rather than a memorised list: 1. **Credentials supplied explicitly in code** when constructing the client. 2. **JVM system properties** (`aws.accessKeyId`, `aws.secretAccessKey`) — Java SDK only. 3. **Environment variables**: `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, and `AWS_SESSION_TOKEN` for temporary credentials. 4. **Web identity token file**: `AWS_WEB_IDENTITY_TOKEN_FILE` plus `AWS_ROLE_ARN`, which triggers `sts:AssumeRoleWithWebIdentity`. This is how EKS IRSA and GitHub Actions OIDC deliver a role. 5. **Shared config and credentials files** (`~/.aws/credentials`, `~/.aws/config`), with `AWS_PROFILE` choosing the profile. A profile here can itself say `role_arn` + `source_profile`, causing an assume-role hop. 6. **Container credentials**: the relative or full URI variable ECS and EKS Pod Identity inject, served by a local agent. 7. **EC2 instance metadata (IMDS)** — the instance profile role, consulted last. The principle behind it: the more explicitly a human stated the credential, the earlier it is tried; the ambient identity the platform attached is the fallback. ## Why the ordering produces the classic bug Because the platform-provided role sits at the bottom, *anything* left above it wins. Real cases: - A `Dockerfile` with `ENV AWS_ACCESS_KEY_ID=...` from a debugging session. The image runs on ECS with a perfectly good task role and never uses it — until the baked key is rotated and every task starts failing at once. - A CI runner that exports keys for one step and leaks them into the whole job's environment, so a later step that should use an OIDC-assumed role runs as the old user. - `AWS_PROFILE=prod` copied into a container's environment from someone's laptop; the file it names does not exist in the image, so the SDK errors out with a profile-not-found instead of falling through to the role. - A `~/.aws/credentials` file baked into a base image, quietly outranking the instance profile on every host that runs it. Each of these looks like an IAM problem — AccessDenied — and is actually a resolution problem. The permission you keep widening is on a role nobody is using. ## Diagnosing it in one command ```bash aws sts get-caller-identity # arn:aws:sts::111122223333:assumed-role/orders-task/abc123 <- the task role, good # arn:aws:iam::111122223333:user/legacy-deploy <- an IAM user, someone left keys around ``` Run it inside the failing container (on ECS, via `aws ecs execute-command`), not on your machine. `aws configure list` complements it by showing which *source* each config value came from (`env`, `shared-credentials-file`, `iam-role`), which pinpoints the shadowing provider. In application code, most SDKs let you log the resolved provider or construct a client with an explicit provider to test a hypothesis. ## Hygiene that prevents it - Never bake credentials or `AWS_PROFILE` into an image; strip them in the Dockerfile if a build stage needs them. - Treat the presence of `AWS_ACCESS_KEY_ID` in a production container as an alarmable condition. - Keep to the default chain in application code. Hard-coding a specific provider makes the binary non-portable across laptop, CI and production, which is the whole point of the chain. - Remember the chain is also *how* the good path works: IRSA, ECS task roles and instance profiles are simply providers late in the same list, which is why correctly configured workloads need zero credential code.
- Why does the SDK consult the instance metadata service last rather than first?Because the ambient platform identity is the fallback, not the override. An operator who explicitly exports credentials or selects a profile expects that to win — for a break-glass session, a cross-account script, or local testing against a different account. Putting the attached role first would make explicit configuration unusable. The cost of that design is exactly the shadowing bug, which is why leftover environment credentials are a production hazard.
- An SDK call inside an EKS pod using IRSA fails with 'unable to locate credentials'. What would you check?Whether the web identity provider ever ran: the pod needs `AWS_WEB_IDENTITY_TOKEN_FILE` and `AWS_ROLE_ARN` in its environment and the projected token file mounted. If they are absent, the pod's ServiceAccount is not the annotated one, or the identity webhook did not mutate the pod because it was created before the annotation. Also check the SDK version is new enough to support the web identity provider.
- Is it ever right to bypass the default chain and construct credentials explicitly in code?Rarely in a service. Legitimate cases are a job that must act as a specific named role and assume it deliberately, tests using a local stub endpoint, or a tool that takes credentials as arguments by design. Otherwise, pinning a provider makes the same binary behave differently across laptop, CI and production and defeats the portability the chain exists to provide.
saying these in an interview costs you the question
- Thinking the attached role always wins over environment variables
- Assuming AccessDenied always means a missing IAM permission
- Baking AWS_PROFILE or keys into a container image
- Believing every SDK language uses a completely different chain
- Debugging with get-caller-identity on the laptop instead of in the container