After you set an EC2 instance's metadata options to require IMDSv2, applications in Docker containers on that host can no longer read instance role credentials from 169.254.169.254, although the same request works on the host itself. What is the cause, and what is the fix?
answer
- not a rate limit, not a lifetime
- it is a packet TTL
- the bridge counts as a hop
- only the token response is affected
- fix the launch template too
basics
~20 sThe IMDSv2 token response is sent with an IP TTL equal to the instance's HttpPutResponseHopLimit, which defaults to 1. Routing from a bridged container to the host consumes that hop, so the token never arrives. Raise the hop limit to 2.
solid answer
~50 sIMDSv2 protects the token from being relayed off the instance by emitting the token response with an IP time-to-live equal to `HttpPutResponseHopLimit`, whose default is 1. A process on the host is one hop from the metadata service and gets the token fine. A container on a bridged Docker network sits behind the host's routing, so the response is forwarded once, the TTL decrements to zero, and the packet is dropped — the container sees the PUT hang or fail while the same request on the host still works. The direct fix is to raise the hop limit to 2 with `modify-instance-metadata-options`, and to set the same value in `MetadataOptions` on the launch template so replacement instances inherit it. Raise it to 2, not higher: every extra hop is another network element that could relay a token. The better long-term fix is to stop containers using the host's role at all and give each workload its own credentials.
code
bash · 4 linesaws ec2 modify-instance-metadata-options \
--instance-id i-0123456789abcdef0 \
--http-tokens required \
--http-put-response-hop-limit 2go deeper
Know that instance metadata is reached at a link-local address and that containers do not automatically inherit the host's access to it once IMDSv2 is enforced.
Explain that HttpPutResponseHopLimit is the IP TTL on the token response, that its default of 1 dies at the container bridge, and that raising it to 2 restores access.
Diagnose from the asymmetry — host works, container does not, only the token step fails, and it started at the hardening change — then fix both the running instance and the launch template, and argue for 2 rather than a comfortable large number.
Take the position that shared host credentials are the real defect: containers should carry their own scoped identity so the instance role stays trivial and the hop limit never has to be relaxed at all.
## The symptom You harden a host by setting `HttpTokens=required`. Immediately, applications running in containers on that host start failing to obtain AWS credentials: SDK calls report that no credential provider succeeded, and a manual `curl -X PUT http://169.254.169.254/latest/api/token` from inside the container times out or returns nothing. The same command run in an SSH session on the host returns a token instantly. Nothing about IAM changed, and the instance profile is still attached. ## Why the container is different The missing piece is `HttpPutResponseHopLimit`. It is not a rate limit and not a token lifetime — it is literally the **IP time-to-live written into the token response packet**. TTL is decremented by each router the packet crosses; at zero it is discarded. That is a deliberate defence. If something outside the instance ever managed to have a token request made on its behalf, the response would have to be routed back out of the instance to reach it — and with a TTL of 1 it dies on the first forwarding step. The token cannot leave the box. A container on Docker's default bridge network is on the far side of exactly such a forwarding step. Its packets are routed and NAT-ed by the host's network stack, so from the metadata service's point of view the container is one hop further away than a process on the host. A TTL of 1 gets the response as far as the host and no further. A container using host networking has no extra hop, which is why the failure looks selective and inconsistent across workloads on the same machine. ## Why only the token step breaks The hop limit is applied to the PUT response, which is why the failure appears exactly when you enforce IMDSv2 and not before. While `HttpTokens` was `optional`, containers were reading metadata with plain IMDSv1 GETs, which never involve the token response — so a fleet can run for years with a hop limit of 1 and expose the problem only on the day someone hardens it. That is the shape of the incident: the change looks like an IAM or a permissions problem, and it is a TTL. ## The fix ```bash aws ec2 modify-instance-metadata-options \ --instance-id i-0123456789abcdef0 \ --http-tokens required \ --http-put-response-hop-limit 2 ``` The attribute can be changed on a running instance and takes effect without a restart. Do the same in the `MetadataOptions` block of the launch template, or the next instance the Auto Scaling group creates arrives back at 1 and the incident repeats — which is the second half of the answer an interviewer is listening for. ## Choose 2, not 64 The valid range runs to 64, and it is tempting to set a comfortable number. Do not. Each additional hop the token response may cross is another network element through which a token could be relayed, which erodes precisely the property that made IMDSv2 worth enforcing. Two is the value that makes one layer of container networking work and nothing more. If a workload needs three, that is a signal about its network topology worth understanding rather than accommodating. ## The better fix Raising the hop limit makes the host's instance role reachable from every container on the host — so every container shares the same credentials and the same blast radius. The stronger design is that containers should not be borrowing the host's identity at all: give each workload its own credentials, so that the container's permissions are scoped to what it does, and the host role can shrink to almost nothing. Container orchestrators on AWS all provide a per-workload credential mechanism for exactly this reason. Where that is available, the correct end state is a hop limit of 1, containers that never touch the instance metadata service, and an instance role narrow enough that reaching it would not be interesting. ## Diagnosing it quickly The tell is the asymmetry: metadata works from the host and not from a container, only the PUT fails, and it started the moment `HttpTokens` became `required`. Compare `curl -v -X PUT` from both sides, check `aws ec2 describe-instances` for the instance's `MetadataOptions`, and confirm the container's network mode. If the container is on host networking and still fails, the hop limit is not your problem and you should be looking at whether the container can reach link-local addresses at all.
- Why is raising the hop limit to a large value a bad idea?The hop limit is the TTL on the token response, and its whole purpose is to keep a token from being routed off the instance. Every extra hop you permit is another network element that could relay one to a remote caller. Two covers one layer of container networking; anything higher trades away the property you enforced IMDSv2 to gain.
- You fix the hop limit on the running instance and the problem returns a week later on new instances. What did you miss?The launch template. `MetadataOptions` set with `modify-instance-metadata-options` applies to that one instance; every instance the Auto Scaling group launches is built from the template and comes back with the default hop limit of 1. The durable fix is to set the metadata options in the launch template — and, fleet-wide, in the account-level instance metadata defaults.
saying these in an interview costs you the question
- Reads the hop limit as the token's time-to-live
- Suggests reverting HttpTokens to optional as the fix
- Sets the hop limit to 64 to be safe
- Assumes host-networked containers hit the same problem
- Fixes the running instance but not the launch template