A Terraform configuration selects an EC2 image with `data "aws_ami"` and `most_recent = true`. Why is that a hazard in production, and what would you do instead?
answer
- read again on every plan
- the newest image is a moving target
- `ami` forces replacement
- same commit, different infrastructure
- pin it as a reviewed input
basics
~20 sBecause the data source is re-read on every plan, most_recent = true resolves to whatever image the vendor published last, so an unrelated apply can silently replace running instances. Pin the image id as a reviewed input instead of resolving it at plan time.
solid answer
~50 s`most_recent = true` with a wildcard name filter means the AMI id is whatever the publisher pushed most recently, and Terraform re-reads the data source on every plan. So the value can change with no commit in your repository. Since `ami` on `aws_instance` forces replacement, the next apply — one triggered by an unrelated change, or by CI on a schedule — plans to destroy and recreate live instances. It also breaks reproducibility: the same commit applied twice can produce different infrastructure, and reverting the code does not revert the image. What I do instead is resolve the image in a promotion step and pass the id in as an input variable, so an image change is a reviewed commit like any other. If a lookup must stay, I pin `owners` so a lookalike community image cannot match, and narrow the filter to an exact version string rather than a wildcard.
code
hcl · 9 linesvariable "ami_id" {
type = string
description = "Pinned base image, bumped by a reviewed commit"
}
resource "aws_instance" "app" {
ami = var.ami_id
instance_type = "t3.small"
}go deeper
Know that most_recent = true returns whichever matching image is newest, that Terraform re-reads it every plan, and that changing an instance's ami replaces the instance rather than updating it.
Explain the chain: re-read on plan, vendor publishes, id changes, ami forces replacement — with no commit in your repository to explain the diff. Name owners as the constraint that stops an untrusted image matching.
Frame it as reproducibility and reviewability: the same commit must produce the same infrastructure. Argue for resolving the image in a promotion step and passing a pinned id in, and explain why ignoring changes to ami is a stopgap that fragments the fleet.
Own the rule for the estate: which inputs may be resolved at plan time and which must be pinned artifacts. Unpinned time-dependent lookups feeding replacement-forcing arguments are a standing outage risk, and the policy should be enforced in review or by a scanner, not left to each team.
## The mechanism ```hcl data "aws_ami" "base" { most_recent = true owners = ["amazon"] filter { name = "name" values = ["al2023-ami-*-x86_64"] } } resource "aws_instance" "app" { ami = data.aws_ami.base.id instance_type = "t3.small" } ``` Three facts combine into the hazard: 1. Terraform reads data sources on **every plan**, not once. There is no pinning, no cache across runs that would keep yesterday's answer. 2. `most_recent = true` tells the provider to sort the matching images by creation date and return the newest. The set of matches changes whenever the publisher releases a new image — typically every few weeks for a maintained distribution. 3. `ami` on `aws_instance` is a **force-replacement** argument. Changing it does not update the instance; it destroys and recreates it. The result: an engineer opens a pull request to change an S3 bucket tag, CI runs `terraform plan`, and the plan destroys and recreates four EC2 instances because a patched image shipped overnight. Nothing in the diff of the pull request explains it. The change was introduced by the outside world. ## Why this is worse than it first looks **It is unreviewable.** The commit that causes the replacement is not the commit that contains it. Code review cannot catch an input that lives in a vendor's account. **It is not reproducible.** Applying the same git SHA in two environments a week apart yields different images, so staging is not a rehearsal for production. The same property makes a rollback ineffective: reverting the Terraform code does not bring the previous image back, because the code never named it. **The blast radius depends on timing.** The change may land during an emergency apply, when someone is targeting an unrelated fix and does not read the whole plan. This is the same failure mode as any implicitly-versioned dependency — it fires at the worst moment because that is when people read plans least carefully. **Without `owners`, it is a security problem too.** Anyone can publish a public AMI, and image names are not reserved. A filter that matches on name alone can, in principle, select an image you did not intend. Always constrain `owners` to the account ids you trust — the distribution vendor's, or your own build account. ## What to do instead **Preferred: resolve once, pass in as an input.** A build or promotion pipeline picks the image (often by querying the same API), writes the id into a tfvars file or a parameter store entry, and opens a pull request. Now the image is a reviewed change with its own commit, its own approval, and its own revert. The Terraform configuration takes `var.ami_id` and is fully deterministic. **Or: make the lookup exact.** If you keep a data source, filter on a fully-qualified image name that includes a version, so the answer is stable until you edit the string. `most_recent = true` is then only a tie-breaker among identical matches, not a policy of "newest wins". **Or: move the rollout somewhere designed for it.** For fleets, put the image in a launch template consumed by an autoscaling group. Changing the image then produces a new launch template version and an instance-refresh you control, rather than a Terraform-driven destroy-and-recreate of named instances. How that rollout is sequenced is a deployment-strategy question; the Terraform-side point is that the image id stops being an argument that forces replacement of long-lived instances. **Partial: freeze existing instances.** A `lifecycle` block ignoring changes to `ami` stops Terraform from acting on the drift, but it is a plaster: existing instances keep the old image forever while any newly created instance gets whatever the lookup returns today, so the fleet quietly becomes heterogeneous. Reach for it only as a stopgap while you move to a pinned input. ## The general lesson The question is really about where non-determinism enters a supposedly declarative system. A data source is an input from outside your repository, and `most_recent` makes that input *time-dependent*. Everything you like about infrastructure as code — review, reproducibility, revert — assumes the inputs are pinned. Every unpinned lookup that feeds a replacement-forcing argument is a scheduled outage waiting for a publisher's release cadence.
- The team insists on keeping the lookup. What is the minimum you would require?Constrain `owners` to trusted account ids so a lookalike public image cannot match, and replace the wildcard with a name filter containing the exact version, so the result only changes when someone edits the string. Then the data source is deterministic per commit and `most_recent` is only breaking ties.
- How would you detect that this pattern already exists in a repository you inherited?Grep for `most_recent` across the configuration and check which arguments consume the result — the dangerous case is a replacement-forcing argument like `ami`. A CI job that plans on a schedule against unchanged code also surfaces it: a plan that proposes changes with no commit behind it is exactly this class of drift.
- Is the same risk present with other data sources?Yes, wherever a lookup can return a different answer over time and feeds a replacement-forcing argument — a "latest" version, an availability zone list that gains a zone, a tag-based lookup that matches a newly created object. The hazard is unpinned time-dependent input, not the AMI data source specifically.
It is the difference between depending on a library version 1.4.2 and depending on latest: the build is fine until the day it is not, and nothing in your repository records what changed.
saying these in an interview costs you the question
- Says Terraform caches the AMI id after the first apply
- Thinks changing `ami` updates the instance in place
- Believes reverting the Terraform code restores the previous image
- Treats `most_recent` as good practice for staying patched
- Omits `owners` and relies on the name filter alone