skip to content

A team bakes AMIs by launching an instance, configuring it by hand and calling CreateImage. What does AWS EC2 Image Builder change about that, and how is an Image Builder pipeline structured?

level: middleimportance: nice to knowfreq 28%

answer

  1. shell history is not a build process
  2. recipe, infrastructure, distribution, pipeline
  3. components are versioned units of work
  4. it tests before it distributes
  5. rebuild when the parent image changes

basics

~20 s

EC2 Image Builder turns hand-baking into a managed, repeatable pipeline: it launches a build instance from a parent image, applies versioned components, runs tests, then distributes the resulting AMI to chosen regions and accounts and tears everything down.

solid answer

~50 s

Hand-baking produces images nobody can reproduce or audit — the steps live in someone's shell history. EC2 Image Builder replaces that with a declared pipeline. An **image recipe** names a parent image plus an ordered list of **components** (versioned build and test steps in AWS's YAML document format) and a block device mapping. An **infrastructure configuration** says what instance type, subnet and instance profile the temporary build instance uses. A **distribution configuration** says which regions the finished AMI is copied to, which accounts or organizations get launch permission, and how it is named. The **pipeline** ties those together and can run on a schedule or on demand, optionally only when the parent image has changed. Image Builder also runs your test components against the built image before distributing, and it can output container images to ECR as well as AMIs. The practical win is a versioned, repeatable, auditable image with an automatic monthly rebuild for OS patches.

go deeper

for a junior

Know that an AMI can be produced by a pipeline instead of by hand, and be able to say why a repeatable build beats configuring an instance manually and snapshotting it.

for a middle

Be ready to name the pieces — recipe, components, infrastructure configuration, distribution configuration, pipeline — and describe the build, test, distribute, tear-down sequence in order.

for a senior

Show the operational angle: scheduled rebuilds when the parent image is patched, test components as a gate before distribution, build instances confined to a private subnet with a least-privilege instance profile, and logs retained for audit.

for a principal

Own image supply chain as a policy: who may publish golden images, how patch cadence is enforced across the fleet, how image provenance is recorded, and how consuming teams discover approved images without hardcoding IDs.

## The problem with hand-baking The manual loop is: launch an instance from a base AMI, SSH in, install and configure things, then call `CreateImage`. It works, and it is how most teams start. It fails at the second month, for reasons that are always the same: - **Not reproducible.** The steps exist in one person's terminal history. Rebuilding on next month's patched base gives a subtly different result. - **Not auditable.** Nothing records what is inside the image or why. - **Not tested.** The image is declared good because the person who built it thinks it is. - **Not distributed.** Copying to other regions and sharing with other accounts is a manual afterthought that drifts. - **The build instance leaks.** Someone forgets to terminate it. ## What Image Builder is EC2 Image Builder is a managed AWS service that runs that same loop for you, declaratively, and cleans up after itself. It is a **service**, not a tool you install: you describe the image and it orchestrates temporary EC2 instances to build and test it. ## The four pieces **Components.** The unit of work — a versioned YAML document in AWS's task-orchestration format, containing named phases with steps (run a script, download a file, reboot and continue, and so on). Components come in two kinds: **build** components that change the image, and **test** components that verify it. AWS publishes a library of managed components (installing the CloudWatch agent, applying updates, running hardening baselines), and you write your own for anything specific. Because components are versioned, an image build records exactly which version of each step produced it. **Image recipe.** Parent image + an ordered list of components + a block device mapping. This is the declarative answer to "what is in this image". Recipes are themselves versioned; you cannot mutate one in place, which is what makes builds reproducible. There is a parallel **container recipe** for producing container images into ECR. **Infrastructure configuration.** How the build runs: instance types to try, VPC subnet and security group, the IAM instance profile the build instance assumes, an SNS topic for notifications, whether to keep the instance on failure so you can debug it, and logging to S3. **Distribution configuration.** What happens to the finished artifact: which regions it is copied into, the AMI name and tags in each, which accounts or organizational units receive launch permission, and any license configurations to attach. **Pipeline.** The object that binds a recipe, an infrastructure configuration and a distribution configuration together, plus a schedule. The schedule can be cron-like, and can be conditioned on whether the parent image has changed — so a monthly "rebuild if the base got patched" pipeline is a configuration setting rather than a script. ## The lifecycle of one run 1. Pipeline triggers (schedule, API call, or an event). 2. Image Builder launches a temporary build instance from the parent image, inside your VPC, with the instance profile you specified. 3. Build components execute in recipe order. 4. The instance is snapshotted into an AMI. 5. A temporary **test** instance is launched from that new AMI and test components run against it. Failing tests stop the run before anything is distributed. 6. The AMI is copied and shared according to the distribution configuration. 7. All temporary instances are terminated. The test stage is the piece hand-baking almost never has, and it is the strongest argument for the service: a bad image is caught before any account can launch it. ## Where it fits, and where it does not Image Builder answers "how is this image produced". It does not answer "how do consumers find the newest one" — you still want the pipeline to publish the resulting AMI ID somewhere consumers resolve, such as an SSM Parameter Store parameter, rather than having teams hardcode IDs. And it does not decide the more important question of what belongs in the image at all versus what an instance does at boot. It is also not the only option — organisations with an existing image-build toolchain often keep it. The interview-relevant point is not brand loyalty; it is that images should be **built by a pipeline, versioned, tested and distributed automatically**, and Image Builder is AWS's managed way to get there without running build infrastructure yourself.

  • What does the test stage of an Image Builder pipeline give you that hand-baking does not?
    It launches a throwaway instance from the freshly built AMI and runs test components against it before anything is distributed. A failed test aborts the run, so a broken image never reaches the accounts that would launch it. Hand-baking has no equivalent gate — the image is distributed the moment it is created, and the first test is production.
  • How should consumers discover the AMI a pipeline just produced?
    Not by hardcoding the ID. Have the pipeline publish the new ID to an SSM Parameter Store parameter per region, or tag images consistently so consumers resolve the newest with a `DescribeImages` owner-and-name filter. Consumers then pin that resolved value for a release rather than resolving on every launch, which keeps deployments reproducible while still region-aware.
  • Does using Image Builder settle how much configuration belongs in the image versus in first-boot user data?
    No — it only makes whatever you decide repeatable. The split is still a judgment call: slow, static, identical-everywhere things belong in the image, while environment-specific configuration, secrets and cluster-join details must stay late-bound. Image Builder makes the baked half auditable; it does not tell you where to draw the line.

saying these in an interview costs you the question

  • Describing hand-baked AMIs as reproducible because the base AMI is fixed
  • Thinking Image Builder replaces the need to version anything
  • Believing the build instance persists and must be managed by you
  • Assuming pipeline output is automatically discoverable by consumers
  • Skipping test components because the build succeeded

context