skip to content

On AWS, how do you decide what gets baked into an EC2 AMI versus what an instance does for itself at first boot through user data?

level: principalimportance: should knowfreq 45%

answer

  1. it is a spectrum, not a binary
  2. slow and static bakes
  3. variable and secret late-binds
  4. scale-out latency is the hard constraint
  5. boot-time dependencies fail when you scale

basics

~20 s

Bake what is slow, static and identical everywhere — OS packages, agents, runtimes, usually the application artifact. Late-bind what varies by environment or must stay fresh — configuration, secrets, cluster identity. The deciding factors are boot latency, reproducibility and how fast you must patch.

solid answer

~50 s

Treat it as a spectrum with two failure modes at the ends. A pure bootstrap image boots fast to build but slowly at runtime, and every scale-out event depends on external package repositories being up and returning the same bytes as last week — which they will not. A fully baked image scales out in seconds and is reproducible, but every change, including a one-line config tweak, becomes an image build and a fleet roll. My rule: bake anything slow, static and identical across environments — the OS, patches, the runtime, monitoring and logging agents, and normally the application artifact itself. Late-bind anything that differs per environment or per instance — configuration pulled from Parameter Store, secrets fetched with the instance's own IAM role, cluster or discovery details. Then let the constraints decide the balance: how fast an Auto Scaling group must add capacity, how quickly you must ship a CVE patch, and whether you can prove which image is running where.

go deeper

for a junior

Know the two options exist — a prepared image versus a first-boot script — and be able to say that installing software at every launch makes instances slower to start.

for a middle

Explain the concrete tradeoffs: boot time, dependence on external package repositories, version drift between instances, and the fact that secrets must never be baked into an image.

for a senior

Argue from operational evidence — measured boot-to-healthy time under scale-out, package mirror outages during incidents, drift between instances in one group — and describe a bootstrap that is idempotent, small and fails loudly through the health check.

for a principal

Own it as a standard: state the rule teams follow, tie it to patch cadence and audit requirements, fund the image pipeline that makes baking viable, and be able to answer at any moment which image version is running in production.

## Two ends of one spectrum Every EC2 instance comes up as *image + first-boot work*. The only question is where the line sits. **Pure bootstrap** (thin base image, everything installed by user data). Attractive because the image never changes and a fix is just a script edit. In production it fails in specific ways: - **Slow scale-out.** Installing packages, pulling a runtime and downloading an artifact takes minutes. Under a traffic spike, capacity arrives after the incident. - **Boot-time external dependencies.** Every launch now depends on a package mirror, a language registry and an artifact store all being available and consistent. An outage in any of them means you cannot add capacity — precisely when you most need to. - **Non-reproducibility.** `dnf install -y nginx` gives a different version this month than last. Two instances in the same Auto Scaling group can differ, which produces bugs that reproduce on one host and not another. - **No atomic rollback.** There is no artifact to roll back to. **Fully baked** (golden image containing everything, user data does nothing). Attractive because boot is fast and deterministic. Its costs: - Every change is an image build, so the change cycle is measured in tens of minutes rather than seconds. - Image sprawl and snapshot storage cost, plus lifecycle policy to manage. - Secrets must never be baked in, so *something* must still happen at boot. - Rolling the fleet for a config change is disproportionate. ## The rule I actually apply **Bake it if it is slow, static and identical everywhere.** OS and patches, the language runtime, system daemons, monitoring and logging agents, and normally the application artifact — a versioned image containing a versioned app is the cleanest deployable unit on EC2. **Late-bind it if it varies or must be fresh.** Environment configuration read at boot from SSM Parameter Store, secrets fetched at runtime with the instance's IAM role, anything derived from the instance's own identity or placement, and anything that must reflect state at launch time rather than build time. **Never bake:** credentials, private keys, environment-specific hostnames, or anything that makes the image usable in only one account. ## The constraints that move the line **Scale-out latency.** If an Auto Scaling group must add capacity in under two minutes, a bootstrap that takes three is not a preference problem, it is a capacity problem. Measure boot-to-healthy, not boot-to-running. **Patch cadence.** Baking centralises patching in the image pipeline: rebuild, test, roll. That is *slower to start* but *auditable and uniform*. Bootstrapping patches whatever the repo has at launch — faster to pick up a fix, impossible to state what is actually deployed. Regulated environments almost always choose the baked, auditable side. **Blast radius of a change.** Baking forces every change through a build-and-roll cycle. That is friction, and friction is the point: it makes changes reviewable and rollbacks atomic. **Fleet size and churn.** With Spot instances or aggressive scaling, launches are constant, and every minute of bootstrap multiplies into real money and real risk. **Team capability.** A baked strategy needs an image pipeline, a distribution mechanism and image lifecycle management. Without those, an organisation baking by hand ends up with unreproducible images and the worst of both worlds. ## Whether the app artifact is in the image This is the sub-decision people argue about. Baking the app gives one immutable artifact per release, trivially rolled back by pointing at the previous AMI. It costs an image build per release, so it demands a fast pipeline. Pulling the artifact at boot from S3 keeps the image stable across releases and shortens the release cycle, at the price of a boot-time dependency and an instance whose contents are not knowable from its AMI ID alone. Both are defensible; what is not defensible is not having decided, so that some services do one and some the other with no stated rule. ## Bootstrap that survives contact with production Whatever stays at boot must be: - **Idempotent**, because retries and re-runs happen; - **Small**, because 16 KB of user data is a hook, not a deployment system; - **Loud on failure** — nothing about a failed user data script stops the instance reaching `running` or being registered with a target group. Have the bootstrap write a health marker only on success, and have the load balancer health check depend on it, so a failed bootstrap fails the instance instead of serving errors; - **Observable** — ship `/var/log/cloud-init-output.log` to CloudWatch Logs so failures are diagnosable after the instance is replaced. ## The answer that lands Do not pick an end of the spectrum. State the rule (slow-and-static bakes, variable-and-fresh late-binds), then name the two or three constraints that would move your line — scale-out latency, patch auditability, release cadence — and say how you would verify it, by measuring boot-to-healthy and by being able to answer "which image is running in production right now".

  • What is the strongest single argument against installing packages in user data at every launch?
    It makes scale-out depend on third-party availability and on repository contents being unchanged. A package mirror outage means you cannot add capacity during the exact incident that demanded it, and a silently updated package means two instances in one Auto Scaling group differ. Bake the dependency and the launch becomes a local, deterministic operation.
  • If nearly everything is baked, why keep user data at all?
    Because some things must not be in an image: secrets, and anything that varies by environment or by the instance itself. A minimal bootstrap fetches configuration from Parameter Store, retrieves secrets with the instance's IAM role, applies instance-specific settings and then writes a health marker. The image stays portable across accounts and environments; the boot script supplies identity.
  • How do you keep a failed bootstrap from putting a broken instance into service?
    Make success explicit. The bootstrap writes a marker file or starts a health endpoint only on its final successful step, and the target group's health check tests that endpoint. A failed script then leaves the instance unhealthy, so it is never registered or is quickly replaced, instead of passing EC2 status checks and serving errors.
  • How does baking change the way you patch a fleet?
    Patching becomes a pipeline event rather than a per-instance one: rebuild the image on the patched parent, run tests, distribute, then roll the fleet by replacing instances. It is slower to start than running an update command everywhere, but it is uniform, auditable and rollback-able — you can state exactly which image every instance is running.

saying these in an interview costs you the question

  • Treating it as a binary choice with one correct answer
  • Baking secrets or environment-specific hostnames into the image
  • Ignoring scale-out latency when defending a heavy bootstrap script
  • Assuming a failed user data script keeps the instance out of service
  • Claiming bootstrapping is reproducible because the script is in version control

context