skip to content

EC2

EC2 is the plain virtual machine underneath much of AWS: instance types, AMIs, EBS-backed volumes, security groups, and auto scaling. You are expected to size an instance and describe how a fleet grows and heals, since EC2 is the baseline every managed compute option gets compared against.

part ofAWSoverview, primer and where to startread it →
on this pageshow

explore

questions

29

What is EC2 user data, at what point in an instance's life does the script run, and as which OS user?

level: juniorimportance: must knowfreq 65%

answer

  1. a text blob, delivered through metadata
  2. cloud-init is what reads it
  3. the first line decides everything
  4. root, not ec2-user
  5. once per instance, not per boot

basics

~20 s

User data is a text blob you attach to an EC2 instance that cloud-init fetches from the instance metadata service at first boot. A script starting with a shebang runs as root, once per instance — not on every reboot.

solid answer

~50 s

User data is a small text payload attached to an EC2 instance — up to 16 KB before base64 encoding. On Linux, cloud-init retrieves it from the instance metadata service at `http://169.254.169.254/latest/user-data` during boot. If the payload starts with a shebang such as `#!/bin/bash`, cloud-init writes it to disk and executes it **as root**, in a non-interactive shell with the working directory at `/` — not as `ec2-user`, and not with a login shell's PATH or environment. By default it runs **once per instance**: cloud-init records the instance ID it already handled, so a reboot, or a stop and start, will not run it again. Output lands in `/var/log/cloud-init-output.log`, which is where you look when the instance comes up but the app is missing. If the payload starts with `#cloud-config` instead, cloud-init interprets it as declarative YAML rather than executing it.

code

bash · 5 lines
bash
#!/bin/bash
set -euxo pipefail
dnf install -y nginx
echo "ok" > /usr/share/nginx/html/health
systemctl enable --now nginx

go deeper

for a junior

Be ready to say what user data is, that the payload needs a shebang, that it executes as root, and that it runs once when the instance is first created rather than on every reboot.

for a middle

Explain the mechanics: cloud-init pulls it from the metadata service, dispatches on the first line, records per-instance state under /var/lib/cloud, and logs output to /var/log/cloud-init-output.log.

for a senior

Show operational judgment — bootstrap failures do not fail the instance, so pair user data with a health signal; keep the script small and idempotent; keep secrets out of it and fetch them at runtime instead.

for a principal

Own the boundary decision: how much belongs in a first-boot script at all versus baked into the image or managed by an agent, and what standard the organisation holds so a failed bootstrap can never quietly serve traffic.

## What user data is When you launch an EC2 instance you can attach a block of text called **user data**. AWS does not interpret it; it merely makes it available to the instance through the **instance metadata service (IMDS)**, an HTTP endpoint reachable from inside the instance at the link-local address `169.254.169.254`, at path `/latest/user-data`. The limit is 16 KB of raw text before base64 encoding (as of 2025), which is why user data is a bootstrap hook, not a delivery mechanism for application artifacts. On most Linux AMIs, including Amazon Linux, Ubuntu and RHEL, the program that reads it is **cloud-init** — a distribution-level boot agent, not an AWS product. AWS supplies the data; cloud-init decides what to do with it. ## How cloud-init decides what the payload is cloud-init dispatches on the first line: - `#!/bin/bash` (or any shebang) → treated as a **script**. cloud-init writes it under `/var/lib/cloud/instance/scripts/` and executes it. - `#cloud-config` → treated as **declarative YAML**: it can create users, write files, install packages, add repositories and so on, without you scripting the steps. - A MIME multipart document → several parts of mixed types in one payload, which is how you combine a `#cloud-config` section with a shell script. A payload with **no** recognised first line is silently ignored. Forgetting the shebang is the single most common reason a user data script "does nothing". ```bash #!/bin/bash dnf install -y nginx systemctl enable --now nginx ``` ## Who runs it, and in what environment The script runs as **root**. There is no `sudo` needed and, conversely, no `ec2-user` home directory, no `~/.bashrc`, no interactive login PATH, and the working directory is `/`. Scripts written and tested in an SSH session frequently break here because they assumed a login shell's environment or a relative path. It also runs **late in boot but before the instance is useful** — after networking is up, which is why package installs work, but potentially before other things you might be counting on. Nothing waits for your script: the instance reaches the `running` state and passes status checks regardless of whether your bootstrap succeeded or failed. ## Once per instance, not once per boot By default the user-scripts stage has **per-instance** frequency. cloud-init keeps state under `/var/lib/cloud/instances/<instance-id>/` and remembers that this instance has already been handled. Consequences: - **Reboot**: does not re-run. - **Stop and start**: same instance ID, so it does **not** re-run either — this surprises people, because the instance did shut down completely. - **New instance from the same AMI and same user data**: runs, because the instance ID is new. This makes user data a *bootstrap* mechanism, not a configuration loop. If you genuinely need something on every boot, either make it a systemd unit installed by the bootstrap, or opt into per-boot frequency in `#cloud-config`: ```yaml #cloud-config cloud_final_modules: - [scripts-user, always] ``` ## Changing user data later On an EBS-backed instance you can change user data with `ModifyInstanceAttribute`, but the instance must be **stopped** first. And because of the once-per-instance rule, changing it and starting the instance again still will not execute the new script — a fact that produces a lot of confused debugging. ## Where to look when it fails - `/var/log/cloud-init-output.log` — the combined stdout and stderr of your script. This is the first place to look. - `/var/log/cloud-init.log` — cloud-init's own module-by-module log, useful when your script never ran at all. ## User data is not a secret Anything that can reach IMDS from inside the instance can read the user data back, and it is visible through the EC2 API to anyone with `ec2:DescribeInstanceAttribute`. Never put passwords, API keys or private keys in it. The correct pattern is to put a **reference** in user data and have the bootstrap fetch the actual secret at runtime using the instance's own IAM credentials.

  • You edit an instance's user data, start it again, and nothing happens. Why?
    The user-scripts stage runs at per-instance frequency, and stopping and starting an instance keeps the same instance ID, so cloud-init sees the work as already done. Either launch a fresh instance, opt into per-boot execution with `cloud_final_modules: [[scripts-user, always]]`, or clear cloud-init's recorded state before rebooting. Editing the text alone changes nothing.
  • Is it acceptable to put a database password in user data so the app can start?
    No. User data is readable by any process on the instance that can reach the metadata service, and by any principal holding `ec2:DescribeInstanceAttribute` on it, and it is often captured in launch configuration and logs. Put an identifier in user data instead — a parameter name or secret ARN — and let the bootstrap fetch the value at runtime with the instance's own IAM credentials.
  • What is the difference between a user data payload starting with #!/bin/bash and one starting with #cloud-config?
    The shebang tells cloud-init to execute the payload as a shell script as root. `#cloud-config` tells it to parse the payload as declarative YAML and hand it to cloud-init modules that create users, write files, install packages or add repositories. A payload with neither marker is recognised by nothing and is quietly discarded.

saying these in an interview costs you the question

  • Saying user data runs on every boot or every reboot
  • Claiming the script runs as ec2-user and needs sudo
  • Omitting the shebang and expecting the script to run anyway
  • Putting credentials in user data because "only the instance can read it"
  • Assuming the instance stays out of service until the script finishes

context

open as a page

In an EC2 Auto Scaling group, what do the minimum, maximum and desired capacity settings each control, and what happens if you manually terminate one of the group's instances?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Desired capacity is the instance count the group tries to run right now; minimum and maximum are the floor and ceiling that scaling can never cross. Manually terminating an instance leaves the group below desired, so it launches a replacement.

open as a page

Decode the EC2 instance type name m7g.2xlarge: what does each part of the name tell you, and how many vCPUs and how much memory does that instance have?

level: juniorimportance: must knowfreq 75%

basics

~20 s

m7g.2xlarge splits into four parts: family m (general purpose, roughly 4 GiB of memory per vCPU), generation 7, processor letter g (AWS Graviton, arm64), and size 2xlarge — 8 vCPUs and 32 GiB of memory.

open as a page

In EC2, what do you get and what do you give up by running an instance as Spot instead of On-Demand, and how is the Spot price actually determined?

level: juniorimportance: must knowfreq 78%

basics

~20 s

Spot Instances run on EC2 capacity AWS currently has idle, at a steep discount off On-Demand, but AWS can reclaim them at any moment with a two-minute warning. Spot prices move gradually with supply and demand; there is no bidding auction.

open as a page

A deployment script that launches EC2 instances hardcodes an AMI ID. It works in us-east-1 and fails in eu-west-1 with InvalidAMIID.NotFound. Why, and what is an AMI actually made of?

level: middleimportance: must knowfreq 55%

basics

~20 s

AMI IDs are scoped to a single AWS region, so an ID registered in us-east-1 does not exist in eu-west-1. Copy the image with CopyImage, which mints a new ID, or look the ID up per region instead of hardcoding it.

open as a page

EC2 Auto Scaling supports target tracking, step, simple, scheduled and predictive scaling policies. How would you choose between them for a web tier, and can they be combined?

level: middleimportance: must knowfreq 70%

basics

~20 s

Target tracking is the default choice: name a metric and the value to hold, and Auto Scaling works out the capacity. Step scaling handles graded reactions to severity, scheduled actions cover known clock-driven load, and predictive scaling pre-scales repeating daily patterns. They compose.

open as a page

An EC2 t3 instance running a web service performs well for a while and then its CPU flatlines at a low fixed percentage while requests queue up. Explain the burstable T-instance CPU credit mechanism behind this, and what unlimited mode changes.

level: middleimportance: must knowfreq 62%

basics

~20 s

Burstable T-family instances earn CPU credits at a fixed hourly rate and spend them to run above a size-specific baseline. In standard mode, an exhausted balance throttles the instance to that baseline; unlimited mode keeps bursting and bills surplus credits instead.

open as a page

On an EC2 instance, how does IMDSv2 differ from IMDSv1 at the request level, and what does setting the instance's metadata option HttpTokens to `required` actually enforce?

level: middleimportance: must knowfreq 58%

basics

~20 s

IMDSv1 answers a plain GET to 169.254.169.254. IMDSv2 requires a PUT to /latest/api/token first, and every later GET must carry that token in an X-aws-ec2-metadata-token header. Setting HttpTokens to required makes the instance reject token-less requests.

open as a page

Compare stopping, hibernating and terminating an EC2 instance: what happens to the instance's RAM contents, its EBS root volume, and any instance-store data in each case?

level: middleimportance: must knowfreq 66%

basics

~20 s

Stopping discards RAM and instance-store data and keeps the EBS root volume. Hibernating first writes RAM to the encrypted root volume and restores it on start. Terminating is permanent and deletes volumes whose DeleteOnTermination flag is true.

open as a page

An EC2 Spot Instance is about to be reclaimed by AWS. What signals do you get beforehand, where does a process running on the instance read them, and what should that process do in response?

level: middleimportance: must knowfreq 68%

basics

~20 s

AWS publishes two signals: an optional earlier rebalance recommendation, and a two-minute Spot interruption notice. Both appear in instance metadata and as EventBridge events. The worker should stop taking new work, checkpoint or requeue in-flight work, and drain from its load balancer.

open as a page

You stop an EC2 instance and start it again a few minutes later. What changes about how clients reach it, and what stays the same?

level: juniorimportance: should knowfreq 62%

basics

~20 s

A stop/start moves the instance onto different host hardware. The auto-assigned public IPv4 address and public DNS name are released and replaced; the instance ID, private IPv4 address and EBS root volume survive. An Elastic IP stays associated.

open as a page

An EC2 Auto Scaling group references a launch template by version. What is the practical difference between pinning it to $Latest, $Default, or a specific version number?

level: middleimportance: should knowfreq 46%

basics

~20 s

Launch template versions are immutable, so changes create new versions. $Latest means every new instance uses the newest version the moment it exists; $Default uses whichever version you explicitly promoted; a pinned number is fully deterministic. Running instances are never altered.

open as a page

EC2 offers three placement group strategies — cluster, spread and partition. What does each one do to where instances physically land, and which kind of workload does each suit?

level: middleimportance: should knowfreq 40%

basics

~20 s

Cluster packs instances close together in one Availability Zone for the lowest network latency. Spread puts each instance on distinct underlying hardware to avoid correlated failure. Partition groups instances into blocks that share no hardware between blocks, for replicated distributed systems.

open as a page

An engineer terminated an EC2 instance and the separate EBS data volume attached to it was deleted along with it. Which EC2 settings decide that outcome, and how would you stop it happening again?

level: middleimportance: should knowfreq 48%

basics

~20 s

Every entry in an instance's block device mappings carries a DeleteOnTermination flag; when it is true, terminating the instance deletes that volume. Set it to false on data volumes, enable termination protection, and keep EBS snapshots as the real safety net.

open as a page

A team needs certainty that twenty EC2 instances of a specific type will be available in one Availability Zone for a launch next month. Which EC2 purchase mechanism actually holds that capacity, and which ones only change the bill?

level: middleimportance: should knowfreq 42%

basics

~20 s

Only an On-Demand Capacity Reservation actually holds EC2 capacity in a specific Availability Zone. Savings Plans and regional Reserved Instances are billing constructs with no capacity guarantee; a zonal Reserved Instance is the one commitment that also reserves capacity.

open as a page

You share a custom EC2 AMI with a second AWS account. They can see the image, but every launch from it fails. What are the usual causes, and how do you share an AMI so it actually launches?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Launch permission on the AMI is not enough when its snapshots are encrypted: the consuming account also needs access to the KMS key. Images encrypted with the AWS-managed EBS key cannot be shared at all — re-encrypt under a customer-managed key first.

open as a page

Instances in an EC2 Auto Scaling group pass their EC2 status checks, but the application on them returns errors and the load balancer has marked them unhealthy. Why has Auto Scaling not replaced them, and what would you change?

level: seniorimportance: should knowfreq 56%

basics

~20 s

The group's health check type is EC2, which only reflects hypervisor and instance status checks — a running instance with a dead application passes. Setting the health check type to ELB makes the group consume the load balancer's verdict and replace failing instances.

open as a page

A scale-in event on an EC2 Auto Scaling group terminates workers that are still processing long-running jobs. How do Auto Scaling lifecycle hooks let you drain that work before the instance goes away?

level: seniorimportance: should knowfreq 45%

basics

~20 s

A terminating lifecycle hook pauses the instance in a Terminating:Wait state and emits an event, giving your drain logic a window to finish or hand off work. You end the wait by calling CompleteLifecycleAction, or extend it with a heartbeat.

open as a page

Your team wants to move an EC2 fleet from x86 instances to AWS Graviton types such as m7g. What actually has to change for the workload to run, and how would you validate the move before committing?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Graviton instances are arm64, so everything compiled has to change: an arm64 AMI, arm64 runtimes, recompiled native code and multi-architecture container images. Interpreted code usually moves untouched; native dependencies and third-party agents are where migrations stall.

open as a page

An EC2 instance shows "1/2 checks passed" in the console. Explain the difference between the system status check and the instance status check, and how your response differs depending on which one failed.

level: seniorimportance: should knowfreq 44%

basics

~20 s

The system status check covers AWS-side host hardware, power and network; the instance status check covers the guest operating system. A failed system check is resolved by moving hosts through stop/start or auto-recovery; a failed instance check needs a reboot or a fix inside the guest.

open as a page

A batch fleet running entirely on EC2 Spot keeps losing most of its instances within the same minute. What makes a whole Spot fleet vanish at once, and how would you configure it so a single reclamation cannot take everything?

level: seniorimportance: should knowfreq 54%

basics

~20 s

The fleet is concentrated in one Spot capacity pool — one instance type and size in one Availability Zone — so a single reclamation event hits every instance. The fix is to spread across many pools and let EC2 pick the deepest ones instead of the cheapest.

open as a page

On AWS, how do you decide what gets baked into an EC2 AMI versus what an instance does for itself at first boot through user data?

level: principalimportance: should knowfreq 45%

basics

~20 s

Bake what is slow, static and identical everywhere — OS packages, agents, runtimes, usually the application artifact. Late-bind what varies by environment or must stay fresh — configuration, secrets, cluster identity. The deciding factors are boot latency, reproducibility and how fast you must patch.

open as a page

How do you decide which EC2 instance family and size to run a service on, and why is a fleet of many small instances not automatically cheaper or safer than fewer large ones?

level: principalimportance: should knowfreq 45%

basics

~20 s

Pick the family from the resource your workload is actually bound by, then the size from measured usage plus headroom. Small sizes carry burstable network and EBS bandwidth and pay fixed per-instance overhead repeatedly; large ones concentrate blast radius and coarsen scaling steps.

open as a page

As the platform lead for an engineering organisation, how do you decide which workloads may run on EC2 Spot and which must never, and how do you bound the damage when Spot capacity disappears across the fleet?

level: principalimportance: should knowfreq 44%

basics

~20 s

Judge each workload by what an unannounced two-minute eviction costs it: restart cost, statelessness, and customer-visible impact. Spot suits retryable, replaceable work; singletons holding state must not use it. Bound the damage with an On-Demand floor sized to the throughput the business cannot lose.

open as a page

A team bakes AMIs by launching an instance, configuring it by hand and calling CreateImage. What does AWS EC2 Image Builder change about that, and how is an Image Builder pipeline structured?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

EC2 Image Builder turns hand-baking into a managed, repeatable pipeline: it launches a build instance from a parent image, applies versioned components, runs tests, then distributes the resulting AMI to chosen regions and accounts and tears everything down.

open as a page

What is the AWS Nitro system, and what does running on Nitro-based EC2 instance types change for the operating system and AMI you boot?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

Nitro is the hardware and hypervisor platform behind current EC2 instances: dedicated cards take over networking, storage and management so nearly all host resources go to the guest. Practically, volumes appear as NVMe devices and the AMI needs ENA and NVMe drivers.

open as a page

In EC2, what is the difference between Dedicated Instances and a Dedicated Host, and in what situation does that difference actually matter?

level: middleimportance: nice to knowfreq 26%

basics

~20 s

Both isolate you from other AWS customers' instances. Dedicated Instances give isolation only — AWS still chooses the hardware. A Dedicated Host allocates you a whole physical server, exposing its socket and core counts and letting you pin instances to it, which is what per-socket software licensing needs.

open as a page

After you set an EC2 instance's metadata options to require IMDSv2, applications in Docker containers on that host can no longer read instance role credentials from 169.254.169.254, although the same request works on the host itself. What is the cause, and what is the fix?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

The IMDSv2 token response is sent with an IP TTL equal to the instance's HttpPutResponseHopLimit, which defaults to 1. Routing from a bridged container to the host consumes that hop, so the token never arrives. Raise the hop limit to 2.

open as a page

A tier behind an EC2 Auto Scaling group takes about eight minutes from launch to serving traffic, but its load can double in under two. How would you decide between warm pools, predictive scaling, and simply carrying more headroom?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Reactive scaling cannot beat an eight-minute boot, so the choice is about predictability and cost. Predictable load gets predictive scaling or scheduled actions; unpredictable spikes need pre-paid capacity — headroom or a warm pool — and the real fix is shortening boot time.

open as a page