An EC2 instance shows "1/2 checks passed" in the console. Explain the difference between the system status check and the instance status check, and how your response differs depending on which one failed.
answer
- two checks, two owners
- one side you cannot reach from inside
- reboot keeps the same hardware
- stop/start changes hardware
- instance store blocks recovery
basics
~20 sThe system status check covers AWS-side host hardware, power and network; the instance status check covers the guest operating system. A failed system check is resolved by moving hosts through stop/start or auto-recovery; a failed instance check needs a reboot or a fix inside the guest.
solid answer
~50 sEC2 runs two independent checks every minute. The **system** status check tests the infrastructure underneath your instance — host hardware, power, and the network and software of the underlying host. You cannot fix any of that from inside the guest, so the remedy is to move the instance to different hardware: stop and start it, or let automatic recovery do it for you. Note that a *reboot* will not help, because a reboot keeps the instance on the same host. The **instance** status check tests the instance itself: a kernel panic, an exhausted filesystem, a corrupt file system, a misconfigured network stack, or an OS that failed to finish booting. That is your side of the shared responsibility, so the response is to reboot, then diagnose through the console output or screenshot and the serial console, and finally to fix whatever the AMI or bootstrap is doing wrong. Both are published to CloudWatch as `StatusCheckFailed_System` and `StatusCheckFailed_Instance`, which is what you alarm on.
code
bash · 6 linesaws cloudwatch put-metric-alarm --alarm-name ec2-system-check-recover \
--namespace AWS/EC2 --metric-name StatusCheckFailed_System \
--dimensions Name=InstanceId,Value=i-0123456789abcdef0 \
--statistic Maximum --period 60 --evaluation-periods 2 --threshold 1 \
--comparison-operator GreaterThanOrEqualToThreshold \
--alarm-actions arn:aws:automate:us-east-1:ec2:recovergo deeper
Know that EC2 runs two checks — one on AWS's infrastructure and one on your instance's operating system — and that "1/2 checks passed" means one of those two halves is broken.
Explain which failure maps to which action: reboot for a guest problem, stop/start for a host problem, and why a reboot cannot help a system-check failure since the instance stays on the same host.
Show the operational path end to end: alarm on StatusCheckFailed_System and StatusCheckFailed_Instance, use console output or the serial console to diagnose the guest, and know that automatic recovery excludes instances with instance store.
Argue the policy: fleet members are replaced rather than recovered, recovery is reserved for genuinely singular instances, and a repeated instance-check failure is an image defect to be fixed at the source rather than an incident to be handled.
## The checks EC2 runs for you Every EC2 instance is probed roughly once a minute by checks that pass or fail independently. The console summarises them as "2/2 checks passed"; the interesting states are the partial ones, because which half failed tells you whose problem it is. - **System status check** — the infrastructure the instance sits on: host hardware, host software, loss of power, loss of network connectivity to the host. A failure here means AWS's side is impaired. - **Instance status check** — the instance's own configuration and guest operating system: it verifies that the instance is reachable and responsive at the OS level. - **Attached EBS status check** — a separate check reporting whether the instance's attached EBS volumes are reachable and able to complete I/O. The corresponding CloudWatch metrics in the `AWS/EC2` namespace are `StatusCheckFailed_System`, `StatusCheckFailed_Instance`, `StatusCheckFailed_AttachedEBS`, and the aggregate `StatusCheckFailed`. Each publishes 0 for pass and 1 for fail. ## System check failed: move the instance Because the fault is in the host, nothing you do inside the guest changes it. The correct action is to get onto different hardware: - **Stop and start** the instance. This releases it from the impaired host and schedules it onto a new one. Everything on EBS survives; instance-store data does not, and the auto-assigned public IPv4 address changes. - **Wait**, if the impairment is a transient host issue AWS is already remediating. Scheduled maintenance and instance retirement notices show up as AWS-scheduled events on the instance. The trap here is the reflex to reboot. `RebootInstances` restarts the guest **on the same host**, so it does exactly nothing for a system-level impairment. Distinguishing reboot from stop/start is the single most common discriminator in this question. ## Automatic recovery EC2 provides automatic recovery for a system-impaired instance: it is migrated onto new hardware while keeping its instance ID, private IPv4 address, any Elastic IP, its EBS volumes and its metadata. Modern instance types have simplified automatic recovery enabled by default, and you can also drive it explicitly with a CloudWatch alarm action: ```bash aws cloudwatch put-metric-alarm --alarm-name ec2-system-check-recover \ --namespace AWS/EC2 --metric-name StatusCheckFailed_System \ --dimensions Name=InstanceId,Value=i-0123456789abcdef0 \ --statistic Maximum --period 60 --evaluation-periods 2 --threshold 1 \ --comparison-operator GreaterThanOrEqualToThreshold \ --alarm-actions arn:aws:automate:us-east-1:ec2:recover ``` The constraint worth remembering: recovery is not available for instances that use instance-store volumes, because there is no way to move those disks. That fact quietly shapes the design — a workload that depends on instance store cannot be auto-recovered and must be rebuildable instead. ## Instance check failed: fix the guest A failed instance check is on your side of the line. Typical causes: a kernel panic, a full root filesystem, a corrupted file system that dropped the instance into a repair prompt, exhausted memory, a misconfigured network interface or firewall rule inside the OS, or a boot that stalled on a bad `/etc/fstab` entry after a volume was detached. The sequence is: reboot first (it clears transient states and is cheap), and if the check does not clear, look at what the instance printed. `GetConsoleOutput` returns the serial console buffer — kernel messages, boot output, panic traces — and `GetConsoleScreenshot` gives you the actual screen, which is how you catch a machine sitting at a filesystem-repair prompt. The EC2 Serial Console gives interactive access to a machine whose network stack is broken, and it works when SSH does not, which is precisely when you need it. As a last resort, stop the instance, detach the root volume, attach it to a healthy instance to repair the filesystem or configuration, and reattach. ## Alarming on it Status checks are the cheapest health signal EC2 gives you and they are per-instance rather than per-application: they will tell you the box is broken and never that the application is. Alarm on `StatusCheckFailed_System` with a recovery action, and on `StatusCheckFailed_Instance` to page or to trigger replacement. For instances in an Auto Scaling group, the group's health check already acts on these signals and replaces the instance, which is generally a better answer than recovering it — a fleet member should be replaced, not repaired. ## The judgment an interviewer wants The strong answer separates three responses: *reboot* fixes a guest problem, *stop/start or recovery* fixes a host problem, and *replace* is what you do when the instance is a fungible member of a fleet. The weak answer reboots everything and hopes. Say out loud that a repeated instance-check failure across replacements is an AMI or bootstrap defect and not an infrastructure event — that is the point at which you stop operating and start fixing the image.
- Why won't rebooting the instance clear a failed system status check?A reboot is an operating-system restart in place — the instance never leaves the host it is on. A system status check failure means that host, its power, or its network is impaired, so restarting the guest on top of it changes nothing. You need a stop and start, or automatic recovery, both of which place the instance on different hardware.
- Which instances cannot use EC2 automatic recovery, and what does that mean for their design?Instances that use instance-store volumes, because those disks are physical drives on the host and cannot move with the instance. Some instance types and placements are excluded too. The design consequence is that such a workload must be rebuildable and replaceable rather than recoverable — its state has to live on EBS, in S3, or in a managed data store.
- You have an instance in an Auto Scaling group failing its instance status check. Should you recover it or let the group replace it?Let it be replaced. A group member is fungible, replacement gives you a known-good instance from the current launch template, and repairing a single machine builds a pet. Recovery is for instances that genuinely cannot be replaced. If the replacements keep failing the same check, stop treating it as an incident — the AMI or bootstrap is defective.
saying these in an interview costs you the question
- Rebooting to clear a failed system status check
- Treating status checks as application health checks
- Believing AWS repairs a guest-OS failure for you
- Assuming every instance supports automatic recovery
- Thinking a stop/start and a reboot are interchangeable