In AWS, when would you run a self-managed NAT instance on EC2 instead of a managed NAT gateway, and what do you take on operationally by doing so?
answer
- managed default versus a host you own
- per-GB processing is the cost driver
- one instance is one point of failure
- the anti-spoofing check must be turned off
- a host can do things an appliance cannot
basics
~20 sChoose a NAT instance only for low-traffic or cost-sensitive environments, or when you need something a NAT gateway cannot do, such as a security group, port forwarding, or traffic filtering. In exchange you own patching, availability failover, and a bandwidth ceiling set by the instance type.
solid answer
~60 sA NAT gateway is the default: AWS runs it, it is redundant inside its Availability Zone, its throughput scales automatically, and there is nothing to patch. You pay an hourly charge plus a per-GB data-processing charge on everything that passes through, which at low traffic volumes can dwarf the cost of a tiny EC2 instance doing the same job. That is the main reason people still build **NAT instances** — a small instance in a public subnet, with `SourceDestCheck` disabled so it is allowed to forward packets that are not addressed to it. The other reason is capability: a NAT instance can carry a security group, act as a bastion, forward specific ports inward, or filter and log egress, none of which a NAT gateway offers. What you take on is real: patching and hardening the host, a throughput ceiling set by the instance type, and your own failover story, because a single instance is a single point of failure that a route table will happily keep pointing at.
code
bash · 9 lines# Required before an EC2 instance can forward traffic for other hosts
aws ec2 modify-instance-attribute \
--instance-id i-0123456789abcdef0 \
--no-source-dest-check
# Verify the attribute took effect
aws ec2 describe-instance-attribute \
--instance-id i-0123456789abcdef0 \
--attribute sourceDestCheckgo deeper
Know that the managed NAT gateway is the normal choice and that a NAT instance is a plain EC2 instance doing the same job by hand. Being able to name one reason for each is enough here.
Explain the four axes — cost shape, bandwidth ceiling, availability model, and what a host can do that an appliance cannot — and mention that the source/destination check must be disabled before an instance can forward at all.
Demonstrate that you price the operational burden, not just the invoice: who patches the host, what the failover time is, and what happens to every private subnet when the single instance stops responding. Be able to describe a credible HA pattern and its recovery window.
Frame it as an egress-architecture decision rather than a device choice: whether the organisation needs an inspected, policy-controlled egress path at all, who owns it, and how that requirement changes the build-versus-managed calculus across every account.
## The default answer, and why it is the default For almost every production VPC, the right egress device is a **NAT gateway**. It is a managed appliance: AWS builds it redundantly within its Availability Zone, scales its throughput for you as traffic grows, patches nothing you can see, and exposes no host for you to compromise. You create it in a public subnet with an Elastic IP and point private route tables at it. That is the whole operational surface. The price is a per-hour charge for each NAT gateway plus a **per-GB data-processing charge on every byte that passes through it, in both directions**, and that second component is what drives people to look for alternatives. It is charged on top of any normal data-transfer-out charge. ## What a NAT instance actually is A NAT instance is an ordinary EC2 instance in a public subnet, with a public address, configured to forward and translate traffic for other subnets. The one AWS-specific step is disabling the instance's **source/destination check**. By default EC2 drops any packet an instance sends or receives whose source or destination is not the instance itself — a sensible anti-spoofing default that makes routing impossible. Turning it off is what converts an instance into a router: ``` aws ec2 modify-instance-attribute \ --instance-id i-0123456789abcdef0 \ --no-source-dest-check ``` The forwarding and address-translation configuration on the host itself is ordinary Linux networking, not an AWS feature. ## The honest comparison **Cost.** A NAT gateway's hourly charge is roughly that of a small always-on instance, and then the per-GB processing charge accrues on top. A `t4g.nano` or `t4g.micro` NAT instance has no per-GB processing charge at all — you pay for the instance and for normal data transfer. For a development account pushing a few gigabytes a month, the instance can be an order of magnitude cheaper. For a production account pushing terabytes, the instance's bandwidth ceiling becomes the binding constraint long before the savings matter. **Bandwidth.** A NAT gateway scales its throughput automatically as traffic grows, with no action from you. A NAT instance is capped by its instance type's network performance, and small burstable types have credit-based network and CPU limits that produce exactly the kind of intermittent, load-dependent slowness that is miserable to diagnose. **Availability.** A NAT gateway is redundant inside its Availability Zone; if the underlying hardware fails, you do not notice. A NAT instance is one instance. If it stops responding, the route table keeps sending traffic to it and every private subnet behind it loses egress. Building real HA means health-checking the instance and having something rewrite the route table entry to a standby — an Auto Scaling group of size one plus a bootstrap that claims the route is the common pattern, and it is genuinely more moving parts than most teams want in the path of all outbound traffic. **Capability.** This is the underrated column. A NAT gateway cannot have a security group attached, cannot forward a port inbound, cannot run a proxy, and gives you no per-connection visibility beyond CloudWatch metrics and flow logs. A NAT instance is a host you control: it can carry a security group, double as a bastion, run an egress proxy that allow-lists destination domains, or log every connection. Teams with a hard egress-control requirement sometimes choose an instance-based path for exactly this reason — though at that point the honest framing is that they are building an egress firewall that happens to also do NAT, not saving money on NAT. **Operational burden.** The NAT instance is a Linux host in the traffic path of everything: it needs patching, hardening, monitoring, and an owner. A NAT gateway needs none of that. ## How to answer Say the default is the NAT gateway and give the two legitimate reasons to deviate — **cost at low volume**, and **capability you cannot get otherwise**. Then name the three things you have signed up for: patching, a bandwidth ceiling, and failover you have to build yourself. A candidate who claims the NAT instance is simply cheaper, with no mention of availability, is telling you they have never had one fail at 3am and take every private subnet's internet access with it. The third option worth mentioning is removing traffic from the NAT path altogether where the destination allows it, which is a different leaf's subject but the correct instinct: the cheapest NAT byte is the one that never traverses the NAT device.
- What exactly does disabling the source/destination check do, and why is it required?By default EC2 discards packets an instance sends or receives whose source or destination address is not the instance's own — an anti-spoofing guard. A router must handle exactly such packets, so forwarding is impossible until you set the instance attribute off. It is a per-instance EC2 setting, not anything inside the guest OS.
- How would you give a NAT instance real high availability?Put it in an Auto Scaling group of size one per Availability Zone, and have the instance claim the private route table's default route on boot. When it dies, the ASG replaces it and the replacement re-points the route. It is workable, but the recovery time is instance-launch time, during which that AZ has no egress.
- Which of the two can you attach a security group to?Only the NAT instance — it is an ordinary instance with an elastic network interface. A NAT gateway accepts no security group at all, so any filtering around it has to live on the instances behind it or on the subnet it serves.
saying these in an interview costs you the question
- Claims a NAT instance is always cheaper without mentioning bandwidth or failover
- Forgets the source/destination check and wonders why forwarding fails
- Thinks a NAT gateway can carry a security group
- Assumes an Auto Scaling group alone gives a NAT instance seamless failover
- Says a NAT gateway needs patching or sizing