skip to content

After a shipment-tracking API moves from VMs into Kubernetes, a carrier firewall that allowlisted the VM's fixed IP rejects its outbound calls; why, and what are your options?

level: seniorimportance: should knowfreq 43%

answer

  1. source address is not the VM
  2. pod and node IPs churn
  3. no egress field in core
  4. reserved NAT or egress gateway
  5. identity beats IP

basics

~20 s

Outbound calls now leave from pod or node addresses that change whenever pods reschedule or nodes are replaced, not from the old fixed IP. Route egress through a NAT or egress gateway with a reserved address, or replace IP trust with authentication such as mTLS.

solid answer

~40 s

A pod gets an IP when it starts, and whether its outbound traffic leaves the cluster with that pod IP or is masqueraded to the node's IP depends on the CNI plugin and network design. Kubernetes itself has no egress-IP field or object. Either way the source is not stable: pods move, and nodes are added, upgraded and replaced, so node IPs churn too. The dependable fix is to send egress through something that owns a reserved address, such as a cloud NAT gateway or a CNI egress-gateway feature, and give the carrier that IP. Pinning the workload to a node pool with static addresses works but is brittle. Long term, move the partner trust to mTLS or signed requests, and treat the IP rule as defence in depth.

go deeper

for a junior

Recall that pods and nodes do not keep fixed addresses, so anything that trusted the VM's IP has to be revisited after a move.

for a middle

Explain which address leaves the cluster in common setups, masqueraded node IP versus routable pod IP, and why neither is stable.

for a senior

Show the options with their failure modes: shared NAT widening trust, static node pools blocking rotation, and the need to test from freshly replaced nodes.

for a principal

Argue for moving partner trust from network location to identity, and decide who owns egress addresses as a platform service.

## Why the source address changes On the VMs, the **shipment-tracking API** called a carrier's label API from one fixed public IP, and the carrier's firewall allowed exactly that address. After the move to a **12-node cluster**, the same calls come from somewhere else, and the somewhere keeps moving: - A **pod IP** is assigned when the pod starts, from whatever range the CNI plugin's IP address management uses. A rollout or reschedule gives the replacement a different IP. - A **node IP** is stable only as long as the node exists. Upgrades and autoscaling replace nodes; on this cluster a rotation drains each node in about **13 minutes** and brings up a new one with a new address. - Kubernetes core defines **no egress-IP object or pod field**. What the outside world sees is decided by the CNI plugin and the network underneath. ## Which address the carrier sees | Egress path | Source IP at the carrier | |---|---| | CNI masquerades cluster-leaving traffic, nodes have public IPs | the node's IP, which changes with node replacement | | Pods use routable IPs, no masquerade | the pod IP, which changes with every new pod | | Nodes are private behind a cloud NAT gateway | the NAT gateway's reserved IP | | CNI egress-gateway feature selects these pods | the gateway node's configured IP | A Service does **not** help here. A `ClusterIP` is a virtual address for traffic *into* the pods, and a `LoadBalancer` Service's external IP is for inbound traffic too; neither is guaranteed to be the source of the pods' outbound connections. ## Options and their trade-offs | Option | Stable? | Cost | |---|---|---| | Cloud NAT gateway with reserved IP for the node subnet | yes | every workload in that subnet shares the IP, so the carrier trusts more than one app | | CNI egress gateway scoped by pod labels | yes | depends on a CNI feature; the gateway becomes a critical path | | Dedicated node pool with static public IPs | only until nodes are replaced | brittle; blocks normal node rotation | | Keep a small forward proxy on a VM during transition | yes | one more thing to decommission later | | mTLS or signed requests with the carrier | independent of IP | needs the carrier's cooperation | The sharing problem in the first row is the one reviewers raise: if the NAT gateway serves the whole cluster, the carrier's allowlist now admits every workload on it, including the GPU model-serving pods. Scope egress by subnet or gateway so only the API's pods use the allowlisted address. ## Internal allowlists move too The same assumption exists inside the company. A database firewall that allowed the three VM IPs now sees node or NAT addresses. Two different controls replace it: - A **NetworkPolicy** with an `egress` rule restricts which destinations the API's pods may reach; for an in-cluster database, an `ingress` rule on the database pods restricts who may connect. Both need a CNI that enforces NetworkPolicy. - A NetworkPolicy does not tell a *remote* system who is calling. An external database still needs credentials or client certificates, because IP alone no longer identifies the service. ## The inbound direction The mirror problem, a partner who calls *in* and whose IP you allowlist, depends on whether the client's source address survives the load balancer and kube-proxy hop. That is a Service traffic-policy question and is handled separately; the migration checklist only needs to flag that inbound allowlists must be re-tested. ## A migration sequence 1. List every external party and internal firewall rule that names the VM IPs. 2. Stand up the egress path (NAT or egress gateway) before the first pod calls out. 3. Send the carrier the new address and run both the old and new IPs in their allowlist through the transition. 4. Test from a pod on a node that was just replaced, not only from a long-lived node. 5. Remove the VM IPs from every allowlist once the VMs are gone, and schedule the move to mTLS.

  • Why is pinning the API to a node pool with static public IPs considered brittle?
    Node upgrades, autoscaling and failures replace machines, and a replacement node rarely inherits the old address. The allowlist then breaks during routine maintenance, and teams start blocking node rotation to protect it, which delays security patches. It also couples a network contract to a scheduling constraint.
  • The API now calls an internal database whose firewall allowed the three VM IPs. What replaces that rule?
    Identity plus policy. The database should authenticate the service with credentials or client certificates, since node and NAT addresses are shared. On the cluster side, a NetworkPolicy egress rule limits which destinations the API pods may reach, and for an in-cluster database an ingress rule limits who may connect, provided the CNI enforces NetworkPolicy.

saying these in an interview costs you the question

  • Pods keep a stable IP, so give the carrier the pod IP
  • A Service's ClusterIP is the source address of the pods' outbound calls
  • A LoadBalancer Service's external IP automatically becomes the egress address
  • Kubernetes has a built-in pod field for setting a fixed egress IP
  • Allowlisting the whole cluster node range is as safe as the old single-IP rule