skip to content

Network Policy & Egress

Most platforms let any workload talk to any other until someone says otherwise; tightening that means allowing by labels rather than addresses, inbound and outbound. Lateral movement is the point.

on this pageshow

questions

5

In a cluster with no network rules written, which workloads can one compromised worker reach, and why?

level: middleimportance: must knowfreq 66%

answer

  1. reachability first, restriction later
  2. nothing filters between workloads by default
  3. one foothold reaches every listening port
  4. lateral movement is the actual threat
  5. selection flips a workload to allowlist

basics

~20 s

A compromised workload can reach every other workload on the cluster network, in any grouping, on any port they are listening on. Most container platforms ship an allow-everything default, so reachability stays universal until a rule narrows it.

solid answer

~50 s

Container platforms almost all start from a flat, permissive network: every workload gets an address, any address can route to any other, and nothing in that path inspects who is talking to whom. So a compromised queue consumer can open a connection to the ledger store, to a neighbouring team's admin endpoint, and to anything else that is listening — including ports that were never published outside the cluster. That is **lateral movement**, and it is the reason network policy exists. The correction is an allowlist written in the platform's own terms: rules that select workloads by `label` and permit named peers, per direction. In the widespread design the rules are additive allowances — a workload that no rule selects keeps the permissive default, and a workload selected by at least one rule for a direction then accepts only what those rules permit in that direction.

go deeper

for a junior

Remember the headline: by default any workload can open a connection to any other, whatever grouping each lives in. Restriction is something a rule adds, never something you get for free.

for a middle

Explain the mechanics: rules select workloads by label, act per direction, and are additive allowances, so a workload no rule selects keeps the open default and two rules on one workload permit the sum of their peers.

for a senior

Show you reason in blast radius. Name what a foothold reaches from where it runs, argue for restricting the highest-value receivers inbound first, and say plainly that reachability control is not caller authentication.

for a principal

Frame the default as a platform-wide risk position with a coverage metric attached: what fraction of workloads any rule selects, how permitted peer sets are kept small as teams ship, and who is accountable when neither is true.

## Reachability is the default; isolation is the exception When a container platform starts a workload, it attaches it to a network on which **every other workload is already reachable**. Each workload is handed an address, routing is arranged so any address can reach any other across hosts, and no component in that path asks whether the conversation was intended. Splitting workloads into separate logical groupings organises names, quotas and access to the platform's own interface — it does **not** put a filter in the packet path. That default is a deliberate design choice, not an oversight. A platform cannot know at install time which workload legitimately calls which; starting closed would mean nothing works until somebody has enumerated the entire dependency graph of every team. So platforms start open and offer a separate mechanism — **network policy** — for narrowing it afterwards. The consequence is that *every* new cluster is, on day one, one flat blast radius. ## What the flat default hands an attacker Picture a shared cluster running a nightly reconciliation batch, a ledger store it reads, and an unrelated queue consumer that has just been compromised through a dependency. With no rules written, that consumer can: - **Dial the ledger store directly**, rather than through the service that was supposed to be the only caller — the store's authorisation is now the only thing in the way. - **Reach ports nobody meant to expose.** "Not published to the outside" is not "not listening": admin surfaces, metrics endpoints and debug listeners are all reachable from inside. - **Reach another team's workloads in another grouping**, because the grouping is an organisational boundary, not a network one. - **Scan cheaply and quietly.** Connection attempts between workloads are ordinary traffic; nothing about them is anomalous by default. - **Do all of this from a trusted position**, where internal callers are commonly assumed to be friendly and are logged less than external ones. The damage is not what the consumer itself was allowed to do. It is everything *reachable* from where it runs. ## Rules are additive allowances The mechanism that fixes this is almost always expressed as allowances rather than prohibitions. A rule names which workloads it **applies to** (by label), a **direction**, and which peers are permitted. What matters is what selection does to the workload: | The workload is… | Inbound | Outbound | |---|---|---| | selected by no rule | everything is allowed | everything is allowed | | selected by an inbound rule only | only the listed peers | everything is allowed | | selected by rules in both directions | only the listed peers | only the listed destinations | Three consequences follow, and all three are asked about: 1. **Selection is what flips the default**, not the existence of rules somewhere in the cluster. Writing your first rule changes nothing for any workload that rule does not select. 2. **The two directions are independent.** Restricting who may call a workload says nothing about where that workload may call out to. 3. **Rules union, they do not intersect.** Two rules selecting the same workload permit the sum of their peers, so a broad rule quietly re-opens what a narrow one closed. Platforms differ in the details — some offer only additive allow rules, others add explicit deny rules with an ordering between them — so in an interview it is worth saying which model you are describing. ## Why your first rule protects less than you expect The common disappointment is writing one careful rule for the ledger store and assuming the cluster is now segmented. It is not: - Every workload the rule does not select still talks to everything. - Anything already inside the permitted peer set is unaffected, so if the permitted peer is compromised, the rule was never in the way. - A rule is a **reachability** control. It does not prove who the caller is, does not encrypt anything, and does not replace the authorisation the receiving workload should be doing itself. A segmented estate is therefore not one clever rule; it is coverage — the fraction of workloads that at least one inbound rule selects — plus small permitted peer sets, maintained as the workloads change.

  • What exactly changes for a workload the moment one rule selects it for a direction?
    For that direction only, it stops accepting everything and accepts just what the selecting rules permit. The other direction is untouched, and every workload the rule does not select keeps the open default. Adding a second rule for the same direction widens the permitted set rather than narrowing it.
  • Does permitting an inbound connection also require an outbound rule for the replies?
    Normally no. Enforcement is commonly connection-aware, so packets belonging to a connection that was already permitted flow both ways without a separate rule. An outbound rule governs connections the workload itself opens, which is why a service that only answers requests can often survive a default-deny outbound posture untouched.
  • Why do platforms ship allow-everything rather than deny-everything?
    Because the platform does not know the dependency graph and cannot invent it. A closed default would mean nothing communicates until every legitimate pair has been enumerated by hand, which makes an empty cluster unusable and pushes an impossible burden onto whoever installs it. The open default trades safety for immediate usability, deliberately.

saying these in an interview costs you the question

  • Says workloads in separate groupings cannot reach each other by default.
  • Thinks a port not published outside the cluster is unreachable from inside it.
  • Assumes the platform denies traffic by default and rules only add exceptions.
  • Believes one rule on one workload segments the whole cluster.
  • Treats an allow rule as proof of which workload is calling.
  • Claims two rules on the same workload narrow each other rather than adding up.
open as a page

A nightly reconciliation batch is switched to default-deny outbound — what stops working first, and why?

level: seniorimportance: must knowfreq 56%

basics

~10 s

Name resolution stops working first. The resolver is itself reached over the network, so unless the outbound allowlist permits it, every lookup fails and each dependency appears to be down rather than blocked.

open as a page

An allow rule written against one replica's address stopped working after that replica moved — what should it have selected instead?

level: middleimportance: should knowfreq 50%

basics

~20 s

Select by label, not by address. A workload's address is a short-lived lease from a per-host pool: a rescheduled replica returns with a different one, so an address-pinned rule silently stops matching, or later matches whoever inherits that address.

open as a page

How would you move a shared cluster carrying dozens of teams from allow-everything to default-deny without causing an outage?

level: principalimportance: should knowfreq 28%

basics

~20 s

Derive the real dependency graph from observed traffic, restrict inbound before outbound, stage it workload by workload starting with the highest-value receivers, and let each team own its own rule. Label hygiene and enforcement coverage are the prerequisites.

open as a page

A network rule was accepted and stored, yet the traffic it forbids still flows — what is missing?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

An enforcing component on the host where the workload runs. A stored rule is only a declaration; something local to each host must translate it into packet filtering; if nothing does, the rule is stored without error and enforces nothing.

open as a page