How would you move a shared cluster carrying dozens of teams from allow-everything to default-deny without causing an outage?
answer
- observe the graph before writing rules
- labels and coverage are the real prerequisites
- inbound first, outbound later
- sequence by value, one workload at a time
- measure coverage, not rule count
basics
~20 sDerive the real dependency graph from observed traffic, restrict inbound before outbound, stage it workload by workload starting with the highest-value receivers, and let each team own its own rule. Label hygiene and enforcement coverage are the prerequisites.
solid answer
~50 sThe hard part is not writing rules; it is knowing the true dependency graph and having labels that mean something. So: observe first, from real traffic over a period long enough to include rare paths, rather than asking teams what they call. Then restrict **inbound before outbound** — inbound shrinks the blast radius immediately and its failures are visible as a caller being refused, whereas outbound restriction breaks the workload itself and disguises the failure as its dependencies being down. Sequence by value, not alphabetically: the stores holding the crown jewels are the biggest win per rule. Ship each rule in the same spec as the workload it protects, owned and reviewed by that team, because a central team writing everyone's rules becomes a bottleneck and goes stale. Finish with a per-grouping backstop so a new workload is covered by default, and measure coverage rather than rule count.
go deeper
Take away the shape of the migration: find out who really talks to whom, restrict who may call the important services first, and change one workload at a time rather than the whole cluster.
Be able to argue the ordering — inbound restriction fails visibly as a refused caller, while outbound restriction breaks the workload and disguises itself as its dependencies being unavailable.
Show the operational plan: observed traffic over a long enough window, an observe-only or mirrored rehearsal, the first run watched end to end, and a break-glass path so an incident does not end with the rules deleted.
Own the parts nobody else can: labels as enforcement input with a gate that rejects workloads missing them, proven enforcement coverage, a per-grouping backstop so new workloads are covered by default, and an honest statement of what segmentation does not provide.
## The prerequisite is a dependency graph you did not invent Every failed rollout of this kind fails the same way: rules were written from what people believed the callers were. Teams describe the paths they designed, not the ones that exist — the reporting job somebody added, the client library that pushes telemetry, the month-end settlement step, the operator's ad-hoc tool. So the first phase is not policy at all, it is observation: collect real connections between workloads for long enough to catch the infrequent paths, and treat that record as the source of truth. Two prerequisites sit alongside it, and both are more likely to sink the effort than any rule: - **Label hygiene.** Selectors turn a naming convention into an enforcement boundary. If labels are inconsistent, a rule either matches nothing or matches too much, and both are silent. Agreeing the small set of labels that rules may select on — and rejecting workloads that ship without them — is the real project. - **Enforcement coverage.** If the installation has no component enforcing policy on every host, the entire programme is theatre: rules will be accepted and nothing will change. Prove that a denied connection is actually denied before writing a second rule. ## Restrict inbound before outbound The two directions differ in both value and risk, and the order is a genuine judgment call worth defending. | | Inbound first | Outbound first | |---|---|---| | What it buys | Immediately shrinks who can reach the valuable receivers | Limits where a compromised workload can send data | | How it fails | A forgotten caller is refused — visible, attributable, quick to fix | The workload itself breaks, and the symptom accuses its dependencies | | Blast radius of a mistake | The one caller you missed | Everything that workload does, including paths that run rarely | Inbound-first is usually right: it delivers most of the lateral-movement reduction, and its failures are legible. Outbound restriction is worth doing afterwards on the workloads that handle the most sensitive data, and it is where you will spend the operational effort. ## Sequence by value, not by inventory - Start with the **receivers that matter** — the ledger store, the credential-issuing service, anything whose compromise is the incident you are trying to prevent. A handful of inbound rules there removes more risk than a hundred rules on stateless frontends. - Enforce **one workload at a time**, with the owning team present for the first real run, including a run that exercises the rare path. - Use an **observe-only mode** if the enforcement implementation offers one; if it does not, get the same effect by applying the rules to a pre-production copy carrying mirrored traffic first. ## Who owns a rule The organisational choice matters more than the technical one. A rule belongs in the same spec as the workload it protects, reviewed by that team like any other change, because they are the ones who will add the dependency that needs a new entry. The platform team's job is different and smaller: 1. Provide the backstop — a rule per grouping that denies unlisted inbound and selects everything in that grouping, so a workload shipped tomorrow is covered rather than forgotten. 2. Own the labels rules may select on, and the check that rejects a workload missing them. 3. Own enforcement coverage and its verification. 4. Own the break-glass: a documented, fast, expiring way to widen a rule during an incident, because the alternative is that someone deletes the rules permanently at three in the morning. ## What to measure, and what this never gives you Rule count is a vanity metric. Track instead the **fraction of workloads selected by at least one inbound rule**, the **number of permitted peers per workload** (which should fall over time), and **denied-connection events**, which double as the signal that enforcement is alive. Silence everywhere is a warning, not a win. Finally, be explicit about the ceiling: this mechanism controls **reachability**. It does not establish who the caller is, does not protect the content of a connection, and does nothing about an attacker who has taken over a workload that is already inside the permitted set. Sold as segmentation it is valuable; sold as authentication it will be trusted for something it cannot do.
- Why is a per-grouping backstop rule worth having once teams own their own rules?Because coverage otherwise decays with every new workload. A backstop that selects everything in a grouping and denies unlisted inbound means a workload shipped without its own rule is protected by default instead of silently open, which turns the team's rule into an explicit widening rather than the only thing standing between the workload and the cluster.
- What would make you stop the rollout and fix something else first?Two findings. If a deliberately forbidden connection still succeeds, nothing is enforcing and every rule written so far is decoration. If workloads routinely ship without the labels rules select on, selectors will match the wrong sets in both directions, silently. Either one means the prerequisite work was skipped, and continuing just accumulates rules that do not do what they claim.
- How do you keep permitted peer sets from growing back to allow-everything?Require a reason and an owner on every entry, especially external address ranges, and review them on a schedule with the observed traffic beside them. Entries with no matching traffic for a long period are the ones to remove. Without that record nobody can tell a load-bearing entry from one added during an incident three years ago, so nothing is ever deleted.
saying these in an interview costs you the question
- Writes rules from what teams say their callers are.
- Restricts outbound across the estate first for a faster win.
- Rolls a default-deny posture out cluster-wide in one change.
- Puts a central team in charge of writing every team's rules.
- Counts rules shipped rather than workloads actually covered.
- Presents segmentation as proof of who the caller is.