skip to content

You relabel a compromised workload to a quarantine label, but its outbound session keeps running — why, and what else must isolation do?

level: seniorimportance: nice to knowfreq 33%

answer

  1. the decision was made at connection setup
  2. tracked state is not re-evaluated
  3. the write still has to reach everywhere
  4. flushing is blunt, not surgical
  5. check the flow records afterwards

basics

~10 s

A label change is evaluated for new connections; established state is not re-checked, and the write takes time to reach every enforcement point. Isolation must flush that state and verify the flows stopped.

solid answer

~50 s

Two things sit between the write and the silence. First, enforcement points generally evaluate policy when a connection is established and then track the connection as state; an already-established session matches existing state and keeps flowing until it is flushed or expires, so an intruder holding a long-lived outbound session is unaffected by the relabel. Second, the write is a control-plane event that has to converge across every enforcement point, and that window is seconds to minutes and is almost never measured. So isolation is three steps, not one: write the label, flush or expire connection state at the enforcement points covering that workload, and confirm from flow records that the conversation actually stopped. The price is that flushing state is not surgical — it also kills legitimate long-lived connections on that workload, and you need the convergence number measured in advance rather than discovered during an incident.

go deeper

for a junior

Know that changing a policy input does not undo decisions already made: an existing connection keeps flowing after a label change, so stopping it takes a separate action.

for a middle

Explain stateful evaluation at connection establishment and control-plane convergence as the two distinct delays, and why only the first explains a session that never stops.

for a senior

Give the full isolation sequence — relabel, flush state, verify from flow records — and be candid about the collateral of a flush and about what relabelling does not revoke.

for a principal

Own the preparation: whether per-workload state expiry exists at all, what a wider flush would cost the estate, and who signs for the convergence figure responders are told to rely on.

## Why the session survives the relabel Attribute-keyed enforcement in a platform dataplane is usually stateful, for the same reason every stateful filter is: evaluating a full policy decision per packet is expensive, so the decision is made when a connection is established and the resulting flow is tracked as state. Subsequent packets match the tracked connection and are forwarded without re-consulting policy. That design has a consequence people meet for the first time during an incident: **changing the inputs to a decision does not revisit decisions already made.** Move a workload to a quarantine label and every *new* connection is denied, immediately and correctly. The session an intruder established twenty minutes ago matches existing state and keeps running — potentially for hours, if it is a long-lived tunnel or a connection with traffic on it that keeps the state fresh. The second delay is convergence. The label write lands in a control plane and must propagate to every enforcement point that covers that workload. In a large fabric this is seconds to minutes. It is rarely instrumented, which means the number quoted to responders during an incident is usually a guess, and the guess is usually optimistic. ## What isolation actually requires Three steps, and the middle one is the one that gets forgotten: 1. **Write the label** — denies new connections once the change converges. 2. **Flush or expire the connection state** for that workload at the enforcement points covering it — this is what stops the traffic that is already running. 3. **Verify** from flow records that the conversation stopped. A flow record carries the five-tuple, byte and packet counts and timestamps, which is exactly enough to answer "did bytes keep moving after the change" — and nothing more. It cannot tell you what was in the session; it can tell you when it ended, which is the question you have. Skipping step 3 is how teams end up believing a host was isolated at 14:02 when data kept leaving until 14:40. ## The price, which is why this is a design decision and not a runbook line Flushing state is blunt. It does not distinguish the intruder's session from the workload's legitimate long-lived connections — a database session, a replication stream, a message-broker connection. On the compromised workload that collateral is usually acceptable; the point is that it is a decision someone makes, and it should be made before the incident rather than at 3am. It also has to be *available*. If your enforcement points offer no way to expire state for a single workload, your only lever is a wider flush, and a wider flush is an outage. Knowing which of those two you have — before you need it — is the actual preparation. And the convergence number has to be measured. Someone should be able to answer "how long after we relabel is the new policy in force everywhere" with a figure from a test they ran on a schedule, not with an assumption. That measurement is a small recurring cost and it is the difference between a responder who knows what they bought and one who hopes. ## The related misconception, worth pre-empting Candidates often extend the relabel into a claim it cannot support: that the workload is now cut off. Relabelling removes what the *label* grants. It does not revoke credentials the workload already holds, tokens it already obtained, or a path that some other rule permits for a different reason. Segmentation is one control; a workload that reached its target using a credential rather than a label is unaffected by anything you do to the label. Saying that out loud — this control does this much and no more — is exactly the honesty this kind of interview is testing for. ## How to answer it Lead with the mechanism: stateful evaluation at establishment, so existing flows are unaffected. Add convergence as the second, smaller delay. Then give the three-step isolation and name the collateral of a state flush and the need for a measured convergence figure. Finish by bounding the control: relabelling denies new connections that the label granted, and nothing else.

  • How would you confirm the isolation actually took effect?
    From flow records covering that workload: look for byte and packet counts on the conversation after the change timestamp. A record carries the five-tuple, counts and timestamps and no payload, so it will not tell you what left, but it answers precisely the question you have — whether anything left at all, and when it stopped.
  • What does relabelling a workload not achieve?
    It does not revoke credentials or tokens the workload already holds, and it does not close paths some other rule permits for an unrelated reason. It removes what that label granted, for new connections, once the change converges. Treating a relabel as full containment overstates a single control's reach.

Revoking someone's guest pass stops them entering again. It does not walk them out of the room they are already standing in.

saying these in an interview costs you the question

  • Assumes a policy change tears down established connections
  • Quotes a convergence time nobody has measured
  • Skips verification and declares the workload isolated
  • Believes a state flush is surgical enough to spare legitimate sessions
  • Treats a relabel as revoking the workload's credentials too

context