skip to content

An allow rule written against one replica's address stopped working after that replica moved — what should it have selected instead?

level: middleimportance: should knowfreq 50%

answer

  1. addresses are leases, labels are claims
  2. churn breaks address-pinned rules
  3. a stale address can be reissued
  4. selectors recompute membership continuously
  5. external peers still need prefix ranges

basics

~20 s

Select by label, not by address. A workload's address is a short-lived lease from a per-host pool: a rescheduled replica returns with a different one, so an address-pinned rule silently stops matching, or later matches whoever inherits that address.

solid answer

~50 s

In a container platform an address is a lease, not an identity. It is taken from the pool of the host the workload happened to land on, and returned when the workload dies; a replacement replica gets whatever is free, often on a different host from a different range. A rule that names `10.42.3.17` therefore stops matching the moment the replica is rescheduled, and the failure is silent — the rule is still stored and still valid, it just selects nothing. The durable selector is a **label**: a claim about what the workload *is*, which the replacement carries because it comes from the same workload spec. The enforcing side re-evaluates membership as workloads appear and disappear. Addresses keep one legitimate use: peers outside the platform, which have no labels, where you allow a published prefix range instead.

code

yaml · 12 lines
yaml
appliesTo:
  labels:
    role: ledger-store
defaultForSelected: deny
inbound:
  - fromLabels:
      role: reconciliation-batch
    toPort: 6100
outbound:
  - toLabels:
      role: audit-sink
    toPort: 6200

go deeper

for a junior

Recall the one-liner: a workload's address is temporary and a replacement gets a different one, so rules should be written against labels that describe what the workload is.

for a middle

Explain both failure directions — the pinned rule stops matching the real caller, and the recycled address may later hand the allowance to an unrelated workload — and note that neither raises an error.

for a senior

Show operational reasoning: diagnose "the rule stopped working" as a selector matching an empty set, and treat label hygiene as security input, since a workload shipped without its label silently escapes the rules meant to cover it.

for a principal

Own the consequence across teams: selectors turn a naming convention into an enforcement boundary, so decide who defines the labels, how a missing one is caught before it ships, and how coarse external prefix ranges are reviewed.

## An address here is a lease, not a name Outside a container platform, pinning a firewall rule to an address is reasonable because addresses are assigned deliberately and change on a change-request timescale. Inside one, the opposite is true. A workload's address is allocated when it starts, usually from a range belonging to the host it was placed on, and released when it stops. Nothing about it is stable: it is not derived from what the workload is, it is not reserved for the next instance of that workload, and it is not chosen by anyone you can ask. Every ordinary event churns it. A replacement during a rolling update, a rescheduled replica after its host was drained, an autoscaler adding and removing copies, a crash and restart — each one can produce a different address, potentially in a different per-host range entirely. ## The two failure directions of an address-pinned rule Pinning to an address does not fail in one way; it fails in two opposite ways, and the second is the dangerous one. | Failure | What happens | How it shows up | |---|---|---| | The rule stops matching | The replica it named no longer exists at that address, so the allowance no longer applies to the real caller | Connections are refused after a routine replacement; looks like an outage with no deployment to blame | | The rule matches the wrong workload | The address is returned to the pool and later handed to an unrelated workload, which inherits the allowance | Nothing looks broken at all — a workload silently holds access it was never granted | Neither failure raises an error. A rule naming an address that currently belongs to nobody is a perfectly well-formed rule; the platform stores it and enforces exactly what it says. ## What a selector actually selects A label-based rule names a **set defined by a property**, not a list of members. The rule says, in effect, "whatever is currently labelled `role: reconciliation-batch` may open connections to whatever is currently labelled `role: ledger-store`, on this port". Two things follow: - **Membership is recomputed as the world changes.** A new replica that carries the label is in the set from the moment it is running; one that is deleted is out of it. Nobody edits the rule. - **The allowance is attached to the role, not the instance.** That matches how the workload is actually described everywhere else — the same labels come from the same workload spec that decided how many replicas to run. The cost of this is that **labels become load-bearing security input**. A typo in a selector is not an error; it simply selects nothing, which for an allow rule means the traffic you meant to permit is denied, and for a broad selector means traffic you never considered is permitted. A workload that ships without the expected label quietly falls outside every rule that was supposed to cover it. ## Where an address is still the right selector Label selection only works for peers the platform knows about. Two cases genuinely need addresses: 1. **Peers outside the platform** — a partner endpoint, a managed data service, an appliance. They have no labels, so an outbound rule names a published prefix range instead. 2. **Callers arriving from outside** whose source range is fixed and published, where an inbound rule may name that range. Both are weaker than label selection and should be written knowing why: a prefix range is coarser than you want, the owner may change it without telling you, and a range covering a shared provider covers every tenant on it, not only your partner. ## What to check when a rule has "stopped working" - Does the selector still match anything? A rule selecting an empty set is the usual cause and looks identical to a correctly restrictive rule. - Did the label on the workload change, or did a new workload ship without it? - Is the peer being named by address where a label was available? - Is the direction the one you meant? Permitting a caller inbound does nothing for the calls that workload makes outward.

  • Why is the silent case — a recycled address — worse than the rule simply breaking?
    Because nothing fails. A broken rule shows up immediately as refused connections after a routine replacement. A reused address hands the allowance to an unrelated workload, which now holds access nobody granted it and no alarm fires, so the mistake can sit for months until someone audits which peers each rule really permits today.
  • How would you permit a partner endpoint that has no labels?
    Name the prefix range the partner publishes in an outbound rule, scoped to the port you actually use, and record why. Accept that it is coarser than a label selector: the range may cover other tenants of the same provider, and the owner can change it without notice, so it needs an owner and a review date rather than being written once.

An address-pinned rule is a guest list written as seat numbers. It works until people move seats — then it turns away the guests you invited and waves in whoever happens to be sitting there now.

saying these in an interview costs you the question

  • Treats a workload's address as a stable identity for that workload.
  • Says a rule selecting nothing is an error the platform reports.
  • Assumes a released address is never reused by another workload.
  • Thinks labels are only documentation and carry no enforcement weight.
  • Claims labels can select peers that live outside the platform.