skip to content

A search tier scales out and the new instances are refused by an address-based allowlist — what should that rule reference instead?

level: middleimportance: must knowfreq 60%

answer

  1. addresses move, membership does not
  2. what does the rule actually name
  3. resolved when the packet arrives
  4. scaling changes membership, not rules
  5. the stateless layer cannot do this

basics

~20 s

The caller's group rather than its addresses. A group-referencing rule on the workload-attached rule set resolves membership when the packet arrives, so instances that appear or disappear are covered with no rule edit and no address list to maintain.

solid answer

~50 s

An allowlist written as addresses is a snapshot of a tier that moves. Every replacement, scale-out and zone rebalance invents addresses the list has never seen, so the rule has to be edited in step with capacity — which nobody does at three in the morning. On the workload-attached stateful layer you can instead write the source as the **caller's group**: the rule says "anything belonging to the search tier's group may reach this port", and membership is resolved as the packet arrives. Scaling changes membership, not the rule. The tempting wrong fix is to widen the allow to the caller's whole subnet range, which silently grants the same access to every other workload that ever lands in that subnet. Note also that the group reference is the *network* boundary only — it decides who may connect, never what the caller is authorized to do once connected.

code

yaml · 13 lines
yaml
# stateful rule set attached to the record store
inbound:
  - protocol: tcp
    portRange: 8443          # the store's listening port
    source:
      group: search-tier     # resolved per packet: new members covered
    action: allow

  - protocol: tcp
    portRange: 8443
    source:
      cidr: 10.20.4.0/24     # every workload placed here, now and later
    action: allow

go deeper

for a junior

Recall why addresses are the wrong thing to write down: instances are replaced and added, so the list is out of date as soon as capacity changes. The rule should name the caller's group instead.

for a middle

Explain that membership is resolved as the packet arrives, which is why scale-out needs no rule edit, and that the stateless subnet layer cannot express this because it only matches address ranges and ports.

for a senior

Recognise the signature — a minority of requests failing, the proportion tracking new capacity — and reject the subnet-wide widening, naming exactly what it grants. Say whether the group reference survives the estate boundary in your topology.

for a principal

Set it as the estate standard: precise group-referencing allows on the workload layer, coarse statements on the subnet layer, address literals only where a group reference genuinely cannot cross a boundary, and never a subnet range used as a trust statement.

## Why an address allowlist rots A customer-facing search tier is exactly the shape that breaks address-based allowlists: it scales with load, its instances are replaced rather than repaired, and its capacity is spread across zones by the platform rather than by you. Each of those events produces addresses that were not in the list when the list was written. The rule is a **snapshot of a moving tier**, and the moment capacity changes it is wrong. The failure has a recognisable signature. Most requests succeed and a minority fail; the proportion tracks how much of the tier is new; the tier looks healthy from its own side because the connection is refused at the far boundary. Under scale-out — that is, under load — the failure rate rises exactly when it hurts most. ## What a group-referencing rule does differently On the workload-attached stateful layer, the source of a rule does not have to be an address range. It can be a **group**: a named collection that workloads belong to. The rule then reads "traffic from anything in the search tier's group may reach this port". Membership is evaluated **when the packet arrives**, not when the rule was written, so: - an instance that joins the group is covered the moment it starts, with no rule edit; - an instance that leaves loses the permission immediately; - the rule survives replacement, scale-out, scale-in and zone rebalancing untouched; - nobody maintains a list of addresses, which means nobody forgets to. The permission is also **directional and specific**: the rule sits on the thing being called and names the caller's group, not the other way round. That is what makes it reviewable — you can read the store's rules and see exactly which tiers may reach it. ## The wrong fix, and why it is tempting The fast repair when pages are firing is to widen the allow to the caller's **whole subnet range**. It works instantly and it is almost always wrong: the permission now belongs to every workload that is in that subnet today and every workload that is ever placed there. A subnet is a placement decision made by whoever allocated the address plan; it is not a statement of trust, and using it as one is how an estate ends up with every tier reachable from every other. | Source written as | What changes when the tier scales | What else is granted | |---|---|---| | A list of instance addresses | the rule is wrong until someone edits it | nothing extra, but it breaks constantly | | The caller's whole subnet range | nothing | every current and future workload in that subnet | | The caller's group | nothing | nothing — membership is the grant | ## Two limits worth knowing First, **the stateless subnet filter cannot do this**. It understands address ranges and ports, and nothing else — there is no membership to resolve at that layer. So group-referencing is a reason to keep precise policy on the workload-attached layer and leave the subnet filter coarse. If you push fine-grained policy down to the stateless layer you are back to maintaining addresses. Second, **a group reference does not always cross an estate boundary**. Platforms differ: on some, a rule in one private address range can reference a group in another network joined to it, and on others it cannot, leaving you with address ranges for cross-estate traffic. Check before designing a topology that assumes it, and expect the cross-estate leg to be coarser than the in-network one. ## What it is not A group-referencing allow decides **who may open a connection**. It does not decide what the caller may then do. Authorization — which caller may read which record, invoke which operation, hold which permission — is a separate mechanism that the platform's access model owns, and a network rule cannot express it. Treating "only the search tier can reach the store" as an authorization control is the classic conflation: it means one compromised member of the group inherits everything the group's network permission grants, with nothing further in the way. ## Saying it in an interview Name the churn first, then the mechanism, then the trap: addresses change because instances are replaced and scaled, a group reference is resolved per packet so membership is the grant, and widening to the subnet range is the fix that stops the pages and quietly grants the whole subnet. Finish with the boundary: this is reachability, not authorization.

  • Why is widening the allow to the caller's subnet range a bad repair?
    Because a subnet is a placement decision, not a trust boundary. The permission then belongs to every workload in that subnet today and every workload placed there later, including ones owned by other teams. It stops the pages and leaves a grant nobody will find until a review asks who can reach the store.
  • Can the stateless subnet filter reference a caller group in the same way?
    No. It matches on address ranges, ports, protocol and direction only — there is no membership for it to resolve. That is a strong argument for keeping precise, per-tier policy on the workload-attached layer and using the subnet filter for coarse whole-subnet statements.
  • Does allowing the search tier's group to reach the store make the store safe?
    No. The rule controls who may open a connection, not what they may do once connected. Authorization is a separate mechanism, and any compromised member of the group inherits the full network reach the group was granted. Treat the group reference as reachability and keep a real authorization control behind it.

saying these in an interview costs you the question

  • Widens the allowlist to the caller's whole subnet range to stop the pages
  • Says instance addresses are stable enough because the tier is long-lived
  • Believes a group reference is just a stored list of addresses
  • Expects group references to work in the stateless subnet filter
  • Thinks reaching the store through a group rule also authorizes the caller