A nightly reconciliation batch is switched to default-deny outbound — what stops working first, and why?
answer
- outbound governs connections the workload opens
- the first connection is a lookup
- denied resolution masquerades as everything being down
- labels inside, prefix ranges outside
- replies to permitted connections still flow
basics
~10 sName resolution stops working first. The resolver is itself reached over the network, so unless the outbound allowlist permits it, every lookup fails and each dependency appears to be down rather than blocked.
solid answer
~40 sRestricting outbound traffic governs the connections a workload *opens*, and the first connection almost every workload opens is a lookup to the platform's resolver. Deny that and nothing resolves, so the batch reports that the ledger store is unreachable and the partner endpoint is unreachable, which sends everyone debugging the wrong things. The allowlist needs, in order: the resolver, the in-platform peers selected by `label`, and the external partner named by its published prefix range and port, because a rule that matches addresses cannot express a hostname. Next to break are the workload's own side-channels — pushing telemetry, fetching a credential from an external store at start-up. What usually does *not* break is replying to permitted inbound connections, since enforcement is commonly connection-aware.
code
yaml · 14 linesappliesTo:
labels:
job: nightly-reconciliation
defaultForSelected: deny
outbound:
- toLabels:
role: cluster-name-resolver
toPort: name-resolution
- toLabels:
role: ledger-store
toPort: 6100
- toAddressRange: 203.0.113.0/24
toPort: 443
reason: settlement partner, published range, review 2027-03go deeper
Hold on to the order of events: a workload looks a name up before it connects to anything, so restricting where it may call out breaks the lookup before it breaks the dependency.
Explain that an outbound rule covers connections the workload opens, that the permitted set becomes an allowlist only once a rule selects the workload, and that in-platform peers are named by label while external ones are named by address range.
Demonstrate you have rolled this out: observe real destinations across several runs, permit the resolver first, enforce on one workload, and recognise a denied outbound connection by its disguise — a hang or refusal that accuses the destination.
Weigh what outbound restriction is worth per unit of operational pain: strong inside the platform where labels exist, weak at the edge where you match somebody else's ranges, and only durable if every entry carries a reason and an owner.
## What restricting outbound actually switches off An outbound rule set governs the connections the workload **initiates**. That is a narrower thing than "all traffic leaving the container", and the distinction decides which failures you should expect. Once the batch is selected by at least one outbound rule, the permissive default is gone for that direction: only destinations some rule permits are reachable, and everything else is dropped. Crucially, the drop happens on the way out, so the workload sees a connection that hangs or is refused — not an error saying a rule denied it. Every symptom therefore arrives disguised as the *destination* being broken. ## The first thing to break is name resolution The batch does not open a connection to the ledger store first. It opens one to the resolver, to turn a name into an address. The resolver is an ordinary network peer, so an outbound allowlist that lists only the ledger store and the partner denies it, and then: - Every lookup fails or times out. - The batch reports the ledger store unreachable **and** the partner unreachable, simultaneously. - The two independent failures look like a platform-wide outage rather than one missing line in a rule. So the first entry in any outbound allowlist is the resolver, permitted on the port it serves. This is the single most common mistake made the first time a team restricts outbound traffic, and it is worth saying out loud in an interview before anything else. ## Then anything whose address you do not control After resolution works, the remaining destinations divide by how they can be named at all: | Destination | How the rule names it | The catch | |---|---|---| | In-platform peer (the ledger store) | By `label`, per port | None — this is what selectors are for | | External peer reached by name (the partner) | By the prefix range the partner publishes, or by a name if the implementation supports name-based rules | A prefix range is coarser than the one endpoint you want; a name-based rule depends on resolution and caching, and can go stale | | External peer with no published range | Cannot be named precisely | You end up permitting far more than you meant, which is worth escalating rather than quietly widening | The honest point here is that outbound restriction is strong inside the platform, where labels exist, and weak at the edge, where you are matching addresses that somebody else controls and may change without notice. ## Then the workload's own side-channels The third wave of breakage is the traffic nobody thinks of as application traffic, because the application did not write it: - Pushing logs, metrics or traces to a collector, when the workload pushes rather than being scraped. - Fetching a credential from an external store at start-up, which turns into a start-up failure rather than a runtime one. - Any client library that phones a service the author never mentioned. ## What does not break Two things reliably surprise people: 1. **Replies to permitted inbound connections keep flowing.** Enforcement is commonly connection-aware, so packets belonging to an already-permitted connection are not re-evaluated as a new outbound connection. A workload that only answers requests can survive a default-deny outbound posture with no rules at all. 2. **Fetching the workload's image is usually unaffected**, because the host agent fetches it under the host's own identity before the workload exists, rather than the workload opening that connection itself. ## Building the allowlist without guessing 1. **Observe before enforcing.** Collect the destinations the batch actually reaches over several runs, including a month-end run, rather than asking the owning team to remember. The list is always longer than the answer you get verbally. 2. **Permit the resolver first**, then the in-platform peers by label, then the external ranges, each scoped to a port. 3. **Enforce on one workload, not the estate.** A nightly batch is a good first candidate precisely because its failure is contained and its dependency list is short. 4. **Watch the first run end to end.** A batch that only touches the partner at the final settlement step will pass the first ninety percent of its run before the missing rule shows up. 5. **Record why each entry exists.** An outbound allowlist with unexplained prefix ranges cannot be pruned later, which is how it grows back to permitting everything.
- The batch now resolves names but the partner call still fails. What is the likely cause?The name resolved to an address outside the range you permitted. Providers rotate endpoints and publish ranges wider than the one address you observed, so a rule built from a single resolved address works until the next rotation. Permit the published range rather than the observed address, scoped to the port, and re-check it on a schedule.
- Why does a service that only answers requests often need no outbound rules at all?Because enforcement is commonly connection-aware: the replies belong to a connection that an inbound rule already permitted, so they are not evaluated as new outbound connections. Such a service needs outbound entries only for the connections it opens itself — name lookups, a credential fetch at start-up, or pushing telemetry.
- How do you tell a denied outbound connection apart from a genuinely unreachable dependency?By correlating with the rule rather than the symptom, since both look like a hang or a refusal. Check whether the destination is reachable from a workload that no outbound rule selects, whether the enforcing component recorded a drop for this workload, and whether the failure started at the moment the rule was applied rather than on a deployment of the dependency.
saying these in an interview costs you the question
- Forgets the resolver and concludes the dependencies are down.
- Assumes replies to a permitted inbound connection need their own outbound rule.
- Thinks an outbound rule can name a hostname wherever addresses are matched.
- Expects the workload to see an error saying a rule denied the connection.
- Builds the allowlist from what the owning team remembers rather than observed traffic.
- Treats one clean run as proof, missing a step that only the month-end run reaches.