skip to content

One inspection stack serves ten tenants: what does your unclassified-traffic rule say, who signs for it, and what hides behind an allow?

level: principalimportance: nice to knowfreq 34%

answer

  1. a default is a risk acceptance
  2. bespoke protocols are unclassified by design
  3. the shared stack shares the blast radius
  4. log-only first, then a dated migration
  5. an exception needs a name and an expiry

basics

~20 s

The default is a risk acceptance, not a technical preference, and a provider cannot accept risk for a customer. Aim at per-tenant defaults: deny for new tenants from onboarding, a dated migration with a log-only phase for the rest, and a named tenant signature on every allow.

solid answer

~50 s

Both global answers are wrong for different reasons. Allow-unclassified everywhere breaks nothing and is exactly the shape of a purpose-built channel, because a bespoke protocol is unclassified by construction — you have written a hole and given it to ten tenants at once. Deny-unclassified everywhere is a defensible posture that will drop bespoke customer applications you have never heard of, at a boundary those customers may not know exists, and the outage lands on them while the decision was yours. The workable answer is per-tenant defaults, and the cost is honest: a policy set, a named owner and a review per tenant. Sequence it — deny from day one for new tenants so breakage surfaces during acceptance testing, and for existing tenants run a log-only phase that produces a per-tenant report of what a deny would have dropped, then migrate on a date they agreed. Whoever wants allow signs for it in writing, with an expiry.

go deeper

for a junior

Know that traffic the classifier cannot name still has to match some rule, and that whether it is allowed or denied is a decision somebody made rather than a property of the device.

for a middle

Explain both failure directions concretely: a permissive default admits bespoke protocols silently, a strict one drops applications nobody has an inventory for. Be able to describe a log-only phase.

for a senior

Show the rollout: report mode first, work the resulting list down to scoped allows with owners, migrate on an agreed date, keep the rollback to a single change made by whoever is on call.

for a principal

Own the ownership question. Be ready to say who signs the risk, what the provider owns instead of the decision, how exceptions expire, and what you escalate when a customer will not engage at all.

## Why this is not a technical decision The unclassified-traffic rule decides what happens to traffic nobody can name. Whichever way it points, somebody bears a consequence they did not choose: with allow, a customer carries a risk they were never quoted; with deny, a customer carries an outage in an application the provider has never heard of. That makes it a **risk acceptance**, and a provider cannot accept risk on a customer's behalf. Getting this right is mostly about finding the person who can sign, and giving them something worth signing. ## The three options, priced **Global allow.** Costs nothing to run and breaks nothing. What it buys the adversary is precise: anything with a bespoke protocol lands in this bucket automatically, so the rule is a hole shaped exactly like a purpose-built channel — and one hole shared by every tenant on the stack. It is also invisible: nothing alerts, the bucket simply grows, and nobody reads it. **Global deny.** A defensible posture, and it will drop applications you cannot enumerate because you do not own any tenant's inventory. The failures arrive as "the transfer to our partner stopped working", days later, from people who did not know a classification engine sat in the path. On a shared boundary the blast radius is every tenant simultaneously, which is the worst property a change can have. **Per-tenant defaults.** The correct answer and the expensive one: a policy set per tenant, an owner per tenant, and a review that recurs. Be ready to say the cost out loud, because a principal answer that pretends the right thing is free is not credible. ## Sequencing, which is most of the real answer 1. **New tenants start at deny.** Onboarding is the only moment where somebody on the customer side is actively testing and expects things to break. A deny introduced during acceptance costs a conversation; the same deny introduced two years later costs an incident. 2. **Existing tenants get a log-only phase first.** Run the deny in report mode and produce, per tenant, the list of endpoint pairs and volumes it *would* have dropped. This is the artefact that makes the conversation possible: it converts "we would like to tighten a rule" into "here are the eleven flows this affects, do you recognise them?" 3. **Work the list down before enforcement.** Each recognised flow becomes a scoped definition or an allow keyed to that endpoint pair — never a blanket allow — with an owner and a review date. 4. **Migrate on a date the tenant agreed**, with a rollback that is one change and can be made by whoever is on call, because the first enforcement window is when you discover what the log-only phase missed. 5. **Whatever is left at the deadline is signed for.** A tenant who wants allow-unclassified keeps it as a written, dated, expiring exception against their name, not as the platform's default. ## Who signs, and what you own The tenant's named security contact signs the risk. The provider owns the **evidence and the reporting**, not the decision: what is in the bucket, how it changed, what a deny would drop, when the exception expires. That split is what keeps the arrangement honest — it means a provider is never in the position of having quietly chosen a customer's posture for them, and it means a tenant who declines to engage still has a dated record of having been asked. If no tenant will engage at all, that is a gap in the service description rather than an engineering failure, and it should be escalated as one. The worst outcome is the common one: the default was set years ago by whoever built the stack, nobody has a name against it, and the bucket has been growing ever since with no reader. ## The tell of a weak answer Picking deny and calling it best practice, with no mention of who absorbs the breakage, no rollout order and no rollback. The security posture is the easy half. The half that is actually being tested is whether you can impose it on ten organisations who never agreed to it, and know which of those ten can say no.

  • The stack cannot hold ten separate policy sets. What do you do instead?
    Then the shared default has to be the safer one, and per-tenant tolerance is carried as scoped allows keyed to specific endpoint pairs rather than as a shared permissive default. A shared allow is a shared blast radius; a shared deny with narrow, owned exceptions at least fails in a direction you can enumerate and report.
  • A tenant refuses to sign either the allow or the deny. How do you resolve it?
    You do not resolve it technically. The contract default applies, the non-response is recorded and reported on the same cycle as everything else, and it is escalated as a service-description gap. What you must not do is quietly pick the permissive option because it generates no complaints — that is choosing their posture for them without telling them.
  • How do you show a tenant the cost of enforcement before enforcement day?
    The log-only report: run the deny in report mode and hand them the endpoint pairs, volumes and schedules it would have dropped. It turns an abstract policy request into a specific list they can recognise, and the flows they cannot explain are the ones worth talking about first.
  • Why is a global allow worse than it looks even though nothing breaks?
    Because nothing breaking is exactly the point of failure. Anything with a bespoke protocol is unclassified by construction, so the rule permits precisely the traffic that was purpose-built, and it does so silently across every tenant on the stack. There is no alert to read and no complaint to trigger a review, so the hole persists by default.

saying these in an interview costs you the question

  • Picks deny globally as best practice with no rollout or rollback
  • Assumes the provider can accept risk on the tenant's behalf
  • Treats a permissive default as safe because nothing breaks
  • Offers one shared policy without pricing per-tenant policy sets
  • Grants exceptions with no named owner and no expiry

context