skip to content

You own a shared Istio mesh used by dozens of teams and want to move it to identity-based default-deny authorization. How would you sequence that work, and what do Istio's policy scoping rules and PERMISSIVE mode mean for the plan?

level: principalimportance: nice to knowfreq 32%

answer

  1. ordering is forced by dependencies
  2. identity is only as fine as the accounts
  3. one prerequisite makes rules enforceable
  4. shadow the decision before enforcing it
  5. one action cannot be overridden downstream

basics

~20 s

Do it in order: give every workload its own service account, reach STRICT peer authentication, observe real callers from telemetry, shadow the rules with dry-run, then apply namespace deny-all with team-owned grants. Identity rules mean nothing while plaintext can still arrive.

solid answer

~50 s

The sequence is forced by dependencies, not preference. **Service accounts first** — mesh identity is per service account, so workloads sharing one are indistinguishable and no policy can separate them. **STRICT next**, because a plaintext caller carries no peer identity: a principal-based rule cannot match it, but any rule written in terms of paths and methods alone will admit it, so authorization written under PERMISSIVE is advisory. **Then observe** — derive each namespace's real caller set from destination-side telemetry over a full business cycle rather than from architecture diagrams. **Then shadow** the intended policies using the dry-run annotation, which records what would have been denied without enforcing it. **Then enforce**, namespace by namespace, deny-all last. On scoping: policies accumulate across mesh, namespace and workload levels, and DENY is evaluated before ALLOW with no exception mechanism — so a mesh-wide DENY is unoverridable by any team beneath it and should be reserved for genuine invariants. Give teams ALLOW policies in their own namespaces as the normal tool.

go deeper

for a junior

Understand that turning on default-deny is a migration with prerequisites, and that the first thing to establish is which callers actually exist rather than which ones the design says exist.

for a middle

Be able to name the ordering and why each step depends on the last, and explain that authorization policies accumulate across scopes while peer authentication policies replace one another.

for a senior

Show how you would use destination-side telemetry and shadow evaluation to make each stage measurable, and how you would give teams a recognisable failure signature and a one-resource rollback.

for a principal

Own the division of ownership — platform holds STRICT, mesh-wide invariants and deny-all defaults, teams hold their own grants — and defend a clear line between coarse mesh authorization and object-level decisions that stay in the application.

## Why the order is not negotiable Each step in this migration is worthless without the one before it, which is what makes the sequencing the actual answer rather than the policies themselves. **1. Service-account hygiene.** Mesh identity is `<trust domain>/ns/<namespace>/sa/<service account>`. If five Deployments in a namespace all run under the namespace default account, they are one identity and no policy can distinguish them. This is unglamorous, touches every team, and must land before anyone writes a rule that names a principal — otherwise you build an authorization model on top of an identity model that cannot express the boundaries you are drawing. **2. STRICT peer authentication.** This is the step most often deferred and it invalidates everything downstream. Under PERMISSIVE, a plaintext request arrives with **no peer identity at all**. A rule matching `source.principals` therefore cannot match it — so far so good. But a rule written only in terms of `to.operation` paths and methods has nothing to check the caller against, and admits the anonymous plaintext request happily. In a large mesh, some of your teams will write rules of exactly that shape. Identity-based authorization is enforceable only once plaintext cannot arrive, which makes STRICT a prerequisite of the security model rather than a hardening extra. **3. Observe the real call graph.** The intended call graph and the actual one differ, and the difference is where outages come from: the reporting job nobody documented, the support tool, another team's direct call. Destination-side request telemetry, grouped by source workload and namespace, gives you the real edges by name. Observe across a full business cycle so weekly and monthly jobs appear. **4. Shadow before enforcing.** Istio supports an experimental dry-run mode: annotate an AuthorizationPolicy with `istio.io/dry-run: "true"` and the proxy evaluates it and reports what the decision *would* have been, without changing what actually happens. That converts a risky flip into a measurable one — you enforce when the shadow denial count for legitimate traffic has been zero for long enough to trust it. The `AUDIT` action serves a related purpose for recording matches. **5. Enforce, narrow to broad.** Turn on per-namespace grants first and add the namespace deny-all last, so that at the moment default-deny takes effect the allow-list is already proven. Do it namespace by namespace; a mesh-wide flip removes your ability to learn cheaply from the first failure. ## What scoping does and does not let you do Authorization policies **accumulate** across the mesh root namespace, the namespace and workload selectors — unlike peer authentication, where the most specific level replaces the others. Every policy selecting a workload participates in one evaluation: CUSTOM, then DENY, then ALLOW. The organisational consequence is the DENY action. There are no priorities and no exception mechanism, so a mesh-wide DENY cannot be carved out by any team's ALLOW policy. That makes it the right tool for a small number of genuine invariants — an administrative path that must never be reachable from outside a namespace — and the wrong tool for everything else, because every future exception becomes a change request to the platform team. The healthy division is: **the platform owns STRICT, the mesh-wide invariants and the deny-all defaults; teams own the ALLOW policies in their own namespaces.** Teams can then express restrictions by narrowing their own grants rather than by asking for a DENY. ## What you still owe the teams - **A failure signature they recognise.** An authorization denial is a 403 with the body `RBAC: access denied`, produced by the proxy, invisible in application logs. If teams do not know that string, every rollout stage generates escalations to you. - **A rollback that is one resource.** Deleting or reverting a policy takes effect as soon as the control plane pushes it. Say so explicitly, because the perceived risk of the change is usually higher than the real one. - **A template.** Most teams need the same two policies — namespace deny-all, plus grants naming principals and paths. Shipping a reviewed template prevents the path-and-method-only rules that would have been unenforceable anyway. ## Where to stop Mesh authorization is coarse: caller identity, method, path, host, port, claims. It cannot express "this user may read this account", because that requires data the proxy does not have. The boundary worth defending is that the mesh enforces **which service, and which end user, may reach which endpoint**, and the application keeps object-level decisions. Pushing further pulls business logic into infrastructure configuration that no test suite covers; stopping short leaves the application trusting its network again.

  • Why can identity-based authorization not be rolled out before STRICT peer authentication?
    Because a plaintext caller has no peer identity. Principal-based rules simply fail to match it, but rules written only over paths and methods will admit it, so the policy set is advisory rather than enforced. In a large mesh some teams will inevitably write the second kind. STRICT is what makes the whole layer meaningful.
  • Why should a platform team be sparing with mesh-wide DENY policies?
    Because DENY is evaluated before ALLOW and has no exception mechanism, so nothing a team writes in their own namespace can carve out a legitimate case. Every exception becomes a platform change request. Reserve DENY for invariants that should never have exceptions, and let teams express restrictions by narrowing their own ALLOW policies.
  • How do authorization policy scoping and peer authentication scoping differ?
    Peer authentication levels replace each other — the most specific matching policy wins and the broader one is ignored. Authorization policies accumulate: every policy selecting the workload, at any level, participates in the same CUSTOM, DENY, ALLOW evaluation. Assuming the replacement rule applies to both is a common and consequential mistake.

saying these in an interview costs you the question

  • Writes authorization policies while the mesh is still permissive
  • Rolls out a mesh-wide deny-all as the first enforced step
  • Assumes a team's ALLOW can override a mesh-wide DENY
  • Leaves workloads sharing one service account before writing principal rules
  • Validates the call graph from diagrams instead of telemetry

context