skip to content

mTLS & Authorization Policy

Istio issues every workload a short-lived SPIFFE identity from istiod's CA and can make mutual TLS the default between pods without touching application code, then gate calls with AuthorizationPolicy. Interviewers ask because this is the clearest example of identity-based service-to-service authz in practice — and because the PERMISSIVE-to-STRICT migration, and what happens to traffic that skips the proxy, are where real deployments break.

on this pageshow

questions

6

In Istio, what do the PeerAuthentication modes STRICT, PERMISSIVE and DISABLE each do to inbound traffic for the workloads they select, and why can you enable mutual TLS without changing application code?

level: juniorimportance: must knowfreq 72%

answer

  1. about inbound connections, not outbound
  2. three modes plus an inherit value
  3. one mode accepts both at once
  4. the migration mode is the middle one
  5. workload beats namespace beats mesh

basics

~20 s

Istio's PeerAuthentication sets what a workload's sidecar accepts on inbound connections: STRICT requires mutual TLS, PERMISSIVE accepts either mutual TLS or plaintext, DISABLE expects plaintext. The proxy performs the TLS, so application code is untouched.

solid answer

~40 s

`PeerAuthentication` is the Istio CRD that configures **inbound** peer authentication for the workloads it selects. `STRICT` means the sidecar only accepts connections that present a valid mesh client certificate; anything plaintext is rejected at the transport layer. `PERMISSIVE` accepts both on the same port — the proxy sniffs the connection and handles mTLS or plaintext accordingly, which is what makes it the migration mode. `DISABLE` turns peer authentication off and expects plaintext. No application change is needed because the sidecar terminates the mesh TLS and forwards a plain connection to the app over loopback, and on the client side the sidecar originates mTLS automatically when the destination is known to support it. Scope follows a precedence chain: a workload-selector policy beats a namespace-wide one, which beats a mesh-wide policy in the root namespace.

code

yaml · 23 lines
yaml
apiVersion: security.istio.io/v1
kind: PeerAuthentication
metadata:
  name: default
  namespace: payments
spec:
  mtls:
    mode: STRICT
---
apiVersion: security.istio.io/v1
kind: PeerAuthentication
metadata:
  name: metrics-exception
  namespace: payments
spec:
  selector:
    matchLabels:
      app: ledger
  mtls:
    mode: STRICT
  portLevelMtls:
    9090:
      mode: DISABLE

go deeper

for a junior

Be able to name the three modes and say in one sentence what each does to an inbound connection, and state plainly that the sidecar does the TLS so application code is unchanged.

for a middle

Explain that the resource is inbound-only, that PERMISSIVE accepts both on the same port by inspecting the connection, and how the workload/namespace/mesh precedence chain resolves.

for a senior

Show you would not flip a namespace to STRICT without first proving from telemetry that nothing plaintext still calls it, and that you know port-level exceptions exist for the cases that cannot move.

for a principal

Own the question of who sets the mesh-wide default and who may override it per namespace, and treat PERMISSIVE as a time-boxed migration state with an owner and an end date rather than a standing posture.

## What the resource actually configures `PeerAuthentication` is a `security.istio.io` CRD that answers one narrow question: **when a connection arrives at this workload's proxy, what transport-level authentication is required?** "Peer" means the calling workload, as opposed to the end user — end-user credentials are a different resource entirely. A minimal namespace-wide policy looks like this: ```yaml apiVersion: security.istio.io/v1 kind: PeerAuthentication metadata: name: default namespace: payments spec: mtls: mode: STRICT ``` Adding `spec.selector.matchLabels` narrows it to specific workloads; omitting the selector applies it to every workload in the namespace. ## The three modes **STRICT** — the proxy's inbound listener only completes a handshake with a peer that presents a valid mesh-issued client certificate. A plaintext TCP connection is closed; the caller typically sees a connection reset rather than an HTTP status, because the rejection happens below HTTP. **PERMISSIVE** — the proxy accepts both on the same port. It inspects the beginning of the connection and routes it either through the mTLS transport path or through a plaintext path. This exists precisely so that a mesh can be turned on incrementally: workloads already in the mesh start talking mTLS to each other, while callers that have not been onboarded keep working. It is the default posture for a fresh installation, and it is a **migration state, not a destination** — a service in PERMISSIVE mode is still reachable by anything that can route a packet to the pod. **DISABLE** — no peer authentication; the sidecar expects plaintext. Real uses are narrow: a port that speaks a protocol the proxy should not wrap, or an endpoint deliberately exposed to a non-mesh scraper. `UNSET` also exists and means "inherit from the next level up". There is also `spec.portLevelMtls`, which overrides the mode for individual ports of the selected workloads — the usual reason a single port needs to stay plaintext without downgrading the whole workload. ## Precedence Three scopes exist, and the most specific wins: 1. **Workload** — a policy with a `selector`, in the workload's own namespace. 2. **Namespace** — a policy with no selector, in that namespace. 3. **Mesh-wide** — a policy with no selector in the mesh's *root namespace* (`istio-system` in a default install). They do not merge or accumulate. A namespace policy replaces the mesh default for that namespace; a workload policy replaces the namespace policy for that workload. This is a frequent source of confusion: someone sets a mesh-wide STRICT policy and is surprised that one namespace still accepts plaintext, because a namespace-level policy is quietly shadowing it. ## Why the application never sees a certificate The proxy is in the network path for the pod, so the mesh TLS session begins and ends at the two proxies. The receiving proxy terminates it and hands the request to the application over the pod's loopback interface as an ordinary plaintext connection. The application keeps listening on a plain socket, its client libraries keep making plain calls, and nothing in the code base learns about keys, chains or handshakes. That is the entire selling point of transport security in the mesh: it is a property of the deployment, not of the code. The consequence people miss is the direction of control. **PeerAuthentication is inbound only.** It says nothing about how this workload calls others. Outbound is decided on the client side: since automatic mutual TLS was introduced, a sidecar originates mTLS by itself when it knows the destination workload has a proxy, and that behaviour is overridden through a `DestinationRule` traffic policy rather than through `PeerAuthentication`. So a policy applied "to service A" protects calls *into* A, not calls *out of* A. ## Where it stops Mutual TLS gives you an authenticated, encrypted channel and a verified peer identity. It does not decide who is *allowed* to call what — every authenticated workload in the mesh can still reach every other one until an authorization policy says otherwise. Turning on STRICT and calling the job done is the single most common misreading of what this resource buys you.

  • If PeerAuthentication only governs inbound traffic, what decides whether this workload's calls out to others use mutual TLS?
    The client side decides. With automatic mutual TLS, the calling sidecar originates mTLS whenever it knows the destination has a proxy, so no configuration is normally needed. To force or forbid it for a destination you set the TLS mode in that destination's traffic policy, not in a PeerAuthentication resource.
  • You have a mesh-wide STRICT policy, but one namespace still accepts plaintext. What is the most likely explanation?
    A more specific policy is shadowing it — either a namespace-level PeerAuthentication in that namespace or a workload-selector policy. The levels do not merge; the most specific matching policy replaces the broader one entirely. Check for policies in that namespace before assuming the mesh-wide one is broken.
  • Why would you ever set portLevelMtls to DISABLE on one port of an otherwise STRICT workload?
    Because something outside the mesh must reach that port directly — a scraper, a health endpoint, or a protocol the proxy should not wrap. It scopes the exception to one port instead of downgrading the whole workload, which keeps the blast radius small and makes the exception visible in the manifest.

PERMISSIVE is the door that still accepts the old brass key while everyone is being issued badges; STRICT is the day the lock is changed and only badges work.

saying these in an interview costs you the question

  • Thinks PeerAuthentication configures outbound TLS to upstreams
  • Says PERMISSIVE encrypts only some of the traffic randomly
  • Believes the application must be rebuilt with TLS libraries
  • Thinks STRICT mode also restricts which services may call
  • Assumes mesh-wide and namespace policies merge together

context

open as a page

You apply your first Istio AuthorizationPolicy with action ALLOW to one workload, and calls that previously worked start returning 403. Explain how Istio evaluates AuthorizationPolicy, including how the CUSTOM, DENY and ALLOW actions relate.

level: middleimportance: must knowfreq 65%

basics

~20 s

Once any ALLOW policy selects a workload, that workload becomes default-deny: anything not matched by an ALLOW rule is rejected with 403. Evaluation runs CUSTOM first, then DENY, then ALLOW, and a DENY match wins outright.

open as a page

In an Istio mesh, where does a workload's mutual-TLS certificate come from, what identity does that certificate carry, and how is it kept fresh?

level: middleimportance: should knowfreq 50%

basics

~20 s

The agent in each pod generates a key and certificate request, authenticates to istiod's built-in CA with the pod's service-account token, and receives a short-lived certificate whose SPIFFE identity encodes the trust domain, namespace and service account. The agent rotates it automatically.

open as a page

An Istio RequestAuthentication resource is applied to a workload with a jwtRules entry for your identity provider, yet requests carrying no token at all still reach the application. Why, and what makes the endpoint actually require a valid token?

level: seniorimportance: should knowfreq 45%

basics

~20 s

RequestAuthentication only defines how to validate a token if one is present; a request with no token is not rejected. To require one, add an AuthorizationPolicy whose rule matches source.requestPrincipals with a wildcard, so unauthenticated requests fail to match any grant.

open as a page

A team flips their namespace's Istio PeerAuthentication from PERMISSIVE to STRICT and part of their inbound traffic immediately starts failing with connection resets. What are the likely causes, and how would you have verified the namespace was ready beforehand?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Something is still calling those workloads in plaintext — a caller with no sidecar, a scraper hitting the pod directly, or a client outside the mesh. Istio's own telemetry records whether each inbound request used mutual TLS, and that is what you check before flipping.

open as a page

You own a shared Istio mesh used by dozens of teams and want to move it to identity-based default-deny authorization. How would you sequence that work, and what do Istio's policy scoping rules and PERMISSIVE mode mean for the plan?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Do it in order: give every workload its own service account, reach STRICT peer authentication, observe real callers from telemetry, shadow the rules with dry-run, then apply namespace deny-all with team-owned grants. Identity rules mean nothing while plaintext can still arrive.

open as a page