skip to content

Istio

Istio is the service mesh most teams mean when they say "mesh": a control plane (istiod) that programs Envoy proxies next to every workload so routing, mTLS and telemetry stop being application code. Interviewers ask about it because it is the place where Kubernetes networking, TLS identity and release control all meet, and because a candidate who has actually run it can say what it costs.

on this pageshow

explore

questions

18

In Istio, what do the PeerAuthentication modes STRICT, PERMISSIVE and DISABLE each do to inbound traffic for the workloads they select, and why can you enable mutual TLS without changing application code?

level: juniorimportance: must knowfreq 72%

answer

  1. about inbound connections, not outbound
  2. three modes plus an inherit value
  3. one mode accepts both at once
  4. the migration mode is the middle one
  5. workload beats namespace beats mesh

basics

~20 s

Istio's PeerAuthentication sets what a workload's sidecar accepts on inbound connections: STRICT requires mutual TLS, PERMISSIVE accepts either mutual TLS or plaintext, DISABLE expects plaintext. The proxy performs the TLS, so application code is untouched.

solid answer

~40 s

`PeerAuthentication` is the Istio CRD that configures **inbound** peer authentication for the workloads it selects. `STRICT` means the sidecar only accepts connections that present a valid mesh client certificate; anything plaintext is rejected at the transport layer. `PERMISSIVE` accepts both on the same port — the proxy sniffs the connection and handles mTLS or plaintext accordingly, which is what makes it the migration mode. `DISABLE` turns peer authentication off and expects plaintext. No application change is needed because the sidecar terminates the mesh TLS and forwards a plain connection to the app over loopback, and on the client side the sidecar originates mTLS automatically when the destination is known to support it. Scope follows a precedence chain: a workload-selector policy beats a namespace-wide one, which beats a mesh-wide policy in the root namespace.

code

yaml · 23 lines
yaml
apiVersion: security.istio.io/v1
kind: PeerAuthentication
metadata:
  name: default
  namespace: payments
spec:
  mtls:
    mode: STRICT
---
apiVersion: security.istio.io/v1
kind: PeerAuthentication
metadata:
  name: metrics-exception
  namespace: payments
spec:
  selector:
    matchLabels:
      app: ledger
  mtls:
    mode: STRICT
  portLevelMtls:
    9090:
      mode: DISABLE

go deeper

for a junior

Be able to name the three modes and say in one sentence what each does to an inbound connection, and state plainly that the sidecar does the TLS so application code is unchanged.

for a middle

Explain that the resource is inbound-only, that PERMISSIVE accepts both on the same port by inspecting the connection, and how the workload/namespace/mesh precedence chain resolves.

for a senior

Show you would not flip a namespace to STRICT without first proving from telemetry that nothing plaintext still calls it, and that you know port-level exceptions exist for the cases that cannot move.

for a principal

Own the question of who sets the mesh-wide default and who may override it per namespace, and treat PERMISSIVE as a time-boxed migration state with an owner and an end date rather than a standing posture.

## What the resource actually configures `PeerAuthentication` is a `security.istio.io` CRD that answers one narrow question: **when a connection arrives at this workload's proxy, what transport-level authentication is required?** "Peer" means the calling workload, as opposed to the end user — end-user credentials are a different resource entirely. A minimal namespace-wide policy looks like this: ```yaml apiVersion: security.istio.io/v1 kind: PeerAuthentication metadata: name: default namespace: payments spec: mtls: mode: STRICT ``` Adding `spec.selector.matchLabels` narrows it to specific workloads; omitting the selector applies it to every workload in the namespace. ## The three modes **STRICT** — the proxy's inbound listener only completes a handshake with a peer that presents a valid mesh-issued client certificate. A plaintext TCP connection is closed; the caller typically sees a connection reset rather than an HTTP status, because the rejection happens below HTTP. **PERMISSIVE** — the proxy accepts both on the same port. It inspects the beginning of the connection and routes it either through the mTLS transport path or through a plaintext path. This exists precisely so that a mesh can be turned on incrementally: workloads already in the mesh start talking mTLS to each other, while callers that have not been onboarded keep working. It is the default posture for a fresh installation, and it is a **migration state, not a destination** — a service in PERMISSIVE mode is still reachable by anything that can route a packet to the pod. **DISABLE** — no peer authentication; the sidecar expects plaintext. Real uses are narrow: a port that speaks a protocol the proxy should not wrap, or an endpoint deliberately exposed to a non-mesh scraper. `UNSET` also exists and means "inherit from the next level up". There is also `spec.portLevelMtls`, which overrides the mode for individual ports of the selected workloads — the usual reason a single port needs to stay plaintext without downgrading the whole workload. ## Precedence Three scopes exist, and the most specific wins: 1. **Workload** — a policy with a `selector`, in the workload's own namespace. 2. **Namespace** — a policy with no selector, in that namespace. 3. **Mesh-wide** — a policy with no selector in the mesh's *root namespace* (`istio-system` in a default install). They do not merge or accumulate. A namespace policy replaces the mesh default for that namespace; a workload policy replaces the namespace policy for that workload. This is a frequent source of confusion: someone sets a mesh-wide STRICT policy and is surprised that one namespace still accepts plaintext, because a namespace-level policy is quietly shadowing it. ## Why the application never sees a certificate The proxy is in the network path for the pod, so the mesh TLS session begins and ends at the two proxies. The receiving proxy terminates it and hands the request to the application over the pod's loopback interface as an ordinary plaintext connection. The application keeps listening on a plain socket, its client libraries keep making plain calls, and nothing in the code base learns about keys, chains or handshakes. That is the entire selling point of transport security in the mesh: it is a property of the deployment, not of the code. The consequence people miss is the direction of control. **PeerAuthentication is inbound only.** It says nothing about how this workload calls others. Outbound is decided on the client side: since automatic mutual TLS was introduced, a sidecar originates mTLS by itself when it knows the destination workload has a proxy, and that behaviour is overridden through a `DestinationRule` traffic policy rather than through `PeerAuthentication`. So a policy applied "to service A" protects calls *into* A, not calls *out of* A. ## Where it stops Mutual TLS gives you an authenticated, encrypted channel and a verified peer identity. It does not decide who is *allowed* to call what — every authenticated workload in the mesh can still reach every other one until an authorization policy says otherwise. Turning on STRICT and calling the job done is the single most common misreading of what this resource buys you.

  • If PeerAuthentication only governs inbound traffic, what decides whether this workload's calls out to others use mutual TLS?
    The client side decides. With automatic mutual TLS, the calling sidecar originates mTLS whenever it knows the destination has a proxy, so no configuration is normally needed. To force or forbid it for a destination you set the TLS mode in that destination's traffic policy, not in a PeerAuthentication resource.
  • You have a mesh-wide STRICT policy, but one namespace still accepts plaintext. What is the most likely explanation?
    A more specific policy is shadowing it — either a namespace-level PeerAuthentication in that namespace or a workload-selector policy. The levels do not merge; the most specific matching policy replaces the broader one entirely. Check for policies in that namespace before assuming the mesh-wide one is broken.
  • Why would you ever set portLevelMtls to DISABLE on one port of an otherwise STRICT workload?
    Because something outside the mesh must reach that port directly — a scraper, a health endpoint, or a protocol the proxy should not wrap. It scopes the exception to one port instead of downgrading the whole workload, which keeps the blast radius small and makes the exception visible in the manifest.

PERMISSIVE is the door that still accepts the old brass key while everyone is being issued badges; STRICT is the day the lock is changed and only badges work.

saying these in an interview costs you the question

  • Thinks PeerAuthentication configures outbound TLS to upstreams
  • Says PERMISSIVE encrypts only some of the traffic randomly
  • Believes the application must be rebuilt with TLS libraries
  • Thinks STRICT mode also restricts which services may call
  • Assumes mesh-wide and namespace policies merge together

context

open as a page

After Istio injection, a Kubernetes pod that declares a single application container shows READY 2/2. What is the second container, and how does the application's traffic end up flowing through it without any change to the application code?

level: juniorimportance: must knowfreq 70%

basics

~20 s

The second container is istio-proxy, an Envoy-based sidecar added by Istio's injection webhook when the pod is created. Injected rules in the pod's own network namespace redirect the pod's inbound and outbound TCP traffic through that proxy, so application code never changes.

open as a page

You apply your first Istio AuthorizationPolicy with action ALLOW to one workload, and calls that previously worked start returning 403. Explain how Istio evaluates AuthorizationPolicy, including how the CUSTOM, DENY and ALLOW actions relate.

level: middleimportance: must knowfreq 65%

basics

~20 s

Once any ALLOW policy selects a workload, that workload becomes default-deny: anything not matched by an ALLOW rule is rejected with 403. Evaluation runs CUSTOM first, then DENY, then ALLOW, and a DENY match wins outright.

open as a page

A newly created pod in a namespace you believe is part of the Istio mesh came up with no istio-proxy container. Explain how Istio decides which pods get a sidecar, and how you would find the reason this one did not.

level: middleimportance: must knowfreq 62%

basics

~20 s

Istio injects through a mutating admission webhook selected by namespace labels — istio-injection=enabled or istio.io/rev=<revision> — with the pod-level sidecar.istio.io/inject label as an override. Check those labels, whether the pod predates them, revision mismatch after an upgrade, host-network pods, and webhook availability.

open as a page

You add an Istio VirtualService that splits traffic 90/10 between subsets named v1 and v2 of a service, and every request to that service immediately starts returning 503. Which resource actually defines those subsets, and what are the two usual causes of the 503?

level: middleimportance: must knowfreq 74%

basics

~20 s

Subsets are defined in a DestinationRule, not in the VirtualService. A weighted split fails with 503 either because no DestinationRule defines those subset names, or because a subset's label selector matches no running pods, leaving it with zero endpoints.

open as a page

An Istio VirtualService for a slow service sets `timeout: 10s` together with `retries: {attempts: 3, perTryTimeout: 5s}`. Explain how those two settings interact, and why a chain of services each configured this way can turn a slowdown into an outage.

level: seniorimportance: must knowfreq 56%

basics

~20 s

The VirtualService timeout is the total budget for the whole request including every retry, while perTryTimeout caps one attempt. Three attempts of five seconds cannot all run inside a ten-second budget, and retries at every hop multiply, so a slow dependency becomes many times its normal load.

open as a page

An Istio VirtualService lists three entries under spec.http: one matching `uri: prefix: /api/v2`, one matching a request header, and one with no `match` block at all. In what order does the proxy evaluate them, and what happens if the entry with no match block is written first?

level: juniorimportance: should knowfreq 60%

basics

~20 s

Istio evaluates a VirtualService's http entries top to bottom and takes the first whose match conditions succeed. An entry with no match block matches every request, so writing it first makes every rule below it unreachable.

open as a page

In an Istio mesh, where does a workload's mutual-TLS certificate come from, what identity does that certificate carry, and how is it kept fresh?

level: middleimportance: should knowfreq 50%

basics

~20 s

The agent in each pod generates a key and certificate request, authenticates to istiod's built-in CA with the pod's service-account token, and receives a short-lived certificate whose SPIFFE identity encodes the trust domain, namespace and service account. The agent rotates it automatically.

open as a page

Istio can program a pod's traffic-capture rules either with an injected `istio-init` container or with the Istio CNI plugin. What does each approach do, and what changes about pod privileges and pod startup when you switch to the CNI plugin?

level: middleimportance: should knowfreq 34%

basics

~20 s

The istio-init container writes the pod's redirection rules from inside the pod, which requires NET_ADMIN and NET_RAW on every workload. The Istio CNI plugin writes the same rules from a node-level DaemonSet during pod network setup, so workload pods need no elevated capabilities.

open as a page

An Istio VirtualService is applied with no `gateways` field set. Which traffic does it govern, and what has to change for that same VirtualService to route requests arriving at an Istio ingress Gateway?

level: middleimportance: should knowfreq 58%

basics

~20 s

With no gateways field, a VirtualService defaults to the reserved value mesh, so it applies only to traffic between sidecars inside the mesh. To take effect at the edge it must list the Istio Gateway resource by name, and its hosts must overlap the hostnames that Gateway serves.

open as a page

An Istio RequestAuthentication resource is applied to a workload with a jwtRules entry for your identity provider, yet requests carrying no token at all still reach the application. Why, and what makes the endpoint actually require a valid token?

level: seniorimportance: should knowfreq 45%

basics

~20 s

RequestAuthentication only defines how to validate a token if one is present; a request with no token is not rejected. To require one, add an AuthorizationPolicy whose rule matches source.requestPrincipals with a wildcard, so unauthenticated requests fail to match any grant.

open as a page

A team flips their namespace's Istio PeerAuthentication from PERMISSIVE to STRICT and part of their inbound traffic immediately starts failing with connection resets. What are the likely causes, and how would you have verified the namespace was ready beforehand?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Something is still calling those workloads in plaintext — a caller with no sidecar, a scraper hitting the pod directly, or a client outside the mesh. Istio's own telemetry records whether each inbound request used mutual TLS, and that is what you check before flipping.

open as a page

An Istio-injected service logs connection failures on its first few outbound calls immediately after each pod starts, then behaves normally. What ordering problem causes this, and what does Istio offer to fix it?

level: seniorimportance: should knowfreq 48%

basics

~20 s

The application container starts before istio-proxy is ready, so early outbound calls hit redirection rules pointing at a proxy that cannot yet serve them. Fix it with holdApplicationUntilProxyStarts, or by running istio-proxy as a Kubernetes native sidecar container, which Kubernetes starts first.

open as a page

An Istio VirtualService http route can carry a `mirror` destination alongside its normal route. What exactly does the proxy send to the mirror target and what does it do with the response, and why is mirroring live production traffic risky?

level: seniorimportance: should knowfreq 42%

basics

~20 s

The proxy sends a fire-and-forget copy of the request to the mirror target with -shadow appended to the Host header, and discards the response entirely, so the client is unaffected. The risk is that the mirrored copy is a real request: it repeats writes and side effects on whatever the mirror target touches.

open as a page

Istio's ambient mode replaces per-pod sidecars with a per-node ztunnel and optional per-namespace waypoint proxies. For an existing sidecar-based Istio mesh, how would you decide whether to move to ambient, and what are you giving up?

level: principalimportance: should knowfreq 38%

basics

~20 s

Decide by what your mesh is actually used for. Ambient's per-node ztunnel gives mTLS and L4 authorization without a proxy in every pod, removing per-pod overhead and mesh-wide upgrade restarts. You give up per-pod isolation and accept a node-level shared component, plus a waypoint hop wherever you need L7.

open as a page

A workload inside an Istio mesh calls an external HTTPS API at api.vendor.example.com. What does adding a ServiceEntry for that host change, and how does the mesh's outboundTrafficPolicy mode affect whether the call works at all?

level: middleimportance: nice to knowfreq 40%

basics

~20 s

A ServiceEntry adds an external host to the mesh's service registry, so VirtualService and DestinationRule rules, timeouts, retries and proper telemetry apply to it. Whether the call works without one depends on outboundTrafficPolicy: ALLOW_ANY passes unknown hosts through, REGISTRY_ONLY blocks them.

open as a page

In a large Istio mesh, each injected istio-proxy's memory footprint and configuration-push time grow as unrelated services are added elsewhere in the cluster. Why does that happen, and what does Istio's `Sidecar` resource do about it?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

By default istiod sends every sidecar the configuration for every service in the mesh, so per-proxy memory and push cost scale with mesh size rather than with a workload's actual dependencies. A Sidecar resource narrows each proxy's visible scope to the namespaces and hosts it really calls.

open as a page

You own a shared Istio mesh used by dozens of teams and want to move it to identity-based default-deny authorization. How would you sequence that work, and what do Istio's policy scoping rules and PERMISSIVE mode mean for the plan?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Do it in order: give every workload its own service account, reach STRICT peer authentication, observe real callers from telemetry, shadow the rules with dry-run, then apply namespace deny-all with team-owned grants. Identity rules mean nothing while plaintext can still arrive.

open as a page