A team flips their namespace's Istio PeerAuthentication from PERMISSIVE to STRICT and part of their inbound traffic immediately starts failing with connection resets. What are the likely causes, and how would you have verified the namespace was ready beforehand?
answer
- a reset, not a status code
- the failure is below HTTP
- the callers you cannot see are the problem
- destination-side telemetry names them
- weekly jobs do not appear in an afternoon
basics
~20 sSomething is still calling those workloads in plaintext — a caller with no sidecar, a scraper hitting the pod directly, or a client outside the mesh. Istio's own telemetry records whether each inbound request used mutual TLS, and that is what you check before flipping.
solid answer
~50 sSTRICT rejects plaintext below HTTP, so the caller sees a reset rather than a status code — that alone tells you the failure is at the transport layer. The causes are all one shape: a client that is not in the mesh. Common ones are workloads in a namespace without proxy injection, jobs or CronJobs that were never injected, monitoring scraping application ports directly rather than through the mesh's metrics endpoint, and anything reaching the pod IP from outside the cluster. Prevention is telemetry, not guesswork: Istio's request metrics on the **destination** side carry a label recording whether the connection was mutual TLS or plaintext, so you group inbound requests to the namespace by that label and by source workload, and you watch it long enough to catch weekly batch traffic — not for an afternoon. Then you move workload by workload, or use a port-level exception for anything that genuinely cannot move.
go deeper
Recognise that STRICT rejects plaintext callers outright, and that a caller with no sidecar is the usual reason something that worked yesterday now fails to connect.
Explain that the rejection happens during the handshake, so it surfaces as a reset rather than an HTTP error, and name the classes of caller that have no certificate.
Show the verification method — destination-side request telemetry grouped by connection security and source workload, observed across a full business cycle — and sequence the change with port-level exceptions and a fast rollback.
Own why the mesh cannot be left permissive indefinitely: identity-based authorization is unenforceable while plaintext can still arrive, so STRICT is a prerequisite for the security model rather than an optional hardening step.
## Read the failure mode first The symptom is diagnostic. STRICT peer authentication is enforced when the TLS handshake happens, which is before any HTTP request exists. So a rejected caller gets a connection reset or a handshake failure — never a 403, never an application error page. If your callers are seeing 403 with the body `RBAC: access denied`, you are looking at authorization, not peer authentication, and the fix is somewhere else entirely. Getting this distinction right in the first minute saves an hour. ## Who is still speaking plaintext Every cause reduces to "a client without a mesh certificate", but the specific instances are worth knowing because they are the ones that survive a casual review: - **Namespaces without injection.** A caller in a namespace that was never labelled for sidecar injection has no proxy and no certificate. Its calls worked fine under PERMISSIVE and die under STRICT. - **Jobs, CronJobs and one-off tooling.** Short-lived workloads are routinely deployed outside the pattern that gets everything else injected, and — critically — a weekly job produces no traffic on the day you look at your dashboard. - **Monitoring scraping application ports.** A Prometheus outside the mesh that scrapes pod ports directly is a plaintext client. The mesh's own answer is to scrape the merged metrics endpoint the agent exposes on port 15020 instead, or to bring the scraper into the mesh. - **Anything reaching the pod IP directly.** Traffic that bypasses the Service, a debugging port-forward, a legacy component on a VM, another cluster. - **A policy shadowing the one you changed.** Peer authentication levels replace rather than merge, so a stray workload-level policy can leave one workload still permissive — or, in the opposite direction, a mesh-wide STRICT you did not know about can be what actually took effect. ## The verification you should have run Istio's standard request metric is recorded on both sides of every call, and the destination-side series carries a label stating the transport security of the connection — mutual TLS or none. That label is the whole answer. Query inbound requests to the target namespace, grouped by that label and by the source workload and namespace, and you get a precise list of who would break, by name. Two disciplines make the difference between a query and a safe migration: **Watch long enough.** Interactive traffic shows up in minutes; the nightly reconciliation job and the monthly report generator do not. A migration window of a full business cycle — a week at minimum — is what turns "no plaintext observed" into evidence rather than luck. **Look at the destination side.** Source-side metrics tell you what your workloads sent, not what unknown clients sent to them. Only the receiving proxy sees every inbound connection, including the ones from things you did not know existed. ## Sequencing the change With the list in hand, the migration is unremarkable: 1. Onboard the callers that can be onboarded — label their namespaces, get them injected, confirm their traffic moves to mutual TLS in the same metric. 2. For anything that genuinely cannot join, use `portLevelMtls` to exempt the specific port rather than leaving the whole workload permissive. The exception is then visible in the manifest and has a small blast radius. 3. Flip STRICT at the smallest scope that makes sense — a workload selector first, then the namespace — rather than mesh-wide in one step. 4. Keep the previous policy ready to reapply. Rollback is a single resource change and takes effect within seconds of the control plane pushing it, which makes this one of the more forgiving changes to attempt during business hours. ## The reason not to stall in PERMISSIVE It is tempting to conclude that PERMISSIVE is fine and STRICT is optional. It is not, and the reason is authorization rather than encryption. A plaintext caller has no peer identity, so identity-based authorization rules cannot match it — but any rule you wrote in terms of paths and methods alone will cheerfully admit it. Every principal-based policy in the namespace is conditional on plaintext being impossible. Until STRICT is on, the authorization layer is an assertion about well-behaved callers rather than an enforced boundary.
- How would you tell from the symptom alone whether the failure is peer authentication or an authorization policy?By what the caller receives. STRICT peer authentication rejects during the TLS handshake, so the client sees a connection reset or handshake failure with no HTTP response at all. An authorization denial happens after the request is parsed and returns HTTP 403 with the body `RBAC: access denied`. Different layer, different fix.
- A monitoring system scrapes application ports directly and would break under STRICT. What are your options?Three, in order of preference: scrape the merged metrics endpoint the agent exposes on port 15020, which is the mesh's intended path; bring the scraper into the mesh so it presents a certificate; or exempt that single port with portLevelMtls set to DISABLE. The last keeps the exception explicit and narrow rather than downgrading the workload.
- Why is it not enough to check source-side mutual TLS metrics before flipping to STRICT?Because source-side series only cover callers that already have a proxy reporting metrics — precisely the ones that are fine. The clients that will break are the ones with no sidecar, so they appear in nobody's source metrics. Only the receiving proxy observes every inbound connection, which is why the destination-side series is the one that matters.
saying these in an interview costs you the question
- Expects a 403 rather than a connection reset
- Checks only source-side mutual TLS metrics
- Observes traffic for an hour and declares the namespace ready
- Flips the whole mesh at once instead of per workload
- Concludes PERMISSIVE is a fine permanent posture