skip to content

Traffic Management CRDs

This is Istio's routing language: VirtualService says which requests go where, DestinationRule defines the subsets and per-connection policy they land on, and Gateway/ServiceEntry mark the mesh's edges. Interviewers ask because weighted routing, fault injection, timeouts, retries and outlier detection are the concrete mechanisms behind every canary or resilience story, and candidates routinely mix up which CRD owns which field.

on this pageshow

questions

6

You add an Istio VirtualService that splits traffic 90/10 between subsets named v1 and v2 of a service, and every request to that service immediately starts returning 503. Which resource actually defines those subsets, and what are the two usual causes of the 503?

level: middleimportance: must knowfreq 74%

answer

  1. two resources, two jobs
  2. weights here, subset definitions there
  3. labels are matched against pods
  4. a destination with no endpoints

basics

~20 s

Subsets are defined in a DestinationRule, not in the VirtualService. A weighted split fails with 503 either because no DestinationRule defines those subset names, or because a subset's label selector matches no running pods, leaving it with zero endpoints.

solid answer

~50 s

The two resources divide the work: the **VirtualService** says *which* requests go where and in what proportion, and the **DestinationRule** defines *what* a subset actually is — a name plus a label selector over the service's pods, optionally with its own `trafficPolicy`. Routing to `subset: v2` with no DestinationRule declaring `v2` produces a route pointing at a destination the proxy has never been told about, and the request fails with 503. The second cause is subtler: the DestinationRule exists, but the subset's `labels` (say `version: v2`) match no pod, so the destination resolves with zero endpoints and you get 503 with no healthy upstream. Both are label-and-name problems rather than routing problems, and `istioctl proxy-config cluster` plus `istioctl analyze` tell them apart in seconds — the first shows no destination at all, the second shows one with an empty endpoint list.

code

bash · 10 lines
bash
# 1. Does the caller's proxy know a destination for subset v2 at all?
istioctl proxy-config cluster deploy/productpage \
  --fqdn reviews.default.svc.cluster.local

# 2. If it does, does that destination have any endpoints?
istioctl proxy-config endpoint deploy/productpage \
  --cluster "outbound|9080|v2|reviews.default.svc.cluster.local"

# 3. Catch a VirtualService referencing an undefined subset
istioctl analyze -n default

go deeper

for a junior

Know that a weighted split needs two resources: the VirtualService carries the weights, and the DestinationRule declares each subset name with the pod labels it selects.

for a middle

Explain both failure modes precisely — an undefined subset name versus a label selector matching no pod — and say why the second leaves a destination that exists but has no endpoints.

for a senior

Demonstrate the diagnosis path: check whether the caller's proxy has a destination for the subset at all, then whether it has endpoints, and narrow the cause to a name, a namespace, an exportTo scope, or pod labels.

for a principal

Own the convention that makes this class of outage rare: a single agreed version label across all Deployments, subsets and rules created together, and validation in the pipeline so a VirtualService referencing an undefined subset never reaches a cluster.

## The division of labour between the two resources Candidates lose this question by assuming one resource does everything. It does not: - **`VirtualService`** — request-level routing. Which requests match, which destination they go to, the `weight` of each destination, plus `timeout`, `retries`, `fault`, `mirror`, `rewrite`. - **`DestinationRule`** — what happens *at* a destination. It names the `host`, declares `subsets` (each a name plus `labels` selecting pods), and carries `trafficPolicy`: `loadBalancer`, `connectionPool`, `outlierDetection`, `tls`. A weighted release therefore always needs **both**. The VirtualService alone is a routing table with entries pointing nowhere. ```yaml apiVersion: networking.istio.io/v1beta1 kind: DestinationRule metadata: name: reviews spec: host: reviews subsets: - name: v1 labels: version: v1 - name: v2 labels: version: v2 ``` ```yaml apiVersion: networking.istio.io/v1beta1 kind: VirtualService metadata: name: reviews spec: hosts: - reviews http: - route: - destination: {host: reviews, subset: v1} weight: 90 - destination: {host: reviews, subset: v2} weight: 10 ``` ## Cause one: the subset name does not exist If no DestinationRule declares `v2` for that host — it was never applied, it landed in the wrong namespace, or its `host` field is spelled differently from the one in the VirtualService — the route references a destination the proxy has no definition for. The request cannot be dispatched and fails with **503**. There is no partial success: because both subsets normally come from the same DestinationRule, a missing or mis-scoped rule usually breaks 100% of traffic, not the 10% you were canarying. That total-failure symptom is a useful signal in itself. Watch the `host` value in particular. Inside a Kubernetes mesh it is resolved relative to the resource's own namespace, so a DestinationRule in namespace `payments` with `host: reviews` refers to `reviews.payments.svc.cluster.local`. If the workload actually lives elsewhere, use the fully qualified name. Cross-namespace consumption is also governed by `exportTo` — a DestinationRule exported only to its own namespace does not exist as far as a caller in another namespace is concerned. ## Cause two: the labels select nothing Here the DestinationRule is correct and the subset exists, but `labels: {version: v2}` matches no running pod. Perhaps the Deployment's pod template labels it `app.kubernetes.io/version: v2`, or `track: v2`, or the v2 pods are not up yet. The destination exists with an **empty endpoint list**, and requests routed to it return 503 with an upstream-unavailable message. The rule to internalise: subset labels are matched against **pod** labels, which come from the Deployment's `spec.template.metadata.labels` — not against the Deployment's own labels and not against Service selectors. A pod must also be a valid endpoint of the Service the host names; a subset cannot select pods the Service does not already back. ## Telling the two apart quickly ```bash istioctl proxy-config cluster deploy/productpage --fqdn reviews.default.svc.cluster.local istioctl proxy-config endpoint deploy/productpage --cluster "outbound|9080|v2|reviews.default.svc.cluster.local" istioctl analyze -n default ``` The first command lists the destinations the caller's proxy knows about, one per subset. **No line for `v2` means cause one.** A line for `v2` whose endpoint listing is empty means **cause two**. `istioctl analyze` catches the common version of cause one directly, reporting a VirtualService that references an undefined subset. ## Why the failure is so abrupt A weighted split is not evaluated leniently. There is no fallback to the other subset when one is undefined, and no health-based skipping of an empty destination — those are different mechanisms with different fields. This is deliberate: silently sending your 10% canary traffic to v1 would make a broken release look successful. Failing loudly is the safer default, and it is why the fix is always in the DestinationRule or the pod labels rather than in the weights. (A third, less common cause of 503 on a newly added DestinationRule lives in its `tls` settings conflicting with the namespace's peer authentication mode. That is mesh security configuration rather than routing, and it is diagnosed with different tools.) ## Per-subset policy Once subsets exist, each may carry its own `trafficPolicy`, which overrides the DestinationRule's top-level one for that subset only. That is how you give a canary tighter connection-pool limits or a different load-balancing setting than the stable version without touching the stable path at all.

  • Where would you put a connection-pool limit that applies only to the canary subset and not to the stable one?
    In the DestinationRule, under `subsets[].trafficPolicy` for that subset. A subset-level `trafficPolicy` overrides the rule's top-level one for that subset alone, so the canary can carry tighter `connectionPool` limits or its own `outlierDetection` while the stable subset keeps the defaults.
  • A DestinationRule with the right subsets exists, but only callers in its own namespace can use it. Why?
    `exportTo` scopes the resource's visibility. A DestinationRule exported only to `.` is invisible to callers elsewhere, so their proxies never learn the subsets and their routes fail exactly as if the rule were missing. The same applies to VirtualService and ServiceEntry.
  • Why does Istio fail the request outright instead of falling back to the other subset in the split?
    Because silent fallback would make a broken canary look healthy — your 10% would quietly be served by v1 and the release would appear to pass. Failing loudly forces the misconfiguration into the open. Skipping unhealthy backends is a separate mechanism with its own fields, not an implicit routing fallback.

saying these in an interview costs you the question

  • Thinks subsets are declared in the VirtualService
  • Expects traffic to fall back to the other subset automatically
  • Matches subset labels against Deployment or Service labels
  • Assumes only the 10% canary share fails
  • Adds a Kubernetes Service per version instead of subsets

context

open as a page

An Istio VirtualService for a slow service sets `timeout: 10s` together with `retries: {attempts: 3, perTryTimeout: 5s}`. Explain how those two settings interact, and why a chain of services each configured this way can turn a slowdown into an outage.

level: seniorimportance: must knowfreq 56%

basics

~20 s

The VirtualService timeout is the total budget for the whole request including every retry, while perTryTimeout caps one attempt. Three attempts of five seconds cannot all run inside a ten-second budget, and retries at every hop multiply, so a slow dependency becomes many times its normal load.

open as a page

An Istio VirtualService lists three entries under spec.http: one matching `uri: prefix: /api/v2`, one matching a request header, and one with no `match` block at all. In what order does the proxy evaluate them, and what happens if the entry with no match block is written first?

level: juniorimportance: should knowfreq 60%

basics

~20 s

Istio evaluates a VirtualService's http entries top to bottom and takes the first whose match conditions succeed. An entry with no match block matches every request, so writing it first makes every rule below it unreachable.

open as a page

An Istio VirtualService is applied with no `gateways` field set. Which traffic does it govern, and what has to change for that same VirtualService to route requests arriving at an Istio ingress Gateway?

level: middleimportance: should knowfreq 58%

basics

~20 s

With no gateways field, a VirtualService defaults to the reserved value mesh, so it applies only to traffic between sidecars inside the mesh. To take effect at the edge it must list the Istio Gateway resource by name, and its hosts must overlap the hostnames that Gateway serves.

open as a page

An Istio VirtualService http route can carry a `mirror` destination alongside its normal route. What exactly does the proxy send to the mirror target and what does it do with the response, and why is mirroring live production traffic risky?

level: seniorimportance: should knowfreq 42%

basics

~20 s

The proxy sends a fire-and-forget copy of the request to the mirror target with -shadow appended to the Host header, and discards the response entirely, so the client is unaffected. The risk is that the mirrored copy is a real request: it repeats writes and side effects on whatever the mirror target touches.

open as a page

A workload inside an Istio mesh calls an external HTTPS API at api.vendor.example.com. What does adding a ServiceEntry for that host change, and how does the mesh's outboundTrafficPolicy mode affect whether the call works at all?

level: middleimportance: nice to knowfreq 40%

basics

~20 s

A ServiceEntry adds an external host to the mesh's service registry, so VirtualService and DestinationRule rules, timeouts, retries and proper telemetry apply to it. Whether the call works without one depends on outboundTrafficPolicy: ALLOW_ANY passes unknown hosts through, REGISTRY_ONLY blocks them.

open as a page