You add an Istio VirtualService that splits traffic 90/10 between subsets named v1 and v2 of a service, and every request to that service immediately starts returning 503. Which resource actually defines those subsets, and what are the two usual causes of the 503?
answer
- two resources, two jobs
- weights here, subset definitions there
- labels are matched against pods
- a destination with no endpoints
basics
~20 sSubsets are defined in a DestinationRule, not in the VirtualService. A weighted split fails with 503 either because no DestinationRule defines those subset names, or because a subset's label selector matches no running pods, leaving it with zero endpoints.
solid answer
~50 sThe two resources divide the work: the **VirtualService** says *which* requests go where and in what proportion, and the **DestinationRule** defines *what* a subset actually is — a name plus a label selector over the service's pods, optionally with its own `trafficPolicy`. Routing to `subset: v2` with no DestinationRule declaring `v2` produces a route pointing at a destination the proxy has never been told about, and the request fails with 503. The second cause is subtler: the DestinationRule exists, but the subset's `labels` (say `version: v2`) match no pod, so the destination resolves with zero endpoints and you get 503 with no healthy upstream. Both are label-and-name problems rather than routing problems, and `istioctl proxy-config cluster` plus `istioctl analyze` tell them apart in seconds — the first shows no destination at all, the second shows one with an empty endpoint list.
code
bash · 10 lines# 1. Does the caller's proxy know a destination for subset v2 at all?
istioctl proxy-config cluster deploy/productpage \
--fqdn reviews.default.svc.cluster.local
# 2. If it does, does that destination have any endpoints?
istioctl proxy-config endpoint deploy/productpage \
--cluster "outbound|9080|v2|reviews.default.svc.cluster.local"
# 3. Catch a VirtualService referencing an undefined subset
istioctl analyze -n defaultgo deeper
Know that a weighted split needs two resources: the VirtualService carries the weights, and the DestinationRule declares each subset name with the pod labels it selects.
Explain both failure modes precisely — an undefined subset name versus a label selector matching no pod — and say why the second leaves a destination that exists but has no endpoints.
Demonstrate the diagnosis path: check whether the caller's proxy has a destination for the subset at all, then whether it has endpoints, and narrow the cause to a name, a namespace, an exportTo scope, or pod labels.
Own the convention that makes this class of outage rare: a single agreed version label across all Deployments, subsets and rules created together, and validation in the pipeline so a VirtualService referencing an undefined subset never reaches a cluster.
## The division of labour between the two resources Candidates lose this question by assuming one resource does everything. It does not: - **`VirtualService`** — request-level routing. Which requests match, which destination they go to, the `weight` of each destination, plus `timeout`, `retries`, `fault`, `mirror`, `rewrite`. - **`DestinationRule`** — what happens *at* a destination. It names the `host`, declares `subsets` (each a name plus `labels` selecting pods), and carries `trafficPolicy`: `loadBalancer`, `connectionPool`, `outlierDetection`, `tls`. A weighted release therefore always needs **both**. The VirtualService alone is a routing table with entries pointing nowhere. ```yaml apiVersion: networking.istio.io/v1beta1 kind: DestinationRule metadata: name: reviews spec: host: reviews subsets: - name: v1 labels: version: v1 - name: v2 labels: version: v2 ``` ```yaml apiVersion: networking.istio.io/v1beta1 kind: VirtualService metadata: name: reviews spec: hosts: - reviews http: - route: - destination: {host: reviews, subset: v1} weight: 90 - destination: {host: reviews, subset: v2} weight: 10 ``` ## Cause one: the subset name does not exist If no DestinationRule declares `v2` for that host — it was never applied, it landed in the wrong namespace, or its `host` field is spelled differently from the one in the VirtualService — the route references a destination the proxy has no definition for. The request cannot be dispatched and fails with **503**. There is no partial success: because both subsets normally come from the same DestinationRule, a missing or mis-scoped rule usually breaks 100% of traffic, not the 10% you were canarying. That total-failure symptom is a useful signal in itself. Watch the `host` value in particular. Inside a Kubernetes mesh it is resolved relative to the resource's own namespace, so a DestinationRule in namespace `payments` with `host: reviews` refers to `reviews.payments.svc.cluster.local`. If the workload actually lives elsewhere, use the fully qualified name. Cross-namespace consumption is also governed by `exportTo` — a DestinationRule exported only to its own namespace does not exist as far as a caller in another namespace is concerned. ## Cause two: the labels select nothing Here the DestinationRule is correct and the subset exists, but `labels: {version: v2}` matches no running pod. Perhaps the Deployment's pod template labels it `app.kubernetes.io/version: v2`, or `track: v2`, or the v2 pods are not up yet. The destination exists with an **empty endpoint list**, and requests routed to it return 503 with an upstream-unavailable message. The rule to internalise: subset labels are matched against **pod** labels, which come from the Deployment's `spec.template.metadata.labels` — not against the Deployment's own labels and not against Service selectors. A pod must also be a valid endpoint of the Service the host names; a subset cannot select pods the Service does not already back. ## Telling the two apart quickly ```bash istioctl proxy-config cluster deploy/productpage --fqdn reviews.default.svc.cluster.local istioctl proxy-config endpoint deploy/productpage --cluster "outbound|9080|v2|reviews.default.svc.cluster.local" istioctl analyze -n default ``` The first command lists the destinations the caller's proxy knows about, one per subset. **No line for `v2` means cause one.** A line for `v2` whose endpoint listing is empty means **cause two**. `istioctl analyze` catches the common version of cause one directly, reporting a VirtualService that references an undefined subset. ## Why the failure is so abrupt A weighted split is not evaluated leniently. There is no fallback to the other subset when one is undefined, and no health-based skipping of an empty destination — those are different mechanisms with different fields. This is deliberate: silently sending your 10% canary traffic to v1 would make a broken release look successful. Failing loudly is the safer default, and it is why the fix is always in the DestinationRule or the pod labels rather than in the weights. (A third, less common cause of 503 on a newly added DestinationRule lives in its `tls` settings conflicting with the namespace's peer authentication mode. That is mesh security configuration rather than routing, and it is diagnosed with different tools.) ## Per-subset policy Once subsets exist, each may carry its own `trafficPolicy`, which overrides the DestinationRule's top-level one for that subset only. That is how you give a canary tighter connection-pool limits or a different load-balancing setting than the stable version without touching the stable path at all.
- Where would you put a connection-pool limit that applies only to the canary subset and not to the stable one?In the DestinationRule, under `subsets[].trafficPolicy` for that subset. A subset-level `trafficPolicy` overrides the rule's top-level one for that subset alone, so the canary can carry tighter `connectionPool` limits or its own `outlierDetection` while the stable subset keeps the defaults.
- A DestinationRule with the right subsets exists, but only callers in its own namespace can use it. Why?`exportTo` scopes the resource's visibility. A DestinationRule exported only to `.` is invisible to callers elsewhere, so their proxies never learn the subsets and their routes fail exactly as if the rule were missing. The same applies to VirtualService and ServiceEntry.
- Why does Istio fail the request outright instead of falling back to the other subset in the split?Because silent fallback would make a broken canary look healthy — your 10% would quietly be served by v1 and the release would appear to pass. Failing loudly forces the misconfiguration into the open. Skipping unhealthy backends is a separate mechanism with its own fields, not an implicit routing fallback.
saying these in an interview costs you the question
- Thinks subsets are declared in the VirtualService
- Expects traffic to fall back to the other subset automatically
- Matches subset labels against Deployment or Service labels
- Assumes only the 10% canary share fails
- Adds a Kubernetes Service per version instead of subsets