skip to content

Your platform runs Linkerd today, and a team asks for a mesh capability Linkerd does not have. How do you decide between working around it, putting the capability somewhere else in the stack, or migrating the mesh to Istio?

level: principalimportance: must knowfreq 50%

answer

  1. capability, not CRD name
  2. edge concerns belong in the ingress tier
  3. name the genuine structural gaps
  4. price the dual-mesh window
  5. you chose the small mesh for a reason

basics

~20 s

Start from what is genuinely missing. Push edge and gateway concerns into an ingress tier, keep Linkerd for in-cluster mTLS and metrics, and treat migration to Istio as justified only by structural gaps such as non-Kubernetes workloads or per-request token policy.

solid answer

~50 s

I would separate three categories. First, requests that are really *edge* concerns — rich routing, WAF, external TLS — belong to an ingress controller or gateway, not the mesh; Linkerd deliberately does not do that job and adding Istio to get it is buying a mesh to solve a gateway problem. Second, requests Linkerd can meet in its own idiom — retries and timeouts, per-route metrics, mTLS-identity-based authorization — are a documentation problem, not a platform problem. Third, the genuine structural gaps: workloads outside Kubernetes, per-request JWT authentication and fine-grained request-level authorization, arbitrary L7 filters via Wasm, or a sidecar-less data plane. Only those justify a migration, because swapping meshes means re-injecting every pod, rebuilding all policy, retraining every on-call engineer, and running two meshes during transition. I would also weigh the reverse cost — Istio's larger operational surface is the thing you chose Linkerd to avoid.

go deeper

for a junior

Know that Linkerd and Istio solve the same problem with different scope, and that switching a mesh affects every pod in the cluster rather than being a configuration change.

for a middle

Be able to say what Linkerd covers well — automatic mTLS, golden metrics, retries and timeouts — and to recognise when a request is really about ingress behaviour rather than about the mesh.

for a senior

Demonstrate that you would look for the capability's proper home before changing platforms, and that you can price a migration concretely: fleet-wide restarts, policy rewrite, invalidated runbooks, dual-mesh traffic.

for a principal

Own the decision rule and the organisational cost. Weigh durability and breadth of the requirement, the operational surface you deliberately chose, distribution and support terms, and whether a scoped second mesh beats a full migration.

## Frame the question correctly first The common failure here is answering a feature-comparison question when you were asked a platform-strategy question. "Istio has X and Linkerd doesn't" is not a decision; it is an input. The decision is about where a capability should live and what changing the mesh costs the organisation. Start by classifying the request. ## Category 1: it is an edge concern wearing a mesh costume A large share of "we need Istio" requests are really requests for gateway behaviour: host and path routing with rich matching, request rewriting, external TLS certificates, WAF rules, authentication of end users at the perimeter, rate limiting by client. Linkerd's deliberate position is that this belongs to an ingress controller in front of the mesh, not to the sidecars. So the answer is to put the capability in the ingress tier — nginx, an Envoy-based ingress controller, or a cloud load balancer — and mesh that controller so traffic joins the mesh at the door. This is the highest-value move available to you, because it satisfies the requirement without touching the mesh at all. ## Category 2: Linkerd can do it, differently Some requests arrive phrased in Istio's vocabulary because that is what the team read about. Retries and timeouts, per-route golden metrics, weighted traffic splitting for a canary, service-to-service authorization based on workload identity — Linkerd has answers for all of these; they just are not called `VirtualService` or `DestinationRule`. Here the work is documentation and enablement, not migration. Say plainly: "you are asking for capability, not for a CRD name." ## Category 3: the real structural gaps These are the ones that would actually justify a mesh change, and you should be able to name them without hedging: - **Workloads outside Kubernetes.** Linkerd is a Kubernetes mesh. If you must mesh VMs or bare-metal services into the same trust and policy domain, that is a structural gap, not a configuration exercise. - **Per-request identity policy.** Validating end-user JWTs at the proxy and authorizing on claims is a different model from authorizing on workload mTLS identity. Linkerd's policy model is centred on workload identity and network scope; if you need per-request end-user policy in the data plane, look hard at whether it belongs at the edge instead — and if it genuinely must be everywhere, that is a real gap. - **Arbitrary L7 extensibility.** Wasm or Lua filters, protocol translation, bespoke in-proxy logic. Linkerd offers no extension point by design. - **Data-plane model.** If the requirement is a sidecar-less data plane for cost or lifecycle reasons, that is an architectural property, not a feature you can bolt on. ## Costing the migration honestly If you land in category three, price it before recommending it: - Every meshed pod must be re-injected and restarted — a fleet-wide rolling operation. - All mesh policy must be rewritten in a different model, and the two models do not map one-to-one. - Both meshes coexist during the transition, and cross-mesh traffic during that window is the hardest part of the plan. - Dashboards, alerts, runbooks and on-call knowledge are all mesh-specific and all get invalidated. - The steady-state operational surface grows. You very likely chose Linkerd because a small, boring mesh had fewer ways to fail at 3am; a migration spends that deliberately. A credible answer proposes a decision rule rather than a verdict: *does more than one team need the missing capability, is there no acceptable home for it outside the mesh, and is the requirement durable?* One team, one quarter, one workaround available means no migration. ## The alternatives to a full swap - **Put a second proxy in the path** for the one workload that needs the exotic behaviour — an Envoy-based gateway in front of that service — while the rest of the fleet stays on Linkerd. Two proxies on one path is ugly, but it is scoped and reversible. - **Move the concern into the application** where it is genuinely business logic that drifted into the platform. - **Run Istio only in the cluster or namespace that needs it**, accepting a heterogeneous platform. Costly in cognitive load, but occasionally the honest answer when one estate has genuinely different requirements. ## Also weigh the non-technical inputs At principal level, distribution and support terms are part of the decision: who publishes the builds you run, how upgrades are supported, what the release cadence and support window are for each project, and whether your organisation needs a commercial backing arrangement. These terms have changed for mesh projects before, so check the current state rather than reciting what was true a couple of years ago. The same applies to feature claims on both sides — data-plane models in this space move quickly, so verify against current releases before you commit a platform decision to a slide. ## What a strong answer sounds like "Tell me the capability, not the CRD. If it is edge behaviour it goes in the ingress tier. If Linkerd expresses it differently, we document it. If it is a structural gap that several teams need durably, I will cost a migration — fleet-wide re-injection, policy rewrite, dual-mesh window, retraining — and compare that against scoping Istio to the one estate that needs it."

  • A team says they need Istio because they want header-based routing to a canary. Is that a valid reason to migrate?
    No. Weighted and header-scoped routing for a canary is available in the ingress tier and, in its own idiom, in Linkerd. Wanting a specific CRD is not a capability gap. I would implement the canary where the traffic decision already happens — usually at ingress — and keep the mesh out of the release-tooling conversation.
  • What would make you accept running both meshes permanently rather than migrating?
    A genuinely separated estate: one cluster or business unit with requirements the other does not share — non-Kubernetes workloads, or regulatory per-request policy — plus separate operational ownership. Permanent heterogeneity is defensible when the boundary is organisational as well as technical; it is not defensible as the residue of an abandoned migration.
  • How would you sequence a Linkerd-to-Istio migration if you decided it was necessary?
    Namespace by namespace, never big-bang. Stand up the new mesh alongside, pick a low-risk namespace with few cross-namespace callers, define how traffic crosses the boundary during the overlap, migrate and observe with both meshes' metrics, then repeat. Keep a rollback path per namespace, and treat the dual-mesh window as the risk to minimise rather than a state to tolerate.
  • How do you stop this decision from being made on feature-list comparisons?
    Insist on a written requirement stated as behaviour under production conditions — what must be true for which traffic, and what breaks if it is not. Feature matrices compare products; requirements compare against your system. Most rows on the matrix turn out to be things you would never enable.

saying these in an interview costs you the question

  • Recommends migrating because the other mesh has more features
  • Treats ingress and gateway needs as mesh requirements
  • Ignores that every pod must be re-injected and restarted
  • Forgets the dual-mesh transition window entirely
  • Compares feature matrices instead of stated requirements

context