Your services already get retries, timeouts and circuit breaking from an in-process resilience library, and public TLS terminates at a shared edge gateway. What does adding a service mesh give you that neither of those does, and what does the library still do better?
answer
- three places policy can live
- north-south versus east-west traffic
- the proxy sees bytes, not intent
- one language versus every language
- fallback needs the call site
basics
~20 sA service mesh applies retries, timeouts, per-hop identity and uniform telemetry to service-to-service traffic through proxies beside each workload, in any language and without code changes. An in-process library still sees call intent — fallbacks, semantics — that a proxy cannot.
solid answer
~50 sThree layers can own cross-cutting traffic concerns, and they cover different traffic. The edge gateway is north-south: it authenticates external clients, terminates the public certificate and applies quotas, but it never sees a call from one internal service to another. A resilience library lives at the call site, so it can do what needs intent — return a cached fallback, bulkhead its own thread pools, know that a charge is not safe to repeat — but you maintain one implementation per language and every policy change is a redeploy of every service. A mesh puts a proxy on both ends of each internal hop, so retries, timeouts, outlier ejection, workload identity and consistent L7 metrics apply uniformly, including to off-the-shelf workloads nobody can rebuild. You pay in extra hops, per-pod resource and a distributed system to operate. Most platforms end up running all three.
go deeper
Be able to say where each piece sits: the gateway at the front door for outside traffic, the library inside your code, the mesh proxy beside each service for internal calls.
Explain concretely what moves when policy leaves the library: same behaviour for every language, changed without a redeploy, applied per hop — and name what stays behind, such as fallback values and in-process pools.
Show judgment about double-retrying, non-idempotent writes and the added hop, and be ready to say when a library plus an edge proxy is genuinely enough for the estate in front of you.
Own the placement rule for the whole platform: which concern lives at which layer, who is allowed to change it, and how you stop three layers from each retrying the same failing call.
## The question behind the question Every distributed system has to solve the same handful of problems on every network call: find an instance, balance across instances, retry the ones that fail, give up after a deadline, authenticate the peer, and emit metrics and traces. The only real design decision is **where that code runs**. There are three answers, and an interviewer is checking whether you understand that they cover different traffic and different failure modes rather than being three brands of the same thing. ## Layer one: the in-process library A resilience library runs inside your application, at the call site. Because it is in the same process, it can see things no network device can see: which method is calling, what a sensible default value would be, whether this particular operation is safe to repeat. ```java // The library sits at the call site, so it can substitute a value on failure. Supplier<Quote> guarded = CircuitBreaker.decorateSupplier(breaker, quotes::fetch); Quote q = Try.ofSupplier(guarded).recover(t -> Quote.cachedOrDefault()).get(); ``` That `recover` branch is the point. A proxy can fail a request; it cannot decide that a stale price from cache is an acceptable answer. Nor can it partition your own thread pools, because it has none of them. The costs are equally concrete: the library exists once per language, so a polyglot estate reimplements it two or three times and the implementations drift; and because it is compiled into the application, changing a timeout means a code change, a build, and a deploy of every service — sometimes dozens of teams' calendars. ## Layer two: the shared edge gateway A gateway is a reverse proxy at the front door with authentication, quotas and routing bolted on. It is the natural home for anything that concerns **external** clients: verifying tokens, terminating the public certificate, throttling per customer, presenting one hostname over many services. Its limit is structural. It sits on the north-south path only. When service A calls service B inside the cluster, that traffic never touches the gateway, so the gateway cannot secure it, retry it, or tell you anything about it. Routing internal calls back out through the edge to gain those properties is a well-known anti-pattern: it doubles latency, makes one box the failure domain for all internal traffic, and still gives you no per-hop identity. ## Layer three: the mesh A service mesh is a **data plane** of proxies deployed next to every workload plus a **control plane** that programs them. Traffic is intercepted transparently — the application dials the ordinary service address and the platform redirects the connection into the local proxy — so enrolment requires no application change. What that buys, specifically: - **Uniformity across languages.** The same retry budget, timeout, outlier ejection and load-balancing algorithm apply to a Java service, a Python service and a vendor image you cannot rebuild. - **Per-hop workload identity.** Each proxy holds a cryptographic identity for its workload, so calls between services are mutually authenticated and authorisable per source workload rather than per network address. - **Consistent east-west telemetry.** Every hop produces the same latency, error and traffic metrics with the same labels, because one implementation emits them. - **Change without redeploy.** Policy is data pushed to proxies, so the platform team can adjust it without touching application code. ## What the mesh cannot do The proxy sees connections and requests, not intent. It does not know whether a POST is idempotent, so mesh-level retries on write paths are a foot-gun unless the application marks them safe. It cannot produce a fallback value. It cannot isolate resources inside your process. And it cannot invent context your application does not emit — if your code drops the trace or deadline header, the proxy has nothing to propagate. It also is not free. Each internal call now crosses two extra proxies, adding latency to every hop and to the tail in particular; each pod carries a proxy's memory and CPU; and your request path now contains a component that can generate its own errors, which means new failure modes to learn and a fleet to keep upgraded. ## How the decision usually lands Small, single-language estates get most of the value from a library and an edge gateway, and a mesh mostly adds operational surface. The arguments that actually justify a mesh are polyglot or unmodifiable workloads, a requirement for encrypted and authorised service-to-service traffic that you cannot get every team to implement, and a platform team that wants to change traffic policy without a fleet-wide code change. Even then, the honest answer to "which one?" is usually **all three, each on the traffic it is actually on**: gateway at the edge, mesh between services, and a thin library for the decisions that need to happen at the call site.
- Does adopting a mesh mean you can retire your API gateway?No. The gateway is on the north-south path and owns external concerns: authenticating untrusted clients, terminating the public certificate, per-customer quotas, exposing one hostname over many services. The mesh covers east-west calls that never reach the edge. Most meshes even implement their own ingress as another proxy, which is a gateway again — so you are choosing who configures it, not removing it.
- If both the library and the mesh retry the same call, what happens?The attempts multiply. Three library attempts each retried three times by the proxy is up to nine requests reaching a dependency that is already failing, and every intermediate hop multiplies again. Pick one owner per hop: usually the mesh for transport-level failures, with the library capped at one, plus a retry budget so retries stay a small percentage of live traffic.
- Which resilience behaviours can a mesh proxy never implement for you?Anything requiring knowledge of the call: returning a cached or default value instead of an error, degrading a page to a partial render, bulkheading your own thread pools or connection pools, and deciding that a specific write is safe to repeat. The proxy can also only forward context your application emits — it cannot originate a deadline or trace your code never set.
saying these in an interview costs you the question
- Claims a service mesh replaces the edge API gateway
- Says the mesh gives fallbacks and bulkheads for free
- Assumes mesh retries are safe on non-idempotent writes
- Calls a mesh zero-cost because no application code changes
- Thinks a per-language library can cover vendor workloads you cannot rebuild