skip to content

Across an organisation, would you put gRPC balancing in every client library or terminate every connection in a sidecar data plane?

level: principalimportance: should knowfreq 42%

answer

  1. not really about the extra hop
  2. count the implementations, not the calls
  3. how does a change reach everyone
  4. a process per workload has a bill
  5. both run at once for a while

basics

~20 s

Client-side balancing costs no extra hop but duplicates the policy in every language and ships only when callers redeploy. A sidecar centralises the policy at the price of a hop, a process per workload, and a new failure domain.

solid answer

~50 s

Both put the decision per call; they differ in where the decision lives and who can change it. **In the client library**: no extra hop, no extra process, and the balancer sees the endpoints directly — but the policy exists once per language, versions drift between them, and changing it means redeploying every caller. **In a sidecar data plane**: one implementation, one place to change policy, uniform per-call telemetry and transport security, and callers that need no balancing code at all — but you now run a process alongside every workload, pay a local hop per call, and own its upgrades and its failures. The honest answer names the decider: how many languages you have, how often policy changes, and who is accountable for the rollout. Either way the caller still owns its deadlines and whether a failed call may be repeated.

go deeper

for a junior

Know that balancing can happen inside the calling process or in a separate local process, and that both end up choosing an endpoint per call.

for a middle

Contrast the concrete costs: an extra hop and a process per workload against one policy implementation per language and connections held by every caller.

for a senior

Argue from the operational side — who diagnoses a failure that now sits between caller and dependency, and how connection counts change when the pool grows.

for a principal

Decide on language count, policy change rate and accountability for rollout, then plan the overlap: both designs running at once, version skew, capacity for the sum, and a config-only rollback.

## The decision is about ownership, not latency Both designs balance per call, and both work. Arguments that reduce to microseconds of extra hop miss what is actually being chosen: **where the policy lives, and who can change it without a deploy**. That is an organisational property, and it is why this is a question for whoever owns the platform rather than for the team writing one service. ## Client-side balancing The caller resolves the pool, holds a subchannel per endpoint and picks one per call. What it buys: - no extra process and no extra hop; the connection goes from the caller to the worker; - the balancer's view of endpoint state is first-hand — it sees connections fail, and it can consume the endpoints' health signals directly; - nothing extra to operate per workload. What it costs: - the policy is implemented once per language, and those implementations differ in defaults, in backoff, and in how faithfully they follow a config; - a policy change reaches production only as fast as your slowest caller redeploys, which for a long tail of services is measured in quarters; - every caller holds a connection to every endpoint, so connection counts multiply with the product of callers and pool size; - getting the configuration to the callers is itself a distribution problem — the resolver and its service config become platform infrastructure whether you planned them or not. ## A sidecar data plane The caller connects to a local process; that process holds the connections to the workers and balances per call. What it buys: - one implementation of the policy for every language in the estate; - policy changes without recompiling or redeploying callers; - one place to enforce transport security and emit per-call telemetry in a shape that is the same everywhere; - callers that carry no balancing configuration at all, which makes a caller written in an unusual language a non-event. What it costs: - a process next to every workload, with its memory, its CPU and its own upgrade campaign across the fleet; - a local hop per call, plus the sidecar's own queuing under load; - a new failure domain sitting between every caller and every dependency, and one whose failures look to the caller like the dependency's; - an endpoint's health signal now reaches the caller second-hand, through the sidecar's own view. ## Comparing them on what actually decides it | Question | Client library | Sidecar data plane | |---|---|---| | How many implementations of the policy? | one per language | one | | How does a policy change ship? | redeploy every caller | roll the data plane | | Extra per-call cost | none | one local hop | | Extra per-workload cost | none | a process | | Connections held | callers times endpoints | sidecar to endpoints | | New failure domain | no | yes | The decision usually falls out of three facts about the organisation: **how many languages** callers are written in, **how often** balancing policy actually changes, and **who is accountable** when a change has to reach everyone. One language and a stable policy make a client-side balancer obviously right. Five languages and a platform team expected to change routing behaviour without a company-wide deploy make the data plane obviously right. ## The migration is the hard part They coexist, and planning that is most of the work: 1. Run the data plane beside a subset of callers first, with the rest still balancing directly — the two paths reach the same workers and must be measured separately. 2. Expect **version skew** for the whole transition: two policies with different backoff, different reactions to an unhealthy endpoint, and different connection counts against the same pool, at the same time. 3. Capacity-plan the pool for the sum, not for either. Connection counts fall as callers move, but not smoothly. 4. Decide the rollback in advance. Pointing a caller back at the pool directly has to be a configuration change, not a code change, or the migration has no cheap exit. ## What neither choice moves Two things stay with the caller under either design, and an answer that forgets them is the weak one: the **deadline** on each call, which only the caller knows how to set, and the decision whether a failed call may be sent again, which depends on what the method does rather than on where it was routed. A data plane can re-route and shed; it cannot invent either of those.

  • What does a sidecar make easy that a client-side policy cannot?
    Changing behaviour for every caller at once. One implementation, rolled by one team, without recompiling or redeploying the callers — plus a single place to enforce transport security and emit per-call telemetry in a shape that is identical across every language in the estate.
  • What stays the caller's job under either choice?
    The deadline on each call and the decision whether a failed call may be sent again. A data plane can re-route, shed and retry-at-the-transport, but it cannot know how long the caller is willing to wait or whether repeating this particular method is safe.
  • How would you keep the migration reversible?
    Make the caller's target a configuration value, so pointing it at the local process or back at the pool is a config change rather than a code change. Then run both paths on real traffic, measure them separately, and capacity-plan the pool for the sum during the overlap.

saying these in an interview costs you the question

  • Treats a sidecar as free, ignoring per-workload memory and upgrade cost
  • Assumes one policy behaves identically across every language implementation
  • Thinks moving balancing into a sidecar removes the need for deadlines
  • Believes the two approaches cannot run side by side during a migration
  • Frames it purely as latency, ignoring who owns the rollout