Across an organisation, would you put gRPC balancing in every client library or terminate every connection in a sidecar data plane?
answer
- not really about the extra hop
- count the implementations, not the calls
- how does a change reach everyone
- a process per workload has a bill
- both run at once for a while
basics
~20 sClient-side balancing costs no extra hop but duplicates the policy in every language and ships only when callers redeploy. A sidecar centralises the policy at the price of a hop, a process per workload, and a new failure domain.
solid answer
~50 sBoth put the decision per call; they differ in where the decision lives and who can change it. **In the client library**: no extra hop, no extra process, and the balancer sees the endpoints directly — but the policy exists once per language, versions drift between them, and changing it means redeploying every caller. **In a sidecar data plane**: one implementation, one place to change policy, uniform per-call telemetry and transport security, and callers that need no balancing code at all — but you now run a process alongside every workload, pay a local hop per call, and own its upgrades and its failures. The honest answer names the decider: how many languages you have, how often policy changes, and who is accountable for the rollout. Either way the caller still owns its deadlines and whether a failed call may be repeated.
go deeper
Know that balancing can happen inside the calling process or in a separate local process, and that both end up choosing an endpoint per call.
Contrast the concrete costs: an extra hop and a process per workload against one policy implementation per language and connections held by every caller.
Argue from the operational side — who diagnoses a failure that now sits between caller and dependency, and how connection counts change when the pool grows.
Decide on language count, policy change rate and accountability for rollout, then plan the overlap: both designs running at once, version skew, capacity for the sum, and a config-only rollback.
## The decision is about ownership, not latency Both designs balance per call, and both work. Arguments that reduce to microseconds of extra hop miss what is actually being chosen: **where the policy lives, and who can change it without a deploy**. That is an organisational property, and it is why this is a question for whoever owns the platform rather than for the team writing one service. ## Client-side balancing The caller resolves the pool, holds a subchannel per endpoint and picks one per call. What it buys: - no extra process and no extra hop; the connection goes from the caller to the worker; - the balancer's view of endpoint state is first-hand — it sees connections fail, and it can consume the endpoints' health signals directly; - nothing extra to operate per workload. What it costs: - the policy is implemented once per language, and those implementations differ in defaults, in backoff, and in how faithfully they follow a config; - a policy change reaches production only as fast as your slowest caller redeploys, which for a long tail of services is measured in quarters; - every caller holds a connection to every endpoint, so connection counts multiply with the product of callers and pool size; - getting the configuration to the callers is itself a distribution problem — the resolver and its service config become platform infrastructure whether you planned them or not. ## A sidecar data plane The caller connects to a local process; that process holds the connections to the workers and balances per call. What it buys: - one implementation of the policy for every language in the estate; - policy changes without recompiling or redeploying callers; - one place to enforce transport security and emit per-call telemetry in a shape that is the same everywhere; - callers that carry no balancing configuration at all, which makes a caller written in an unusual language a non-event. What it costs: - a process next to every workload, with its memory, its CPU and its own upgrade campaign across the fleet; - a local hop per call, plus the sidecar's own queuing under load; - a new failure domain sitting between every caller and every dependency, and one whose failures look to the caller like the dependency's; - an endpoint's health signal now reaches the caller second-hand, through the sidecar's own view. ## Comparing them on what actually decides it | Question | Client library | Sidecar data plane | |---|---|---| | How many implementations of the policy? | one per language | one | | How does a policy change ship? | redeploy every caller | roll the data plane | | Extra per-call cost | none | one local hop | | Extra per-workload cost | none | a process | | Connections held | callers times endpoints | sidecar to endpoints | | New failure domain | no | yes | The decision usually falls out of three facts about the organisation: **how many languages** callers are written in, **how often** balancing policy actually changes, and **who is accountable** when a change has to reach everyone. One language and a stable policy make a client-side balancer obviously right. Five languages and a platform team expected to change routing behaviour without a company-wide deploy make the data plane obviously right. ## The migration is the hard part They coexist, and planning that is most of the work: 1. Run the data plane beside a subset of callers first, with the rest still balancing directly — the two paths reach the same workers and must be measured separately. 2. Expect **version skew** for the whole transition: two policies with different backoff, different reactions to an unhealthy endpoint, and different connection counts against the same pool, at the same time. 3. Capacity-plan the pool for the sum, not for either. Connection counts fall as callers move, but not smoothly. 4. Decide the rollback in advance. Pointing a caller back at the pool directly has to be a configuration change, not a code change, or the migration has no cheap exit. ## What neither choice moves Two things stay with the caller under either design, and an answer that forgets them is the weak one: the **deadline** on each call, which only the caller knows how to set, and the decision whether a failed call may be sent again, which depends on what the method does rather than on where it was routed. A data plane can re-route and shed; it cannot invent either of those.
- What does a sidecar make easy that a client-side policy cannot?Changing behaviour for every caller at once. One implementation, rolled by one team, without recompiling or redeploying the callers — plus a single place to enforce transport security and emit per-call telemetry in a shape that is identical across every language in the estate.
- What stays the caller's job under either choice?The deadline on each call and the decision whether a failed call may be sent again. A data plane can re-route, shed and retry-at-the-transport, but it cannot know how long the caller is willing to wait or whether repeating this particular method is safe.
- How would you keep the migration reversible?Make the caller's target a configuration value, so pointing it at the local process or back at the pool is a config change rather than a code change. Then run both paths on real traffic, measure them separately, and capacity-plan the pool for the sum during the overlap.
saying these in an interview costs you the question
- Treats a sidecar as free, ignoring per-workload memory and upgrade cost
- Assumes one policy behaves identically across every language implementation
- Thinks moving balancing into a sidecar removes the need for deadlines
- Believes the two approaches cannot run side by side during a migration
- Frames it purely as latency, ignoring who owns the rollout