skip to content

How does a load-balanced Feign client behave when the chosen instance is down, and how do you make it resilient?

level: principalimportance: should knowfreq 35%

answer

  1. default: no auto-retry, error surfaces
  2. spring-retry -> maxRetriesOnNextServiceInstance routes around
  3. GET-only retries unless retryOnAllOperations (idempotency!)
  4. circuitbreaker.enabled + fallback/fallbackFactory
  5. cached instance list -> stale node can still be picked

basics

~20 s

By default a single failing instance just surfaces the error. Add spring-retry so the load balancer retries on another instance, tune per-attempt timeouts, and wrap the client with a Resilience4j circuit breaker plus a fallback for graceful degradation.

solid answer

~40 s

Out of the box, if SCLB picks a dead instance the call fails with a connection error and nothing retries automatically. To make it resilient you layer several mechanisms. First, enable load-balancer retries by adding spring-retry (and spring.cloud.loadbalancer.retry.enabled=true): FeignBlockingLoadBalancerClient will then retry, and crucially can retry on the next instance, routing around a bad node. Second, set tight connect/read timeouts so a slow node fails fast. Third, add a circuit breaker: Spring Cloud OpenFeign integrates with Resilience4j via spring.cloud.openfeign.circuitbreaker.enabled=true, and you supply a fallback (fallback or fallbackFactory) for graceful degradation. The load balancer also relies on the DiscoveryClient's instance list, which is cached and updated periodically, so a just-crashed instance may still be selected until the registry and cache converge — retries bridge that gap.

code

java · 24 lines
java
// application.yml
// spring:
//   cloud:
//     loadbalancer:
//       retry:
//         enabled: true
//         maxRetriesOnSameServiceInstance: 0
//         maxRetriesOnNextServiceInstance: 2   # route around a dead node
//     openfeign:
//       circuitbreaker:
//         enabled: true

@FeignClient(name = "order-service", fallbackFactory = OrderFallbackFactory.class)
interface OrderClient {
    @GetMapping("/orders/{id}")
    OrderDto get(@PathVariable long id);
}

@Component
class OrderFallbackFactory implements FallbackFactory<OrderClient> {
    @Override public OrderClient create(Throwable cause) {
        return id -> OrderDto.unavailable(id); // graceful degradation
    }
}

go deeper

for a junior

Know that a failing call surfaces an error and resilience needs extra setup.

for a middle

Add spring-retry for next-instance retries and a fallback for degradation.

for a senior

Combine timeouts, retries, and a Resilience4j circuit breaker; respect idempotency.

for a principal

Reason about instance-list caching/staleness, health checks, retry idempotency, and layered defense-in-depth across the failure modes.

**Default behavior on a bad instance.** Spring Cloud LoadBalancer selects an instance from the (cached) `DiscoveryClient` list. If that instance is down, the HTTP call throws a **connect exception**; by default **there is no automatic retry** and the caller sees the failure. Load balancing spreads load but does not by itself provide failover. **1) Load-balancer retries (route around failures).** Add **`spring-retry`** to the classpath and enable `spring.cloud.loadbalancer.retry.enabled=true` (on by default when spring-retry is present). Then `FeignBlockingLoadBalancerClient`/`RetryLoadBalancerInterceptor` retries. Key knobs under `spring.cloud.loadbalancer.retry.*`: - `maxRetriesOnSameServiceInstance` — retries against the **same** instance. - `maxRetriesOnNextServiceInstance` — retries against the **next** instance — this is what routes around a dead node. - `retryableStatusCodes` — which HTTP status codes trigger a retry (beyond exceptions). - `retryOnAllOperations` — by default only idempotent (GET) requests retry; set true to retry POST etc. (dangerous unless the endpoint is idempotent). **Idempotency caveat:** retrying non-idempotent calls (POST that charges a card) risks duplicate side effects. Keep `retryOnAllOperations=false` unless the server dedupes. **2) Fail fast with timeouts.** Retries only help if a stuck instance fails quickly. Tight `connectTimeout`/`readTimeout` (per attempt) let the retry to the next instance happen promptly instead of after a 10s hang. **3) Circuit breaker + fallback.** Enable `spring.cloud.openfeign.circuitbreaker.enabled=true`; Feign then wraps calls with Spring Cloud CircuitBreaker (typically **Resilience4j**). When failure rates cross a threshold the breaker **opens**, short-circuiting calls to fail fast and let the downstream recover. Provide a **fallback** for graceful degradation: `@FeignClient(name="order-service", fallback = OrderFallback.class)` (fallback must be a Spring bean implementing the client interface) or `fallbackFactory` when you need the triggering exception. You can further tune the Resilience4j instance (sliding window, failure-rate threshold, wait duration) via `resilience4j.circuitbreaker.instances.*`. **4) Discovery-list staleness.** SCLB doesn't query the registry on every call; it uses a **cached** instance list (`spring.cloud.loadbalancer.cache.*`, default TTL ~35s) refreshed periodically, and the registry itself has propagation/heartbeat lag (e.g. Eureka lease renewals). So a crashed instance can still be **selected** for a short window — another reason retries to the next instance matter. Health checks (`spring.cloud.loadbalancer.configurations=health-check`) can proactively drop unhealthy instances from selection. **Putting it together (defense in depth).** Tight timeouts (fail fast) + next-instance retry (route around) + circuit breaker with fallback (shed load, degrade gracefully) + health-checked, freshly-cached instance lists. Each layer covers a different failure mode; relying on load balancing alone is the common principal-level mistake.

  • Why is retrying only enabled for GET requests by default, and when would you change that?
    GETs are idempotent, so a retry can't cause duplicate side effects. Non-idempotent calls (POST/PATCH) could double-charge or double-create on retry. Only set retryOnAllOperations=true when the endpoint is idempotent or deduplicates (e.g. via an idempotency key).
  • A node crashed one second ago but Feign still routes to it. Why, and what mitigates it?
    SCLB uses a cached instance list and the registry has heartbeat/propagation lag, so the dead instance lingers in the pool briefly. maxRetriesOnNextServiceInstance routes the failed call to a healthy node, and enabling the health-check load-balancer configuration proactively excludes unhealthy instances.

saying these in an interview costs you the question

  • Believing Feign automatically retries on another instance out of the box
  • Thinking load balancing alone provides failover/resilience
  • Enabling retryOnAllOperations without considering idempotency
  • Assuming SCLB re-queries the registry on every single call

context