How does a load-balanced Feign client behave when the chosen instance is down, and how do you make it resilient?
answer
- default: no auto-retry, error surfaces
- spring-retry -> maxRetriesOnNextServiceInstance routes around
- GET-only retries unless retryOnAllOperations (idempotency!)
- circuitbreaker.enabled + fallback/fallbackFactory
- cached instance list -> stale node can still be picked
basics
~20 sBy default a single failing instance just surfaces the error. Add spring-retry so the load balancer retries on another instance, tune per-attempt timeouts, and wrap the client with a Resilience4j circuit breaker plus a fallback for graceful degradation.
solid answer
~40 sOut of the box, if SCLB picks a dead instance the call fails with a connection error and nothing retries automatically. To make it resilient you layer several mechanisms. First, enable load-balancer retries by adding spring-retry (and spring.cloud.loadbalancer.retry.enabled=true): FeignBlockingLoadBalancerClient will then retry, and crucially can retry on the next instance, routing around a bad node. Second, set tight connect/read timeouts so a slow node fails fast. Third, add a circuit breaker: Spring Cloud OpenFeign integrates with Resilience4j via spring.cloud.openfeign.circuitbreaker.enabled=true, and you supply a fallback (fallback or fallbackFactory) for graceful degradation. The load balancer also relies on the DiscoveryClient's instance list, which is cached and updated periodically, so a just-crashed instance may still be selected until the registry and cache converge — retries bridge that gap.
code
java · 24 lines// application.yml
// spring:
// cloud:
// loadbalancer:
// retry:
// enabled: true
// maxRetriesOnSameServiceInstance: 0
// maxRetriesOnNextServiceInstance: 2 # route around a dead node
// openfeign:
// circuitbreaker:
// enabled: true
@FeignClient(name = "order-service", fallbackFactory = OrderFallbackFactory.class)
interface OrderClient {
@GetMapping("/orders/{id}")
OrderDto get(@PathVariable long id);
}
@Component
class OrderFallbackFactory implements FallbackFactory<OrderClient> {
@Override public OrderClient create(Throwable cause) {
return id -> OrderDto.unavailable(id); // graceful degradation
}
}go deeper
Know that a failing call surfaces an error and resilience needs extra setup.
Add spring-retry for next-instance retries and a fallback for degradation.
Combine timeouts, retries, and a Resilience4j circuit breaker; respect idempotency.
Reason about instance-list caching/staleness, health checks, retry idempotency, and layered defense-in-depth across the failure modes.
**Default behavior on a bad instance.** Spring Cloud LoadBalancer selects an instance from the (cached) `DiscoveryClient` list. If that instance is down, the HTTP call throws a **connect exception**; by default **there is no automatic retry** and the caller sees the failure. Load balancing spreads load but does not by itself provide failover. **1) Load-balancer retries (route around failures).** Add **`spring-retry`** to the classpath and enable `spring.cloud.loadbalancer.retry.enabled=true` (on by default when spring-retry is present). Then `FeignBlockingLoadBalancerClient`/`RetryLoadBalancerInterceptor` retries. Key knobs under `spring.cloud.loadbalancer.retry.*`: - `maxRetriesOnSameServiceInstance` — retries against the **same** instance. - `maxRetriesOnNextServiceInstance` — retries against the **next** instance — this is what routes around a dead node. - `retryableStatusCodes` — which HTTP status codes trigger a retry (beyond exceptions). - `retryOnAllOperations` — by default only idempotent (GET) requests retry; set true to retry POST etc. (dangerous unless the endpoint is idempotent). **Idempotency caveat:** retrying non-idempotent calls (POST that charges a card) risks duplicate side effects. Keep `retryOnAllOperations=false` unless the server dedupes. **2) Fail fast with timeouts.** Retries only help if a stuck instance fails quickly. Tight `connectTimeout`/`readTimeout` (per attempt) let the retry to the next instance happen promptly instead of after a 10s hang. **3) Circuit breaker + fallback.** Enable `spring.cloud.openfeign.circuitbreaker.enabled=true`; Feign then wraps calls with Spring Cloud CircuitBreaker (typically **Resilience4j**). When failure rates cross a threshold the breaker **opens**, short-circuiting calls to fail fast and let the downstream recover. Provide a **fallback** for graceful degradation: `@FeignClient(name="order-service", fallback = OrderFallback.class)` (fallback must be a Spring bean implementing the client interface) or `fallbackFactory` when you need the triggering exception. You can further tune the Resilience4j instance (sliding window, failure-rate threshold, wait duration) via `resilience4j.circuitbreaker.instances.*`. **4) Discovery-list staleness.** SCLB doesn't query the registry on every call; it uses a **cached** instance list (`spring.cloud.loadbalancer.cache.*`, default TTL ~35s) refreshed periodically, and the registry itself has propagation/heartbeat lag (e.g. Eureka lease renewals). So a crashed instance can still be **selected** for a short window — another reason retries to the next instance matter. Health checks (`spring.cloud.loadbalancer.configurations=health-check`) can proactively drop unhealthy instances from selection. **Putting it together (defense in depth).** Tight timeouts (fail fast) + next-instance retry (route around) + circuit breaker with fallback (shed load, degrade gracefully) + health-checked, freshly-cached instance lists. Each layer covers a different failure mode; relying on load balancing alone is the common principal-level mistake.
- Why is retrying only enabled for GET requests by default, and when would you change that?GETs are idempotent, so a retry can't cause duplicate side effects. Non-idempotent calls (POST/PATCH) could double-charge or double-create on retry. Only set retryOnAllOperations=true when the endpoint is idempotent or deduplicates (e.g. via an idempotency key).
- A node crashed one second ago but Feign still routes to it. Why, and what mitigates it?SCLB uses a cached instance list and the registry has heartbeat/propagation lag, so the dead instance lingers in the pool briefly. maxRetriesOnNextServiceInstance routes the failed call to a healthy node, and enabling the health-check load-balancer configuration proactively excludes unhealthy instances.
saying these in an interview costs you the question
- Believing Feign automatically retries on another instance out of the box
- Thinking load balancing alone provides failover/resilience
- Enabling retryOnAllOperations without considering idempotency
- Assuming SCLB re-queries the registry on every single call