skip to content

Resilience & Stability

Patterns that keep a service usable when its dependencies fail or its load spikes: retry, circuit breaker, bulkhead, timeout, fallback, throttling, load leveling and health monitoring. Together they are the standard answer to how you stop one failure from cascading.

part ofResilience & cloud-native patternsoverview, primer and where to startread it →
on this pageshow

questions

page 2 of 2

A currency-conversion microservice occasionally fails. A developer adds a fallback that returns a hardcoded exchange rate of 1.0 (treating the foreign amount as if it were already in the local currency) whenever the call fails, so the checkout flow never breaks. What's wrong with this specific choice of default value, and what would be a safer approach?

level: middleimportance: should knowfreq 55%

basics

~20 s

Pretending the exchange rate is always 1.0 can massively over- or under-charge customers without anyone noticing — it fails silently in a way that looks like success. A safer default is to block or clearly flag the transaction instead of guessing.

open as a page

A Kubernetes deployment defines livenessProbe with initialDelaySeconds: 5, periodSeconds: 10, and failureThreshold: 3 for a Java service whose JVM takes about 45 seconds to finish class-loading and warm its caches before it can respond. What happens when this deployment starts, and how would you fix it?

level: middleimportance: should knowfreq 65%

basics

~20 s

The health check starts too early and fails because the app isn't ready yet, so Kubernetes thinks it's broken and restarts it before it ever finishes starting up. Fix: give it a separate 'still starting' probe with a longer allowance, rather than just delaying the regular checks.

open as a page

What does an application give up by adopting queue-based load leveling for a request that used to be handled synchronously? Name and explain at least two concrete costs.

level: middleimportance: should knowfreq 55%

basics

~20 s

The caller no longer gets an instant answer back, it has to check later. And under real load, the answer can now take much longer to arrive because the request has to wait its turn in the queue.

open as a page

In a multi-tenant platform where all tenants share one backend capacity pool, what techniques let you throttle for per-tenant fairness so a single heavy tenant can't degrade service for everyone else?

level: middleimportance: should knowfreq 55%

basics

~10 s

Instead of one big shared limit for everybody, give each customer their own smaller limit. That way one customer sending a lot of traffic can only use up their own share, not everyone else's.

open as a page

If an API gateway has a 1-second end-to-end budget for a request that fans out to three parallel downstream calls plus one sequential call after they return, how should it split that 1-second budget across the calls, and what's the risk of splitting it evenly without regard to the call graph shape?

level: middleimportance: should knowfreq 45%

basics

~20 s

Calls that run in parallel can each get most of the budget since they finish together, but calls that run one after another must each get a smaller slice so their total doesn't exceed the time left. Splitting evenly regardless of shape wastes budget or starves later steps.

open as a page

What are the real costs of adopting per-dependency bulkheads across a service with many downstream calls, and under what circumstances would you deliberately choose NOT to bulkhead a particular dependency?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Bulkheads cost extra setup, extra resources (each pool needs its own capacity, some of which sits idle), and extra things to configure and monitor correctly. For a dependency that's low-risk, low-traffic, or extremely reliable, that overhead often isn't worth it — you'd only add a bulkhead where a failure would actually do real damage if left unisolated.

open as a page

When wiring a fallback for an open circuit breaker — say, returning a cached value or a default response instead of calling the failing dependency — what should that fallback avoid doing, and what can go wrong if it's implemented carelessly?

level: seniorimportance: should knowfreq 62%

basics

~20 s

A fallback is the backup answer given when the breaker blocks a call. It should be fast, safe, and not depend on the thing that just failed. Done badly, the fallback can itself be slow, call another struggling service, or quietly hide a real problem from the people who need to know about it.

open as a page

How do circuit breaker implementations differ across Resilience4j (Java), Hystrix (Netflix's now-legacy library), and Polly (.NET) in how you configure and wire a breaker into a call, and why did the industry largely move away from Hystrix's approach?

level: seniorimportance: should knowfreq 52%

basics

~30 s

All three let you wrap a risky call so it can fail fast and recover automatically. Hystrix was the older, heavier one from Netflix, built around wrapping calls in special 'command' objects. Resilience4j is a newer, lighter Java library that wraps a plain function call instead. Polly is the .NET equivalent, using a similar lightweight wrapping style. Hystrix is no longer actively developed, so most new Java projects use Resilience4j instead.

open as a page

A team adds a fallback so that whenever the primary recommendation service is slow, requests fall through to a secondary 'simple trending list' service instead. During a major primary outage, the secondary service — which was sized only for occasional traffic — collapses under the full redirected load, and now both are down. What went wrong with this fallback design, and how would you fix it?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The backup service wasn't built to handle all the traffic at once, so when everyone switched to it, it broke too. Fix it by making sure the fallback can actually handle full load, or by rate-limiting/shedding instead of just dumping everything on it.

open as a page

A platform team wants to add synthetic monitoring — scripted checks that periodically exercise a service's health endpoints and key user flows from outside the cluster — and wire failures directly into automated remediation such as auto-restart, auto-scale, or traffic failover. What can go wrong if this is built without safeguards, and how do you guard against it?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Synthetic monitoring means scripts periodically pretend to be a user and check the app still works from the outside. If a single failed check triggers an automatic fix, like restarting servers, a flaky check or a small real problem can trigger a big automatic overreaction, like restarting everything at once, and make things worse instead of better.

open as a page

A downstream HTTP service returns a 429 Too Many Requests response with a Retry-After: 30 header. How should a well-behaved client's retry logic use this header, and what should it do if the header is absent on a 429 or 503 response?

level: seniorimportance: should knowfreq 55%

basics

~20 s

The Retry-After header tells the client exactly how long to wait, so the client should honor that value instead of using its own backoff timer. If the header isn't present, the client should fall back to its normal exponential-backoff-with-jitter schedule.

open as a page

What makes the supervisor itself a production risk in the Scheduler Agent Supervisor pattern, and how would you design around that risk?

level: seniorimportance: should knowfreq 40%

basics

~20 s

If there's only one supervisor and it crashes, stuck steps stop getting fixed even though everything else works. Fix it by running more than one supervisor safely, usually with leader election so only one is active at a time.

open as a page

How should a team decide where to set a throttling threshold relative to a service's SLOs, and what goes wrong in production if that threshold is set too aggressively versus too loosely?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Set the limit based on real measured capacity, with some safety margin, not a guess. Too strict and you block normal traffic for no reason; too loose and the throttle never actually kicks in before things break.

open as a page

How should you decide what value to set a downstream call's timeout to, and what specifically goes wrong if a caller's timeout is set longer than the callee's own internal processing timeout (or its thread-pool queue wait)?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Base the timeout on the callee's real latency (like its 99th-percentile response time) plus margin, not a round number picked by guessing. If the caller waits longer than the callee itself would ever take, the caller just blocks pointlessly once the callee has already given up or hung.

open as a page

You're setting resilience standards for a platform where dozens of teams each own services calling many downstream dependencies. How would you design an organization-wide bulkhead strategy — covering partition boundaries, defaults, and how bulkheads interact with circuit breakers and timeouts — so individual teams don't each have to re-derive correct isolation from scratch?

level: principalimportance: should knowfreq 30%

basics

~20 s

Instead of asking every team to figure out bulkheads themselves, you build sane defaults into shared infrastructure (like a service mesh) that isolate every outgoing call automatically, give teams an easy way to tighten limits for their riskiest dependencies, pair it with required timeouts and circuit breakers, and monitor pool saturation everywhere so problems get caught before they cause outages.

open as a page

You're designing an API where a client uploads a document and expects a synchronous HTTP response within two seconds containing the fully processed result. Why would introducing a queue-based load-leveling pattern between the API and the processing worker be a poor fit here, and what would you do instead?

level: principalimportance: should knowfreq 40%

basics

~20 s

A queue adds unpredictable waiting time, which clashes with a hard two-second promise, since the request might sit behind other work in line. For this kind of tight, synchronous deadline, it's usually better to size the service to handle real load directly, or fall back to sync-with-a-time-budget and only queue as an overflow path.

open as a page

When would you deliberately avoid the Scheduler Agent Supervisor pattern in favor of something simpler, and what does it cost you to adopt it unnecessarily?

level: principalimportance: should knowfreq 35%

basics

~20 s

Skip it for short, simple operations that fit in one process and can just retry in place - the pattern's durable state, watchdog loop, and idempotency requirements are real infrastructure you shouldn't pay for unless the operation genuinely spans processes and needs to survive a crash mid-flight.

open as a page

In a production system running many replica instances of the same service, each with its own in-process circuit breaker guarding calls to a shared downstream dependency, what operational pitfalls can emerge from that per-instance breaker design, and how would you mitigate them?

level: principalimportance: nice to knowfreq 38%

basics

~20 s

If every copy of your service has its own separate circuit breaker watching the same downstream dependency, they can trip and recover at different, uncoordinated times, or all probe recovery at once and re-overload the dependency. Fixes include adding randomness to timing, or moving the breaker logic to a shared layer like a service mesh so it's coordinated.

open as a page

Across a large platform with dozens of teams independently adding fallbacks and stale-cache defaults to their services, what organizational and observability practices would you put in place so that graceful degradation doesn't quietly erode overall product quality or hide real reliability problems over time?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Make sure every fallback is tracked and visible, not silent; review how often they're used company-wide; and set rules for which kinds of data are allowed to use a fallback at all, so teams don't quietly guess on important things.

open as a page

A globally distributed service uses DNS-based health checks, such as Route 53 health checks, to automatically fail traffic away from an unhealthy region. What makes this pattern riskier than an in-cluster readiness probe, and how would you design against flapping and split-brain?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

DNS-based failover switches which region's address gets handed out based on health checks, but DNS answers get cached everywhere, in browsers, ISPs, and apps, for minutes, so recovery is slow and uneven. If the health check itself flaps, users can end up split across two different regions at once, which is riskier than a fast, uncached, in-cluster check.

open as a page

AWS's own SDKs default to a formula sometimes called 'decorrelated jitter': sleep = min(cap, random_between(base, previous_sleep * 3)). How does this differ mechanically from full jitter (random(0, min(cap, base * 2^attempt))), and why might it be preferred for long retry sequences?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Full jitter recalculates a random delay from scratch each time based on the attempt number. Decorrelated jitter instead bases each new random delay on the previous one, which tends to spread retries out more smoothly over a long sequence of attempts.

open as a page

In a large distributed system with an edge/gateway layer, individual backend services, and per-tenant concerns all in play, how would you design a throttling policy across these layers, and what goes wrong if throttling is only applied at one of them?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Put limits at more than one point — at the front door, inside each service, and per customer — instead of just one place. If you only limit at one spot, problems from a different spot slip through and can still overload something downstream.

open as a page

A synchronous HTTP request triggers a message being placed on a queue for asynchronous background processing, and the message includes the original request's deadline as a field. Why is that deadline mostly meaningless for the queue consumer to enforce as a 'give up and fail fast' bound the way it would be for a synchronous downstream call, and what should the consumer actually do with it?

level: principalimportance: nice to knowfreq 25%

basics

~20 s

Once work is on a queue, no one is blocked waiting for it the way a synchronous caller waits on a socket, so racing the original deadline doesn't 'fail fast' for anyone. The consumer should instead use it to detect and discard work that's already too stale to matter, not to bound how long its own processing takes.

open as a page

showing 31–53 of 53