Resilience & Stability
Patterns that keep a service usable when its dependencies fail or its load spikes: retry, circuit breaker, bulkhead, timeout, fallback, throttling, load leveling and health monitoring. Together they are the standard answer to how you stop one failure from cascading.
part ofResilience & cloud-native patternsoverview, primer and where to startread it →on this pageshowhide
explore
- Retry and Backoff6 questions
- Circuit Breaker6 questions
- Bulkhead6 questions
- Timeout & Deadline Propagation6 questions
- Fallback & Graceful Degradation6 questions
- Throttling & Rate Limiting6 questions
- Queue-Based Load Leveling5 questions
- Health Endpoint Monitoring6 questions
- Scheduler Agent Supervisor6 questions
- AI Engineerrole
- API Designskill
- Backend Developerrole
- Data Engineerrole
- DevOps / SRE Engineerrole
- Forward Deployed Engineerrole
- Full Stack Developerrole
- Game Developerrole
- Java Backend Developerrole
- Kotlin Backend Developerrole
- Server-Side Game Developerrole
- Software Architectrole
- Software Design & Architectureskill
- System Designskill
questions
page 2 of 2A currency-conversion microservice occasionally fails. A developer adds a fallback that returns a hardcoded exchange rate of 1.0 (treating the foreign amount as if it were already in the local currency) whenever the call fails, so the checkout flow never breaks. What's wrong with this specific choice of default value, and what would be a safer approach?
basics
~20 sPretending the exchange rate is always 1.0 can massively over- or under-charge customers without anyone noticing — it fails silently in a way that looks like success. A safer default is to block or clearly flag the transaction instead of guessing.
A Kubernetes deployment defines livenessProbe with initialDelaySeconds: 5, periodSeconds: 10, and failureThreshold: 3 for a Java service whose JVM takes about 45 seconds to finish class-loading and warm its caches before it can respond. What happens when this deployment starts, and how would you fix it?
basics
~20 sThe health check starts too early and fails because the app isn't ready yet, so Kubernetes thinks it's broken and restarts it before it ever finishes starting up. Fix: give it a separate 'still starting' probe with a longer allowance, rather than just delaying the regular checks.
What does an application give up by adopting queue-based load leveling for a request that used to be handled synchronously? Name and explain at least two concrete costs.
basics
~20 sThe caller no longer gets an instant answer back, it has to check later. And under real load, the answer can now take much longer to arrive because the request has to wait its turn in the queue.
In a multi-tenant platform where all tenants share one backend capacity pool, what techniques let you throttle for per-tenant fairness so a single heavy tenant can't degrade service for everyone else?
basics
~10 sInstead of one big shared limit for everybody, give each customer their own smaller limit. That way one customer sending a lot of traffic can only use up their own share, not everyone else's.
If an API gateway has a 1-second end-to-end budget for a request that fans out to three parallel downstream calls plus one sequential call after they return, how should it split that 1-second budget across the calls, and what's the risk of splitting it evenly without regard to the call graph shape?
basics
~20 sCalls that run in parallel can each get most of the budget since they finish together, but calls that run one after another must each get a smaller slice so their total doesn't exceed the time left. Splitting evenly regardless of shape wastes budget or starves later steps.
What are the real costs of adopting per-dependency bulkheads across a service with many downstream calls, and under what circumstances would you deliberately choose NOT to bulkhead a particular dependency?
basics
~20 sBulkheads cost extra setup, extra resources (each pool needs its own capacity, some of which sits idle), and extra things to configure and monitor correctly. For a dependency that's low-risk, low-traffic, or extremely reliable, that overhead often isn't worth it — you'd only add a bulkhead where a failure would actually do real damage if left unisolated.
When wiring a fallback for an open circuit breaker — say, returning a cached value or a default response instead of calling the failing dependency — what should that fallback avoid doing, and what can go wrong if it's implemented carelessly?
basics
~20 sA fallback is the backup answer given when the breaker blocks a call. It should be fast, safe, and not depend on the thing that just failed. Done badly, the fallback can itself be slow, call another struggling service, or quietly hide a real problem from the people who need to know about it.
How do circuit breaker implementations differ across Resilience4j (Java), Hystrix (Netflix's now-legacy library), and Polly (.NET) in how you configure and wire a breaker into a call, and why did the industry largely move away from Hystrix's approach?
basics
~30 sAll three let you wrap a risky call so it can fail fast and recover automatically. Hystrix was the older, heavier one from Netflix, built around wrapping calls in special 'command' objects. Resilience4j is a newer, lighter Java library that wraps a plain function call instead. Polly is the .NET equivalent, using a similar lightweight wrapping style. Hystrix is no longer actively developed, so most new Java projects use Resilience4j instead.
A team adds a fallback so that whenever the primary recommendation service is slow, requests fall through to a secondary 'simple trending list' service instead. During a major primary outage, the secondary service — which was sized only for occasional traffic — collapses under the full redirected load, and now both are down. What went wrong with this fallback design, and how would you fix it?
basics
~20 sThe backup service wasn't built to handle all the traffic at once, so when everyone switched to it, it broke too. Fix it by making sure the fallback can actually handle full load, or by rate-limiting/shedding instead of just dumping everything on it.
A platform team wants to add synthetic monitoring — scripted checks that periodically exercise a service's health endpoints and key user flows from outside the cluster — and wire failures directly into automated remediation such as auto-restart, auto-scale, or traffic failover. What can go wrong if this is built without safeguards, and how do you guard against it?
basics
~20 sSynthetic monitoring means scripts periodically pretend to be a user and check the app still works from the outside. If a single failed check triggers an automatic fix, like restarting servers, a flaky check or a small real problem can trigger a big automatic overreaction, like restarting everything at once, and make things worse instead of better.
A downstream HTTP service returns a 429 Too Many Requests response with a Retry-After: 30 header. How should a well-behaved client's retry logic use this header, and what should it do if the header is absent on a 429 or 503 response?
basics
~20 sThe Retry-After header tells the client exactly how long to wait, so the client should honor that value instead of using its own backoff timer. If the header isn't present, the client should fall back to its normal exponential-backoff-with-jitter schedule.
What makes the supervisor itself a production risk in the Scheduler Agent Supervisor pattern, and how would you design around that risk?
basics
~20 sIf there's only one supervisor and it crashes, stuck steps stop getting fixed even though everything else works. Fix it by running more than one supervisor safely, usually with leader election so only one is active at a time.
How should a team decide where to set a throttling threshold relative to a service's SLOs, and what goes wrong in production if that threshold is set too aggressively versus too loosely?
basics
~20 sSet the limit based on real measured capacity, with some safety margin, not a guess. Too strict and you block normal traffic for no reason; too loose and the throttle never actually kicks in before things break.
How should you decide what value to set a downstream call's timeout to, and what specifically goes wrong if a caller's timeout is set longer than the callee's own internal processing timeout (or its thread-pool queue wait)?
basics
~20 sBase the timeout on the callee's real latency (like its 99th-percentile response time) plus margin, not a round number picked by guessing. If the caller waits longer than the callee itself would ever take, the caller just blocks pointlessly once the callee has already given up or hung.
You're setting resilience standards for a platform where dozens of teams each own services calling many downstream dependencies. How would you design an organization-wide bulkhead strategy — covering partition boundaries, defaults, and how bulkheads interact with circuit breakers and timeouts — so individual teams don't each have to re-derive correct isolation from scratch?
basics
~20 sInstead of asking every team to figure out bulkheads themselves, you build sane defaults into shared infrastructure (like a service mesh) that isolate every outgoing call automatically, give teams an easy way to tighten limits for their riskiest dependencies, pair it with required timeouts and circuit breakers, and monitor pool saturation everywhere so problems get caught before they cause outages.
You're designing an API where a client uploads a document and expects a synchronous HTTP response within two seconds containing the fully processed result. Why would introducing a queue-based load-leveling pattern between the API and the processing worker be a poor fit here, and what would you do instead?
basics
~20 sA queue adds unpredictable waiting time, which clashes with a hard two-second promise, since the request might sit behind other work in line. For this kind of tight, synchronous deadline, it's usually better to size the service to handle real load directly, or fall back to sync-with-a-time-budget and only queue as an overflow path.
When would you deliberately avoid the Scheduler Agent Supervisor pattern in favor of something simpler, and what does it cost you to adopt it unnecessarily?
basics
~20 sSkip it for short, simple operations that fit in one process and can just retry in place - the pattern's durable state, watchdog loop, and idempotency requirements are real infrastructure you shouldn't pay for unless the operation genuinely spans processes and needs to survive a crash mid-flight.
In a production system running many replica instances of the same service, each with its own in-process circuit breaker guarding calls to a shared downstream dependency, what operational pitfalls can emerge from that per-instance breaker design, and how would you mitigate them?
basics
~20 sIf every copy of your service has its own separate circuit breaker watching the same downstream dependency, they can trip and recover at different, uncoordinated times, or all probe recovery at once and re-overload the dependency. Fixes include adding randomness to timing, or moving the breaker logic to a shared layer like a service mesh so it's coordinated.
Across a large platform with dozens of teams independently adding fallbacks and stale-cache defaults to their services, what organizational and observability practices would you put in place so that graceful degradation doesn't quietly erode overall product quality or hide real reliability problems over time?
basics
~20 sMake sure every fallback is tracked and visible, not silent; review how often they're used company-wide; and set rules for which kinds of data are allowed to use a fallback at all, so teams don't quietly guess on important things.
A globally distributed service uses DNS-based health checks, such as Route 53 health checks, to automatically fail traffic away from an unhealthy region. What makes this pattern riskier than an in-cluster readiness probe, and how would you design against flapping and split-brain?
basics
~20 sDNS-based failover switches which region's address gets handed out based on health checks, but DNS answers get cached everywhere, in browsers, ISPs, and apps, for minutes, so recovery is slow and uneven. If the health check itself flaps, users can end up split across two different regions at once, which is riskier than a fast, uncached, in-cluster check.
In a large distributed system with an edge/gateway layer, individual backend services, and per-tenant concerns all in play, how would you design a throttling policy across these layers, and what goes wrong if throttling is only applied at one of them?
basics
~20 sPut limits at more than one point — at the front door, inside each service, and per customer — instead of just one place. If you only limit at one spot, problems from a different spot slip through and can still overload something downstream.
A synchronous HTTP request triggers a message being placed on a queue for asynchronous background processing, and the message includes the original request's deadline as a field. Why is that deadline mostly meaningless for the queue consumer to enforce as a 'give up and fail fast' bound the way it would be for a synchronous downstream call, and what should the consumer actually do with it?
basics
~20 sOnce work is on a queue, no one is blocked waiting for it the way a synchronous caller waits on a socket, so racing the original deadline doesn't 'fail fast' for anyone. The consumer should instead use it to detect and discard work that's already too stale to matter, not to bound how long its own processing takes.
showing 31–53 of 53