How do you design graceful degradation under a partial outage — deciding per endpoint what a fallback should actually return?
answer
- classify: critical-path vs enrichment
- degrade reads (cache/empty/partial + flag)
- fail-fast unsafe writes — no fake success
- fail-closed for security/auth
- observable: log cause, Micrometer, alerts
basics
~20 sClassify each dependency by how critical it is. For non-critical/enrichment calls, return cached or default data and mark the response degraded. For critical writes, fail fast rather than fake success. Never let a fallback hide data loss or return misleading 'success'.
solid answer
~50 sGraceful degradation is a per-endpoint design decision, not a blanket setting. Start by classifying dependencies: **critical-path** (payment, auth) vs **enrichment** (recommendations, ratings, avatars). Enrichment calls should degrade — return last-known-good from a cache, an empty/default value, or a partial response with a `degraded=true` flag so the UI can show 'live data unavailable'. Critical writes should usually **fail fast** with a controlled 503, because a fallback that fakes success can cause silent data loss or double-charges. Build responses so partial success is representable (e.g., return the order with a null 'estimated delivery' rather than failing the whole call). Make degradation **observable**: log the cause (use Feign `fallbackFactory` or the fallback's Throwable), emit metrics/alerts, and expose breaker state via Actuator. Guard against fallback-of-a-fallback and stale-cache pitfalls. The principle: degrade reads, fail-fast unsafe writes, and never return data that misleads the caller.
code
java · 22 lines// A degraded-aware response envelope keeps the contract stable and honest.
public record ProductView(
Product product,
List<Item> recommendations, // enrichment — may be empty when degraded
boolean recommendationsDegraded) {}
@Service
public class ProductViewService {
private final ProductRepository products; // critical: no fallback -> fail fast
private final RecommendationClient reco; // enrichment: degrade
public ProductView view(String sku) {
Product p = products.findBySku(sku) // if this fails, let it 500/503
.orElseThrow(() -> new ProductNotFound(sku));
try {
return new ProductView(p, reco.recommend(sku), false);
} catch (Exception e) { // breaker OPEN / timeout / etc.
log.warn("reco degraded for {}: {}", sku, e.toString());
return new ProductView(p, List.of(), true); // partial, honestly flagged
}
}
}go deeper
Know that fallbacks should return safe defaults and not crash.
Distinguish enrichment (degrade) from critical (fail fast) and flag degraded responses.
Design partial-success DTOs, avoid stale-cache/fallback-of-fallback traps, and wire observability.
Set an org-wide degradation policy + criticality matrix, standard degraded envelope, fail-closed security stance, and treat fallback-rate as an operational signal.
## Graceful degradation = a product decision expressed in code **Graceful degradation** means the system keeps serving *reduced* functionality during a partial outage instead of failing wholesale. The hard part isn't the mechanism (fallbacks) — it's deciding, per endpoint, *what a correct degraded answer is*. A wrong fallback can be worse than an error. ## Step 1 — classify dependencies by criticality - **Critical path**: without it the operation is meaningless or unsafe — payment authorization, inventory decrement, authentication. - **Enrichment / non-critical**: improves the response but isn't required — recommendations, product ratings, related items, avatars, delivery ETA. This classification drives the fallback policy. ## Step 2 — choose a degraded response per class **Reads / enrichment → degrade:** - **Last-known-good cache**: serve the previously cached value (best UX). Watch staleness — tag with an age or `stale=true`. - **Empty / neutral default**: `emptyList()` for recommendations, a generic placeholder. - **Partial response**: return the core object with the missing part nulled/flagged (`estimatedDelivery=null`, `ratingsAvailable=false`), rather than failing the entire request. Design DTOs so partial success is *representable*. - **`degraded=true` flag**: let the client render 'live data temporarily unavailable' instead of showing wrong data as if it were fresh. **Writes / critical path → usually fail fast:** - A fallback that returns a fake 'order placed' when the payment service is down causes **silent data loss, lost revenue, or double-processing**. Prefer a controlled `503 Service Unavailable` (or a domain `ServiceUnavailableException` mapped to 503) with a Retry-After, so the caller/queue retries later. - If the write can be **deferred**, an alternative is to enqueue it (outbox / message) and return 'accepted' — but that's a real async design, not a fake fallback. ## Step 3 — make degradation observable - **Never swallow silently.** Use Feign `fallbackFactory` (or the fallback's `Throwable`) to **log the cause** at WARN with correlation id. - **Metrics + alerts**: Resilience4j publishes to Micrometer — track fallback rate, breaker state transitions, slow-call rate. Alert when a breaker stays OPEN. - **Expose state**: `resilience4j` Actuator endpoints / health indicators show breaker states for ops. - A spike in fallbacks is a *symptom of an outage in progress* — treat it as a first-class signal. ## Step 4 — avoid the classic traps - **Fallback-of-a-fallback**: don't call another remote inside a fallback without its own breaker/timeout; you just move the failure. - **Stale cache masking a long outage**: bound cache age; a week-old price is dangerous. - **Masking 4xx as degraded**: a 404/400 is a client/logic error, not an outage — don't silently degrade it (see Feign 404 trap). - **Fallback hides root cause from the caller who needed to know**: e.g., an auth check that 'degrades' to allow-through is a security hole. Fail closed for security-relevant calls. - **Inconsistent contract**: the degraded DTO must be schema-compatible with the normal one or clients break. ## Step 5 — organizational consistency At scale, set a **standard**: factory-based fallbacks, a shared `Degradable<T>` response envelope with a `degraded` flag, a naming/metrics convention, and a documented criticality matrix reviewed in design. This turns ad-hoc fallbacks into a coherent resilience policy. ## The one-line principle **Degrade reads, fail-fast unsafe writes, fail-closed security, and never return data that misleads the caller — and always make the degradation observable.** ## When this matters High-fan-in services, checkout/order flows, dashboards aggregating many sources — anywhere a partial outage is likely and 'all-or-nothing' is unacceptable to the business.
- When is failing fast strictly better than returning a fallback?For unsafe writes and security decisions. A fallback that fakes 'payment succeeded' or 'authorized' causes silent data loss, revenue loss, or a security breach. There, fail fast (503 + Retry-After, or fail-closed for authz) so the caller/queue retries or the request is correctly denied — a degraded *lie* is worse than an honest error.
- How do you keep degradation observable rather than a silent hole?Log the cause in the fallback (Feign `fallbackFactory` gives the Throwable) at WARN with a correlation id; publish Resilience4j metrics to Micrometer (fallback rate, breaker state, slow-call rate); expose breaker state via Actuator health; and alert when a breaker stays OPEN. A rising fallback rate is your early outage signal.
saying these in an interview costs you the question
- Returning a fake 'success' from a fallback on a critical write, hiding data loss.
- Degrading a security/authorization check to 'allow' (should fail closed).
- Serving unbounded stale cache during a long outage.
- Swallowing the exception in the fallback with no logging or metrics.
- A degraded DTO with an incompatible schema that breaks clients.