Resilience & cloud-native patterns
The catalogued solutions for building applications on cloud infrastructure, grouped into resilience, data management, messaging and deployment. They are worth learning as named patterns because interviewers and design docs use the names as shorthand for a whole set of trade-offs.
on this pageshowhide
guide
overview
~1 minCloud design patterns are the working vocabulary of distributed-system interviews. When a candidate says "put a circuit breaker in front of it" or "make it a saga", the interviewer expects the rest to follow: which failure the pattern addresses, what it costs, and what new problem it brings with it. The name is shorthand for a bundle of trade-offs, and the round tests whether you know what sits behind it. Junior questions ask what a pattern is for; senior and principal questions ask when it backfires, how two patterns interact, and what you would watch in production to know it is doing its job. The subject splits into four sections. [Resilience and stability](/topics/found-cloud-design-patterns-resilience) covers keeping a service usable when a dependency is slow or down or traffic spikes: timeouts, retries, circuit breakers, bulkheads, fallbacks, throttling, load leveling and health checks. [Data management](/topics/found-cloud-design-patterns-data-management) covers making reads fast and spreading data out: caching, precomputed views, sharding, secondary index tables, and keeping large payloads off paths that cannot carry them. [Messaging and integration](/topics/found-cloud-design-patterns-messaging) covers asynchronous work and the gateway tier: worker pools on a queue, prioritisation, processing pipelines, gateways, translation layers and the saga for business transactions that span services. [Deployment and operational](/topics/found-cloud-design-patterns-deployment) covers topology and runtime control: helper processes beside the application, repeatable deployment units, external configuration, feature flags and the gatekeeper. Start with resilience, and inside it with [timeouts](/topics/found-cloud-design-patterns-timeout) and [retries](/topics/found-cloud-design-patterns-retry-backoff), because nearly every other pattern assumes a call can fail, stall or be repeated. Messaging comes next, since queues and at-least-once delivery change what "done" means. Data management follows, where the recurring question is how far a copy may drift from its source. Deployment comes last, because it arranges the pieces the other sections define.
primer
A few ideas run under every section; with them in place, most questions below read as one of them applied to a specific pattern. - **A remote call has three outcomes, not two.** Besides success and failure there is *unknown*: the request may or may not have taken effect when the caller gave up. Timeouts keep the unknown case bounded in time; retries turn it into a possible duplicate. That is why idempotency comes up in the retry, messaging and saga sections alike, and why "just retry" is rarely a complete answer. - **Failures spread through shared resources.** A cascade starts when a slow dependency holds threads, connections or queue slots that healthy work also needs. Timeouts, bulkheads, circuit breakers and load shedding each cut a different path of that spread; a strong answer says which path a given pattern cuts and which it leaves open. - **Every copy is a consistency decision.** Caches, materialized views, index tables, read models and regional replicas each buy read speed or locality and pay with staleness and a second write that can fail on its own. The interview question is nearly always the same: how stale may it get, how do you notice drift, and how do you repair it. - **Asynchrony moves a problem rather than removing it.** A queue absorbs bursts and decouples availability, but delivery is typically at-least-once, ordering is limited, and the caller no longer sees the result directly. The design then needs duplicate-tolerant consumers, a way to report completion, and a place for messages that keep failing. - **Across service boundaries, compensation replaces rollback.** When each service owns its store, a single atomic transaction across them is usually off the table. You sequence local steps and define a business-level reversal for each, accepting that intermediate states are visible to others. - **Cross-cutting work moves out of the application.** Gateways, sidecars, external configuration and flags pull routing, security, configuration and release control out of service code. That buys consistency and fewer copies of the same logic, and costs an extra hop, another component to operate and a shared point of failure. - **Patterns compose, and they interfere.** Retries stacked at several layers multiply load; a breaker changes what a retry sees; a health check that tests a shared dependency can turn one outage into a fleet-wide one. Senior questions tend to live at these seams rather than inside a single pattern.
- Transient fault
- A failure expected to clear on its own within a short time, such as a dropped connection or a brief overload; the only kind of fault a retry is meant to address.
- Idempotency
- The property that applying an operation more than once has the same effect as applying it once; what makes retries and redelivered messages safe.
- Exponential backoff
- A retry schedule in which the wait before each attempt grows by a constant factor, usually up to a cap, giving a struggling dependency room to recover.
- Jitter
- Randomness added to retry delays so that many clients failing at the same moment do not retry at the same moment.
- Cascading failure
- A failure that spreads from one component to its callers because they wait on it, retry against it, or run out of shared resources while doing so.
- Deadline
- An absolute point in time by which a whole request must finish, passed down a call chain so each hop knows how much time remains.
- Load shedding
- Deliberately rejecting or dropping lower-priority work when a system is saturated, so the work it keeps can still complete in time.
- At-least-once delivery
- A messaging guarantee that every message is delivered, possibly more than once; the common default, which makes consumers responsible for tolerating duplicates.
- Poison message
- A message that fails processing every time it is delivered; without a retry limit and a side queue it blocks or loops a consumer indefinitely.
- Compensating transaction
- A business-level action that counteracts an earlier completed step, such as a refund for a charge, used where a distributed rollback is not available.
- Eventual consistency
- A guarantee that copies of data converge once updates stop arriving, with no promise about what a reader sees before then.
- Read model
- A copy of data shaped for a specific query and kept up to date from the write side; also called a projection or materialized view.
- Shard key
- The attribute that decides which partition holds a record; it fixes which queries stay local and where load concentrates.
- Hot partition
- A shard receiving a disproportionate share of traffic because of key skew or access pattern, which caps throughput regardless of how many shards exist.
- Blast radius
- The set of users, requests or components affected when one part fails; bulkheads, stamps and cells exist to keep it small.
The four sections describe different layers of the same system, and most patterns show up in more than one of them under a different role. ### Around a single call The resilience patterns stack around each outbound call. A timeout is the innermost limit; retries wrap it; a circuit breaker decides whether to attempt the call at all; a bulkhead caps how much of the caller the call may occupy; a fallback decides what to return when everything inside has given up. Throttling and load leveling sit at the other end of the same call, protecting the service that receives it. [Health endpoints](/topics/found-cloud-design-patterns-health-monitoring) are what the platform and load balancer read to decide whether an instance receives traffic or gets restarted, so they belong to the same conversation about failure. ### Across the queue The queue in [load leveling](/topics/found-cloud-design-patterns-load-leveling) is the same queue that [competing consumers](/topics/found-cloud-design-patterns-competing-consumers) drain, that a priority scheme splits, and that asynchronous request-reply exposes to a client. The [saga](/topics/found-cloud-design-patterns-saga) builds on all of this: it needs reliable messaging between steps, idempotent participants, and the compensation idea from the core ideas above. The data section feeds on messaging too, since a CQRS read model is a materialized view kept current by events, and a claim check pairs a queue with storage because the payload does not fit the message. ### Around the whole deployment Gateway offloading and the sidecar or ambassador answer the same question from two placements: one shared component at the edge, or one helper beside each instance. An ambassador is also a natural home for the retries and breakers from the resilience section. Deployment stamps and geodes apply the sharding idea to entire deployments rather than rows. External configuration and feature flags are the runtime switches that fallbacks, kill switches and gradual rollouts rely on.
- Timeout & Deadline Propagation →
Most other resilience patterns assume a call can stall; learn to bound it and to pass a remaining-time budget down a call chain first.
- Retry and Backoff →
Introduces idempotency and the risk of amplifying load, two ideas that return in messaging, sagas and every retry discussion afterwards.
- Circuit Breaker →
The standard answer to a dependency that stays down, and the clearest example of failing fast instead of waiting on a lost cause.
- Queue-Based Load Leveling →
Brings the queue into the picture: buffering bursts, decoupling availability, and the shift from immediate results to eventual completion.
- Cache-Aside →
The most common data pattern and the gentlest introduction to staleness, invalidation and the second write that can fail.
- Saga Pattern →
Combines messaging, idempotency and compensation into one design; take it once the earlier steps are solid.
Adding retries to an operation that is not idempotent; after a timeout the outcome is unknown, and a blind retry can apply the change twice.
Retrying at every layer of a call chain, so attempts multiply per hop and one user request becomes a burst against the deepest dependency.
Retrying on a fixed short delay with no jitter, so every client that failed together hits the recovering dependency together.
Leaving a call with no timeout, or one longer than the caller's own budget, so a single slow dependency can hold threads until the caller fails too.
Wiring a restart decision to a check that tests a shared database, so one database outage restarts every instance instead of just taking them out of rotation.
Describing a queue-based design as exactly-once; say at-least-once and explain what the consumer does when the same message arrives twice.
Serving a fallback or cached response with no limit on how stale it may be and no signal to the caller that the answer is degraded.
Proposing a saga without saying what each compensation does once its effect has been seen outside the system, such as a sent email or a shipped parcel.
Choosing a shard key for convenience rather than access pattern; a steadily increasing key under range sharding sends every new write to one shard.
Naming a pattern as the whole answer; each one here adds a component, a hop or a consistency window, and the follow-up asks which.
The same handful of choices recurs across all four sections. Naming the one you are making, and what would change your mind, is usually worth more than the pattern's name. - **Fail fast versus keep trying.** Retries recover from transient faults; breakers, deadlines and load shedding give up early to protect the caller and the dependency. The deciding facts are whether the fault will likely clear within the caller's budget and whether the operation is safe to repeat. - **Freshness versus load.** Caches, views and replicas cut read cost by serving data that may be behind. The question is how stale a given screen or decision can tolerate, and whether you bound it by expiry, by invalidation on write, or both. - **Synchronous versus asynchronous.** A direct call gives the caller an immediate, simple answer and couples its availability to the dependency. A queue decouples them and absorbs bursts, at the price of duplicates, ordering limits and a separate path for reporting results. - **Centralised versus per-instance.** A gateway applies a policy once at the edge; a sidecar applies it beside every instance. An edge policy is easier to reason about and becomes a shared bottleneck; a per-instance one scales with the fleet and multiplies what has to be deployed and upgraded. - **Isolation versus utilisation.** Separate pools, shards, stamps and machines contain failures and noisy neighbours; consolidation raises utilisation and lowers cost. Security boundaries, very different scaling profiles and the acceptable blast radius decide where to stop consolidating.
Several shapes recur across the sections under different names; recognising one is often how you place an unfamiliar question. - **Bound the damage.** Timeouts bound time, bulkheads bound concurrency, quotas bound resources, stamps bound blast radius. - **Send a pointer, not the payload.** Claim check, valet key and index table all move a small reference through the constrained path and keep the heavy data, or the heavy access, somewhere better suited to it. - **Precompute for the read.** Cache-aside, materialized views and CQRS read models all trade write-time or background work for cheap reads, and all inherit the same staleness and rebuild questions. - **Put a proxy in the path.** Gateway routing, aggregation and offloading, the ambassador, the gatekeeper and the anti-corruption layer each place a component between two parties to translate, protect or simplify the conversation. - **Buffer, then drain at your own pace.** Load leveling, competing consumers, priority queues and asynchronous request-reply share one structure: accept work quickly, store it, and process it at the rate the back end can sustain. - **Split by key into independent units.** Sharding splits rows, bulkheads split resource pools, deployment stamps and geodes split whole deployments. The recurring questions are the key, uneven load, and requests that span units. - **Decide at runtime, not at build time.** External configuration and feature flags turn values and behaviour into data that can change without a redeploy, which is powerful and is also a new way to break a running fleet.
explore
- Resilience & Stability53 questions
- Retry and Backoff6 questions
- Circuit Breaker6 questions
- Bulkhead6 questions
- Timeout & Deadline Propagation6 questions
- Fallback & Graceful Degradation6 questions
- Throttling & Rate Limiting6 questions
- Queue-Based Load Leveling5 questions
- Health Endpoint Monitoring6 questions
- Scheduler Agent Supervisor6 questions
- Data Management39 questions
- Cache-Aside5 questions
- Materialized View5 questions
- Sharding Pattern6 questions
- Valet Key5 questions
- Claim Check6 questions
- Index Table6 questions
- Event Sourcing & CQRS (bridge)6 questions
- Messaging & Integration52 questions
- Competing Consumers6 questions
- Priority Queue6 questions
- Pipes and Filters6 questions
- Gateway Aggregation6 questions
- Gateway Offloading5 questions
- Gateway Routing6 questions
- Anti-Corruption Layer (bridge)6 questions
- Asynchronous Request-Reply5 questions
- Saga Pattern6 questions
- Deployment & Operational44 questions
- Sidecar and Ambassador5 questions
- Deployment Stamp6 questions
- Geode6 questions
- External Configuration Store5 questions
- Feature Flags6 questions
- Static Content Hosting5 questions
- Compute Resource Consolidation5 questions
- Gatekeeper6 questions
- AI Engineerroleanchors this topic
- API Designskillanchors this topic
- Backend Developerroleanchors this topic
- Data Engineerroleanchors this topic
- DevOps / SRE Engineerroleanchors this topic
- Full Stack Developerroleanchors this topic
- Java Backend Developerroleanchors this topic
- Kotlin Backend Developerroleanchors this topic
- Software Design & Architectureskillanchors this topic
- System Designskillanchors this topic
- Forward Deployed Engineerrole
- Game Developerrole
- Redisskill
- Server-Side Game Developerrole
- Software Architectrole
questions
188 · 4 sectionsWhat is the bulkhead pattern in software resilience, and why would a service give each downstream dependency its own pool of threads or connections instead of sharing one pool across all dependencies?
basics
~20 sBulkhead means giving each external service its own separate pool of workers/connections, like separate lifeboats on a ship. If one dependency gets slow or breaks, it only uses up its own pool — it can't eat all the workers other dependencies need.
In the circuit breaker pattern used to protect a caller from a failing downstream dependency, what are the three states a breaker cycles through, and what makes it flip from closed to open?
basics
~20 sA circuit breaker watches calls to another service. Normally it's 'closed' and lets calls through. If too many fail, it 'opens' and stops calling that service for a while, failing fast instead. After a timeout it lets a few test calls through ('half-open') to check if the service recovered.
A checkout page normally shows personalized product recommendations pulled from a separate recommendation service. If that service is slow or completely down, what should the checkout page do instead of failing the whole page, and what is the general name for this strategy?
basics
~20 sSkip or replace the broken part instead of crashing everything — show checkout without recommendations, or with generic ones, so the customer can still buy. This is called graceful degradation, and the substitute response is a fallback.
What is a liveness probe versus a readiness probe in a container orchestrator like Kubernetes, and what does each one control?
basics
~20 sA liveness probe checks if an app is stuck and needs restarting. A readiness probe checks if it's ready to handle traffic right now. Failing liveness means restart the container; failing readiness means stop sending it requests, but leave it running.
A web app writes directly to a downstream image-processing service. During a marketing campaign, request volume spikes to 50x normal for ten minutes and the service falls over. Why would putting a queue between the web app and the image-processing service help, instead of just calling it directly?
basics
~20 sA queue lets the web app drop off work instantly and move on. Workers pull jobs from the queue at a steady pace they can handle, so a flood piles up safely instead of crashing the service.
In the cache-aside (lazy-loading) caching pattern, walk through step by step what happens when application code requests a key that is not currently in the cache.
basics
~20 sThe app checks the cache first. If the value isn't there (a miss), the app reads it from the real database, saves a copy in the cache, then returns it to the caller. Next time it's a fast cache hit.
A message queue advertises a 256 KB maximum message size, but a service needs to move a 50 MB video file through an event-driven pipeline. What is the standard fix, and why don't teams just send the file directly in the message body?
basics
~10 sSave the big file to storage like S3, then send only a small pointer (its location or ID) through the queue. The receiver downloads the real file from storage when it needs it.
In a cloud-native application, why might a team split the write side and read side of a service into separate models using CQRS (Command Query Responsibility Segregation) instead of using one shared model for both?
basics
~20 sBecause writing data and reading data back have different needs. CQRS keeps them as two separate models so each can be built, scaled, and hosted the way that suits it best, instead of one design compromising for both jobs.
A key-value store like Azure Table Storage or DynamoDB lets you look up an item fast only by its primary key (partition key + sort key). Your application also needs to find a customer's orders by their email address, which is not part of that key. Name a simple technique for making that lookup fast without switching databases, and say what it costs you.
basics
~20 sBuild a second, smaller table that maps the field you want to search by (email) to the ID of the matching row in the main table. Look there first, then fetch the real record. It costs extra storage and effort to keep both tables in sync.
What is a materialized view, and why might a system use one instead of running the same expensive query against source tables every time a client asks for the data?
basics
~20 sA materialized view is a saved copy of a query's result, stored ahead of time. Instead of recalculating joins and aggregates on every request, the app just reads the pre-built copy, which is much faster but can be slightly out of date.
When a new service needs to integrate with a legacy system that has messy, inconsistent naming and data structures, what problem does putting an Anti-Corruption Layer between them solve?
basics
~10 sAn Anti-Corruption Layer is a translator between two systems. It converts the old system's weird data and terms into clean data your new system understands, so the mess doesn't spread into your new code.
A client calls a REST endpoint that kicks off a report that takes 2 minutes to generate. Instead of holding the connection open, the server responds immediately with HTTP 202 Accepted and a URL the client can check later. What is this approach called, and why use it instead of just blocking the connection until the report is ready?
basics
~20 sThe server says 'got it, working on it' right away instead of making the client wait. It hands back a link the client can check later to see if the work is done. This keeps the connection short and frees the client to do other things while it waits.
In a messaging system, what is the Competing Consumers pattern, and why might a team run several consumer instances reading from the same queue instead of just one?
basics
~20 sSeveral worker programs all watch the same queue and grab the next job when free. This lets jobs get done faster (more workers = more parallel work) and keeps things running if one worker crashes.
A mobile app screen needs data from five different backend microservices to render fully. Instead of having the app call all five services directly over the network, what does putting a gateway aggregation pattern in front of them do, and why does that help?
basics
~20 sIt puts one gateway in front of many services. The gateway makes all the backend calls itself and sends the client one combined response, so the client needs only one request instead of many slow round trips.
In an API gateway architecture, why would a team terminate TLS at the gateway instead of having every backend service handle its own TLS handshake, and what does that decision cost them?
basics
~10 sThe gateway holds the certificate and does the encryption handshake once, so backend services get plain, already-decrypted traffic instead of every service needing its own certificate and crypto work.
In cloud deployments, what does 'compute resource consolidation' mean, and what cost problem is it meant to solve?
basics
~10 sIt means running several small tasks or services together on one bigger machine instead of each getting its own, so the machine's capacity is actually used and you pay for less idle hardware.
What is the deployment stamp (scale unit) pattern, and why would a team run many independent copies of the same application instead of one large shared deployment?
basics
~20 sA deployment stamp is a complete, self-contained copy of an app and its database, deployed once per group of customers. Instead of building one giant system for everyone, you stamp out many identical smaller copies, each handling its own slice of users.
A team currently bakes database URLs, connection-pool sizes, and API timeout values directly into a Docker image at build time. What problems does this cause once they need to run that same image across dev, staging, and production, and how does moving that configuration into an external configuration store fix it?
basics
~20 sIf settings are baked into the image, you need a different image per environment, and changing a value means rebuilding and redeploying. An external store keeps those settings outside the artifact so one image runs everywhere, each instance just reads its own values from the store at startup.
What is a feature flag, and how does using one to gate a new code path let a team deploy code to production without releasing it to users yet?
basics
~20 sA feature flag is an on/off switch in code that lets you turn a feature on for some or all users without redeploying. It separates 'the code is live on servers' from 'users can see/use it.'
In the Gatekeeper cloud design pattern, what does the dedicated gatekeeper host do to a client's request before it reaches a backend service, and why does it run separately from that backend?
basics
~10 sA gatekeeper is a stripped-down front-door server that checks and cleans every incoming request before it's allowed through, so bad or malformed input never touches the real backend systems directly.