skip to content

Deployment & Operational

Topology and operations patterns: sidecar and ambassador, deployment stamps, geodes, external configuration, feature flags, static content hosting, compute consolidation and the gatekeeper.

part ofResilience & cloud-native patternsoverview, primer and where to startread it →
on this pageshow

questions

page 2 of 2

If the feature-flag evaluation service itself becomes unreachable or slow, what should a client SDK do, and how does the answer differ between a flag guarding a nice-to-have feature versus a flag acting as a kill switch for a risky dependency?

level: seniorimportance: should knowfreq 55%

basics

~20 s

If the flag system can't be reached, the app should fall back to a safe default instead of crashing or hanging, usually the last value it knew, cached locally. For most features, safe means off; for a kill switch protecting against a broken dependency, safe might mean staying off (fail closed) so the risky code never runs by accident.

open as a page

In what circumstances would you decide NOT to add a dedicated Gatekeeper host in front of a backend service, even though the service accepts requests from outside its own trust boundary?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Skip it when the extra server, delay, and upkeep aren't worth it — for example, low-stakes data, a small team that can't run two systems well, or when a simpler tool like a managed API gateway with good input checks already covers the risk.

open as a page

Many production geode deployments route a given user or account consistently to the same 'home' geode under normal conditions, even though every geode is technically capable of serving any request. Why do teams add this kind of routing stickiness on top of an active-active geode design, and what do they give up when a user's home geode becomes unavailable and traffic reroutes elsewhere?

level: seniorimportance: should knowfreq 35%

basics

~20 s

It makes each user's experience predictable and consistent, since their requests always hit the same copy of the data instead of racing replication delays between regions. The trade-off is that when their home region goes down and they get sent somewhere else, they might briefly see slightly older data until things catch up.

open as a page

A platform team is deciding whether to give every service mTLS and retries via a sidecar proxy or via a shared in-process client library. What concrete trade-offs should drive that decision?

level: seniorimportance: should knowfreq 55%

basics

~20 s

A sidecar works with any programming language and can be upgraded without touching app code, but it costs extra memory/CPU per instance and adds a tiny delay. A library is lighter and faster but only works in one language and needs every app to be rebuilt to upgrade it.

open as a page

A CDN caching static assets is configured to include a custom request header in its cache key, and shortly after, some users start receiving other users' cached error pages or unexpected content for a shared static URL. What kind of failure is this, and how does it happen mechanically?

level: seniorimportance: should knowfreq 45%

basics

~20 s

This is cache poisoning: the CDN stored a wrong or attacker-influenced response under a cache key that many different users then hit, so instead of one broken response for one weird request, everyone gets served the bad cached copy.

open as a page

A CDN serving static assets from many edge points of presence reports a cache hit ratio of only 40%, meaning most requests still reach the object storage origin. What mechanisms would you look at to raise that ratio, and what's the role of an origin shield in this?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Low hit ratio usually means too many different edge locations each independently asking the origin for the same file, or the cache expiring too fast. An origin shield adds one extra caching layer between the edges and the origin so only one edge has to fetch a file the first time, and every other edge gets it from the shield instead.

open as a page

You're designing a placement/bin-packing strategy for a shared fleet running thousands of heterogeneous workloads (mixed CPU-bound, memory-bound, and I/O-bound; mixed criticality). What are the key design decisions you'd need to make to maximize density without creating systemic contention risk?

level: principalimportance: should knowfreq 35%

basics

~10 s

You'd group workloads by resource profile and priority, cap what each can use, watch real usage instead of guesses, and keep critical stuff separated from risky/bursty stuff — then keep adjusting as workloads change.

open as a page

An organization introduces a centralized external configuration store that all 200 of its microservice instances read from at startup and poll periodically. Walk through what happens to the fleet if that store becomes unavailable or briefly serves corrupted data, what mitigations and versioning/rollback strategy you'd put in place, and under what circumstances you'd advise a team not to adopt a fully centralized config store at all.

level: principalimportance: should knowfreq 44%

basics

~30 s

If the store goes down or serves bad data, every dependent service can be affected at once, so it's a single point of failure for the whole fleet. Mitigations: cache last-known-good config locally, version every change so you can instantly roll back, replicate/harden the store, and gate risky changes with canary rollout. For a small number of services, a fully centralized store can be overkill - simpler options (env vars, files in the deploy pipeline) may suffice.

open as a page

When a feature flag has multiple targeting rules, for example always off for free-tier users, always on for internal staff, and a 50% rollout for everyone else, how does a flag evaluation engine typically resolve conflicts between rules, and what architectural trade-offs come from evaluating those rules client-side in a browser or mobile SDK versus server-side?

level: principalimportance: should knowfreq 40%

basics

~20 s

Targeting rules are usually checked in order, top to bottom, and the first one that matches a user wins, so rule order matters a lot. Doing that check in the browser or app is fast but exposes your rules and upcoming features to anyone who inspects the app; doing it on your server keeps rules private but adds a network step.

open as a page

When designing the internal channel between a gatekeeper host and its trusted host, what design choices most affect whether the pattern actually caps the blast radius of a gatekeeper compromise, and how does pairing this with a short-lived, scoped access token (as in the Valet Key pattern) change the calculus?

level: principalimportance: should knowfreq 25%

basics

~20 s

How you build the pipe between the front-door server and the real worker matters as much as having two servers at all — a narrow, structured pipe with short-lived, limited permissions keeps a break-in small, while a wide-open one lets it through anyway.

open as a page

A platform team runs an active-active geode deployment across four regions on top of a globally-distributed database that supports writes originating from any region. They need to ship a backward-incompatible schema change to a core, heavily-used table. Walk through why this is riskier than the same change in a single-region deployment, and describe a rollout strategy that avoids taking any geode offline or serving inconsistent responses during the migration.

level: principalimportance: should knowfreq 30%

basics

~20 s

In one region, you flip a switch and everyone's on the new version at once. Across four regions, you roll out gradually, so for a while some regions run old code and some run new - meaning the database has to work for both at the same time, or things break for whichever region hasn't updated yet.

open as a page

A platform team running Istio at 3,000-pod scale sees intermittent 503s during rolling deploys and a slow but steady rise in per-node memory pressure over months. Walk through the sidecar-specific failure modes that could explain each symptom.

level: principalimportance: should knowfreq 35%

basics

~20 s

Two separate problems: pods dying before their sidecar finishes forwarding the last requests causes brief errors during deploys, and every pod carrying its own extra proxy process, multiplied by thousands of pods, slowly eats memory across the cluster.

open as a page

When would you deliberately avoid the deployment stamp pattern, and how do teams that do adopt it handle product features that inherently need a view across every stamp, like global search or cross-customer reporting?

level: principalimportance: nice to knowfreq 25%

basics

~20 s

Skip stamping when you're small — running many full stacks for a handful of customers just wastes money and adds complexity you don't need yet. For features needing a global view across all stamps (search, reporting), teams usually copy data out of every stamp into one separate system built just for that.

open as a page

Under what circumstances would offloading assets to object storage plus a CDN actually be the wrong call, and what would you use instead?

level: principalimportance: nice to knowfreq 40%

basics

~20 s

If content changes per user or per request, is accessed rarely, or must never leave certain servers or regions for legal reasons, caching it broadly on a CDN either doesn't help or actively causes problems, so you'd keep it dynamic or restrict it instead.

open as a page

showing 31–44 of 44