skip to content

Operations & Deployment

Running many services in production: finding each other, staying up under failure, being observable, being deployed safely, and being configured without a redeploy. Operations is where the microservices tax is actually paid.

part ofMicroservices architectureoverview, primer and where to startread it →
on this pageshow

questions

page 2 of 2

A Kubernetes-deployed microservice exposes both a liveness probe and a readiness probe, plus a /metrics endpoint scraped by Prometheus. What's the functional difference between the two probes, and what happens to the service's traffic and pod status if only one of them is misconfigured to always report healthy?

level: principalimportance: should knowfreq 55%

basics

~20 s

Liveness checks whether the app should be restarted because it's stuck; readiness checks whether it should currently receive traffic. If readiness is broken and always says "yes," a struggling instance keeps getting new requests it can't handle while it's failing.

open as a page

You're designing service discovery for a multi-region microservices platform. How do CAP-theorem trade-offs (e.g. Consul's CP/Raft model vs Eureka's AP model) and load-balancing strategy choices factor into deciding between a client-side discovery library, a service mesh, or relying on Kubernetes-native DNS discovery per cluster?

level: principalimportance: should knowfreq 40%

basics

~20 s

Pick based on what you need: strict-consistency registries (Consul/Raft) can refuse to answer during a network partition, while availability-first ones (Eureka) may hand out stale info but never go fully silent. At multi-region scale, most teams end up with per-region discovery plus a mesh or global routing layer to handle traffic across regions, rather than one giant global registry.

open as a page

How does the sidecar container pattern let a microservices team add cross-cutting operational behavior (like traffic encryption, retries, or metrics collection) without changing each service's application code, and what does it cost?

level: seniorimportance: nice to knowfreq 35%

basics

~20 s

A sidecar is a helper container deployed alongside a service's container in the same pod that handles shared plumbing like security or metrics, so the service's own code doesn't need to implement that plumbing itself.

open as a page

In a GitOps pipeline, a team wants to commit their Kubernetes Secret manifests directly into the same git repository that drives deployment (for example via Bitnami's 'Sealed Secrets' controller), including in a repo with broad read access. How can a Secret be safely committed to git, and what does this pattern protect against versus what does it not protect against?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Normally you can't put a real password in git because anyone who can read the repo can read it. Sealed Secrets lets you encrypt the password first with a key only your cluster knows, so the encrypted version is safe to commit — only your cluster can turn it back into the real password.

open as a page

'Progressive delivery' combines canary-style traffic control, feature flags, and automated analysis into one release process. As a principal engineer choosing a deployment strategy per service, when would you deliberately NOT use canary or progressive delivery, even though the tooling is available?

level: principalimportance: nice to knowfreq 45%

basics

~20 s

Sometimes a fancy gradual rollout isn't worth it - for tiny low-traffic services, batch jobs, or changes where a database migration already commits you either way, a simpler rolling deploy with good tests is safer and cheaper than building a whole canary pipeline.

open as a page

A principal engineer is asked whether a 15-service platform running on Kubernetes should adopt a sidecar-based service mesh like Istio. What factors would make this a bad idea, and what does a 'sidecar-less' or 'ambient' mesh architecture change about that calculus?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

For a small number of services, a full mesh can cost more in complexity and per-pod overhead than the traffic-management and mTLS benefits are worth; newer ambient-mesh designs try to get similar benefits without a sidecar in every pod.

open as a page

showing 31–36 of 36