skip to content

Operations & Deployment

Running many services in production: finding each other, staying up under failure, being observable, being deployed safely, and being configured without a redeploy. Operations is where the microservices tax is actually paid.

part ofMicroservices architectureoverview, primer and where to startread it →
on this pageshow

questions

page 1 of 2

Why should application configuration (like database URLs, timeouts, or feature flags) be kept outside the compiled application artifact instead of hardcoded, especially when the same service runs in dev, staging, and production?

level: juniorimportance: must knowfreq 75%

answer

  1. build once deploy everywhere
  2. 12-factor config
  3. env vars vs hardcoded
  4. config != code
  5. secrets need extra protection

basics

~20 s

Because the same built package has to run in different places (laptop, test, live servers) that need different settings. If settings are baked into the code, you'd have to rebuild the whole app just to change a URL, and test settings could leak into production.

solid answer

~30 s

Externalizing config follows the 12-factor app principle: build one artifact once and deploy that same artifact everywhere, letting environment-specific values (DB connection strings, API keys, feature flags, timeouts) come from outside the code — environment variables, mounted files, or a config service. This decouples the deploy pipeline from configuration changes: you can flip a flag or rotate a credential without rebuilding. It reduces blast radius versus conditional branches inside code, and it keeps secrets out of source control.

go deeper

for a junior

Should know config shouldn't be hardcoded and can name env vars or a config file as the mechanism; doesn't need config-server internals.

for a middle

Should articulate the build-once-deploy-everywhere rationale and know Spring profiles or Kubernetes ConfigMaps as concrete mechanisms.

for a senior

Should discuss the secrets-vs-config distinction, fail-fast validation, and config drift as an operational risk across many services.

for a principal

Should reason about config as a cross-cutting platform capability — governance, audit trail, and the trade-off of centralizing vs distributing config ownership across dozens of teams.

## What "externalized configuration" means **Externalized configuration** means any value that changes between deploys or environments without changing the application's behavior in code — connection strings, credentials, thread-pool sizes, timeouts, feature flags, log levels — lives outside the compiled artifact and is injected at runtime. The **12-factor app** methodology's "Config" factor says: build one immutable artifact (a jar, a container image) and supply config at runtime via - environment variables, - mounted config files/ConfigMaps, - or a dedicated config service the app queries at startup. Concretely, a Spring Boot service reads defaults from `application.yml`, and Spring's property resolution order lets environment variables (like `DATABASE_URL`), JVM system properties, or a Config Server override those defaults per profile (`dev`, `staging`, `prod`) — so a Kubernetes ConfigMap or Secret mounted as env vars determines behavior without touching the built jar. ## Why it exists This exists because microservices multiply the number of places configuration lives. If values are hardcoded or baked in at build time, every environment needs its own build, which breaks the promise of "build once, promote the same artifact through the pipeline" and increases the risk that what you tested isn't actually what you ship. Externalizing config also separates **who can change config** (ops, SRE, feature owners) from **who can change code** (gated by code review and CI), enabling faster, lower-risk operational changes — like rotating a credential or tuning a timeout under load — without a full build-and-deploy cycle for every tweak. ## The trade-off, in both directions The trade-off runs both directions. Externalizing adds **indirection**. You now need a place to store, version, and secure config: - env vars in orchestrator specs, - a git-backed config repo, - a config server. And if that store is inconsistent across environments you get "works on my machine" bugs at the config layer instead of the code layer. **Over-externalizing** — turning every constant into a flag — creates config sprawl that's hard to audit, so mature setups add a config schema or validation step to keep it sane. And treating secrets as just another config value is itself a risky trade-off: **secrets need extra protection** (encryption, access auditing, rotation) that ordinary config doesn't, so lumping them into the same plaintext pipeline is a common anti-pattern. ## Failure modes in production In production, several failure modes recur. 1. **A missing or misspelled environment variable** can silently fall back to a default that's wrong for production (say, pointing at an in-memory test database) rather than causing a hard, obvious failure — this is why fail-fast validation at startup matters. 2. **Environment drift** is another: staging's ConfigMap slowly diverges from production's over months of ad hoc tweaks, until a staging-only bug appears in prod, or vice versa. 3. **Type mismatches** — a config value parsed as the string "true" in one environment's loader but expected as a boolean in another — can pass validation in one place and silently misbehave in another. 4. **Config dumped at startup** for debugging purposes can accidentally leak secrets into logs if secrets weren't segregated from ordinary config in the first place. ## Where it shows up A concrete real-world pattern: Kubernetes ConfigMaps and Secrets mounted as environment variables or volumes are now the standard runtime mechanism for this, often paired with a centralized store like **Spring Cloud Config** or **HashiCorp Consul** for per-service, per-environment key-value configuration that's centrally managed, versioned, and audited. This lets a platform team change a shared timeout value for many services in staging with a single commit, and roll it forward to production through the normal promotion pipeline, with zero code changes and zero rebuilds — exactly the operational agility that hardcoded configuration would make impossible.

  • What's the difference between a config value and a secret, and why might you treat them differently?
    A config value like a timeout or a feature flag isn't sensitive if exposed, whereas a secret like a database password grants access if leaked, so secrets need encryption at rest, access auditing, and rotation that ordinary config doesn't. Most shops route secrets through a dedicated store (Vault, a cloud secrets manager) rather than plain ConfigMaps or plaintext env-var files, and restrict who or what can read them. Mixing the two in one plaintext pipeline is a common security anti-pattern.
  • How would you validate that a required config value is present before a service starts serving traffic?
    Fail fast: on startup the service should assert that all required properties are present and well-typed and refuse to start, or fail its readiness probe, rather than starting with a null or default and misbehaving mysteriously on the first request. Spring Boot supports this via `@ConfigurationProperties` classes combined with JSR-303 validation annotations enforced at boot time.
  • If two services need slightly different timeout values in staging vs prod, where should that difference live?
    In profile- or environment-scoped config layers, such as Spring's `application-staging.yml` / `application-prod.yml` or separate namespace ConfigMaps, rather than in code branches like `if (env == "prod")`. That keeps the built artifact identical across environments and makes the difference visible and auditable in one place.

Like a stage play using the same script (the built app) in every city, just with different set dressing for each theater (env-specific config) — you don't rewrite the script per city, you just change what's on stage.

saying these in an interview costs you the question

  • Hardcodes environment checks like if (env == 'prod') in application code
  • Stores DB passwords in the same plaintext file as ordinary timeouts
  • Doesn't mention build-once-deploy-everywhere or an equivalent idea
  • Thinks rebuilding the app is the normal way to change a URL
  • No fail-fast validation for missing required config

context

open as a page

Why do teams building a microservices system typically package each service as a container image rather than installing it directly onto a shared virtual machine?

level: juniorimportance: must knowfreq 75%

basics

~10 s

A container bundles a service with everything it needs to run, so it starts the same way everywhere and doesn't fight with other services sharing a machine over conflicting libraries or versions.

open as a page

In a blue-green deployment, two identical production environments (blue = live, green = idle) exist side by side. Explain how traffic cutover works and what happens if the new version needs to be rolled back.

level: juniorimportance: must knowfreq 75%

basics

~20 s

You run two copies of the app - one live, one idle with the new version. Once the idle one is tested and ready, you flip a switch (router/load balancer) so all traffic goes to it. If something breaks, you flip back instantly.

open as a page

In a microservice that calls a downstream payment service over HTTP, what problem does wrapping that call in a circuit breaker solve, and how does the breaker decide when to stop letting calls through?

level: juniorimportance: must knowfreq 85%

basics

~20 s

A circuit breaker watches for repeated failures calling another service and, once it sees too many, stops sending new requests for a while so the failing service isn't overloaded and your own service doesn't get stuck waiting.

open as a page

In a microservices architecture, what is service discovery and why can't services just call each other using hardcoded IP addresses?

level: juniorimportance: must knowfreq 85%

basics

~10 s

Service discovery lets services find each other's current network location automatically, instead of using fixed IP addresses that change when instances scale, restart, or move.

open as a page

In a sidecar-based service mesh like Istio or Linkerd, what is a sidecar proxy and how does it change the way one microservice's network traffic reaches another?

level: juniorimportance: must knowfreq 70%

basics

~10 s

A sidecar is a small helper program that runs next to each service and handles all its network traffic — sending, receiving, encrypting — so the service's own code doesn't have to.

open as a page

A team runs a Spring Cloud Config Server backed by a Git repository, serving configuration to 40 microservices. Walk through how a configuration change made in a git commit reaches a running service, and what has to happen for services to pick up the change without a restart.

level: middleimportance: must knowfreq 70%

basics

~20 s

The config values live in a separate git repo. A central 'config server' reads that repo and hands values to each service on request. Editing the repo alone doesn't reach running services — they either re-fetch on next restart or must be explicitly told to refresh live.

open as a page

In Kubernetes, how do a Deployment and a Service work together to let a microservice be released and scaled without breaking the services that call it?

level: middleimportance: must knowfreq 80%

basics

~10 s

A Deployment manages the running copies of a service and replaces them gradually during updates, while a Service gives callers one stable address that always points at whichever copies are currently healthy.

open as a page

Describe how a canary deployment progressively shifts traffic to a new service version, and what signals should automatically trigger a rollback.

level: middleimportance: must knowfreq 85%

basics

~20 s

You send a new version to a small slice of real users first (like 5%), watch error rates and latency, and only widen the slice if it looks healthy. If it looks bad, you send everyone back to the old version.

open as a page

A service calls three downstream dependencies (inventory, pricing, and recommendations) using a shared thread pool for all outbound HTTP calls. Recommendations starts responding slowly. Why can that alone stall inventory and pricing calls too, and what does the bulkhead pattern do about it?

level: middleimportance: must knowfreq 65%

basics

~20 s

If all outgoing calls share one pool of worker threads, a slow dependency can hog every thread, leaving none free for calls to healthy dependencies too. A bulkhead gives each dependency its own separate, limited pool so one slow dependency can't starve the rest.

open as a page

A client calls a downstream service and gets a connection timeout. It retries three times with exponential backoff and jitter before giving up. Why is the jitter necessary in addition to the exponential growth, and what can go wrong if retries are added without a cap on total attempts or a retry budget?

level: middleimportance: must knowfreq 80%

basics

~10 s

Adding random jitter to retry delays stops many clients from all retrying at exactly the same moment and slamming the recovering service again; without limits, retries can multiply traffic and make an outage worse.

open as a page

Compare client-side and server-side service discovery: who queries the registry, who performs load balancing, and what does the calling service need to know about in each pattern?

level: middleimportance: must knowfreq 80%

basics

~10 s

Client-side: the calling app looks up healthy instances itself and picks one. Server-side: the caller just calls a fixed address, and a separate component (load balancer/proxy) looks up instances and forwards the request.

open as a page

How do service registries like Consul and Eureka determine that an instance is unhealthy and stop routing traffic to it? Contrast active health checks with heartbeat/lease-based detection.

level: middleimportance: must knowfreq 75%

basics

~20 s

Consul actively pings each instance, and Eureka's instances periodically say 'I'm alive.' If pings fail or the 'I'm alive' messages stop arriving within a set time window, the registry marks that instance as down and stops handing out its address.

open as a page

In Istio's architecture, what specifically does the control plane (istiod) do versus what the data plane (the Envoy sidecars) does, and how does configuration get from one to the other?

level: middleimportance: must knowfreq 65%

basics

~10 s

The control plane is the 'brain' that decides the rules for routing and security; the data plane is all the sidecar proxies that actually carry traffic following those rules.

open as a page

How does a service mesh like Istio establish mutual TLS (mTLS) between two services automatically, without either service's code doing any TLS handshake itself, and what is PERMISSIVE mode for?

level: middleimportance: must knowfreq 60%

basics

~10 s

The sidecars on both ends do the encryption handshake for the services automatically, using certificates the mesh hands out and rotates — the app code just sends plain traffic to its own sidecar.

open as a page

You need a feature-toggle system used by 100+ microservices to gradually roll out a new checkout flow. What toggle types would you consider, where should toggle state live, and what should happen to a request if the toggle-evaluation service is unavailable mid-request?

level: seniorimportance: must knowfreq 60%

basics

~20 s

Use an on/off switch that can target a percentage of users, store its state somewhere fast and locally readable rather than deep inside one remote service, and make sure that if the switch-checking system goes down, requests fall back to a safe default instead of failing or exposing the risky new feature to everyone.

open as a page

When a microservices platform runs dozens of small services as containers under an orchestrator, what production failure modes tend to show up that a smaller, single-deployable system wouldn't hit, and how do teams mitigate them?

level: seniorimportance: must knowfreq 70%

basics

~10 s

With many independently-run services, a bad rollout, resource fights between services, or mismatched versions across services can each cause outages a single app never would; teams mitigate with limits, gradual rollouts, and contract testing.

open as a page

How do feature flags let a team decouple 'deploying code' from 'releasing a feature,' and what operational costs does relying heavily on flags introduce?

level: seniorimportance: must knowfreq 80%

basics

~20 s

Feature flags are on/off switches in the code. You can ship new code to production turned off, then flip it on for some or all users later without a new deployment - and flip it back off fast if it breaks.

open as a page

When a single user-facing request fans out across five microservices, what mechanism lets a tool like Jaeger reconstruct the full call tree and show which of those five services caused the added latency, and what has to happen at each service boundary for that to work?

level: seniorimportance: must knowfreq 75%

basics

~20 s

Each call in the chain gets tagged with a shared trace ID, and each step gets its own ID linked back to whoever called it. Tools like Jaeger stitch these into one timeline showing where time actually went.

open as a page

When a platform team configures retry and timeout policy in a service mesh's data plane (e.g. Istio's VirtualService/DestinationRule) instead of in each service's application code, what do they gain, what do they give up, and where can this go wrong?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Moving retries/timeouts into the mesh means one shared config controls every service's behavior instead of each team writing its own — easier to standardize, but harder for a single service to fine-tune its own special cases, and blind retries can make an outage worse.

open as a page

What's the practical difference between storing a database password as a Kubernetes Secret (base64-encoded, mounted into a pod) versus fetching it at runtime from HashiCorp Vault using dynamic, short-lived credentials? What operational problem does each approach create?

level: middleimportance: should knowfreq 65%

basics

~20 s

A Kubernetes Secret is like a hidden config file with the password stored as-is (just encoded, not encrypted by default), and it doesn't expire. Vault instead hands out a temporary password that auto-expires, so a leak stops working soon. Vault is safer but more complex to run.

open as a page

Why does scaling a microservices system out (e.g., via Kubernetes Horizontal Pod Autoscaling) per-service give different cost/performance characteristics than scaling a monolith, and what new operational overhead does per-service scaling introduce?

level: middleimportance: should knowfreq 65%

basics

~10 s

You can add more copies of just the busy service instead of the whole app, saving resources, but now you have many separate things to size, monitor, and tune instead of just one.

open as a page

In a rolling deployment, instances of a service are updated a few at a time while old and new versions both serve traffic simultaneously. What risks does this mixed-version window create, and how do you mitigate them?

level: middleimportance: should knowfreq 70%

basics

~20 s

You replace old servers with new ones in small batches instead of all at once, so there's no downtime, but for a while some users hit the old version and some hit the new one - they both need to work together correctly.

open as a page

A central config server outage blocks several services that are mid-deploy from starting, because they can't fetch their configuration at bootstrap. How would you architect configuration delivery so that a config-server outage doesn't cascade into a platform-wide outage?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Don't make every service's ability to even start depend on one central system being up right now. Give each service a fallback: a locally cached copy of its last-known-good config it can start from if the central source is unreachable.

open as a page

Design an automated rollback system for a service deployment pipeline: what should trigger it, and what can cause an automated rollback itself to fail or make things worse?

level: seniorimportance: should knowfreq 65%

basics

~10 s

The pipeline watches health metrics right after a deploy, and if things look bad (errors spike, requests fail), it automatically switches back to the previous version without waiting for a human to notice.

open as a page

Three microservices each write their own log lines to separate log streams while handling one incoming HTTP request. What has to be threaded through that request for an engineer to later pull up every log line from all three services that belongs to that one request, and how does this relate to (but differ from) distributed tracing?

level: seniorimportance: should knowfreq 60%

basics

~20 s

A shared ID gets attached to a request and passed to every service it touches. If each service logs that ID on every line, you can search for the ID and pull all related lines from every service, in order.

open as a page

What are the main production failure modes of service discovery systems - such as stale registry entries, a registry outage, or a 'thundering herd' on registry recovery - and how do teams mitigate them?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Registries can hand out addresses for instances that already died (stale data), can go down themselves and block all lookups, or can get hammered by every service reconnecting at once after an outage. Mitigations include client-side caching with fallback, retries, and staggered reconnects.

open as a page

How does Kubernetes implement service discovery under the hood - what roles do the Service object, CoreDNS, and kube-proxy each play, and how does this differ from a registry-based tool like Consul or Eureka?

level: seniorimportance: should knowfreq 70%

basics

~20 s

Kubernetes gives every Service a stable name and virtual IP. CoreDNS resolves the name to that IP, and kube-proxy sets up networking rules on each node so traffic to that IP gets sent to one of the healthy Pods behind it - no app-level registry client needed.

open as a page

A service mesh is often sold on giving you 'free' observability — uniform metrics, distributed tracing, and access logs for every service without code changes. What exactly does the mesh actually provide here, what's the catch, and what production failure mode commonly trips up teams relying on it?

level: seniorimportance: should knowfreq 50%

basics

~20 s

The mesh's sidecars automatically record how long calls took and whether they succeeded, for every service, without any app code — but they can't fully trace a request's journey unless the app helps pass along a few tracing headers.

open as a page

When would you deliberately choose NOT to containerize and orchestrate a microservice under something like Kubernetes, even though the rest of your system uses that model?

level: principalimportance: should knowfreq 40%

basics

~20 s

When the overhead of running an orchestrator (people, tooling, complexity) costs more than what you get from it — e.g., a tiny team, very few services, or a workload that doesn't fit the container/orchestrator model well.

open as a page

showing 1–30 of 36