skip to content

Service Discovery

Instances come and go, so callers need a way to find healthy ones: client-side or server-side discovery backed by a registry such as Consul, Eureka or Kubernetes DNS. Health checks are what keep discovery from confidently returning dead instances.

part ofMicroservices architectureoverview, primer and where to startread it →
on this pageshow

questions

6

In a microservices architecture, what is service discovery and why can't services just call each other using hardcoded IP addresses?

level: juniorimportance: must knowfreq 85%

answer

  1. registry = phone book of live instances
  2. register on startup / evicted via health check on failure
  3. ephemeral IPs from autoscaling and rescheduling
  4. client-side vs server-side lookup

basics

~10 s

Service discovery lets services find each other's current network location automatically, instead of using fixed IP addresses that change when instances scale, restart, or move.

solid answer

~40 s

In microservices, instances come and go dynamically due to autoscaling, deployments, and failures, so IPs can't be hardcoded. Service discovery solves this with two pieces: a registry that tracks which instances are alive and where, and a discovery mechanism that lets callers look up healthy instances at call time. Instances (or a platform agent) register on startup and deregister (or get evicted via failed health checks) on shutdown/failure. Callers then either query the registry directly (client-side discovery) or go through an infrastructure component like a load balancer or DNS that queries the registry on their behalf (server-side discovery). This decouples service consumers from the physical location of dependencies, letting the platform scale, reschedule, and heal without every caller needing manual reconfiguration.

go deeper

for a junior

Can explain the core problem (dynamic IPs) and that a registry tracks who's alive, even if fuzzy on registration mechanics.

for a middle

Names at least one concrete registry (Consul, Eureka, Kubernetes Services) and describes register/deregister plus health checks.

for a senior

Compares client-side vs server-side discovery and discusses staleness/consistency trade-offs of the registry.

for a principal

Reasons about discovery as part of overall platform reliability - ties it to deployment strategy, multi-region topology, and registry availability guarantees.

## Why a hardcoded address stops working Service discovery exists because the network location of a microservice instance is no longer a fixed, hand-assigned fact - it's a rapidly changing piece of runtime state. In a monolith or a handful of statically deployed VMs, hardcoding an IP or hostname into a config file was fine because that address rarely changed. In a microservices system running on an autoscaler, a container orchestrator, or spot/preemptible infrastructure, instances of a given service are created and destroyed constantly: - **autoscaling** adds and removes instances in response to load; - **deployments** roll new instances in and retire old ones; - **health-based self-healing** kills and replaces failing instances. Each of these events can hand out a brand-new IP address. If a caller's configuration hardcodes a specific list of IPs, that list starts going stale within minutes of being written, and every subsequent change requires someone to manually edit and redeploy every caller's configuration - a process that doesn't scale past a handful of services and is guaranteed to lag reality. ## The registry - the live list of instances The mechanism that solves this has two halves. The first is a **registry**: a piece of infrastructure - a dedicated tool like `Consul` or Netflix `Eureka`, or a built-in platform feature like Kubernetes' `Service/Endpoints` objects - that maintains a live, queryable list of which instances of which services currently exist and are healthy. Instances participate in keeping this list accurate in one of two ways: 1. by actively **registering** themselves on startup and **deregistering** on graceful shutdown; 2. or more robustly, by **being monitored**: the registry (or an agent) periodically checks each instance's health (an HTTP ping, a TCP connect, or waiting for a heartbeat/renewal signal from the instance itself) and removes any instance that stops responding or stops heartbeating within a configured window. This health-check layer matters because clean deregistration on shutdown can't be relied upon - processes crash, nodes get killed by the underlying infrastructure, and network partitions cut instances off without warning, so the registry needs an independent way to notice an instance has gone bad rather than trusting it to say so itself. ## Discovery itself - the two broad patterns The second half is discovery itself: how a caller turns 'I need to talk to service X' into an actual address it can open a connection to. There are two broad patterns. | Pattern | Who resolves the address | What happens | |---|---|---| | **Client-side discovery** | the calling application | it queries the registry directly (often via a small library), gets back a list of currently healthy instances, and picks one itself using some load-balancing strategy before connecting directly - cutting out a network hop but requiring every caller to speak the registry's protocol | | **Server-side discovery** | an intermediary | the caller just sends its request to a fixed, well-known address - a load balancer, an API gateway, or in Kubernetes, a Service's stable virtual IP - and that intermediary is the one that actually queries the registry and routes the request onward, keeping the caller completely unaware that discovery is even happening | ## The staleness trade-off The trade-off inherent in any discovery system is the gap between reality and what the registry currently believes: there is always some window between an instance actually failing and the registry detecting and removing it, bounded by how frequently health checks or heartbeats run. During that window, a caller can still be handed a dead instance's address, which is why production systems pair discovery with client-side resilience rather than assuming the registry is instantaneously accurate: - short connect timeouts; - retries against a different instance; - circuit breakers. Widening the check interval reduces registry/network load but slows failure detection; shrinking it speeds up detection at the cost of load and a higher chance of false positives from a single transient blip. ## Kubernetes as the everyday case A concrete, ubiquitous real-world instance of this is Kubernetes itself: every Deployment's Pods get created and destroyed constantly by the scheduler, autoscaler, and rolling updates, yet application code just calls a stable Service DNS name like `orders.default.svc.cluster.local` - Kubernetes handles tracking which Pods are currently ready and routing traffic to them, so the developer never hardcodes a Pod IP. This is exactly the problem service discovery is built to solve: letting the platform churn instances freely underneath, while callers keep working against a stable name instead of a fragile, manually maintained list of addresses.

  • What happens if a service instance crashes without deregistering itself - how does the system recover?
    Most registries rely on active health checks or heartbeats rather than trusting a clean deregistration, since crashes rarely give the process a chance to unregister. If an instance misses its heartbeat or fails an HTTP/TCP health check for a configured number of cycles, the registry marks it unhealthy and removes it from the pool of addresses returned to callers, typically within seconds to tens of seconds.
  • Is DNS enough for service discovery on its own?
    Plain DNS with A records is a weak fit because it wasn't built for fast churn - TTLs are often minutes, and DNS doesn't natively carry health status, so callers can keep hitting dead instances until caches expire. Platforms usually pair DNS with short TTLs and an active health-checking control plane (like Kubernetes Services + kube-proxy, or Consul's DNS interface) to make it viable.

It's like a restaurant delivery app: instead of memorizing a chef's home address, you look up 'nearest open restaurant serving X' every time you order, because chefs move, close for the night, or open new locations.

saying these in an interview costs you the question

  • says you just hardcode IPs and update them manually when things change
  • doesn't mention health checks / failure detection at all
  • thinks service discovery is only about DNS
  • confuses service discovery with API gateway routing rules

context

open as a page

Compare client-side and server-side service discovery: who queries the registry, who performs load balancing, and what does the calling service need to know about in each pattern?

level: middleimportance: must knowfreq 80%

basics

~10 s

Client-side: the calling app looks up healthy instances itself and picks one. Server-side: the caller just calls a fixed address, and a separate component (load balancer/proxy) looks up instances and forwards the request.

open as a page

How do service registries like Consul and Eureka determine that an instance is unhealthy and stop routing traffic to it? Contrast active health checks with heartbeat/lease-based detection.

level: middleimportance: must knowfreq 75%

basics

~20 s

Consul actively pings each instance, and Eureka's instances periodically say 'I'm alive.' If pings fail or the 'I'm alive' messages stop arriving within a set time window, the registry marks that instance as down and stops handing out its address.

open as a page

What are the main production failure modes of service discovery systems - such as stale registry entries, a registry outage, or a 'thundering herd' on registry recovery - and how do teams mitigate them?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Registries can hand out addresses for instances that already died (stale data), can go down themselves and block all lookups, or can get hammered by every service reconnecting at once after an outage. Mitigations include client-side caching with fallback, retries, and staggered reconnects.

open as a page

How does Kubernetes implement service discovery under the hood - what roles do the Service object, CoreDNS, and kube-proxy each play, and how does this differ from a registry-based tool like Consul or Eureka?

level: seniorimportance: should knowfreq 70%

basics

~20 s

Kubernetes gives every Service a stable name and virtual IP. CoreDNS resolves the name to that IP, and kube-proxy sets up networking rules on each node so traffic to that IP gets sent to one of the healthy Pods behind it - no app-level registry client needed.

open as a page

You're designing service discovery for a multi-region microservices platform. How do CAP-theorem trade-offs (e.g. Consul's CP/Raft model vs Eureka's AP model) and load-balancing strategy choices factor into deciding between a client-side discovery library, a service mesh, or relying on Kubernetes-native DNS discovery per cluster?

level: principalimportance: should knowfreq 40%

basics

~20 s

Pick based on what you need: strict-consistency registries (Consul/Raft) can refuse to answer during a network partition, while availability-first ones (Eureka) may hand out stale info but never go fully silent. At multi-region scale, most teams end up with per-region discovery plus a mesh or global routing layer to handle traffic across regions, rather than one giant global registry.

open as a page