In a microservices architecture, what is service discovery and why can't services just call each other using hardcoded IP addresses?
answer
- registry = phone book of live instances
- register on startup / evicted via health check on failure
- ephemeral IPs from autoscaling and rescheduling
- client-side vs server-side lookup
basics
~10 sService discovery lets services find each other's current network location automatically, instead of using fixed IP addresses that change when instances scale, restart, or move.
solid answer
~40 sIn microservices, instances come and go dynamically due to autoscaling, deployments, and failures, so IPs can't be hardcoded. Service discovery solves this with two pieces: a registry that tracks which instances are alive and where, and a discovery mechanism that lets callers look up healthy instances at call time. Instances (or a platform agent) register on startup and deregister (or get evicted via failed health checks) on shutdown/failure. Callers then either query the registry directly (client-side discovery) or go through an infrastructure component like a load balancer or DNS that queries the registry on their behalf (server-side discovery). This decouples service consumers from the physical location of dependencies, letting the platform scale, reschedule, and heal without every caller needing manual reconfiguration.
go deeper
Can explain the core problem (dynamic IPs) and that a registry tracks who's alive, even if fuzzy on registration mechanics.
Names at least one concrete registry (Consul, Eureka, Kubernetes Services) and describes register/deregister plus health checks.
Compares client-side vs server-side discovery and discusses staleness/consistency trade-offs of the registry.
Reasons about discovery as part of overall platform reliability - ties it to deployment strategy, multi-region topology, and registry availability guarantees.
## Why a hardcoded address stops working Service discovery exists because the network location of a microservice instance is no longer a fixed, hand-assigned fact - it's a rapidly changing piece of runtime state. In a monolith or a handful of statically deployed VMs, hardcoding an IP or hostname into a config file was fine because that address rarely changed. In a microservices system running on an autoscaler, a container orchestrator, or spot/preemptible infrastructure, instances of a given service are created and destroyed constantly: - **autoscaling** adds and removes instances in response to load; - **deployments** roll new instances in and retire old ones; - **health-based self-healing** kills and replaces failing instances. Each of these events can hand out a brand-new IP address. If a caller's configuration hardcodes a specific list of IPs, that list starts going stale within minutes of being written, and every subsequent change requires someone to manually edit and redeploy every caller's configuration - a process that doesn't scale past a handful of services and is guaranteed to lag reality. ## The registry - the live list of instances The mechanism that solves this has two halves. The first is a **registry**: a piece of infrastructure - a dedicated tool like `Consul` or Netflix `Eureka`, or a built-in platform feature like Kubernetes' `Service/Endpoints` objects - that maintains a live, queryable list of which instances of which services currently exist and are healthy. Instances participate in keeping this list accurate in one of two ways: 1. by actively **registering** themselves on startup and **deregistering** on graceful shutdown; 2. or more robustly, by **being monitored**: the registry (or an agent) periodically checks each instance's health (an HTTP ping, a TCP connect, or waiting for a heartbeat/renewal signal from the instance itself) and removes any instance that stops responding or stops heartbeating within a configured window. This health-check layer matters because clean deregistration on shutdown can't be relied upon - processes crash, nodes get killed by the underlying infrastructure, and network partitions cut instances off without warning, so the registry needs an independent way to notice an instance has gone bad rather than trusting it to say so itself. ## Discovery itself - the two broad patterns The second half is discovery itself: how a caller turns 'I need to talk to service X' into an actual address it can open a connection to. There are two broad patterns. | Pattern | Who resolves the address | What happens | |---|---|---| | **Client-side discovery** | the calling application | it queries the registry directly (often via a small library), gets back a list of currently healthy instances, and picks one itself using some load-balancing strategy before connecting directly - cutting out a network hop but requiring every caller to speak the registry's protocol | | **Server-side discovery** | an intermediary | the caller just sends its request to a fixed, well-known address - a load balancer, an API gateway, or in Kubernetes, a Service's stable virtual IP - and that intermediary is the one that actually queries the registry and routes the request onward, keeping the caller completely unaware that discovery is even happening | ## The staleness trade-off The trade-off inherent in any discovery system is the gap between reality and what the registry currently believes: there is always some window between an instance actually failing and the registry detecting and removing it, bounded by how frequently health checks or heartbeats run. During that window, a caller can still be handed a dead instance's address, which is why production systems pair discovery with client-side resilience rather than assuming the registry is instantaneously accurate: - short connect timeouts; - retries against a different instance; - circuit breakers. Widening the check interval reduces registry/network load but slows failure detection; shrinking it speeds up detection at the cost of load and a higher chance of false positives from a single transient blip. ## Kubernetes as the everyday case A concrete, ubiquitous real-world instance of this is Kubernetes itself: every Deployment's Pods get created and destroyed constantly by the scheduler, autoscaler, and rolling updates, yet application code just calls a stable Service DNS name like `orders.default.svc.cluster.local` - Kubernetes handles tracking which Pods are currently ready and routing traffic to them, so the developer never hardcodes a Pod IP. This is exactly the problem service discovery is built to solve: letting the platform churn instances freely underneath, while callers keep working against a stable name instead of a fragile, manually maintained list of addresses.
- What happens if a service instance crashes without deregistering itself - how does the system recover?Most registries rely on active health checks or heartbeats rather than trusting a clean deregistration, since crashes rarely give the process a chance to unregister. If an instance misses its heartbeat or fails an HTTP/TCP health check for a configured number of cycles, the registry marks it unhealthy and removes it from the pool of addresses returned to callers, typically within seconds to tens of seconds.
- Is DNS enough for service discovery on its own?Plain DNS with A records is a weak fit because it wasn't built for fast churn - TTLs are often minutes, and DNS doesn't natively carry health status, so callers can keep hitting dead instances until caches expire. Platforms usually pair DNS with short TTLs and an active health-checking control plane (like Kubernetes Services + kube-proxy, or Consul's DNS interface) to make it viable.
It's like a restaurant delivery app: instead of memorizing a chef's home address, you look up 'nearest open restaurant serving X' every time you order, because chefs move, close for the night, or open new locations.
saying these in an interview costs you the question
- says you just hardcode IPs and update them manually when things change
- doesn't mention health checks / failure detection at all
- thinks service discovery is only about DNS
- confuses service discovery with API gateway routing rules