Where does Eureka sit on the CAP spectrum, and how do its design choices (peer replication, caching, self-preservation) reflect that? When might you not use it?
answer
- AP: available + partition-tolerant, not consistent
- peers replicate best-effort, no quorum/leader
- vs Consul/Zookeeper = CP (consensus)
- correctness pushed to client: LB retry + circuit breaker
- skip for k8s DNS / service mesh / CP needs
basics
~20 sEureka is an AP system: during a network partition it stays available and returns possibly-stale data rather than blocking to guarantee consistency. Everything about it — peer replication without consensus, client caches, self-preservation — favors availability over a perfectly consistent registry.
solid answer
~50 sEureka deliberately chooses availability and partition-tolerance over consistency (AP), unlike consensus-based registries like Consul or Zookeeper (CP). Its design embodies that: server nodes form a flat peer-to-peer cluster that replicates registrations best-effort with no leader election or quorum, so any node keeps serving reads/writes during a partition and nodes may temporarily diverge. Clients cache the registry locally and poll for deltas, so they keep routing even if all servers are unreachable. Self-preservation stops eviction during suspected partitions to avoid emptying the registry. The trade-off is pervasive staleness — you may route to instances that are gone — so the architecture pushes correctness to the client via load-balancer retries and circuit breakers. You'd skip Eureka when the platform already provides discovery (Kubernetes Services/DNS), when you need strong consistency (CP registry), or for a service mesh where the sidecar/control plane owns discovery.
code
java · 21 lines// HA peer-aware Eureka servers (AP topology) — each node peers with the other
// --- node A (application.yml) ---
// spring.application.name: eureka-server
// eureka:
// client:
// register-with-eureka: true # register self with peers
// fetch-registry: true # fetch peers' registries
// service-url:
// defaultZone: http://eureka-b:8761/eureka/
// --- node B ---
// eureka.client.service-url.defaultZone: http://eureka-a:8761/eureka/
// Replication between peers is asynchronous + best-effort: NO leader election,
// NO quorum. During a partition each node keeps serving; views may diverge and
// reconcile after the partition heals. That is the deliberate AP choice.
// Client resilience is where correctness lives (not the registry):
@Bean
ReactorLoadBalancer<ServiceInstance> loadBalancer() { /* Spring Cloud LoadBalancer w/ retry */ return null; }
// + Resilience4j @CircuitBreaker on the outbound call to shed dead instances.go deeper
Know Eureka favors staying available over being perfectly up-to-date.
State that Eureka is AP and contrast at a high level with a CP registry.
Explain how caching, best-effort peer replication, and self-preservation all encode the AP choice and push correctness to clients.
Reason about topology (peers per zone), consistency trade-offs vs Consul/Zookeeper, and when platform-native discovery or a service mesh makes Eureka the wrong tool.
## CAP framing The **CAP theorem** says that under a network **partition (P)** a distributed system must choose between **consistency (C)** and **availability (A)**. Eureka is unambiguously **AP**: when nodes can't talk, it keeps **serving and accepting registrations** and tolerates **divergent, stale views**, rather than blocking to preserve a single correct answer. Netflix built it for cloud environments where partitions are routine and a registry that goes *unavailable* is worse than one that's *slightly wrong*. ## How the design encodes AP 1. **Peer-to-peer replication, no consensus.** Eureka servers are configured as **mutual peers** (each lists the others in `service-url.defaultZone`). A registration to one node is **asynchronously replicated** to peers on a **best-effort** basis — there is **no leader, no quorum, no Raft/Paxos**. Consequence: nodes can hold **different registries** for a while (eventual consistency), and a write never blocks on agreement. Contrast **Zookeeper/Consul**, which use consensus and will refuse writes / minority reads during a partition to stay consistent (**CP**). 2. **Client-side caching + polling.** Clients hold a full local registry and refresh deltas (~30s). This removes the server from the request hot path and means clients **keep routing even if every Eureka server is down** — pure availability, at the cost of staleness. 3. **Self-preservation.** On a heartbeat drop below the renewal threshold (0.85), the server **stops evicting** rather than risk emptying a healthy registry during a partition — availability over correctness again. 4. **Lease + heartbeat, timer eviction.** Even removal is lazy (up to ~90s + sweep). Nothing in the system is real-time or transactional. ## The systemic consequence Because *every* layer is eventually consistent, **the registry is never ground truth** — a consumer must assume any given instance might already be gone. The architecture therefore **delegates correctness to the caller**: **Spring Cloud LoadBalancer** retries another instance, and a **circuit breaker (Resilience4j)** sheds load from failing endpoints. If you assume the registry is authoritative, you build fragile systems. ## HA topology Production runs **multiple Eureka servers as peers**, typically one per zone/region, each with `register-with-eureka: true` and pointing at the others. Clients list all zones. There's no split-brain resolution because there's nothing to resolve — divergence is acceptable and reconciles when the partition heals. ## When NOT to use Eureka - **Kubernetes-native platforms:** K8s **Services + cluster DNS** (and readiness probes/endpoints) already provide discovery and load balancing; running Eureka on top is usually redundant. Spring Cloud Kubernetes exposes discovery via the K8s API instead. - **Strong-consistency needs:** if you truly need a consistent view (e.g. leader/config coordination), a **CP** store (Zookeeper, Consul, etcd) fits better. - **Service mesh:** with Istio/Linkerd, the **sidecar + control plane** own discovery and routing; app-level Eureka is unnecessary. - **Very small/simple deployments:** static config or DNS may be enough; a registry cluster is operational overhead. - **Maintenance posture:** Netflix Eureka 1.x is in maintenance; some teams pick Consul or platform-native discovery for that reason. ## Gotchas at the architecture level - Don't disable self-preservation in prod to 'get accurate data' — you reintroduce the partition-wipe failure mode. - Sizing the renewal threshold and heartbeat intervals interacts with cluster size; autoscaling changes the expected-renewals baseline. - Cross-region replication is best-effort; design for regional isolation, not a single global consistent registry. ## Bottom line Eureka is a pragmatic **AP** registry: choose it when **availability during partitions** and simple operations matter more than a perfectly consistent instance list, and when the platform doesn't already give you discovery — and always build clients that are correct despite stale data.
- How does Eureka differ from Consul or Zookeeper as a registry?Eureka is AP: peer replication is best-effort with no consensus, so it stays available and eventually consistent during partitions. Consul and Zookeeper are CP: they use consensus (Raft/ZAB) and will sacrifice availability (refuse writes / minority reads) during a partition to keep a strongly consistent view.
- In a Kubernetes deployment, do you still need Eureka?Usually not. Kubernetes Services + cluster DNS plus readiness probes already provide name-based discovery and load balancing. Spring Cloud Kubernetes can expose discovery via the K8s API. Running Eureka on top is typically redundant operational overhead unless you're mixing non-K8s workloads.
saying these in an interview costs you the question
- Calling Eureka a CP or strongly-consistent system.
- Claiming Eureka uses leader election / quorum / Raft like Zookeeper.
- Treating the registry as an authoritative real-time source of truth.
- Recommending Eureka on top of Kubernetes without noting native discovery makes it redundant.