skip to content

Two Eureka servers are configured as peers. An instance registers against server A, but a client reading server B does not see it. How does replication between Eureka peers actually work, and what would you check?

level: seniorimportance: nice to knowfreq 34%

answer

  1. no leader, no quorum, full mesh
  2. each server replays writes itself
  3. marked replication, never forwarded on
  4. renewals repair missing registrations
  5. one shared hostname breaks it

basics

~20 s

Each Eureka server forwards writes directly to every configured peer over HTTP, asynchronously and best-effort, marked so the peer does not forward them again. There is no leader and no quorum, so peers converge eventually rather than immediately.

solid answer

~50 s

A Eureka server is also a Eureka client of its peers. When it accepts a registration, renewal, cancellation or status change, it replays that write to every peer it has configured, as a batched async HTTP call flagged as a replication so the receiver applies it without forwarding it on. There is no leader election, no quorum, and no rollback if a peer call fails — the cluster is eventually consistent by design, which is the same availability-first choice as self-preservation. Divergence heals mostly through heartbeats: a replicated renewal for an instance a peer has never heard of comes back as a not-found, which prompts a re-registration on that peer. So for the missing entry I would check three things: whether each server lists every *other* server individually by a resolvable address rather than one shared virtual hostname, whether server B is simply serving a cached read that has not turned over yet, and whether B was recently restarted and is still inside its startup registry sync.

go deeper

for a junior

Know that every Eureka server keeps a full copy of the registry and that servers copy changes to each other, so two servers can briefly disagree.

for a middle

Explain that replication is direct, asynchronous, best-effort and one hop only, and that there is no leader or quorum anywhere in the design.

for a senior

Diagnose with it: check peer URLs are individual and resolvable, rule out the server's read cache before blaming replication, test connectivity in both directions, and allow for startup sync on a restarted node.

for a principal

Be able to defend the choice — full replication with no consensus keeps every node writable during a partition, which suits a registry whose clients already cache and already fail over, and is disqualifying for anything that must be a store of record.

## The peer model Eureka's clustering is deliberately the least sophisticated thing that could work. There is **no leader, no consensus protocol, and no quorum**. Every server holds a full copy of the registry, accepts writes from any client, and forwards each write to its peers. Two nodes, five nodes, one per availability zone — the topology is a fully connected mesh described by configuration. A Eureka server is itself a Eureka client: it can register with its peers and it fetches their registries. That is what makes bootstrapping work at all. ## What gets replicated, and how The four write operations — **register, renew, cancel, status change** — are each replayed to every configured peer as an HTTP call, batched and issued asynchronously. Two properties follow: **It is best-effort.** The write to the local registry succeeds and the client is told so *before* peers have acknowledged anything. If a peer call fails or times out, there is no rollback and no retry queue that guarantees delivery. The cluster diverges and is expected to reconverge. **It is not transitive.** A replicated request is marked as a replication, and a server applies such a request without re-forwarding it. That prevents an infinite forwarding loop in a mesh, but it also means the originating server is solely responsible for reaching *all* peers. This is not gossip — a peer that server A cannot reach does not learn the change from server C. ## How divergence heals The self-healing mechanism is the heartbeat itself. Renewals are replicated too, so an instance registered only on A produces a stream of replicated renewals arriving at B. When B receives a renewal for an instance it has no lease for, it answers not-found, and that response tells the origin to send a full registration instead. Over a heartbeat interval or two, a missed registration repairs itself. A missed *cancel* is repaired more crudely — by the lease simply expiring on the peer that never heard about it, assuming that peer is not in self-preservation. One consequence worth knowing: each peer stamps leases with **its own** clock and computes **its own** self-preservation threshold from the renewals it personally receives. Peers can therefore hold slightly different views of expiry, and one peer can be suppressing eviction while another is evicting normally. ## What to check when a peer is missing an entry **1. Peer URLs.** Each server must list the *other* servers by an address that resolves to that specific node. The classic misconfiguration is pointing every server at one shared virtual hostname or load-balancer address, so that servers end up replicating to themselves or to an arbitrary node, and registrations land wherever the balancer happened to route. List peers individually. **2. Read caching, not replication.** Before concluding replication failed, remember that a Eureka server serves reads from a cached response that refreshes roughly every 30 seconds. An entry can be perfectly replicated and still be invisible in B's responses for that long. Give it a minute before hunting. **3. Startup sync.** A freshly restarted server first tries to fill its registry from its peers. If that sync comes back empty it deliberately waits before serving traffic, so that it does not answer with a nearly empty registry and, worse, start evicting things it never learned about. A server queried during that window looks broken and is not. **4. Connectivity in both directions.** Because replication is per-origin, a one-way network problem produces exactly this asymmetry: registrations made against A appear on B, but registrations made against B never reach A. Test both directions rather than assuming a symmetric failure. **5. Which server did the client actually use?** Clients are configured with a list of server URLs and use one of them. "The instance is missing" is often "the instance registered against the peer the caller does not read from, and has not converged yet". ## The honest summary Eureka's replication buys availability — any server can accept any write while any other is unreachable — at the cost of any consistency guarantee at all. It is the correct trade for a registry whose consumers already cache stale data and already have to survive a wrong address. It would be the wrong trade for a store of record, which is exactly why registries built on a consensus protocol behave so differently under the same partition.

  • Why must a replicated write be flagged so the receiving peer does not forward it again?
    Because the peers form a fully connected mesh. If every server forwarded every write it received, a single registration would bounce between nodes indefinitely, multiplying traffic without adding information. The flag makes replication exactly one hop, which is also why the originating server must reach every peer itself.
  • Can two Eureka peers report different registries at the same moment, and is that a bug?
    They can, and it is not. There is no synchronous replication and no quorum, so peers converge only eventually. Clients are expected to tolerate that, exactly as they tolerate their own cached copy being stale — the address list is a hint that gets corrected when a connection fails.
  • What breaks if all Eureka servers are placed behind a single load-balanced hostname?
    Both peering and client behaviour become non-deterministic. A server replicating to that hostname may reach itself or an arbitrary peer, and a client's registration lands on whichever node the balancer chose, so retries and reads can hit a node that has not converged. Peers should be addressed individually, and clients given the full list.

saying these in an interview costs you the question

  • Describes Eureka peers as electing a leader
  • Assumes a write is acknowledged only after peers confirm it
  • Thinks peers gossip changes onward to each other
  • Puts all Eureka servers behind one load-balanced hostname
  • Expects two peers to always return identical registries

context