skip to content

Your JVM services discover each other through Eureka today, and the platform is moving onto an orchestrator that already tracks which instances of each service are ready. How would you decide whether Eureka stays?

level: principalimportance: should knowfreq 44%

answer

  1. two registries, one slower
  2. what does it know the platform does not
  3. reach, policy, migration time
  4. self-preservation versus rolling deploys
  5. addresses get reused in a cluster

basics

~20 s

Decide by asking what Eureka still knows that the platform does not. If everything lands in one cluster, the orchestrator's own readiness tracking makes Eureka a slower second source of truth and it should go; keep it only for reach beyond the cluster or for per-caller routing policy.

solid answer

~60 s

The question is whether you now have **two registries disagreeing about the same fleet**. An orchestrator already knows which instances passed their readiness gate and removes them the moment they stop passing; Eureka learns the same facts later, through heartbeats and caches, and self-preservation can make it actively refuse to remove an instance the platform has already taken out of rotation. Running both means the slower, less-informed source is the one your callers route on. So if all services land in one cluster, I would delete Eureka and let the platform own membership. I would keep it in three cases: reach — callers must find instances outside the cluster, on VMs or in another cluster; policy — you genuinely depend on per-caller load balancing such as zone affinity that the platform's flat routing does not express; and time — a long migration where some services are not moved yet. In that case I would run both, shorten the lease, guarantee deregistration on shutdown, and set a date to remove it rather than letting it become permanent.

go deeper

for a junior

Understand the shape of the question: an orchestrator already knows which instances are ready, so a separate registry may be duplicating a job the platform now does.

for a middle

Be able to name the concrete frictions — slower removal through leases and caches, self-preservation firing during rollouts, and deregistration racing the termination grace period.

for a senior

Argue the keep-or-delete case from operational evidence: measure how long callers hold dead addresses under each mechanism, and identify which of reach, routing policy or metadata your system actually depends on.

for a principal

Own the retirement as a programme, not an opinion. Set the criterion in advance, sequence callers rather than registries, and make deleting the transitional system an owned deliverable with a date, because the real failure mode is that it never gets removed.

## Reframe the question before answering it An interviewer is not asking whether Eureka is good. They are asking whether you can recognise that **a capability moved into the platform**, and whether you can retire something that still works. The frame is: what does Eureka know that the platform does not, and is that knowledge worth a second registry? ## What changes underneath you On a VM fleet, Eureka was the only thing that knew where instances lived, and self-registration was the only mechanism available. Under an orchestrator, membership is already tracked as a first-class fact: the platform starts the instance, gates it behind a readiness signal, knows its address, and removes it from the routable set the instant it stops being ready or begins terminating. That is strictly more information than a heartbeat conveys, and it arrives strictly sooner. ## Three concrete frictions when both run **Two sources of truth with different eviction semantics.** The platform removes an instance on failed readiness in seconds. Eureka removes it on lease expiry plus an eviction sweep plus cache turnover, in minutes. Your callers route on the slower one, so the platform's fast, accurate removal buys you nothing at the point where it matters. **Self-preservation fights the orchestrator.** Rolling deploys and autoscaling produce exactly the pattern that trips self-preservation: many instances stopping in a short window. The registry then refuses to evict anything, including instances the platform has already terminated, precisely during the deploy when addresses are churning fastest. **Deregistration becomes a race with the termination grace period.** A pod receiving a stop signal must send its cancel to the registry and finish in-flight work before it is forcibly killed. Get that ordering or that grace window wrong and every rollout leaves a trail of leases pointing at addresses that no longer exist — and in a cluster, addresses get **reused**, so a stale entry can point at a completely different workload rather than at nothing. ## What Eureka still legitimately offers **Reach beyond one cluster.** A registry is just a list; it does not care where a member runs. If callers must find instances across two clusters, or in a cluster and on legacy VMs during a multi-year migration, a single registry spanning both is a real answer and the platform's in-cluster membership is not. **Per-caller routing policy.** Client-side discovery lets each caller apply its own selection — zone affinity, weighting, filtering on instance metadata your application defines. Flat platform routing typically expresses none of that. **Application-level metadata.** Eureka carries arbitrary key/value data about an instance — version, feature capability, region — that callers can filter on. If your routing logic genuinely reads that, moving it costs work. **Independence from cluster networking.** Callers hold addresses and connect directly, with no dependency on the platform's own resolution path being healthy. ## The decision, stated plainly - **Single cluster, no external callers, no per-caller policy → delete it.** Two registries is worse than either alone, and the one you would be keeping is the slower and less-informed of the two. Retire it; do not keep it "just in case", because a registry nobody trusts still gets consulted at 3am. - **Hybrid or multi-cluster reach, or real dependence on client-side policy → keep it, deliberately.** Then make the platform's readiness the authority on whether an instance is usable, keep the Eureka lease short, guarantee cancellation on shutdown inside the grace period, and be honest that during a partition self-preservation is going to serve stale addresses. - **Mid-migration → run both, with an end date.** Keep the registry authoritative only for the population that has not moved, migrate callers service by service, and treat "Eureka removed" as an explicit deliverable with an owner. The failure mode here is organisational, not technical: transitional infrastructure survives when nobody is accountable for deleting it. ## How you would know you were right Measure the window between an instance stopping and the last caller ceasing to dial it, before and after. If Eureka's removal is consistently the long pole, that number *is* the argument for retiring it — and if it is not, you have found a genuine reason it is still earning its place.

  • What is the specific danger of a stale Eureka entry in a cluster, versus on a VM fleet?
    Address reuse. On a static fleet a stale address usually just refuses the connection and the caller fails over. In a cluster the address can be recycled onto an entirely different workload, so a stale entry may resolve to something that answers — wrongly. That turns a clean failure into a silent misroute.
  • If you keep Eureka during a migration, what stops it from becoming permanent?
    A named owner and a removal date treated as a deliverable, plus a shrinking list of services still registered. Transitional infrastructure survives on ambiguity: as long as it works and nobody is accountable for deleting it, it stays and quietly becomes a dependency again.
  • A team argues Eureka should stay because it survives a control-plane outage. How do you weigh that?
    It is a real property — cached client registries keep working when the registry is gone. But weigh it against what else stops working in that outage: if the control plane is down you are likely not scheduling, scaling, or deploying either. Discovery surviving alone rarely changes the blast radius enough to justify a second registry.
  • How would you migrate callers off Eureka without a flag day?
    One caller at a time, not one registry at a time. Keep instances registering in both systems for the transition, switch each calling service's resolution to the platform, and verify against the two lists that they agree for that service before moving the next. Stop double-registering only once no caller reads Eureka.

saying these in an interview costs you the question

  • Keeps Eureka because rewriting callers feels risky, with no other reason
  • Claims Eureka detects failures faster than the orchestrator
  • Ignores that two registries can disagree about the same fleet
  • Treats a service mesh as the automatic answer to every discovery question
  • Overlooks that cluster addresses are recycled onto other workloads

context