skip to content

Your platform already gives every service a stable internal DNS name that resolves to a health-checked virtual IP. A team proposes running a dedicated service registry alongside it. How would you decide whether that registry is redundant here or genuinely justified?

level: principalimportance: nice to knowfreq 28%

answer

  1. one source of truth for membership
  2. what the VIP structurally cannot do
  3. multiplexed connections pin to one backend
  4. off-platform callers need one namespace
  5. a registry becomes tier-0

basics

~20 s

Decide by naming the capability the virtual IP structurally cannot provide — off-platform members, instance-aware per-request balancing, or routing metadata. Absent one of those, a second registry mainly adds a second, disagreeing answer to who is healthy.

solid answer

~60 s

I would make them name the capability gap, because the platform's name-plus-VIP already covers most of discovery: one source of truth for membership, health-gated inclusion, and server-side balancing with no library in any caller. A second registry duplicates that source of truth, and two systems that both claim to know who is healthy will disagree precisely during an incident. So the question is what the VIP structurally cannot do. Three answers are legitimate. Callers or callees living outside the platform that must appear in the same namespace. Per-request, instance-aware balancing — long-lived multiplexed connections pin to one backend through a VIP, so spreading load needs the caller to see instances individually and often to prefer its own zone. And instance-level metadata that routing depends on, such as version or shard ownership. If none of those applies, it is redundant. If one does, I would still price it honestly: the registry becomes a tier-0 dependency with quorum, capacity and upgrade ownership, and I would require one authoritative source with the other derived from it.

go deeper

for a junior

Know that a virtual IP in front of a pool already hides which instances exist, so callers use one stable name — and that adding a registry on top is a decision with costs, not an automatic upgrade.

for a middle

Be able to contrast the two: a virtual IP balances per connection with no client library, while a registry hands the caller a list of instances and lets it choose per request, at the price of a library and a service to run.

for a senior

Show that you would look for the specific capability gap — off-platform members, skew from multiplexed connections, metadata-driven routing — and that you would define what callers do when the registry is unreachable before shipping it.

for a principal

Own the invariant: one authoritative source of membership. Be prepared to refuse a second registry on the grounds that duplicated truth costs more than the hop it saves, and to say what evidence would change your mind either way.

## Start from the invariant, not the products The invariant worth defending is **one source of truth for membership**. Discovery exists to answer "who should receive traffic for this service right now". A platform that publishes a name resolving to a health-checked virtual IP already answers it, and answers it in a way every caller can consume with no library at all. Introducing a second system that answers the same question is not additive; it creates two answers that must be reconciled, and the reconciliation always comes due at the worst moment. So the review question is not "is a registry good". It is: **what does the caller need that a stable address in front of a health-checked pool structurally cannot give it?** That reframing does most of the work, because it converts an architectural preference into a checkable list. ## What the platform's name-plus-VIP already covers - Membership gated on health, maintained by one component, with a single place to look during an incident. - Server-side selection: the caller connects to one address and the proxy decides which backend serves it. - Zero client surface. No library, no version skew across languages, nothing to upgrade in dozens of services. - A single failure domain to reason about, that the platform team already operates. Against that, a registry's costs are concrete: a component whose unavailability can stop every caller from finding anything, a client library in every language you use, a new set of failure semantics (what does a caller do when the member list comes back empty?), and — the expensive one — a second opinion about health. ## The gaps that genuinely justify one **Population.** If some callers or callees live off the platform — legacy virtual machines, another cluster, a second cloud, an acquisition — they cannot be expressed in the platform's own service abstraction. A registry that spans both worlds gives one namespace instead of two half-namespaces plus a manual bridge. This is the strongest and most common justification. **Instance-aware, per-request balancing.** This is the strongest *technical* argument and the one candidates most often miss. Balancing through a VIP happens per connection. A protocol that multiplexes many requests over one long-lived connection therefore pins all of those requests to whichever backend the connection landed on, and load skews badly — most visibly right after a scale-up, when new backends receive almost nothing. Client-side discovery lets the caller hold instances individually, spread requests across them, prefer its own availability zone, and react to a slow instance per request rather than per connection. If the workload has that shape, no amount of VIP tuning fixes it. **Routing metadata.** If selection depends on facts about the instance — build version, shard ownership, capability flags, canary weight — you need somewhere to attach them and something that acts on them. A VIP membership list generally exposes health and nothing else. **Non-traffic uses.** Sometimes the registry is wanted for lookups, tagging or configuration distribution rather than for routing. That may be a fair reason to run one, but it is a different justification and should be argued on its own terms rather than smuggled in as a discovery decision. ## Reasons that do not survive scrutiny "Client-side discovery is faster because it removes a hop" — measure the hop before spending a tier-0 dependency on it; it is usually a small fraction of the call. "We want to avoid DNS TTLs" — a health-checked VIP has already moved membership out of DNS, so the TTL complaint is about a problem the platform solved. "A registry is best practice" — best practice for systems that lack a platform abstraction, which is precisely the condition being questioned. And "we will run both, they will agree" is the assumption that fails in every incident review. ## If you approve it, what you require - **One authoritative source.** Either the registry derives its membership from the platform, or the platform's members are registered from one place. Never two independent registration paths for the same instance. - **Defined behaviour under registry unavailability.** Callers must serve the last known-good member list rather than treating an empty answer as "no instances" — failing static beats failing closed. - **A convergence objective.** State how quickly a membership change must reach callers and measure it, so the thing you bought the registry for is actually being delivered. - **Ownership.** Quorum sizing, capacity, upgrades and the client library across every language, named and staffed. ## Knowing later that it was wrong Six months on, the signals are unambiguous: the two membership views disagree during incidents; nobody consumes the instance metadata that justified the purchase; teams have quietly added a fallback to the platform name; and the registry appears in postmortems as a contributing factor rather than a mitigation. The counter-signal is equally clear — measurable load evening out across backends, cross-environment callers working without a bridge, and routing decisions that could not have been expressed otherwise.

  • What is the strongest purely technical argument for a registry when a health-checked virtual IP already exists?
    Per-request, instance-aware balancing. Long-lived multiplexed connections are balanced once, at connect time, so through a VIP all their requests pin to a single backend and load skews — worst right after a scale-up. Client-side discovery lets the caller spread per request and prefer its own zone.
  • If you approve it, what would you require before it carries production traffic?
    One authoritative membership source with the other derived from it, defined caller behaviour when the registry is unavailable — serve the last known-good list, never treat empty as no instances — a measured convergence objective, and named ownership for quorum, capacity, upgrades and the client library.
  • How would you know six months later that it was the wrong call?
    The two views disagree during incidents; the instance metadata that justified it is unused; teams have added quiet fallbacks to the platform name; and the registry shows up in postmortems as a contributing cause. Those together mean you bought a failure domain and no capability.

saying these in an interview costs you the question

  • A dedicated registry is always better than platform DNS
  • Running both is fine because they will agree
  • Client-side discovery is faster, so it should be the default
  • A registry is just infrastructure and cannot cause an outage
  • We need a registry because the platform's DNS has TTLs

context