skip to content

Service Names & Resolution

A stable name and virtual address in front of replicas that are replaced constantly, and the platform's own resolution behind it. Asked because callers must survive every replica changing.

on this pageshow

questions

4

A scoring service's replicas are replaced many times a day; what does its stable service name resolve to, and why does that survive the churn?

level: middleimportance: must knowfreq 74%

answer

  1. one name outlives every replica
  2. the address belongs to the declaration
  3. only the member list changes
  4. readiness decides who is in the set
  5. a member is picked per new connection

basics

~20 s

It resolves to a virtual address allocated to the service declaration itself, not to any replica. Replica churn changes only the member set behind that address, so the name and address a caller holds stay valid for as long as the service exists.

solid answer

~40 s

Declaring a service in front of a replica set makes the platform allocate two long-lived things that belong to the declaration rather than to any replica: a name resolvable inside the platform, and, in the usual mode, a virtual address that the name resolves to. Nothing listens on that virtual address; dispatch rules programmed on every host recognise traffic aimed at it and rewrite the destination to one address from the current member set. That member set is the only part that churn touches — as replicas are created and destroyed, addresses enter and leave the set while the name and the virtual address stay put. So the caller holds a fact with a long lifetime, and the platform keeps a short-lived list current on its behalf.

go deeper

for a junior

Remember the shape: callers address a service by a name, not by a replica's address, because replica addresses are handed out and taken back constantly.

for a middle

Be able to say what is allocated to the declaration (a name and a virtual address), what is maintained behind it (a set of ready members), and that a member is picked when a connection is opened.

for a senior

Show where this leaks in production: the pick is per connection, so pooled callers pin a member, and a caller that stores a member address has undone the abstraction.

for a principal

The judgment is about what you make the default for every team — one indirection the platform owns, versus callers holding member lists and owning the churn themselves.

## Why a replica's own address is not an answer A replica gets an address when it starts and gives it up when it stops. On a platform that replaces replicas many times a day — a new build rolled out, a host drained, a crashed process rescheduled elsewhere — any replica address a caller has written down is a fact with a lifetime of minutes. The service abstraction exists to hand callers something that outlives what it points at: **one name, declared once, that stays valid while the replicas behind it are created and destroyed**. ## What the name resolves to When you declare a service in front of a set of replicas, the platform allocates two things that belong to the *declaration*, not to any replica: - a **stable name**, resolvable inside the platform by every workload allowed to reach it; - in the usual mode, a **virtual address** that the name resolves to. The virtual address is deliberately strange: it is not configured on any interface, and no process anywhere is listening on it. It is a token. Dispatch rules programmed on every host recognise traffic aimed at it and rewrite the destination to the address of one current member. Both the name and the virtual address stay valid for the life of the service declaration — scaling from three replicas to thirty, replacing every replica, or landing the workload on entirely different hosts does not move them. ## The member set behind it Behind the virtual address the platform maintains a **member set**: the addresses of the replicas that match the service's label selector *and* currently report ready. The platform keeps that set current as replicas appear and disappear, and pushes each change out to the dispatch rules on every host. The set, not the name, is what churn touches. The division of labour is worth stating plainly: - the **caller** holds a name and nothing else; - the **platform** holds the mapping from name to virtual address, and from virtual address to the current members; - the **replicas** hold nothing about each other at all. ## What happens to one call 1. The caller resolves the service name through the platform's own resolution and gets the virtual address. 2. The caller opens a connection to that address. 3. Dispatch rules on the caller's own host pick one member from the current set and rewrite the destination to that member's address. 4. Connection tracking keeps that established flow pointed at the same member until the connection closes. 5. A replica is replaced: the set loses one address and gains another, and the *next* connection picks from the new set. The caller's name, its resolved address and its code are untouched. Step 3 is the one candidates skip, and it is where the value is: the choice of replica is made by the platform, per connection, on the machine the caller is running on — not by the caller, and not by a single box somewhere in the middle. ## The other mode, where the name is the set Platforms also offer a mode with no virtual address, in which the name resolves to the member addresses themselves and each replica may additionally carry its own stable per-instance name. There is then no platform dispatch, and the caller picks a member itself. | | stable name plus virtual address | direct instance addressing | |---|---|---| | the name resolves to | one address that outlives replicas | the current member addresses | | who picks a member | dispatch rules on each host | the caller | | replica churn is | invisible to the caller | visible in every answer it resolves | | suits | ordinary request traffic | reaching one particular member | Designs genuinely differ in where that dispatch runs and how the rules are programmed, and some platforms put a hop in the path that picks per request rather than per connection. What does not differ is the shape: a long-lived name in front of a short-lived set. ## What the abstraction does not buy you - It does not re-pick for a connection that is **already established**. The pick happens at connect time, so a pooled connection held for hours stays on the member it first landed on. - It does not make a member correct. A replica that reports ready and then goes subtly wrong keeps receiving its share. - It does not retry your failed call, and it promises nothing about the order members are chosen in. - It does not help a caller that goes around it. A caller that resolves a *member* address once and stores it has re-created exactly the problem the name existed to solve. That last point is the practical test of whether someone understands the abstraction: the stable thing must be the thing you hold.

  • Does the virtual address change when the service is scaled from three replicas to thirty?
    No. Scaling changes only the member set behind the address — twenty-seven addresses are added to it as those replicas start reporting ready. The name and the virtual address were allocated to the service declaration and are unaffected by how many replicas sit behind them, which is precisely why callers need no redeploy when a service is scaled.
  • What breaks if a caller resolves the name once at start-up and holds that answer for a week?
    In the virtual-address mode, very little: the answer it cached is the service's own address, which is still valid. In the mode where the name resolves to member addresses directly, almost everything: the addresses it cached belong to replicas that are long gone, and an address can even have been handed to an unrelated workload since.
  • Nothing is listening on the virtual address — so what happens to a packet sent to it?
    Dispatch rules on the sending host match it, choose one address from the current member set, and rewrite the destination before the packet leaves. The reply is rewritten back so the caller sees answers from the address it addressed. The virtual address therefore never needs to exist on any interface.

A team's published extension rings whoever is on shift: the number never changes, the person answering is whoever is currently available, and a call already answered stays with that person until it ends.

saying these in an interview costs you the question

  • Says the name resolves to one chosen replica's own address
  • Believes callers must re-resolve after every replica replacement
  • Thinks the virtual address is a real interface on some host
  • Assumes a single central proxy handles every call in the cluster
  • Confuses the stable name with the choice of balancing algorithm
open as a page

A replica stops reporting ready but keeps running — what happens to its membership of the set behind the service's virtual address?

level: middleimportance: should knowfreq 58%

basics

~20 s

Its address is removed from the member set, so no new connections are dispatched to it, while the process itself keeps running untouched. When it reports ready again the address is added back silently, with no caller involvement and no event a caller sees.

open as a page

A queue consumer's pooled connections keep reaching a replaced replica of a scoring service — why did the stable name not protect it?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Because a member is chosen when a connection is established, not when a request is sent. A pool holds its connections for hours, so the choice made at open time is never revisited, and the abstraction only helps callers that open new connections.

open as a page

Teams keep asking to address individual replicas directly instead of through the service's virtual address — how do you decide which requests to grant, and what do you require of the ones you do?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Grant it only where the caller genuinely must reach a particular member rather than any member: replicas forming a group among themselves, a caller that must reach the instance holding a given piece of state, or per-instance collection. Everything else keeps the virtual address.

open as a page