Teams keep asking to address individual replicas directly instead of through the service's virtual address — how do you decide which requests to grant, and what do you require of the ones you do?
answer
- a member, or one particular member
- default stays behind the name
- the caller inherits the platform's job
- membership must be re-read, not stored
- write the justification beside the declaration
basics
~20 sGrant it only where the caller genuinely must reach a particular member rather than any member: replicas forming a group among themselves, a caller that must reach the instance holding a given piece of state, or per-instance collection. Everything else keeps the virtual address.
solid answer
~50 sThe test is whether the caller needs a *specific* member or merely *a* member. Ordinary request traffic needs any member, so it belongs behind the stable name and the platform's dispatch. Legitimate exceptions are narrow: replicas that must find each other by name to form a group, a caller that must reach the one instance holding a particular piece of state, and per-instance collection or administrative paths. Where you grant it, the caller has taken over the platform's job, so require it to re-read membership rather than store the first answer, to spread its own calls, to drop members that vanish, to treat a listed member as possibly not ready, and to own retries. Keep the virtual address in place for the same service so the exception does not spread into the normal request path.
go deeper
Know that the default is to address a service by its stable name, and that reaching one specific replica is an unusual request with a reason behind it.
Be able to name what changes when the name resolves to members instead of a virtual address: no platform dispatch, and churn visible to the caller.
Argue the narrow cases that justify it and list what the caller now owns — membership, spread, dead members, retries and connection lifetime.
Set the standard: default to the indirection, grant exceptions against a written justification, keep both addresses in place, and watch for callers disagreeing about membership during rollouts.
## Two different questions hide behind one request "Let me address replicas directly" is asked for two very different reasons, and the whole decision is telling them apart. - **"I need a particular member."** The work is bound to one instance: peers forming a group among themselves, a caller that must reach the instance that holds a given piece of state, a collection path that must read each member separately rather than a random one, or an administrative call aimed at one instance under investigation. Here the virtual address is genuinely the wrong tool — it deliberately hides which member you reach, which is exactly what this caller cannot have. - **"I want better behaviour than the platform's dispatch gives me."** The team wants an ordering, an affinity, or a spread they think they can do better. This is not an addressing requirement at all; it is a routing or balancing requirement wearing an addressing costume, and granting it hands one team a permanent copy of a problem the platform already owns. The first is a small, checkable set. The second should almost always be refused, because the cost lands on the team's future on-call rather than on the person asking today. ## What the caller signs up for Granting direct addressing means the caller has taken over the job the platform was doing. Make that explicit and reviewable, because it is invisible in the code afterwards: - **Re-read membership.** The set changes constantly, so the caller must resolve again rather than keep the first answer for the life of the process. - **Treat a listed member as possibly unusable.** Depending on the mode, the answer may include members that are not currently ready, so the caller needs its own handling rather than assuming everything it sees can serve. - **Spread its own calls.** With no dispatch in the path, an evenly loaded replica set is now the caller's responsibility. - **Drop members that are gone**, and not retry into an address that has stopped answering — remembering that a released address can later belong to a completely different workload. - **Own retries and timeouts**, including the question of which work is safe to repeat. - **Re-open connections periodically**, because pinning is now entirely in its hands. If a team cannot say how it does those six things, the honest answer is that it is asking for direct addressing and expecting the platform's guarantees anyway. ## Keep both, and keep the exception narrow The two modes are not mutually exclusive, and the standard worth setting is to have both in front of the same replicas: 1. The stable name and virtual address stay, and remain the only address used by ordinary request traffic from other services. 2. The per-instance names exist alongside them and are used *only* by the narrow path that justified them. 3. The justification is written down where the service is declared, so a reviewer a year later can see it was a decision rather than a habit. Without point three, direct addressing spreads by copy-paste. Somebody's template that legitimately needed per-instance names becomes the starting point for a service that did not, and the estate accumulates callers holding member lists with no idea why. | | virtual address for everything | direct addressing for the narrow case | |---|---|---| | who tracks membership | the platform | the caller | | who spreads load | the platform | the caller | | replica churn is | invisible | visible and must be handled | | worth it when | the caller needs any member | the caller needs one named member | ## The failure mode to watch for afterwards The estate-level symptom of granting this too freely is uniform and slow: several callers that each hold their own idea of who exists, disagreeing during every rollout. One holds a member that is gone, another has not noticed a member that arrived, and errors cluster around deploys of a service whose own logs look healthy. Because each caller solved it differently, there is no single place to fix it — which is the strongest practical argument for keeping the indirection the platform owns as the default, and treating every exception as something a team must justify and then maintain. A final note on register: this is a standard you set, not a mechanism you recite. The good answer names the narrow legitimate cases, states what the caller inherits, and says how you stop the exception becoming the norm.
- A team argues direct addressing is faster because it skips a hop — is that a good reason?Rarely. The dispatch is a destination rewrite performed on the caller's own host, not a trip through a separate box, so the saving is small and hard to measure. Against it sits a caller that must now track membership, spread calls and handle churn. Grant it for needing a specific member, not for that.
- How do you stop a granted exception spreading to services that never justified it?Make it visible where the service is declared, with the reason attached, and keep the stable name in place for the same replicas so ordinary traffic has a correct default to use. Then review new uses against the narrow list: peers forming a group, reaching the instance holding a piece of state, and per-instance collection.
saying these in an interview costs you the question
- Grants direct addressing to get better load spreading
- Assumes a directly addressed set still hides replica churn
- Thinks the caller inherits nothing beyond one extra lookup
- Removes the stable name once per-instance names exist
- Treats a resolved member list as valid for the process lifetime