A service reaches its dependency only through a DNS name that resolves to several instance addresses, with no registry and no proxy in between. Which parts of service discovery does that give you, and which does it not?
answer
- naming, not health
- present or absent, nothing between
- TTL is a hint, not a contract
- one address, held by the pool
- SRV adds port and weight
basics
~20 sDNS-only discovery supplies naming and resolution: a stable name mapping to a set of addresses. It supplies no health state, no push updates and no real per-request balancing, and its TTL is a hint that clients routinely ignore.
solid answer
~50 sDNS gives you the naming half of discovery and little of the rest. You get a stable logical name mapping to a set of addresses, so callers stop hard-coding IPs and the fleet can change without a code change — and it works from any language, any platform, even across organisational boundaries. What you do not get is health: a record is simply present or absent, so an instance that is listening but broken keeps its record until some external actor edits the zone. You do not get push — a removal only takes effect once every cache along the path expires, and TTL is advisory, not a contract. You do not get balancing: the caller normally uses the first address that connects and keeps it for the life of a pooled connection. And you get almost no metadata — no version, zone or shard, and no port unless you move to `SRV` records. That is why DNS usually fronts a virtual IP rather than the instances themselves.
code
bash · 5 lines# address records: a set of endpoints, and nothing about their state
dig +short api.internal.example A
# SRV adds priority, weight, port and target for a named service
dig +short _grpc._tcp.api.internal.example SRVgo deeper
Be able to say that a DNS name maps to one or more addresses so callers never hard-code an IP, and that the name is what stays stable when instances are replaced.
Explain the mechanics: what a TTL actually instructs a cache to do, why several address records is not load balancing, and why nothing in a DNS answer conveys whether the instance is healthy.
Show production judgement about staleness — who caches the answer, how long traffic keeps flowing to a withdrawn address, and why you would front the instances with a virtual IP rather than publishing them directly.
Own the trade: DNS reaches every caller and every platform at the cost of coarse, unpushable membership. Be ready to argue where that ceiling is acceptable and where the organisation should pay for a richer discovery layer.
## The four jobs discovery has to do It helps to split "service discovery" into four separate jobs, because every model does some of them and not others: 1. **Naming** — a stable identifier a caller can hold in configuration that outlives any individual instance. 2. **Resolution** — turning that identifier into one or more network addresses. 3. **Membership** — deciding which instances currently *deserve* traffic, which means knowing that they are alive and ready. 4. **Selection** — picking which of them serves this particular request. DNS does jobs 1 and 2 very well, job 3 crudely, and job 4 not at all. Every DNS-only design succeeds or fails on whether the missing jobs matter for that dependency. ## What DNS genuinely gives you A name is a real decoupling: `api.internal.example` can point anywhere, and the zone is a single, auditable source of truth for where "the API" lives. Multiple address records for one name express a set, so the fleet's size is invisible to callers. Support is universal — no client library, no sidecar, no language binding, and it reaches callers you do not control, including third parties and other organisations. `SRV` records extend this: an entry of the `_service._proto.name` form carries a target host, a **port**, a priority and a weight, so a caller can discover where to connect rather than assuming a port by convention. And caching is designed in, which is why the system scales to the whole internet. ## What it does not give you **Health.** There is no field for it. A record exists or it does not, and withdrawing it requires something outside DNS to notice the instance is unhealthy and edit the zone. Until that happens, a process that accepts connections but fails every request keeps receiving traffic. **Push.** DNS is pull with a TTL hint. In the best case a caller re-resolves once the TTL elapses. In practice the TTL is advisory at every hop: caching resolvers may floor or extend it, runtimes keep their own address caches that never look at the record's TTL, and a caller holding an established keep-alive connection does not resolve anything at all. From the zone owner's side, the time until traffic actually stops is effectively unbounded. **Balancing.** Authoritative servers and resolvers commonly rotate the order of an answer set, which is where the phrase "DNS round robin" comes from — but rotation is not balancing. Clients typically connect to the first usable address and keep it; some stub resolvers re-sort address sets by destination-selection rules, which defeats rotation entirely; and nothing in the path considers load, latency or connection count. A ten-record answer routinely produces a badly skewed distribution. **Metadata.** Version, availability zone, shard ownership, canary weight, protocol capability — there is nowhere to put them that routing can act on. `SRV` gives you priority and weight only, and `TXT` conventions are exactly that, conventions. ## The characteristic failure An instance dies. Its record is still published, so callers keep opening connections to it and failing. Someone withdraws the record — and traffic still arrives, because each caching layer keeps the old answer for its own interval, and existing pooled connections never re-resolve. Teams respond by cutting the TTL, discover that failover is still slow and unpredictable, and now also pay a much higher query rate. The lesson is that **TTL bounds when a cooperative cache asks again, not when traffic stops.** ## Where DNS-only is still right — and the usual fix It is a good fit for stable dependencies, third-party endpoints, callers you do not control, cross-platform reach, and any case where a few tens of seconds of staleness is harmless. The strongest common design is not DNS-to-instances at all but **DNS-to-a-virtual-IP**: the name resolves to one address that almost never changes, and membership, health checking and selection all live behind that proxy or load balancer. That is server-side discovery, and it deliberately reduces DNS's role to the one job it does best. The third model, client-side discovery, replaces the DNS lookup with a query to a registry that returns instances plus their state and metadata; the caller then balances per request. It buys back health, push and selection at the cost of a library in every caller and a registry to operate. ## Answering it in an interview Name the four jobs, place DNS on each of them, then say the two sentences that show judgement: a record carries no health, and a TTL is a hint rather than a guarantee. Finish with where you would put a VIP or a registry instead — that is the part interviewers are listening for.
- What changes if that DNS name resolves to a single virtual IP in front of the instances instead of to the instances themselves?It becomes server-side discovery. Resolution collapses to one address that rarely changes, so caller-side caching stops mattering for membership, and the proxy owns health checking and per-request selection. You pay an extra network hop and make that proxy a thing you must scale, observe and fail over.
- Why does lowering the TTL to a few seconds not reliably speed up failover?Because TTL only asks cooperative caches to re-query. Caching resolvers may floor it, runtimes keep independent address caches that never read the record's TTL, and a reused keep-alive connection re-resolves nothing at all. You get a much higher query rate and only some of the intended speed-up.
- When is DNS-only discovery clearly the right choice?When the dependency is stable, the callers are diverse or outside your control, and tens of seconds of staleness is acceptable — third-party endpoints, cross-platform or cross-organisation calls, and anything where adding a client library to every caller costs more than the staleness does.
saying these in an interview costs you the question
- DNS round robin spreads load evenly across the returned addresses
- Setting a low TTL guarantees clients fail over within that time
- If an instance becomes unhealthy DNS stops returning its record
- A client that receives several addresses tries all of them
- Discovery is solved as soon as every service has a hostname