skip to content

A service is registered in Consul under the name `web`. What does a DNS A-record lookup of `web.service.consul` give you, what does an SRV lookup of the same name add, and how would you narrow the answer to instances carrying a particular tag?

level: juniorimportance: should knowfreq 60%

answer

  1. the .consul zone on port 8600
  2. address only, or address plus port
  3. tags go in front of the name
  4. datacenter sits between service and consul

basics

~20 s

An A lookup of web.service.consul returns one address per healthy instance, in randomised order. An SRV lookup additionally returns each instance's registered port and node name. Prefixing a tag — primary.web.service.consul — filters to instances registered with that tag.

solid answer

~40 s

Consul's agent answers DNS on port 8600, and `web.service.consul` is the discovery name for the service registered as `web`. An **A** query returns one address per instance that Consul currently considers healthy, shuffled so naive clients spread load; it tells you *where* but not *on which port*, so it only works when every instance listens on a well-known port. An **SRV** query for the same name returns the port and target node alongside the weight and priority, which is what you want when instances take a dynamic port — a container mapped to an ephemeral host port, for example. Tags are a prefix filter: `primary.web.service.consul` returns only instances registered with the tag `primary`, and `web.service.dc2.consul` reaches another datacenter. Health filtering happens before the answer is built, so failing instances never appear.

code

bash · 11 lines
bash
# All healthy instances (addresses only)
dig @127.0.0.1 -p 8600 web.service.consul +short

# Same instances, with the registered port
dig @127.0.0.1 -p 8600 web.service.consul SRV +short

# Only instances registered with the tag "primary"
dig @127.0.0.1 -p 8600 primary.web.service.consul +short

# The same service in another federated datacenter
dig @127.0.0.1 -p 8600 web.service.dc2.consul +short

go deeper

for a junior

Be able to say the name form out loud — web.service.consul for addresses, an SRV query when you need the port, and a tag as a prefix — and know the agent answers on port 8600.

for a middle

Explain why SRV is mandatory for dynamically-ported instances, what the zero TTL is protecting against, and how the datacenter segment fits into the name.

for a senior

Expect to be asked what breaks in production: resolver-side caching that defeats the zero TTL, UDP answer limits hiding instances, and DNS shuffling being no substitute for load-aware balancing.

for a principal

Own the decision of whether DNS is the right discovery surface at all for a given fleet, versus the HTTP API or a proxy that consumes it, and what that choice costs in staleness and in per-agent query load.

## The DNS interface, and why it exists Consul deliberately exposes discovery through DNS as well as through its HTTP API, because DNS is the one lookup mechanism every language, library and legacy binary already speaks. An application that can be pointed at a hostname needs no Consul SDK at all: it resolves `web.service.consul` and connects. Each Consul agent — including the client agent running on the application's own node — serves this zone on **port 8600** by default (UDP and TCP). In production, node resolvers are usually configured to forward the `.consul` domain to the local agent, or the agent's DNS port is bound to a loopback address that the resolver stub points at, so applications can use ordinary names. ## The name grammar The canonical discovery form is: ``` [tag.]<service>.service[.<datacenter>].consul ``` - `web.service.consul` — all healthy instances of `web` in the local datacenter. - `primary.web.service.consul` — only those instances registered with the tag `primary`. Tags are arbitrary strings set at registration time (`"tags": ["primary", "v2"]`), commonly used for role or version. - `web.service.dc2.consul` — the same service in the datacenter named `dc2`, when the datacenters are federated. Consul also serves the RFC 2782 style name, `_web._tcp.service.consul`, for resolvers that construct SRV names that way. ## A records versus SRV records An **A** (or AAAA) query returns just addresses: ``` $ dig @127.0.0.1 -p 8600 web.service.consul web.service.consul. 0 IN A 10.0.1.14 web.service.consul. 0 IN A 10.0.1.22 ``` That is enough only if the port is a constant you already know. The moment instances get dynamic ports — a scheduler mapping a container to whatever host port is free — an address alone is useless. An **SRV** query carries the port: ``` $ dig @127.0.0.1 -p 8600 web.service.consul SRV web.service.consul. 0 IN SRV 1 1 21985 node1.node.dc1.consul. ``` The target is the node's name inside the `.consul` domain, and Consul returns the node's address as an additional A record in the same response, so a well-behaved resolver needs only one round trip. Note what SRV does **not** carry: tags, health state, or metadata. Tags are a query-side filter, and health is applied before the answer is assembled. ## What the answer set already excludes The records you get back are not the full registration list. Consul builds the DNS answer from instances whose checks are acceptable, so an instance whose check has gone critical, or whose node has dropped out of the gossip pool, simply is not in the response. This is why DNS-based discovery in Consul "just works" for a rolling restart: the dying instance disappears from resolution without anyone editing config. ## Practical caveats that bite people **TTLs are zero by default.** Consul returns records with a TTL of 0 so that resolvers do not cache a dead instance. Some resolvers and language runtimes cache anyway — the JVM's DNS caching is the classic example — which reintroduces staleness Consul was trying to avoid. If you *want* caching, `dns_config.service_ttl` lets you set a non-zero TTL per service. **UDP answers are limited.** A DNS response over UDP has a size ceiling, so Consul caps how many records it returns over UDP (`udp_answer_limit`) and can set the truncate bit (`enable_truncate`) so a client retries over TCP. With a large pool, do not assume DNS shows you every instance. **Ordering is randomised, not balanced.** Shuffling gives a crude spread across instances, but it is not load-aware: nothing in a DNS answer reflects how busy an instance is. **No health detail travels.** DNS gives you a yes/no. When you need to know *why* an instance is out, or want the check output, you go to the HTTP health API instead.

  • Why does Consul return DNS records with a TTL of zero, and when would you raise it?
    A zero TTL tells resolvers not to cache, so an instance that fails a check disappears from resolution immediately rather than after a cache expiry. You raise it — via `dns_config.service_ttl` — when query volume against the agent is heavy and you can tolerate a few seconds of staleness, typically for large, stable pools where instances rarely churn.
  • An application resolves `web.service.consul` once at startup and reuses the address forever. What goes wrong?
    It pins itself to whichever instance answered first and never learns about scaling, replacement or failure — the zero TTL is wasted because the process cached the result itself. Fix it by resolving per connection or per pool refresh, or by putting a proxy in front that re-resolves, rather than relying on process-lifetime resolution.

An A record is a street address; an SRV record is the address plus the apartment number. If everyone in the building is on the same floor you can get by with the street address — until they are not.

saying these in an interview costs you the question

  • Thinking an A record includes the service's port
  • Assuming DNS returns every registered instance, healthy or not
  • Writing the tag after the service name instead of before
  • Believing DNS ordering balances load by instance utilisation
  • Expecting health-check output to be visible over DNS

context