skip to content

An Envoy cluster with type: LOGICAL_DNS points at a DNS name that resolves to many backend addresses, yet almost all traffic lands on one backend. What does LOGICAL_DNS do that STRICT_DNS does not, and which would you use here?

level: seniorimportance: should knowfreq 35%

answer

  1. type decides how endpoints are learned
  2. one logical host versus every address
  3. load balancing needs more than one host
  4. connection reuse pins traffic further
  5. internal fleet wants STRICT_DNS or EDS

basics

~20 s

LOGICAL_DNS keeps a single logical host built from one resolved address, so the load balancer has exactly one endpoint to choose. STRICT_DNS keeps every address returned by DNS as a separate host and balances across all of them, which is what a multi-backend name needs.

solid answer

~50 s

A cluster's `type` decides how its endpoint list is built, and the two DNS types differ sharply. **STRICT_DNS** resolves the name periodically and maintains *every* returned address as its own host in the cluster; the `lb_policy` then spreads requests across them and each has its own connection pool and health state. **LOGICAL_DNS** deliberately keeps only a single logical host — it uses the first address from each resolution and re-resolves in the background, letting existing connections continue while new ones use the freshest address. That design exists for very large external names, where holding thousands of rotating addresses would be wasteful, and where the client only ever needs one working connection. The cost is exactly your symptom: with one endpoint, load balancing has nothing to balance, and with connection reuse — HTTP/2 especially — the traffic pins to one backend for a long time. For an internal service name that legitimately resolves to your fleet, use `STRICT_DNS`, or `EDS` if you have a control plane feeding endpoints.

go deeper

for a junior

Know that an Envoy cluster's type field says how its backend list is built, and that a DNS-based cluster resolves a hostname rather than listing IP addresses.

for a middle

Explain the difference between STRICT_DNS keeping every resolved address as a host and LOGICAL_DNS keeping a single logical host, and why the latter leaves the load-balancing policy with nothing to do.

for a senior

Diagnose uneven upstream traffic by checking cluster membership before touching the lb_policy, account for connection reuse pinning traffic further, and pick the type deliberately for internal fleets versus large external names.

for a principal

Own the standard for how endpoint membership is learned across a fleet — DNS refresh versus a control plane feeding EDS — and weigh propagation speed, control-plane availability and blast radius rather than leaving each team to pick a cluster type ad hoc.

## The cluster type field A cluster answers one question: which endpoints are in this upstream group and how do we learn about them? The `type` field names the mechanism: - **`STATIC`** — endpoints are literal IP:port pairs in `load_assignment`. No resolution at all. - **`STRICT_DNS`** — a hostname is resolved on a timer and *all* returned addresses become hosts. - **`LOGICAL_DNS`** — a hostname is resolved, but the cluster holds a single logical host. - **`EDS`** — endpoints are delivered dynamically by a management server. - **`ORIGINAL_DST`** — the destination comes from the connection's original destination address, for transparent proxying. Everything else about the cluster — `lb_policy`, `connect_timeout`, upstream TLS, protocol options — is orthogonal to the type. But the type decides how many things the load balancer has to choose between, which is why it can quietly nullify your load-balancing configuration. ## STRICT_DNS in detail Envoy asynchronously resolves the configured name at `dns_refresh_rate` (five seconds by default), or according to the record TTL if `respect_dns_ttl` is enabled. Every address in the answer becomes a distinct host in the cluster, with its own connection pool and its own health status. Hosts that disappear from a later answer are removed; new ones are added. Two names that resolve to the same address still produce two hosts, because identity is per resolved entry. This is what you want when the name genuinely represents a set of interchangeable backends: `LEAST_REQUEST` or `ROUND_ROBIN` then has real choices, an unhealthy backend can be ejected individually, and connection pools are per host. ## LOGICAL_DNS in detail LOGICAL_DNS was built for a different problem: talking to a huge external service whose name resolves to a large, rapidly rotating address set. Materialising every address would mean a big host table churning constantly, for no benefit — the client just needs *a* working connection to *the* service. So the cluster keeps one logical host. Resolution still happens in the background, and the address used for **new** connections is refreshed, but existing connections are not torn down when the answer changes. The load balancer sees one host, so `lb_policy` is effectively inert. ## Why your traffic pinned to one backend Put those together with connection reuse and the symptom is inevitable: - Only one endpoint exists, so every request goes to whatever address that logical host currently holds. - With HTTP/2 to the upstream, one connection carries many concurrent streams, so even long-lived load stays on the single connection to that one address. - With HTTP/1.1 and keep-alive, the pool's connections were all established against the address that was current when they were created. Nothing is broken; the cluster is doing precisely what LOGICAL_DNS promises. The mismatch is between the type and the intent. ```yaml clusters: - name: service_backend connect_timeout: 0.25s type: STRICT_DNS # every resolved address becomes a host lb_policy: LEAST_REQUEST load_assignment: cluster_name: service_backend endpoints: - lb_endpoints: - endpoint: address: socket_address: { address: backend.internal, port_value: 8080 } ``` ## Choosing between them - **Internal service names resolving to your own fleet** → `STRICT_DNS`, so balancing and per-endpoint health work. - **Large external endpoints with volatile address sets** → `LOGICAL_DNS`, to avoid a churning host table. - **You run a control plane that knows the real membership** → `EDS`, which removes DNS from the path entirely and gives you locality and per-endpoint metadata. - **A fixed, known set** → `STATIC`. ## Related sharp edges `dns_lookup_family` controls whether Envoy uses A records, AAAA records, or both — a cluster set to one family against a name that only publishes the other resolves to nothing and every request fails, which looks like an outage rather than a config error. And with `STRICT_DNS`, a name resolving to a very large answer set produces a correspondingly large host table, so a low refresh rate against a big set costs real CPU. Whatever the type, remember DNS answers are cached by Envoy on its own schedule, so a backend rotation is visible only as fast as the refresh interval allows. ## How to confirm it in production The direct check is the cluster's membership: how many hosts does the cluster actually have? One host in a cluster you believed was balancing over ten is the whole answer, and it points straight at the type field rather than at the load-balancing policy you were about to start tuning.

  • Why does Envoy offer LOGICAL_DNS at all if it defeats load balancing?
    Because for very large external services the address set is huge and rotates constantly, and materialising every address would mean a big, churning host table that buys nothing — the client only needs one working connection. LOGICAL_DNS keeps a single logical host, refreshes the address used for new connections in the background, and lets existing connections continue undisturbed.
  • How does EDS change the picture compared with STRICT_DNS?
    EDS takes DNS out of the path: a management server supplies the endpoint set directly, so membership changes propagate as fast as the control plane pushes rather than as fast as a refresh timer allows, and endpoints can carry locality and metadata that DNS cannot express. The trade is that you now depend on a control plane being available and correct.
  • A STRICT_DNS cluster resolves to nothing and every request fails, though the name resolves fine from a shell. What would you check first?
    `dns_lookup_family`. If the cluster is restricted to one address family and the name only publishes records of the other, Envoy gets an empty answer and the cluster has no hosts, so requests fail with no healthy upstream. Your shell resolver has no such restriction, which is why the name looks fine there.

saying these in an interview costs you the question

  • Thinks LOGICAL_DNS balances across all resolved addresses
  • Blames the lb_policy when the cluster has one host
  • Assumes DNS changes take effect immediately
  • Believes cluster type and load balancing are the same setting
  • Expects per-endpoint health with a single logical host

context