Your service runs in two federated Consul datacenters, and you want callers to use the remote datacenter only when no healthy instance exists locally. Which Consul discovery mechanisms would you weigh for that, and what does each one cost you?
answer
- federation forwards, it does not replicate
- the remote datacenter is just a name segment
- one name, a server-side policy behind it
- the failover trigger is your health signal
- invisible failover is a debugging problem
basics
~20 sPlain federated DNS gives you an explicit remote name — web.service.dc2.consul — but no automatic failover; the caller must choose. A prepared query with a failover policy makes Consul do the fallback server-side and expose it as a single name under .query.consul.
solid answer
~60 sWAN federation by itself is addressability, not failover: joining the datacenters' server tiers over the WAN gossip pool lets you resolve `web.service.dc2.consul`, but a caller resolving `web.service.consul` never falls back on its own. To get automatic fallback you register a **prepared query** whose failover policy names the datacenters (or asks for the nearest N), resolved through `<name>.query.consul`. Consul evaluates it server-side: healthy local instances if any exist, otherwise the remote set. The costs are real. The caller can no longer tell from the name where it landed, so latency and egress become invisible in the client's own telemetry. The failover decision now hinges entirely on local health being *correct* — a broken local check flips the whole datacenter's traffic across the WAN. Each datacenter keeps its own Raft, so a remote lookup crosses the WAN and inherits its latency and its partitions. Often the better answer is not to fail over at the discovery layer at all, but a tier above it, where the decision is observable and deliberate.
go deeper
Know that a federated datacenter is addressed by putting its name into the query — web.service.dc2.consul — and that this is an explicit choice, not a fallback.
Explain that each datacenter keeps its own catalog and Raft, that the WAN pool only forwards requests, and that a prepared query is the mechanism that adds server-side failover.
Argue the operational consequences: opacity to the caller, flapping with no damping, and a shallow health check becoming the lever that moves a datacenter's entire load.
Own the placement of the failover decision itself — discovery layer versus a tier above it — and defend it on observability, blast radius, granularity and how much you trust the health signal that fires it.
## What federation gives you, and what it does not WAN federation joins the *server* tiers of two or more datacenters into a shared WAN gossip pool. Each datacenter keeps its own Raft peer set, its own leader, and its own catalog — nothing is replicated across the link. What federation buys is *forwarding*: a query naming another datacenter is routed to that datacenter's servers and answered there. The visible surface is a name: ``` web.service.dc1.consul # explicitly local web.service.dc2.consul # explicitly remote web.service.consul # this agent's own datacenter ``` Notice what is missing: there is no fallback. `web.service.consul` returns an empty answer when nothing local is healthy. Whoever wants failover has to implement it — either in the caller, or in the discovery layer. ## Option one: let the caller choose The caller resolves the local name, and on an empty answer resolves the remote one. This is explicit and completely observable — the client's own code and telemetry record which datacenter it chose. The costs are that every caller must implement the same logic identically, in every language you use, and the fallback decision is made per-caller with no coordination, so behaviour during a partial local outage is a matter of luck rather than policy. ## Option two: a prepared query with a failover policy A prepared query is a named, server-side query definition registered against the catalog. Its service block can restrict to healthy instances, filter by tag, sort by proximity, and — the part that matters here — carry a **failover** policy that names specific datacenters or asks for the nearest few. Callers resolve one name: ``` web-failover.query.consul ``` Consul evaluates it at query time: healthy local instances if there are any, otherwise instances from the first datacenter in the failover list that has some. Every caller gets identical behaviour by construction, the policy changes in one place, and nothing needs redeploying to change it. The trade-offs are the interesting part: **Opacity.** The name no longer tells anyone where the traffic went. A caller that suddenly has 90 ms of extra latency has no local signal saying "you are now talking to the other region". You have to build that observability deliberately. **Health becomes a global lever.** Failover fires on local health being empty. A misconfigured check, a dependency that every local instance shares, or an over-aggressive timeout can turn *all* local instances critical at once and send a datacenter's full traffic across the WAN — into a remote datacenter sized for its own load. That is a capacity event caused by a health-check bug. **Nothing prevents flapping.** Local health recovering and failing repeatedly moves traffic back and forth with no damping. Discovery-layer failover has no notion of hysteresis or of a human deciding to fail back. **WAN dependence.** The remote lookup itself crosses the link. A partition that takes out local instances *and* the WAN leaves you with nothing, and cross-datacenter data transfer usually has a bill attached. ## Option three: do not fail over at the discovery layer The alternative worth naming in a principal-level answer is keeping each datacenter's discovery strictly local and moving the failover decision up a tier — to whatever already steers traffic between regions before it reaches a datacenter at all. That decision is coarse-grained, observable, and typically reversible by a human, and it fails a whole region rather than one service at a time. Its cost is granularity: you cannot fail over a single misbehaving service without moving everything. ## How to decide Ask three questions: 1. **Is the remote instance actually usable?** Cross-datacenter latency, plus the data locality the service needs, decides this before anything else. Failing over a chatty service to a database it cannot reach quickly is worse than serving errors. 2. **How much do you trust local health?** Automatic failover is only as good as the signal that triggers it. If your checks are shallow or share a dependency, an automatic global lever is dangerous. 3. **Who should notice?** A prepared query hides the event; a caller-side or tier-above decision surfaces it. During an incident, an invisible failover is a debugging problem, not a save. A defensible position: use prepared queries where the remote datacenter is genuinely equivalent and the health signal is deep and per-instance; use an explicit tier-above decision where failing over is a big, rare, human-supervised event. What you should not do is add automatic cross-datacenter failover to the discovery layer simply because Consul supports it.
- Why does a prepared query change what your client-side telemetry can tell you?The client resolves a single name and never learns which datacenter answered, so its metrics show a latency change with no attributable cause. You have to add that visibility deliberately — surfacing which set the query resolved to, and correlating client latency with the query's own behaviour — otherwise the first sign of a cross-region failover is an unexplained latency shift during an incident.
- What makes automatic cross-datacenter failover dangerous when health checks are shallow?Failover fires when the local healthy set is empty, so any fault that makes every local instance critical at once — a shared dependency, a check that probes something common, an over-tight timeout — moves an entire datacenter's traffic across the WAN. The remote side is sized for its own load, so a health-check bug becomes a capacity incident. Deep, per-instance checks are the prerequisite for trusting the lever.
- Does WAN federation replicate one datacenter's catalog into another?No. Each datacenter keeps its own Raft peer set, leader and catalog; the WAN pool only lets server tiers find each other and forward requests. A cross-datacenter lookup is answered by the remote datacenter's own servers over the WAN link, which is why it inherits that link's latency and fails when the link does.
saying these in an interview costs you the question
- Assuming WAN federation replicates catalogs between datacenters
- Thinking a plain service name falls back to a remote datacenter
- Adding automatic failover without deep per-instance health checks
- Ignoring that the caller cannot see which datacenter answered
- Treating cross-datacenter latency and egress cost as free