skip to content

Consul

Consul is HashiCorp's service networking product: a registry with health checking, a KV store, and a built-in mesh, run as an agent on every node. Interviewers ask about it because it is the classic non-Kubernetes answer to "how do services find each other", and because its three faces — discovery, config, and Connect — get confused with one another.

on this pageshow

questions

18

In Consul's service mesh, what is an intention, what identity does it match on, and what happens to a call between two mesh services when no intention matches it?

level: juniorimportance: must knowfreq 68%

answer

  1. authorization between two named services
  2. identity, not source IP address
  3. read from the mTLS certificate's SPIFFE SAN
  4. checked by the destination's sidecar
  5. no match falls back to ACL default_policy

basics

~20 s

A Consul intention is an authorization rule that permits or denies traffic from one mesh service to another. It matches the service identity carried in the sidecar's mTLS certificate rather than a source IP address. When nothing matches, Consul falls back to the ACL default_policy.

solid answer

~50 s

An intention is Consul's service-to-service authorization rule, written as a `service-intentions` config entry whose `Name` is the destination service and whose `Sources` list callers with `Action = "allow"` or `"deny"`. The decision is made on **identity**, not address: every sidecar holds a leaf certificate whose URI SAN is a SPIFFE ID naming the service, so when `web`'s proxy opens an mTLS connection to `db`'s proxy, the destination proxy reads the caller's service name out of that certificate and applies the matching intention. Enforcement happens at the destination sidecar, using the intention set Consul has already pushed to it, so changes take effect in seconds with no restart. If no intention matches the pair, Consul defers to the ACL `default_policy`: `deny` gives you a deny-by-default mesh, `allow` lets everything through. Teams that want default deny either run ACLs with `default_policy = "deny"` or write an explicit `*` → `*` deny intention.

code

hcl · 13 lines
hcl
Kind = "service-intentions"
Name = "db"

Sources = [
  {
    Name   = "web"
    Action = "allow"
  },
  {
    Name   = "*"
    Action = "deny"
  }
]

go deeper

for a junior

Be ready to say plainly that an intention allows or denies one service calling another, and that the caller is identified by the certificate its sidecar presents rather than by an IP address.

for a middle

Explain the config entry shape — destination in Name, callers in Sources — how the SPIFFE identity is read from the peer certificate, and that the default when nothing matches follows the ACL default_policy.

for a senior

Show that you have operated this: precedence rules over overlapping intentions, enforcement living on the destination proxy so changes apply without restarts, and the gap where traffic reaching the app port directly is never governed at all.

for a principal

Own the posture question. Argue for ACLs with default_policy = "deny" from day one, describe how you would migrate an allow-all mesh to default deny without an outage, and say who authors intentions as services multiply.

## What an intention actually is In Consul's service mesh, an **intention** is an authorization rule between two *services* — not two hosts, not two IP ranges. It is expressed as a config entry: ```hcl Kind = "service-intentions" Name = "db" # the DESTINATION service Sources = [ { Name = "web", Action = "allow" }, { Name = "*", Action = "deny" } ] ``` You write it with `consul config write intentions.hcl`, or with the older `consul intention create -allow web db` CLI. The entry is named after the destination, and every caller you care about appears in `Sources`. ## Where the identity comes from When a service joins the mesh, Consul issues its sidecar proxy a short-lived X.509 leaf certificate signed by the mesh CA. The certificate's URI Subject Alternative Name is a SPIFFE-style identifier of the form: ``` spiffe://<trust-domain>.consul/ns/<namespace>/dc/<datacenter>/svc/web ``` Both sides present a certificate — the connection between two sidecars is always mutual TLS. So when `web`'s sidecar dials `db`'s sidecar, the receiving proxy does not have to guess who is calling from a source address that may belong to a NAT gateway, a shared node, or a recycled pod IP. It reads `svc/web` out of the verified peer certificate. That is the entire reason intentions are more useful than a firewall rule in a dynamic environment: the identity travels with the workload, and it survives rescheduling, autoscaling and IP reuse. ## Where enforcement happens Enforcement is on the **destination** side. The inbound listener of the destination's sidecar terminates the mTLS connection and checks the caller's identity against the intentions it holds. Consul distributes the relevant intention set to each proxy in advance, so the check is a local decision — there is no per-request round trip to a Consul server. Two practical consequences follow. First, adding or removing an intention takes effect within seconds and requires no restart of the application or the proxy. Second, a proxy that has temporarily lost contact with the Consul servers keeps enforcing the last set it received rather than failing open. ## The default when nothing matches This is the part candidates most often get wrong. Consul does not have a fixed built-in default. If no intention matches a source/destination pair, the result follows the ACL system's `default_policy`: - `default_policy = "deny"` — the mesh is deny-by-default; only explicitly allowed pairs connect. This is the production posture. - `default_policy = "allow"` — which is what you get in a quick dev setup with ACLs disabled — everything is permitted unless an intention denies it. If you cannot yet turn on strict ACLs, you can still get default-deny semantics by writing a catch-all `*` source with `Action = "deny"` and layering specific allows above it. ## Precedence, not "deny always wins" When several intentions could apply, Consul resolves them by **precedence**, computed from how specific the source and destination names are. An exact service name outranks a wildcard, and a more specific match wins regardless of whether it says allow or deny. So an explicit `web → db` allow beats a `* → db` deny; deny does not automatically override. Candidates who assume "the deny rule always wins" are surprised when a broad deny fails to close a hole an exact allow left open. ## L4 by default, L7 when the protocol says so By default an intention is a connection-level (L4) decision: the whole TCP connection is either accepted or refused, and a refused caller sees the connection closed rather than an application error. If the destination's protocol has been declared as HTTP-based, an intention can instead carry `Permissions` with per-request rules (path, method, header), and a request that fails them gets an HTTP 403 while the connection itself stays up. ## What intentions do not do They do not encrypt anything — the mTLS between sidecars does that, and it is on regardless of whether the call is allowed. They also govern only traffic that actually enters a sidecar: anything that reaches the application's own listening port directly bypasses them entirely, which is why mesh workloads should bind their app port to localhost or be protected at the network layer as well.

  • If both a `web` → `db` allow and a `*` → `db` deny exist, which one applies?
    The allow. Consul resolves overlapping intentions by precedence computed from how specific the source and destination names are, and an exact name outranks a wildcard. Deny does not win by virtue of being a deny — it only wins when its match is at least as specific. That is why a broad catch-all deny is a floor, not an override.
  • A service inside the mesh is still reachable from a machine that has no sidecar. Why did the intention not stop it?
    Intentions are enforced by the sidecar's inbound listener, so they only govern traffic that arrives there. A client that connects to the application's own port directly never meets the proxy and is never asked for a certificate. The fix is to bind the application to localhost (or an interface only the proxy can reach) so the sidecar is the sole ingress path.
  • How quickly does deleting an intention take effect on running traffic?
    Within seconds, and without restarting the application or the proxy. Consul pushes the intention set to each sidecar ahead of time, so the update is a config push followed by a local decision change. Existing connections that were already allowed at L4 are not necessarily torn down, so a deny is best thought of as blocking new connections rather than instantly severing established ones.

A firewall rule is a guest list of street addresses; an intention is a guest list of names, checked against photo ID at the destination's door.

saying these in an interview costs you the question

  • Says intentions match on the caller's source IP address
  • Assumes a deny intention always overrides a matching allow
  • Thinks intentions are what encrypts the traffic
  • Believes a missing intention always means the call is denied
  • Claims an intention change requires restarting the proxies

context

open as a page

A service registered in Consul's service mesh has a sidecar proxy running, but its outbound calls still go straight to the destination's real address and never get mTLS. How is an application normally expected to address an upstream in Consul's mesh, and what does transparent proxy mode change?

level: middleimportance: must knowfreq 52%

basics

~20 s

By default Consul does not intercept outbound traffic. Each upstream is declared in the sidecar registration with a local_bind_port, and the application must dial 127.0.0.1 on that port. Transparent proxy mode instead installs iptables rules so all outbound traffic is redirected into the sidecar.

open as a page

In a Consul cluster, what is the difference between a client agent and a server agent, and why does the standard deployment put an agent on every node instead of having applications talk directly to the servers?

level: middleimportance: must knowfreq 72%

basics

~20 s

Server agents hold the catalog and replicate it through Raft with an elected leader. Client agents hold no catalog: they run local health checks, carry local registrations, take part in gossip, and forward requests to servers — which is why every node runs one.

open as a page

Two instances of a service are registered in Consul and one instance's HTTP check has gone critical. Explain what a DNS lookup and the `/v1/health` and `/v1/catalog` HTTP endpoints each return now, and how that state differs from the instance having been deregistered.

level: middleimportance: must knowfreq 68%

basics

~20 s

DNS and /v1/health/service/<name>?passing return only the healthy instance. /v1/catalog/service/<name> returns both, because the catalog lists registrations regardless of health. A critical instance is still registered and returns automatically on recovery; a deregistered one is gone until something registers it again.

open as a page

You render an nginx upstream list from Consul with consul-template and reload nginx from the template's `command`. After an incident the rendered file came out with an empty upstream and every request failed. Walk through consul-template's render-and-reload cycle, and how you would make it safe.

level: seniorimportance: must knowfreq 58%

basics

~20 s

consul-template holds a blocking query per dependency in the template, re-renders when one changes, writes the file atomically and then runs the command. It will happily render an empty list if Consul answers with zero results, so guard the render, validate the config, and only then reload.

open as a page

A service is registered in Consul under the name `web`. What does a DNS A-record lookup of `web.service.consul` give you, what does an SRV lookup of the same name add, and how would you narrow the answer to instances carrying a particular tag?

level: juniorimportance: should knowfreq 60%

basics

~20 s

An A lookup of web.service.consul returns one address per healthy instance, in randomised order. An SRV lookup additionally returns each instance's registered port and node name. Prefixing a tag — primary.web.service.consul — filters to instances registered with that tag.

open as a page

A service reads its configuration from Consul's KV store with `GET /v1/kv/config/app/db_url` and finds the JSON response's Value field is a string like `cG9zdGdyZXM6Ly8...` rather than the URL it wrote. Why is that, and which query parameters return the raw value or a whole prefix in one call?

level: juniorimportance: should knowfreq 52%

basics

~20 s

Consul's KV HTTP API wraps every entry in JSON and base64-encodes the Value so arbitrary bytes survive the encoding. Decode it, or append ?raw to get the value verbatim; ?recurse returns every key under a prefix in one response.

open as a page

Consul compiles a destination service's `service-router`, `service-splitter` and `service-resolver` config entries into a discovery chain. In what order do those three stages apply, and which of them require the destination's protocol to be declared as HTTP?

level: middleimportance: should knowfreq 40%

basics

~20 s

Consul applies them router first, then splitter, then resolver. The router picks a route from L7 request attributes and the splitter divides traffic by weight, so both need an HTTP-family protocol declared in service-defaults; the resolver, which selects subsets and failover targets, also works for plain TCP.

open as a page

An agent must react within a second when a key changes in Consul's KV store, without polling the API in a loop. Explain how a blocking query on `GET /v1/kv/config/app` works, what the `X-Consul-Index` response header is for, and the two index values a client has to handle defensively.

level: middleimportance: should knowfreq 50%

basics

~20 s

A blocking query is a long poll. You send the last X-Consul-Index value back as ?index=N with ?wait, and Consul holds the connection open until that path's index advances or the wait expires, then you repeat with the new index.

open as a page

Two instances of a job read the Consul KV key `config/app/limits`, each modify it, and each write it back with `PUT /v1/kv/config/app/limits`. How does a check-and-set write using the entry's ModifyIndex stop one from silently clobbering the other, and how does the API report a CAS that did not apply?

level: middleimportance: should knowfreq 42%

basics

~20 s

Pass the ModifyIndex you read as ?cas=<index> on the PUT. Consul applies the write only if the key's current ModifyIndex still matches; otherwise it rejects it — returning HTTP 200 with a body of false, not an error status.

open as a page

Consul's service mesh issues each sidecar its own certificate. What does that certificate carry, how long is it valid by default, and what happens across the fleet when you rotate the mesh root or move the CA to the Vault provider?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Each sidecar gets a short-lived leaf certificate whose URI SAN is a SPIFFE identity naming the service, valid for LeafCertTTL — 72 hours by default — and renewed automatically. Rotating the root signs a new intermediate and re-issues leaves; cross-signing is what keeps connections working during the rollover.

open as a page

Services inside a Consul service mesh must call a third-party API and a legacy database that will never run a sidecar. What does a Consul terminating gateway do for that traffic, and what control do you keep once the traffic leaves it?

level: seniorimportance: should knowfreq 26%

basics

~20 s

A Consul terminating gateway is a mesh member that accepts mTLS from sidecars and forwards to destinations that are not in the mesh, optionally originating TLS to them. Intentions still govern who may reach each destination, but past the gateway mesh identity is gone.

open as a page

After a batch of instances registered in Consul is terminated abruptly, their entries linger in the catalog in a critical state and never go away. Why does Consul keep them, and which registration and check settings make it remove them automatically?

level: seniorimportance: should knowfreq 42%

basics

~20 s

A failing check never deletes a registration — only an explicit deregistration does. Set deregister_critical_service_after on the check so the agent removes a service that stays critical, deregister in a shutdown hook, and stop agents with a graceful leave rather than a kill.

open as a page

A five-server Consul datacenter loses three of its servers, so there is no leader. What continues to work for service discovery, what stops, and how do the `stale` and `consistent` read modes change the answer?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Without a quorum there is no leader, so all writes fail: registrations, deregistrations and check-state updates. Reads depend on mode — stale reads are answered by any surviving server from its own replicated state, so DNS keeps resolving, while default and consistent reads fail.

open as a page

A cron-style job runs on three nodes and must execute on exactly one. Explain how acquiring a Consul KV key with a session — `PUT /v1/kv/service/job/leader?acquire=<sessionID>` — elects a leader, what happens to that lock when the session is invalidated, and what LockDelay protects against.

level: seniorimportance: should knowfreq 35%

basics

~20 s

Each node creates a session and tries to acquire the same key; exactly one acquire returns true and that node is leader. If the session is invalidated — the agent fails, a check goes critical, or the TTL lapses — Consul releases or deletes the key, and LockDelay blocks re-acquisition briefly so the old holder cannot still be acting.

open as a page

Your platform already runs Consul for service discovery and now needs a service mesh. What would make you turn on Consul's own mesh rather than adopt Istio, and where does Consul's model cost you by comparison?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Consul's mesh wins when workloads span VMs and several clusters, because one catalog and one identity domain already cover them and intentions are simple to reason about. It costs you a Raft server cluster to operate and a smaller L7 policy vocabulary than Istio's.

open as a page

Your service runs in two federated Consul datacenters, and you want callers to use the remote datacenter only when no healthy instance exists locally. Which Consul discovery mechanisms would you weigh for that, and what does each one cost you?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Plain federated DNS gives you an explicit remote name — web.service.dc2.consul — but no automatic failover; the caller must choose. A prepared query with a failover policy makes Consul do the fallback server-side and expose it as a single name under .query.consul.

open as a page

Consul's KV store holds the runtime configuration for every service on your platform, and all of the teams' tooling authenticates with a single shared token. What is your plan for limiting who can write where, and for getting data back after someone runs a recursive delete on the wrong prefix?

level: principalimportance: nice to knowfreq 22%

basics

~20 s

Scope tokens to prefixes under a default-deny ACL system, one token per workload rather than one shared token, and keep the desired state outside Consul. Snapshots restore the whole cluster, not one prefix, so schedule prefix-level exports for recovery.

open as a page