In a Consul cluster, what is the difference between a client agent and a server agent, and why does the standard deployment put an agent on every node instead of having applications talk directly to the servers?
answer
- one binary, a mode flag
- only one tier runs Raft
- checks execute where the service runs
- local state syncs upward, not downward
- gossip, not polling, spots a dead node
basics
~20 sServer agents hold the catalog and replicate it through Raft with an elected leader. Client agents hold no catalog: they run local health checks, carry local registrations, take part in gossip, and forward requests to servers — which is why every node runs one.
solid answer
~50 sBoth run the same `consul` binary; the mode is a startup flag. **Servers** (typically three or five) form the Raft peer set: they elect a leader, store the service catalog durably, and answer queries. **Clients** store almost nothing. A client agent joins the LAN gossip pool, holds the registrations made against it, executes the health checks for services on its own node, syncs that local state up to the catalog through anti-entropy, and proxies HTTP and DNS requests to the servers via RPC. Running one everywhere buys three things: applications talk to `127.0.0.1:8500` or the agent's DNS port and never need to know a server address; health checks execute next to the workload, so a check reflects the node's own view rather than a server's; and gossip spreads failure detection across the fleet instead of making the servers poll everything. It also keeps the small Raft peer set from becoming the connection sink for the whole fleet.
go deeper
Know that both tiers run the same binary, that servers hold the data and clients do not, and that your application talks to the agent on localhost rather than to a server.
Explain Raft on the servers, gossip membership across everyone, and anti-entropy syncing a node's local service and check state up into the catalog — including which side wins a conflict.
Show you have operated it: what happens when the agent dies but the workload survives, what a partition between a node and the servers freezes, and why check locality changes what the check actually measures.
Own the placement decision — dedicated server nodes, agent-per-node versus agentless access to server APIs, and how each choice moves connection load, check load and blast radius around the fleet.
## One binary, two modes Consul ships as a single agent binary. Started in server mode it joins the Raft peer set; started without that flag it runs as a client. Every node in the datacenter runs an agent — servers on a handful of dedicated nodes, clients everywhere else — and all of them join the same LAN gossip pool. ## What the servers own The **server agents** are the stateful tier. They: - Elect a leader through **Raft** and replicate every catalog write to a quorum before acknowledging it. Three or five servers is the norm; the cluster tolerates the loss of one out of three or two out of five. - Store the **catalog**: which nodes exist, which services each node registered, and the state of every check. - Answer queries from clients over the server RPC port, and serve the HTTP and DNS interfaces themselves too. - Participate in the WAN gossip pool when datacenters are federated, so cross-datacenter requests can be forwarded server to server. ## What the clients own A **client agent** is deliberately cheap. It: - Joins **LAN gossip** (the Serf membership protocol), which is how the cluster detects that a node has died without any server polling it. - Accepts **registrations** for services on its own node — from a config file in the agent's config directory, or from `PUT /v1/agent/service/register` on the local HTTP API. - **Runs the health checks** for those local services on their configured interval. - Performs **anti-entropy**: the agent's local state is authoritative for the services on its node, and it periodically reconciles that state up into the servers' catalog. If a server's catalog disagrees with the agent about a node's services, the agent wins. - **Forwards** HTTP and DNS requests it cannot answer to a server over RPC. The client keeps no durable copy of the catalog. Restart it and it re-registers what its config directory and API-registered state say, then resyncs. ## Why an agent on every node **Locality of registration.** A service registers with the agent on its own machine, so registration is a loopback call that cannot fail because of a network partition to the servers. The agent then owns the job of getting that fact to the catalog. **Locality of health checking.** This is the part candidates most often miss. The check runs *on the node the service runs on*. An HTTP check against `http://localhost:8080/health` genuinely tests what a local caller would experience, and a script check can inspect the host itself. If the servers ran all checks centrally, the check load would grow with the fleet and every check would measure the server-to-node path rather than the service. **A stable local endpoint.** Applications and tooling talk to `127.0.0.1:8500` (HTTP) and the agent's DNS port (8600 by default). They never carry a list of server addresses, and a server replacement is invisible to them. **Scalable failure detection.** Gossip distributes failure detection across the whole membership rather than concentrating it. When a node stops gossiping, its `serfHealth` check goes critical and its services drop out of discovery — with no health check of the service itself involved. **Protecting the Raft tier.** The servers are the scarce, consensus-bound resource. Funnelling every application's DNS lookups and registrations through client agents keeps connection counts and check execution off them. ## The failure modes this shape produces - **The agent dies but the workload lives.** Gossip marks the node failed and every service it registered disappears from discovery, even though the processes are fine. Discovery health is agent health. - **The node is partitioned from the servers.** Local checks keep running and the local API keeps answering, but anti-entropy cannot sync, so the catalog holds a view that is frozen at the moment of the split. - **Servers lose quorum.** Writes stop; what reads can still be served depends on the consistency mode requested. ## Running without client agents It is possible to point an application straight at a server's HTTP API, and some containerised setups do exactly that. You lose local checks, you lose the loopback registration path, you make the servers hold the fleet's connections, and you have to give every application a server address list. It is a deliberate trade, not the default.
- If a client agent restarts, what happens to the services that were registered against it through the HTTP API?Anything registered only via the API is in-memory state and is lost on restart unless the agent is configured to persist it or the application re-registers on startup. Registrations that live in the agent's config directory come back automatically because the agent reloads them. This is why long-lived registrations are usually written as config or re-asserted by the application at boot.
- What does anti-entropy mean in Consul, and which side wins a disagreement?Anti-entropy is the periodic reconciliation between an agent's local view of its own node — its services and check states — and the servers' catalog. The agent is authoritative for its own node, so if an operator deletes one of its services from the catalog directly, the next sync puts it back. Removing a service for good means deregistering it at the agent.
- Why are Consul server counts almost always odd — three or five rather than four?Raft needs a strict majority to commit, so a four-server cluster tolerates exactly one failure, the same as three, while adding a node's worth of replication cost and a larger quorum to reach. Odd sizes give the best failure tolerance per server. Five is chosen over three when you want to survive two simultaneous server losses, at the cost of slower writes.
saying these in an interview costs you the question
- Saying client agents cache or replicate the catalog
- Thinking health checks execute on the servers
- Assuming applications must know the server addresses
- Believing more servers always means more fault tolerance
- Treating the client agent as optional overhead with no role