You have roughly 5,000 servers to keep configured and you also need ad-hoc commands across them. What does adopting Salt's persistent minion agent buy you over a purely SSH-push approach, and what does it cost?
answer
- connection cost per fleet, not per host
- persistent agent buys events and latency
- salt-ssh trades the bus for no daemon
- master is a root-everywhere blast radius
- syndic tiers when one master is not enough
basics
~20 sThe agent buys constant-time fan-out and event-driven reaction: connections are established once, so a command is one publish rather than 5,000 SSH handshakes, and minions can push events the master reacts to. It costs an agent lifecycle, key management and a master that is root-everywhere.
solid answer
~50 sThe persistent minion changes the cost model. An SSH-push tool pays connect, key exchange and auth per host, per run, and parallelises with forks; Salt pays it once at minion start, so publishing to 5,000 hosts is a single message and completion is bounded by the slowest minion rather than by connection scheduling. It also unlocks things SSH push structurally cannot do: minions raise events, so beacons plus the reactor give you closed-loop remediation, and the minion-side scheduler converges hosts without anything reaching in. The costs are real: you must install, upgrade and monitor a daemon on every host, manage accepted keys, keep the master reachable, and accept that a compromised master executes as root fleet-wide — a much sharper blast radius than distributing SSH access. Salt hedges with `salt-ssh`, an agentless mode using a roster, which gives you the state system without the bus, at SSH speed.
code
yaml · 8 linesweb1:
host: 10.0.1.10
user: deploy
sudo: True
web2:
host: 10.0.1.11
user: deploy
sudo: Truego deeper
Be able to state the basic difference: Salt normally runs a minion daemon that stays connected, while a push tool opens a connection per host each run. Know that salt-ssh exists as Salt's agentless mode.
Explain why the persistent connection makes fan-out roughly constant work, and what depends on it — events, beacons, the reactor, the minion scheduler — so you can say concretely what salt-ssh gives up.
Argue both sides with operational detail: agent lifecycle and version skew, key acceptance during provisioning, a dead minion being indistinguishable from a dead host, and batching so a fleet-wide job does not stampede the master and the mirrors.
Own the decision and its governance. Weigh the root-everywhere master against distributed SSH credentials, decide whether environments share a master, plan syndic or multi-master topology before the fleet forces it, and be comfortable recommending a mixed estate rather than one model everywhere.
## Frame the question as a cost model, not a feature list The honest comparison is about where the connection cost lives and what capabilities that unlocks. Everything else follows. **Push over SSH.** Every run: open TCP, negotiate, authenticate, ship the payload, execute, collect, close — per host. Parallelism is forks or threads on the control node, so 5,000 hosts means tuning fork counts and living with a long tail. There is no persistent channel, so the control node only knows what a host is doing while it is holding a connection to it. **Persistent agent.** `salt-minion` connects out to the master once and stays connected. A job is a single publish; each minion evaluates the target itself and returns its result. Fan-out is close to constant work for the master, and the wall-clock time is dominated by the slowest minion rather than by connection setup multiplied by fleet size. At 5,000 hosts that difference stops being an optimisation and becomes the difference between a usable and an unusable tool for ad-hoc work. ## What the persistent connection unlocks This is the half candidates usually miss. A live bidirectional channel is not just faster — it makes the minion an active participant: - **Events flow upward.** Beacons on the minion emit events; the master's reactor acts on them. That is closed-loop remediation with detection at the edge, and it has no equivalent in a model where the control node must initiate everything. - **The minion can converge itself.** The minion-side scheduler runs `state.apply` on an interval with a splay, so drift correction does not depend on anyone reaching in — useful for hosts that are frequently offline or behind restrictive networks. - **No inbound access needed.** Minions dial out, so hosts behind NAT or with no inbound SSH still participate. With push, connectivity to every host is a hard prerequisite. - **Real-time inventory.** `grains.items` across the fleet answers "which hosts are on this kernel" in seconds, because the channel is already open. ## What it costs 1. **Agent lifecycle.** A daemon on every host has to be installed at build time, upgraded, and monitored — and a minion that is dead is invisible, so "no return" now has two meanings: the host is down, or the agent is. Version skew between master and minions is a real operational chore. Push tools genuinely have less to go wrong here, which is their main selling point. 2. **Key management.** Keys must be accepted, and that step is where you decide whether an unknown host may join. Auto-accept is convenient and a bad idea in most estates; without it, provisioning has to include a key-acceptance step. 3. **Blast radius.** This is the argument that carries weight at a senior level. The master can execute arbitrary commands as root on every minion, with no per-host credential in the way. Whoever controls the master — or its reactor rules, or the pillar tree — controls production. It deserves the controls you would put on a CI signing key: restricted access, audited change to the state and reactor repositories, and a hard look at who can publish fleet-wide. A push model distributes that authority into SSH credentials, which is more auditable per host and harder to fully centralise. 4. **Master availability and scale.** The master is now on the critical path for both ad-hoc work and, if you use the reactor, remediation. At 5,000 minions you plan for it: multiple masters, or `salt-syndic` to tier masters under a top-level master so publishes fan out through a tree instead of one process. ## Salt's own hedge: `salt-ssh` Salt does not force the choice. `salt-ssh` runs the same state system over SSH with no minion installed, using a **roster** instead of accepted keys: ```yaml web1: host: 10.0.1.10 user: deploy sudo: True ``` You get states, pillar and execution modules without a daemon — and you give up exactly what the bus provided: the constant-time fan-out, events, beacons and the reactor, and the minion scheduler. It is the right tool for appliances you cannot install software on, for bootstrapping the minion itself, and for a small number of hosts where the agent is not worth it. ## The answer that lands Say that at 5,000 hosts with a genuine ad-hoc requirement, the agent's constant-time fan-out and the event bus are worth real money, and that the price is a daemon fleet plus a master that is a root-everywhere trust concentration you must govern deliberately. Then note that a mixed estate is normal: minions where you own the image, `salt-ssh` for the edges. Refusing to name the security cost is what makes this answer sound junior.
- What exactly do you give up by choosing salt-ssh over installing minions?The bus. You keep states, pillar and execution modules, but lose constant-time fan-out — every run is an SSH connection per host again — and lose everything that depends on a live channel: events, beacons, the reactor, and the minion-side scheduler. It is the right call for appliances, for a handful of hosts, or for bootstrapping the minion itself.
- How would you reduce the blast radius of the master without giving up the agent model?Treat the master as production-change infrastructure: restrict who can publish, require review on the state, pillar and reactor repositories since those are executable authority, and separate environments onto different masters rather than sharing one across prod and dev. Keep an audit trail of published jobs through the job cache or an external returner, and prefer targeted publishes over habitual `'*'`.
- At 5,000 minions, what breaks first and what do you do about it?The single master — publish throughput, the file server under a fleet-wide highstate, and its own availability as a single point on the critical path. Batch large jobs so minions do not stampede, stagger scheduled convergence with a splay, and scale the topology with multiple masters or a syndic tier rather than a bigger box.
saying these in an interview costs you the question
- Says the agent is always better without naming a cost
- Ignores that the master executes as root fleet-wide
- Assumes salt-ssh keeps beacons and the reactor working
- Treats one master as sufficient at any fleet size
- Argues the agent is only about avoiding SSH keys