How does Salt's master/minion architecture work, and why can a command like `salt '*' cmd.run 'uptime'` come back from thousands of hosts in seconds?
answer
- one publish, many listeners
- minions dial out, master never in
- target evaluated on the minion
- 4505 publish, 4506 return
- offline minions simply miss the job
basics
~20 sSalt minions hold a persistent ZeroMQ connection out to the master. A command is published once onto that bus, every minion that matches the target expression evaluates it locally and runs it, then returns its own result, so fan-out time barely grows with host count.
solid answer
~50 sSalt runs two daemons: `salt-master` and `salt-minion`. The minion dials **out** to the master and keeps a persistent connection — 4505 to subscribe to published jobs, 4506 to return results — so no inbound port is opened on the managed host. When you run `salt '*' cmd.run 'uptime'`, the master publishes **one** message containing the target expression and the function; every connected minion receives it, decides for itself whether it matches, and the matching ones execute locally and send results back on 4506. That is why fan-out is roughly constant work for the master instead of one connection per host. Targeting can be a glob, a grain (`-G 'os:Ubuntu'`), a pillar (`-I`), an explicit list (`-L`), a regex (`-E`) or a compound expression (`-C`). Crucially, remote execution is first-class: `cmd.run`, `pkg.install` and `test.ping` need no state file at all, and applying config is just another execution module, `state.apply`.
code
bash · 5 linessalt-key -L
salt-key -A -y
salt '*' test.ping
salt -G 'os:Ubuntu' cmd.run 'uptime'
salt -b 10 '*' state.applygo deeper
Know the two daemons by name and that the minion connects out to the master over a persistent connection. Be able to say what salt '*' test.ping does and that '*' is a target expression, not a shell glob over files.
Explain the publish/return split across 4505 and 4506, that each minion matches the target itself, and why that makes fan-out cost roughly independent of fleet size. Show you know the common targeting flags and that execution modules work without any state file.
Demonstrate operational judgment: batch large jobs so you do not stampede mirrors, use --async plus job lookup for long runs, and know that offline minions silently miss a publish so convergence needs a scheduled run, not a one-off command.
Own the topology and the risk. Decide between multi-master and syndic tiers for throughput and failure domains, define who may publish to the whole fleet, and treat the master as a root-everywhere trust concentration that needs the same controls as a CI signing key.
## Two daemons, one long-lived connection Salt is built around `salt-master` and `salt-minion`. The important direction is that the **minion connects out to the master**, not the other way round. That means a managed host needs no inbound listening port and works fine behind NAT; the master needs two ports reachable: - **4505** — the publish port. Every minion holds a subscription here and receives jobs the master publishes. - **4506** — the request/return port. Minions post job results back, and fetch files and pillar data over it. Trust is established with public keys. On first contact a minion sends its public key; it stays pending until an operator accepts it (`salt-key -L` to list, `salt-key -a <id>` or `-A` to accept). Until then the minion receives nothing. The minion's identity is its **minion id**, defaulting to the FQDN and settable in `/etc/salt/minion`. ## Why fan-out is fast The cost model is the whole point of the architecture. A push tool that reaches hosts over SSH pays a TCP connect, a key exchange and an authentication round trip **per host**, then parallelises with forks or threads — so 5,000 hosts is 5,000 connections that someone has to schedule. Salt pays that cost once, at minion start, and keeps the socket. Publishing a job is a **single** message onto the bus carrying the target expression, the function name and its arguments: ```bash salt '*' cmd.run 'uptime' ``` Every connected minion receives that one message and **evaluates the target against itself**. Non-matching minions drop it; matching minions run the function locally with their own copy of the code, then return results individually over 4506. Master-side work is one publish plus a stream of returns, so latency is dominated by the slowest minion rather than by connection setup multiplied by fleet size. ## Targeting Because matching happens on the minion, the target expression is just data on the wire: - glob on minion id (the default): `salt 'web*' test.ping` - grain: `salt -G 'os:Ubuntu' pkg.upgrade` - pillar: `salt -I 'role:database' state.apply` - list: `salt -L 'web1,web2' service.restart nginx` - regex: `salt -E '^web\d+' test.ping` - compound: `salt -C 'web* and G@os:Ubuntu' test.ping` - nodegroup (named in the master config): `salt -N frontend test.ping` One caveat interviewers like: grains are supplied by the minion, so grain targeting is a convenience, not a security boundary. ## Remote execution is not a side feature In several configuration-management tools the only real verb is "converge this host to the config". In Salt, **remote execution exists independently of any state file**. `test.ping`, `cmd.run`, `pkg.install`, `service.restart`, `disk.usage`, `network.interfaces` are execution modules you call directly, ad hoc, and they are how most operators first use Salt. Desired-state configuration is layered on top as just another execution module — `state.apply` (with no argument it applies the highstate assembled from `top.sls` in `file_roots`; with an argument it applies one SLS). This is why Salt is often described as an orchestration and remote-execution framework that happens to ship a configuration-management system. ## Jobs, asynchrony and returns Every publish gets a **job id (JID)**. By default the CLI blocks and prints returns as they arrive, giving up on stragglers after the timeout (`-t`). With `--async` the CLI prints the JID immediately and you inspect the job later: ```bash salt --async '*' state.apply salt-run jobs.lookup_jid 20240101120000123456 salt '*' saltutil.running ``` Returners can additionally ship results to an external store instead of, or as well as, the master job cache. ## Where it bites in production - **Thundering herd.** Publishing a heavy job to the whole fleet makes every minion start at once — hammering package mirrors or the master's file server. Use `-b`/`--batch` (`salt -b 10 '*' state.apply`) to run in waves. - **Fire and forget.** A publish is not queued for minions that are offline; they simply never see that job. Re-target them later, or use the minion-side scheduler for recurring runs. - **Master scaling and blast radius.** One master is both a throughput ceiling and a security concentration — whoever controls it executes as root everywhere. Salt offers multi-master and `salt-syndic` tiers for the throughput half of that problem.
- If targeting is evaluated on the minion, what stops an untrusted minion from running jobs meant for another host?Every minion still authenticates with its accepted public key, and the master encrypts payloads such as pillar data per minion, so a minion only gets data targeted at it. But the *published* job itself is broadcast, and grains are minion-supplied — which is why grain targeting is treated as convenience, not authorization, and why pillar is the place for anything secret.
- What is the practical difference between running `salt '*' state.apply` and `salt '*' cmd.run 'some-script.sh'`?`state.apply` runs the declarative state system: it compiles SLS into a dependency-ordered set of states, checks current versus desired for each, and reports what changed — so re-running it is idempotent. `cmd.run` just executes a command every time and reports it as changed unless you constrain it with `unless`, `onlyif` or `creates`.
- How do you keep a fleet-wide `state.apply` from overwhelming the master and your package mirrors?Batch it. `salt -b 10 '*' state.apply` runs ten minions at a time and only starts the next as the previous finish, and a percentage such as `-b 5%` scales with fleet size. For scheduled convergence, stagger with the minion-side scheduler and a splay so minions do not all wake at the same second.
saying these in an interview costs you the question
- Says the master SSHes into each minion to run commands
- Thinks the master evaluates the target and contacts matches individually
- Assumes offline minions get the job queued and replayed later
- Believes Salt can only apply state files, not ad-hoc commands
- Treats grain-based targeting as a security boundary