skip to content

Why declare a per-host metrics agent as a one-copy-per-host workload rather than a replicated service with the copy count set to the number of hosts?

level: middleimportance: should knowfreq 50%

answer

  1. coverage, not quantity
  2. the count is derived from host membership
  3. two on one host, none on another
  4. a new host gains a copy by itself
  5. the filter is the only dial you have

basics

~20 s

Because the requirement is coverage, not quantity. A one-copy-per-host workload derives its count from the set of matching hosts and guarantees exactly one copy on each. A copy count is just a number, and nothing stops it landing two copies on one host and none on another.

solid answer

~50 s

The two declarations express different invariants. A replicated service says `this many copies exist somewhere`; the scheduler is free to place them wherever they fit, so with the count set to the host count you can easily get two copies on a roomy host and none on a busy one — and a host that joins tomorrow gets nothing until someone edits the number. A one-copy-per-host workload says `every host matching this filter has exactly one copy`. The count is derived, not declared: a host joining gains a copy automatically, a host leaving loses one, and a host that stops matching the filter has its copy removed. For anything whose whole job is to observe or serve the host it sits on, coverage is the correctness property and quantity is a coincidence that usually holds.

code

pseudocode · 12 lines
pseudocode
# one copy per host: the target is recomputed from host membership
for each host in cluster:
    if matches(host, workload.hostFilter) and not hasCopy(host, workload):
        placeCopy(host, workload)

for each copy in copiesOf(workload):
    if copy.host not in cluster or not matches(copy.host, workload.hostFilter):
        removeCopy(copy)

# replicated service: the target is a number someone wrote down
while countRunning(service) < service.copies:
    placeCopy(anyHostWithRoom(service), service)   # may be a host that already has one

go deeper

for a junior

The requirement is coverage: every host gets exactly one copy. A plain copy count only promises a total, and totals can pile two copies onto one host while another gets none.

for a middle

Explain that the per-host count is derived from host membership on every pass of the platform's loop, so joins and departures are handled automatically, while a declared count goes stale the moment the estate changes.

for a senior

Show the costs too: a fleet-wide reservation footprint, an update that touches every host and briefly drops coverage there, and a host filter that must describe a property rather than enumerate today's hosts.

for a principal

Set the rule for what is allowed to run on every host at all. Each per-host workload is a tax on the entire estate and a shared failure domain with everything serving beside it, so the bar for adding one belongs to whoever owns platform capacity.

## A derived count against a declared one The distinction is not about how many copies end up running — on a good day both give you the same number. It is about **what the platform is holding true**. | | Replicated long-running service | One-copy-per-host workload | |---|---|---| | Invariant | this many copies are running **somewhere** | **each matching host** has exactly one copy | | Where the count comes from | you declare it | derived from host membership | | Two copies on one host | allowed, and common | prevented by the contract | | A host with none | allowed | a gap the platform fills | | New host joins | nothing happens | a copy appears | | Host leaves or stops matching | count is refilled elsewhere | that copy is simply removed | A metrics and log-shipping agent reads the host it runs on: its capacity, its running workloads, the output streams landing there. Its correctness property is therefore **coverage** — every host observed, exactly once. Quantity is not the requirement; it is a number that happens to equal the requirement while nothing moves. ## Why the copy-count version goes wrong Three failure modes, in the order you usually meet them: 1. **Uneven placement.** The scheduler places copies where they fit. Nothing in a plain copy count says "one per host", so a host with spare capacity can take two and another can take none. The result is a host with no observation at all — and a duplicate stream from another, which double-counts everything measured per host. 2. **Membership drift.** Hosts join and leave routinely: capacity changes, a host is replaced, a failed one is cycled out. Every one of those events makes a declared count stale, and the correction is a human editing a number. Between the event and the edit you are blind on some host. 3. **Silent partial coverage.** Nothing complains. The declared number of copies *is* running, so the platform is satisfied; the gap only shows up as a host missing from a dashboard, which is exactly the kind of absence nobody notices. The one-copy-per-host contract removes all three by construction, because the platform derives the target from the host set on every pass of its loop rather than from a number written down once. ## What you give up This contract is narrower on purpose, and a candidate who only lists its advantages has not used it. - **You cannot choose the number.** There is no meaningful way to run "a few" of these; the only dial is which hosts match the filter. - **Its footprint scales with the estate.** Every host pays the agent's reservation, so an agent that is slightly too heavy is slightly too heavy hundreds of times over. - **Replacement is per host.** Rolling a new version across the estate touches every host, and on the host being updated the coverage is briefly gone — a gap that is inherent rather than accidental. - **It competes with the workloads it observes.** The agent shares its host's capacity with whatever is serving there, so its ceiling matters more than its importance suggests. ## Choosing the filter is the real decision Because the count is derived from matching hosts, the filter is the whole control surface. Common shapes: all hosts; only hosts in a particular role or hardware class; every host except those reserved for a special purpose. The trap is a filter written around today's estate — if it enumerates hosts rather than describing a property of them, you have reinvented the stale declared count with extra steps. ## When a replicated service is right anyway Not everything that exists on many hosts wants this contract. If the workload does not care where it runs — a stateless request handler, a queue worker, an aggregator that receives data pushed to it — then quantity genuinely is the requirement and a declared count is the honest declaration. The test is simple: **does this workload's job depend on which host it is on?** If yes, coverage is the invariant. If no, you want a number, and pinning one copy per host just makes the fleet's size dictate your capacity for no reason.

  • What happens to a one-copy-per-host workload when a host is added to the cluster?
    The platform's next pass sees a matching host with no copy and places one, with no change to any declaration. That is the point of the contract: membership is the input, so the estate growing or shrinking is handled by the same loop that keeps coverage true, rather than by someone remembering to raise a number.
  • Is there any case where two copies of a per-host agent on one host would be acceptable?
    Not under this contract — it exists precisely to prevent that, because a doubled agent double-reports everything it measures per host and doubles the load it puts on the host it is observing. If you genuinely need two different collectors on each host, they are two workloads with two filters, not one workload with the guarantee relaxed.

Fire extinguishers are mounted one per floor, not "twelve per building". If you buy twelve and let someone put them wherever there is wall space, you can end up with two on one floor and none on another — and the building still reports twelve.

saying these in an interview costs you the question

  • Says setting the copy count to the host count gives the same guarantee
  • Thinks the scheduler naturally spreads copies one per host without being told
  • Believes a per-host workload can be scaled by raising a number
  • Forgets that hosts join and leave, so a declared count goes stale
  • Ignores that every host pays the agent's reservation
  • Treats double coverage on one host as harmless duplication