skip to content

Per collector, per site, or per fleet - how fine should a several-hundred-host fleet's credential identity be cut?

level: seniorimportance: nice to knowfreq 28%

answer

  1. scope at the unit of action
  2. would you stop one host alone?
  3. twenty sites, four hundred hosts
  4. enrolment and retirement decide survival
  5. the destination may offer one identity

basics

~20 s

Cut at the unit you would actually act on. If you would never take one collector offline without taking its whole site offline, per-site identity gives most of the containment for a fraction of the identities, and per-host granularity is paying for a distinction you will never use.

solid answer

~50 s

Granularity is a trade between the radius you leave standing and the identities you must operate forever. Work it with numbers: 400 collectors across 20 sites is 20 hosts per site, so a credential per site cuts the radius by a factor of twenty while costing 20 identities instead of 400 - five percent of the enrolments, withdrawals and grants of the per-host cut. The right unit is the one at which you would actually take action during an incident and at which you would actually reconfigure. Two constraints bound the choice: the destination must be able to hold and withdraw that many identities separately, and someone must remove an identity when a host is retired, or the fleet quietly accumulates holders nobody tracks. Per-host is right where each host is independently valuable and independently suspect; below that, finer cuts buy resolution nobody uses.

go deeper

for a junior

Recall that more identities means smaller damage per leak and more work to run, and that the choice is a trade rather than a rule.

for a middle

Explain the arithmetic: how many holders share each value, how many identities that implies, and what each intermediate cut buys against the finest one.

for a senior

Show the operational judgement - the unit you would actually act on, what the destination can support, and why enrolment and retirement decide whether the model lasts.

for a principal

Set the standard for the estate: the granularity teams must reach, the automation that makes it survivable, and the cases where you accept a coarser cut and record why.

## The unit of action, not the unit of hardware The instinct is that finer is always better, so the ideal is one identity per host, per process, per run. The useful test is different: **what is the smallest unit you would actually take action on?** Scoping finer than your unit of action buys a distinction you will never exercise, and you pay its operational cost every day regardless. Two actions define the unit: - **Withdrawal.** Would you really stop one collector while its nineteen neighbours keep uploading? If the site's data is useless with one host missing, the honest unit is the site. - **Investigation.** When the record names a holder, what is your next move? If it is *quarantine the whole site and rebuild it*, then site resolution was all you needed. ## The arithmetic, stated with its assumptions Assume 400 collectors spread over 20 sites, so about 20 per site. | cut | holders per credential | identities to operate | radius of one leak | record resolves to | |---|---|---|---|---| | per fleet | 400 | 1 | the whole estate's uploads | the value only | | per site | 20 | 20 | one site | the site | | per collector | 1 | 400 | one host | the host | The middle row is the one teams under-use. It is a **twentyfold reduction in radius for five percent of the identity count** of the finest cut, and it converts an un-shrinkable suspect set into a group of twenty. If a fleet cannot get to per-host today, per-site is not a consolation prize - it is most of the benefit at a small fraction of the price. ## What the destination has to support Fine cutting is only available if the receiving system can hold many identities and withdraw them independently. Some destinations accept exactly one login for a partner and nothing finer; there, the holder axis is simply not offered, and the cuts you can still make are by right and by environment. Say this out loud when you answer - a plan that assumes an identity per host against a destination that has one account is not a plan. ## The costs that scale with identity count 1. **Enrolment.** Every new host needs an identity before it can upload, which means the identity must be created as part of building the host rather than by a human afterwards. 2. **Withdrawal on retirement.** Every retired host leaves an identity behind. One nobody removed is a holder you no longer track - a hole on the very axis you paid to close. 3. **Grants.** Hundreds of identities need their rights kept equal and kept narrow; divergence between them is invisible until an incident finds the one that was different. 4. **Diagnosis.** *This host cannot upload* becomes a per-identity question rather than a fleet-wide one, which is better for blast radius and worse for the support load. The first two are the ones that decide whether the cut survives. A per-host identity model that depends on a person remembering to create and remove identities degrades toward sharing within a year, because the pressure to reuse an existing value when a host is rebuilt at three in the morning is enormous. ## How to choose, in one paragraph Start from the incident you expect. If a single collector can be compromised independently and is worth investigating on its own, cut per host and make enrolment and withdrawal part of building and destroying the machine, not a human step. If collectors are identical, disposable and always handled as a group, cut per site or per role and spend the saved effort on the right axis instead, where the cut is cheaper and the reduction is often larger. Whichever you choose, say what you left standing: the group that still shares one value is the group you will not be able to tell apart when it matters.

  • What makes a per-host identity model decay back into sharing?
    Manual enrolment. If creating an identity for a rebuilt host is a human step, someone will paste a neighbour's value at three in the morning and the fleet drifts back toward one credential. The model survives only when the identity is created as part of building the machine and withdrawn as part of destroying it.
  • The destination will only hold one account for your organisation. What scoping is left?
    The holder axis is unavailable, so cut the other two: narrow the rights that one value carries to exactly what the fleet does, and keep separate values for separate environments if the destination offers even that. Then be explicit in the risk record that containment is fleet-wide and attribution is impossible, because that is now a property of the integration, not a choice you can revisit quietly.

saying these in an interview costs you the question

  • Argues finer granularity is always better regardless of what it costs
  • Plans per-host identity without a way to create and remove them automatically
  • Ignores whether the destination can hold more than one identity
  • Treats an intermediate cut such as per site as a failure rather than a reduction
  • Forgets that retired hosts leave identities nobody tracks