skip to content

How do you run a host benchmark profile fleet-wide when it needs privileged access on every target?

level: principalimportance: nice to knowfreq 34%

answer

  1. push versus pull topology
  2. concentrated credentials, one box
  3. reads only, short-lived, scoped sudo
  4. assessing is not remediating
  5. unassessed hosts are not compliant hosts

basics

~20 s

Choose between a central runner holding credentials to every host and a local agent that ships only results, scope the privilege to reads, keep assessing separate from remediating, and measure coverage against an inventory so unassessed hosts never read as compliant.

solid answer

~50 s

The profile has to read root-owned state - file modes on `/etc/shadow`, `sshd_config` contents, mount options - so the real question is the shape of that access. Push (a central runner with SSH credentials everywhere) is easy to version and operate but concentrates fleet-wide root in one box that is usually less scrutinised than what it assesses. Pull (a local agent shipping only results) removes the inbound credential at the cost of agent lifecycle and version skew. Whichever you pick, narrow the privilege to the reads the profile needs, issue short-lived credentials, and keep the assessing identity separate from the remediating one so you have not built a fleet-wide remote-execution channel and called it compliance. Then fix the reporting: coverage is measured against an authoritative inventory, unassessed is its own category, and the results store is sensitive because it maps your weaknesses.

go deeper

for a junior

Understand why this kind of profile needs privileged access at all: it reads root-owned configuration and system state, which an unprivileged account simply cannot see.

for a middle

Be able to contrast a central runner connecting outward with an agent running locally and shipping results, and name one concrete cost on each side rather than declaring a winner.

for a senior

Show how you would harden the runner itself, scope its credentials to the reads the profile needs, and build coverage reporting that reconciles against an inventory instead of against the run's own output.

for a principal

Own the ownership split and the metric definition. Decide who chooses the controls, who operates the reach, and who remediates, and be ready to defend why unassessed hosts are reported as unknown even though it makes the quarterly number look worse.

### Why this is hard A host benchmark profile is not a network scan. To assert that `/etc/shadow` is mode `0000`, that `sshd_config` sets `PermitRootLogin no`, that `/tmp` is mounted `nodev,nosuid`, or that an audit daemon is enabled, the runner must read root-owned configuration and query system state on the machine itself. Multiply that by the whole estate and you have designed something with privileged read access to every host you own. The design question is not *how do I get root everywhere* — it is *how do I get the minimum access this needs, in a shape whose compromise is survivable*. ### The two topologies **Push / agentless.** A central runner holds credentials — SSH keys, WinRM, cloud instance-connect — and connects outward to each target. Attractive because there is nothing to install and the profile version is unambiguous: one runner, one profile, one run. The cost is concentration. That runner is now a box that can become root on every host in the fleet, and it holds long-lived credentials to prove it. It is, by construction, the highest-value target in your infrastructure, and it usually gets less scrutiny than the workloads it assesses. **Pull / local.** A small agent or a scheduled local execution runs the profile on the host, as a local privileged process, and ships only the result outward. No inbound credentials, no fleet-wide key material, and the blast radius of a compromised result collector is a report rather than root. The costs are real too: an agent to package, version and keep alive on every host; a self-reported result you have to trust; and profile-version skew across a fleet that does not upgrade uniformly. Neither is correct in the abstract. What a strong answer does is name the constraint that decides it — an estate with a working configuration-management agent already on every host has most of the pull model paid for; a fleet of appliances you cannot install software on forces the push model and forces you to invest in credential brokering instead. ### Reducing the privilege regardless of topology - **Scope the credential, not just the account.** Most checks are reads. A sudo allowlist of the specific commands and file reads a profile needs is far narrower than unrestricted root, and it is auditable. - **Make the credential short-lived.** Brokered, expiring credentials issued per run beat a key that has been in the runner's home directory for two years. - **Separate reading from fixing.** The runner that assesses should not be the mechanism that remediates. Once one identity can both read every host and change every host, you have built a fleet-wide remote-execution channel and called it compliance. - **Isolate the runner.** Its own hardening, its own change control, its own access review — held to the standard it enforces on everything else, which is an argument you should be able to make out loud. ### Ownership Three roles, and conflating them is where programmes stall. Security owns the **profile content** — which controls, which thresholds, which exemptions are acceptable. The platform team owns the **runner and its reach** — that it executes, on schedule, everywhere, with credentials that are managed. The teams that own the hosts own the **remediation**. The failure mode when one group owns all three is that the profile quietly narrows until it always passes, because the same people are graded on the result and choose the questions. ### The number that lies The report says 100% pass. It is computed over the hosts that returned a result. The hosts that matter most — the forgotten one built by hand two years ago, the appliance nobody can log into, the node in the region the runner has no route to — are exactly the ones missing from the denominator. So: - reconcile every run against an authoritative inventory, not against its own output; - report **assessed / not assessed** as a first-class category, and never let a missing host read as a compliant one; - treat an unexplained drop in host count as an incident in the assessment pipeline, because a runner that silently stops reaching a subnet looks exactly like a fleet that got healthier. ### And the results themselves The output is a ranked, machine-readable list of every weakness on every host you own, with remediation instructions attached. It deserves the access controls and retention policy you would give any other crown-jewel dataset. Handing broad read access to the compliance dashboard because "it is only reports" gives an attacker who reaches it a better map of your estate than they could build themselves.

  • Which is safer, a central runner with keys to every host or an agent on each host?
    Neither in the abstract; the constraint decides. Push concentrates credentials, so it is defensible only if the runner is hardened, its credentials are brokered and short-lived, and its access is reviewed like a privileged system. Pull removes the inbound credential but adds an agent to package, version and keep alive everywhere, and you must handle profile-version skew. If a managed agent is already on every host, most of the pull model is already paid for.
  • Your fleet report shows 100% pass. What is the first thing you check?
    The denominator. That figure is computed over hosts that returned a result, and the hosts most likely to be non-compliant are the ones the runner never reached - built by hand, in an unrouted subnet, or with a broken agent. Reconcile the run against an authoritative inventory, report assessed versus unassessed as a first-class number, and treat a silent drop in host count as an incident in the pipeline rather than good news.
  • Why keep the assessing identity separate from the remediating one?
    Because an identity that can both read and change every host is a fleet-wide remote-execution channel, and it will be a more attractive target than anything it protects. Assessment needs reads; remediation needs writes and should go through the same change control as any other production change. Keeping them apart also stops the tempting shortcut where a failing control is 'fixed' by the same job that reports on it, with no record anyone can review.
  • Who should own the profile content itself?
    Security owns which controls and thresholds apply and which exemptions are acceptable; the platform team owns that the runner executes everywhere with managed credentials; host owners own remediation. When one group owns all three, the profile quietly narrows until it always passes, because the people graded on the result also choose the questions.

A central runner with root everywhere is a master key cut for every door in the building, kept by whoever happens to run the inspection schedule.

saying these in an interview costs you the question

  • Gives the runner unrestricted root on every host and stops there
  • Reports coverage as a share of the hosts that responded
  • Assumes hosts missing from the report are compliant
  • Uses the same identity to assess and to remediate
  • Treats the findings store as low-sensitivity operational data

context