skip to content

An auditor asks you to defend '94% endpoint log coverage' — which denominator and evidence do you present?

level: principalimportance: should knowfreq 40%

answer

  1. the denominator is the argument
  2. reconcile two independent asset lists
  3. numerator is arrivals, not installs
  4. a contract counting enrolments buys enrolments
  5. residual is a named register with owners

basics

~20 s

Name the denominator first — which authoritative asset list the figure is over — and prove the numerator from events indexed in a stated window, not agents enrolled. Present the missing six percent as a named, owned list.

solid answer

~50 s

A coverage percentage is an argument about two numbers, and the denominator is the one people skip. The CMDB, the identity directory, DHCP and DNS records and network-authentication logs each yield a different host count for the same estate, and hosts appearing in one and not another are precisely the uncovered ones — so I reconcile at least two independent inventories and declare which list the figure is over. The numerator must be hosts whose events actually landed inside a stated window, not agents installed; otherwise the metric rewards enrolment. That matters most where a provider co-manages the estate and its contract counts enrolled devices, because you are paying for a configuration they control rather than for data you receive. So: publish reporting-in-24h over a reconciled denominator, hold the residual as a named exception register with an owner and review date, and move the provider's obligation onto delivery measured in your SIEM.

go deeper

for a junior

Know that any coverage percentage needs a stated denominator and a stated window, and that 'covered' should mean events arrived, not that an agent was installed.

for a middle

Be able to explain why the CMDB, the identity directory and DHCP records give different host counts, and why the hosts in the gaps between them are the least likely to be reporting.

for a senior

Show how you construct the figure end to end — reconciled denominator, reporting-in-window numerator, exception register — and what you would fix first when the two numbers diverge.

for a principal

Own the metric as an incentive. Decide what the contract measures, whose system is the source of truth for it, what target is honest, and how the residual is governed rather than rounded away.

## The number is an argument, not a measurement '94% coverage' is meaningless until three things are fixed: what the denominator counts, what the numerator counts as covered, and over what window. Auditors who know the domain go straight to the first; the ones who do not will accept a number that is quietly indefensible, which is worse. ## Choosing and defending a denominator Every estate has several plausible host counts, and they never agree: - the **CMDB or asset register** — authoritative on paper, usually stale, and the place decommissioned hosts go to live forever; - the **identity directory** — good for domain-joined machines, blind to anything that never joins; - **DHCP and DNS records** — captures whatever actually appeared on the network, including things nobody registered, and churns constantly; - **network authentication or NAC** — captures what was admitted to the wire; - the **cloud provider's own inventory**, for instances that live for hours. The hosts that exist in one list and not another are not a rounding error; they are the population most likely to be uncovered, because nobody owns them. So the defensible construction is: pick the list you can name as authoritative for the scope being audited, reconcile it against at least one independent list, and disclose both the choice and the size of the disagreement. If two systems differ by eight hundred hosts, that disagreement is itself a finding for the asset owner and belongs in the report next to the percentage. The one edit never to make is removing unreachable or silent hosts from the denominator. It raises the number every time and reduces its meaning to zero. ## Choosing a numerator There is a ladder of increasingly honest definitions of 'covered': 1. **Licensed** — you bought capacity. Says nothing. 2. **Installed / enrolled** — a configuration record exists. This is what most consoles report and what most contracts count. 3. **Heartbeating** — the agent checks in. Proves the agent lives, not that security data arrives. 4. **Reporting in the last 24 hours** — events attributable to that host were indexed. This is the minimum defensible definition. 5. **Reporting on every expected channel** — the host is delivering all the record types the detections require, not just one. This is the definition that survives contact with an incident. Publish at rung four or five, and say which. Rung two is the one that looks best and means least. ## The incentive problem with a co-managing provider Where a provider runs the platform or the estate, the metric is not just a report — it is a contractual obligation, and obligations shape behaviour. An SLA written on enrolled devices is an SLA on a configuration the provider fully controls; they can hit 98% enrolment while your reporting figure sits at 91%, and both statements are true. The provider is not cheating; the contract asked for the wrong thing. The fix is contractual, and it is a negotiation rather than an email: - move the obligation to **reporting within a window on named channels**, over a denominator both sides agree in writing; - define the **measurement source as your SIEM**, not the provider's console, so the evidence is on your side of the boundary; - keep enrolment as a leading indicator that helps them manage the work, not as the obligation; - agree how disputed hosts are handled — a joint reconciliation cadence, and a rule for what happens to a host neither party can find. Expect resistance, because delivery depends on network paths and host states the provider does not wholly own. That is a real objection and it is answered by scoping — excluding stages genuinely outside their control — not by reverting to enrolment. ## The residual The last few percent are dominated by hosts nobody owns, short-lived instances, kit in transit and equipment with genuine constraints, and the cost curve turns vertical while risk reduction does not. Chasing 100% is not the defensible position; a **named exception register** is: one row per host or class, with an owner, a reason, a compensating measure where one exists, and a review date. What you are proving to an auditor is not that the gap is zero but that the gap is *known, owned and shrinking* — a register that is smaller and better-attributed than last quarter is a stronger answer than a higher percentage with no register behind it. ## Two honest caveats to volunteer Volunteer them, because being asked for them is worse. First, the window: a figure measured over 24 hours will differ from one measured over an hour, and laptops that travel legitimately fall in and out. State the window and the treatment of known-offline hosts, with an expiry so they cannot sit in a benign bucket forever. Second, the scope of the claim: a covered host means a delivered log source. It is not a promise that anything malicious on that host would be detected — that is a different measurement with a different owner, and conflating the two is how a coverage number becomes a false assurance to a board. ## What an interviewer is listening for That you interrogate the denominator before defending the percentage; that your numerator is delivery rather than configuration; that you see the incentive a contract metric creates; and that you present the residual as an owned list rather than as rounding.

  • Two asset systems disagree by 800 hosts. Which one do you report against?
    Report against the list you can defend as authoritative for the audited scope, and disclose the reconciliation. In practice I take the union for the denominator — a host that exists in either system is a host — and raise the disagreement itself as a finding for the asset owner, because 800 machines nobody agrees exist is a bigger exposure than the percentage they move.
  • The provider meets a 98% enrolment SLA while your reporting figure is 91%. What changes?
    The measurement in the contract. Enrolment is a configuration they control completely; reporting is the thing you are actually buying. I would move the SLO to reporting within a window on named channels, over a denominator agreed in writing, measured in our SIEM rather than their console, and keep enrolment as a leading indicator. Where delivery genuinely depends on stages they do not control, scope those out explicitly rather than reverting to enrolment.
  • Why not simply drive coverage to 100%?
    Because the last few percent are hosts nobody owns, ephemeral instances and equipment with real constraints, and the cost curve turns vertical while the risk reduction flattens. The defensible position is a stated target plus an exception register with an owner, a reason, a compensating measure and a review date per entry — and evidence that the register is shrinking. A rounded-up number with nothing behind it fails the first serious question.

saying these in an interview costs you the question

  • Quotes a coverage percentage without naming its denominator
  • Counts enrolled agents as covered hosts
  • Drops unreachable hosts from the denominator
  • Presents the residual as rounding rather than a named list
  • Accepts a provider figure measured on the provider's own console

context