skip to content

Surviving the Trip

A record that never leaves the host, arrives four hours late, or lands under a vendor field name no rule knows is worth nothing later. Interviewers probe the plumbing behind the search bar.

on this pageshow

explore

questions

12

In SOC log-source coverage, what is the difference between an enrolled host and a reporting host?

level: juniorimportance: must knowfreq 72%

answer

  1. configuration fact versus observed fact
  2. console says deployed, index says nothing
  3. last-seen per source, event time and index time
  4. count arrivals, not installations
  5. silence is missing data, not calm

basics

~20 s

Enrolled means a host is registered in the collection console, which is a configuration fact. Reporting means its events actually arrived in the SIEM inside a recent window, which is an observed fact. Count coverage from arrivals.

solid answer

~50 s

Enrolment is what someone configured; reporting is what the pipeline can prove. An agent can be installed, licensed and shown as deployed while sending nothing: the host is powered off, the audit policy producing the records was never enabled, the log path moved after an upgrade, the shipper's queue is wedged, or a firewall now blocks the collector. A dashboard counting installed agents is counting intent. The figure that matters is per source: when did the SIEM last index an event from it, and does the last hour look like this source's normal hour? I would build coverage from a last-seen table joined to an asset denominator, publish both numbers, and work the gap as a queue with named owners. Silence from a source is never evidence it was quiet; it is evidence you have no data.

go deeper

for a junior

Be ready to say plainly that an agent being installed and events actually arriving are two different facts, and that coverage should be counted from the second one.

for a middle

Explain how the gap opens in practice — powered-off hosts, audit policy never enabled, a rotated log file, a blocked collector port — and how a per-source last-seen table is built and read.

for a senior

Show how you turn the enrolled-versus-reporting gap into a worked queue with owners, add a per-source volume check on top of last-seen, and reconcile against an asset inventory you did not write.

for a principal

Own which single figure leaves your team and what it commits you to. Publishing enrolment as coverage creates an incentive to enrol rather than to collect, and that incentive outlives whoever set the metric.

## Two different facts A log-source coverage number is built from two claims that people routinely conflate. - **Enrolled** — a configuration statement. An agent is installed, a forwarder is configured, a cloud audit trail is switched on, a device is pointed at a syslog collector. The evidence lives in a management console, a deployment tool, or a licence count. - **Reporting** — an observed statement. Events attributable to that source were indexed by the SIEM within a stated window, at a volume consistent with what that source normally produces. Enrolment is a promise. Reporting is delivery. Only the second is evidence, because only the second was produced by the pipeline you will actually search during an investigation. ## Why the gap opens The distance between the two numbers is never zero, and the causes are mundane rather than exotic: - the host is powered off, in a drawer, on a shelf, or was rebuilt and never re-enrolled; - the agent is running but the underlying audit source is not — Windows audit policy not configured for the categories the rules need, `auditd` rules never loaded, a cloud audit trail enabled in one region only; - the log file the shipper tails was rotated to a new inode, or moved when a package upgraded; - the network path changed: a new egress rule, a route, an expired certificate on the collector's TLS listener; - back pressure — the shipper's disk queue filled and it is dropping, or an index-side filter or volume cap is discarding the events after they arrive; - the source sends *some* channels and not others, so it looks alive while the specific data your detections need has stopped. None of these change the console. All of them change what you can search. ## Measuring reporting honestly The basic construct is a **last-seen table**: one row per source, carrying the most recent event time and the most recent index time, plus the volume in the last hour and in the same hour last week. - **Event time** is the timestamp inside the record. It can be wrong — a skewed host clock, or a backlog being replayed after an outage will produce old event times arriving now. - **Index time** is when the SIEM actually received and stored it. This is what tells you whether the path works right now. Compare them. Current event times with a stale index time means the pipeline has stalled or is lagging. Current index time with event times eleven days old means a backlog is draining. Alert on index time; investigate with both. Last-seen alone catches total silence only. A source can keep one channel flowing while another dies, so its last-seen stays fresh while half its telemetry is gone. That is why per-source, per-channel **volume** against that source's own baseline — with an absolute floor for sources that are normally sparse — is the stronger check. ## The denominator A percentage needs a bottom half. The CMDB, the identity directory, DHCP or DNS lease records, and network-authentication logs each yield a different host count for the same estate, and the hosts that exist in one list and not another are exactly the ones nobody is collecting from. Whatever list you pick, name it when you publish the number, and reconcile against at least one independent source. Resist the temptation to remove silent hosts from the denominator. That is the one edit that always raises the percentage and always lowers the truth. ## What silence proves The direction of the claim matters. Records arriving prove activity happened. Records **not** arriving prove only that you have no records. It does not prove the host was idle, it does not prove nothing malicious ran, and it does not prove the host still exists. A host that has been silent for six days is an unexplained gap to be worked, not a clean host to be ticked off — and it is also not, on that evidence alone, a compromised one. ## What an interviewer is listening for That you distinguish configuration from observation; that you know coverage is computed from arrivals over a named denominator; that you can list two or three concrete ways an installed agent stops delivering; and that you say out loud that absence of data is missing evidence rather than good news.

  • Should a source's last-seen check use the event's own timestamp or the time the SIEM indexed it?
    Keep both and compare them. Event time comes from the host and can be skewed by a wrong clock or by a backlog replaying old records; index time tells you when data actually landed. Current event times with a stale index time means the path has stalled; a current index time carrying eleven-day-old event times means a queue is draining. Alert on index time, then use both to explain what happened.
  • A laptop has been powered off for three weeks. Is that a coverage gap?
    It is a gap in your number even though nothing was missed while the machine was off, and from the SIEM alone you cannot tell 'off' from 'silenced'. Reconcile against an independent signal — directory last sign-in, DHCP or DNS records, MDM check-in — and move confirmed-offline hosts into a separate bucket with an expiry date. Do not delete them from the denominator to make the percentage look better.
  • Why is a per-source event count better than a simple last-seen check?
    Last-seen only catches total silence. A source that keeps emitting one channel while another dies still shows a fresh last-seen, so losing process telemetry while authentication events keep flowing is invisible. Comparing per-channel volume to that source's own baseline, with an absolute floor for normally sparse sources, catches partial loss — which is the more common and more dangerous failure.

saying these in an interview costs you the question

  • Counts installed agents and calls the result coverage
  • Treats no alerts from a host as a clean host
  • Assumes a green agent status proves events are indexed
  • Uses one global last-seen rather than per-source last-seen
  • Deletes silent hosts from the denominator to raise the percentage

context

open as a page

Why normalise vendor logs onto a shared schema like OCSF or ECS?

level: juniorimportance: must knowfreq 62%

basics

~10 s

A shared schema gives every product one field name for the same idea, so a single detection, search or pivot works across all feeds instead of being rewritten once per vendor dialect.

open as a page

Windows Security Event ID 1102 fired on a compromised workstation — does that mean the logs for that period are gone?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Event ID 1102 records that the Windows Security log was cleared, and is written after the clear itself. Anything already forwarded to the collector survives there, so the cleared window is usually still readable centrally.

open as a page

Your collector shows a syslog source as healthy while no security events arrive — how is that possible?

level: middleimportance: should knowfreq 56%

basics

~20 s

Health and security data travel as separate channels. The agent's heartbeat keeps flowing while the security channel stops: the audit source is off, the tailed file rotated, or a parser drops the records. Health proves the agent lives, not delivery.

open as a page

A CEF mapping caps process command lines at 1023 characters — what does that destroy?

level: middleimportance: should knowfreq 48%

basics

~20 s

Everything past the cap is gone with no error, so executions that differ only in their tail arrive as byte-identical strings. They dedupe into one repetitive-looking event, and the bytes that would have distinguished them never left the mapper.

open as a page

An authenticated vulnerability scan floods a host's Security channel — how can that erase the intrusion window before the forwarder ships it?

level: middleimportance: should knowfreq 48%

basics

~10 s

A Windows event channel is a fixed-size file that overwrites its oldest records first. If the write rate beats the forwarder's drain rate, records are destroyed before being shipped, silently and with no marker.

open as a page

A site's syslog volume fell to zero eleven days ago and nobody noticed — tampering or a benign change?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Treat it as unexplained rather than calm. Pin the cut to the minute, scope what stopped, demand a change record whose timestamp matches, corroborate from a surface off that path. Those eleven days are a blind period.

open as a page

After a merger, two proxy feeds share one normalised action field with different meanings — what breaks?

level: seniorimportance: should knowfreq 38%

basics

~10 s

Every rule, dashboard and verdict reading that field silently averages two vocabularies. One feed's deny means the request was blocked; the inherited feed's deny means a monitor-mode policy matched and the traffic still completed.

open as a page

A branch office's endpoint events arrive hours after their host timestamps — what does that break during an intrusion investigation?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Late arrival breaks scheduled detections that scan a rolling window of event time, skews any timeline that mixes fast and slow sources, and makes the live picture of a host stale while you are deciding what to do about it.

open as a page

An auditor asks you to defend '94% endpoint log coverage' — which denominator and evidence do you present?

level: principalimportance: should knowfreq 40%

basics

~20 s

Name the denominator first — which authoritative asset list the figure is over — and prove the numerator from events indexed in a stated window, not agents enrolled. Present the missing six percent as a named, owned list.

open as a page

Investigators blame a 40-minute gap in a host's forwarded events on the intruder, but your deployment job stopped the agent — how do you settle it?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

A change record is a claim, not evidence. Corroborate it with the agent service's own stop and start records, the same gap on other hosts in that deployment ring, and the host's local channel, which kept recording while shipping stopped.

open as a page

A shared-schema field rename would break forty live detections — how do you ship it?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Publish both names for a fixed, announced window, migrate the consumers you can enumerate, then remove the old name on a stated date. A silent cut breaks content nobody warned; a permanent alias quietly becomes the schema.

open as a page