skip to content

Your collector shows a syslog source as healthy while no security events arrive — how is that possible?

level: middleimportance: should knowfreq 56%

answer

  1. health and data are different channels
  2. the agent is alive, its input is not
  3. generation, rotation, transport, parser, filter
  4. measure the message, not the messenger
  5. synthetic event that must arrive

basics

~20 s

Health and security data travel as separate channels. The agent's heartbeat keeps flowing while the security channel stops: the audit source is off, the tailed file rotated, or a parser drops the records. Health proves the agent lives, not delivery.

solid answer

~50 s

A source's health status is generated by the shipper about itself, often over a different port, pipeline and index than the security data. So every failure between a record's creation and its indexing leaves the heartbeat untouched: the audit source stopped producing (audit policy not applied, `auditd` rules flushed), the log file rotated to a new inode and the tailer kept the old handle, an upgrade moved the path, the parser began dropping records it could no longer normalise, an index-side filter or volume cap discarded them, or a firewall blocks the syslog port while the agent's check-in over 443 still succeeds. The fix is to measure the message, not the messenger: per-source, per-channel volume against that source's own hourly baseline with an absolute floor, plus a scheduled synthetic event that must be searchable within an agreed window.

go deeper

for a junior

Know that a green source-health tile describes the agent, not the data, and that the real question is whether events from that source are searchable right now.

for a middle

Be able to walk the path from record creation to indexed event and name where it dies silently: audit source off, rotated file, blocked port, failing parser, index-side filter.

for a senior

Demonstrate the check you would build — per-source per-channel volume against a comparable window with an absolute floor, plus an end-to-end synthetic event — and how you avoid it firing every weekend.

for a principal

Decide what source health means contractually and who is accountable for it when a platform team, a network team and a provider each own one stage of the path and none owns delivery.

## Two channels, one dashboard When a source-health view says a source is fine, ask what produced that claim. Almost always it is one of: - the agent's own **heartbeat** or check-in to its management plane; - a transport-level fact — an established TCP session on the collector; - a **last-seen** field computed over the health index rather than over the security data. All three can be perfectly true while the security channel is dead, because they are produced by different code, over a different path, into a different index. The heartbeat says the agent is alive. It does not say the agent has anything to send, that what it sends is parseable, or that what it sends survives to the searchable index. ## The failures that mute data without killing the agent Walk the path from creation to search, and each stage has a silent failure: 1. **Generation.** The record is never produced. Windows audit policy was not configured for the categories the detections need; `auditd` rules were flushed by a config management run; a cloud audit trail is enabled in one region and not another; verbose logging was turned down after a disk incident. The agent has nothing to read and reports itself healthy, because it is. 2. **Reading.** The file rotated to a new inode and the tailer is still holding the old handle; a package upgrade moved the path or changed the syslog facility; permissions changed so the shipper can no longer read the file it is watching. 3. **Transport.** An egress rule, a route change, or an expired certificate on the collector's TLS listener stops the security stream, while the agent's management check-in over a different destination and port still succeeds — which is why the console stays green. 4. **Normalisation.** A vendor version bump changed the record layout, the parser now fails, and failed records are dropped rather than quarantined. Volume goes to zero with no error visible to the SOC. 5. **Indexing.** An index-side filter, a routing rule, or a licence volume cap discards the events after ingest. They were received and are not searchable. Each of these produces the same symptom on the dashboard: green. ## Building a check that actually proves delivery Two constructs, and they answer different questions. **Per-source, per-channel volume with a floor.** For each source and each channel, compare the count in the last hour to that source's own baseline for the comparable window — the same hour last week is usually better than the previous hour, because it survives night, weekend and shift patterns. Add an **absolute floor** for sources that are normally sparse: a source that produces four events an hour has a normal variance that includes zero, so a purely statistical check will never fire on it. The floor is what catches it. Also evaluate per source, not in aggregate. A global events-per-minute graph is dominated by the loudest few sources; an entire class of quiet-but-important sources can die inside the noise band of one busy one and never move the total. **A synthetic source event.** On a schedule, cause the source itself to emit a uniquely tagged, benign record — a scheduled task, a `logger` line, a service account performing a harmless read — and alert when that record is not searchable in the SIEM within an agreed window. This is the only check that exercises the whole path end to end: generation, reading, transport, parsing, indexing, search. Keep the tag distinctive, document it, and make sure it is excluded from detection logic so it never becomes an alert of its own. ## Grading the source, not the agent A useful source status has three states rather than two: **delivering** (volume within expectations for every expected channel), **degraded** (one channel or a large fraction of volume missing), and **silent** (below the floor). Degraded is the state most tools omit, and it is the one that hides the longest, because the source keeps a fresh last-seen the whole time. Finally, decide who is woken and when. A single sparse source falling below its floor is a ticket. A whole class of sources — every domain controller, every host at one site, every device behind one collector — going quiet at once is a different event, because it means something structural changed on a shared path, and it is worth waking someone for at 02:00 even though nothing has alerted. ## What an interviewer is listening for That you never accept a component's self-report as proof of delivery; that you can name at least three stages where data dies silently; and that your proposed check measures the data itself, ideally end to end with a synthetic event.

  • How would you build a synthetic source check that genuinely proves delivery?
    Make the source itself emit a uniquely tagged benign record on a schedule — a scheduled task, a `logger` line, a service account doing a harmless read — and alert when that record is not searchable in the SIEM within an agreed window. It exercises generation, agent, network, parser and index in one shot, which no component's self-report does. Keep the tag distinctive, document it, and exclude it from detection logic.
  • Why can't a single global events-per-minute graph catch this?
    Aggregate volume is dominated by a handful of very loud sources. A whole class of quiet but important sources can go to zero inside the normal variance of one busy one and never move the total. Silence has to be evaluated per source and per channel against that source's own history, with an absolute floor for sources that are naturally sparse.
  • What breaks a per-source volume baseline in practice?
    Legitimate rhythm and legitimate change. Nights, weekends, holidays, shift patterns and batch windows make volume swing by design, so compare like windows — the same hour last week — rather than the previous hour. And decommissioning, migrations or a new site shift the baseline permanently, so the check needs a change feed or an ageing baseline, otherwise it either screams every Sunday or quietly relearns a real collapse as normal.

saying these in an interview costs you the question

  • Trusts an agent's self-reported status as proof of delivery
  • Watches total SIEM volume instead of per-source volume
  • Thinks an established TCP session means events are indexed
  • Assumes heartbeat and security records travel the same path
  • Has no check that catches partial loss of one channel

context