skip to content

A laptop reconnects and flushes six hours of buffered EDR events - what does that do to your hunt?

level: seniorimportance: should knowfreq 40%

answer

  1. absence of evidence, with a timestamp on it
  2. the data set for an hour keeps growing
  3. who was even reporting
  4. re-run keyed on arrival, not event time
  5. size the re-sweep from the lag tail

basics

~20 s

It invalidates the negative result. The sweep covered the hours those records describe, but the records had not arrived yet, so nothing found meant nothing had arrived. Re-run the sweep over what has been received since.

solid answer

~50 s

A hunt's negative result is only as good as what had reached the platform when it ran. An offline agent buffers locally and flushes on reconnect, so records with event times inside your window can land hours after your query. The fix is to re-sweep by arrival rather than by event time: search records received since your last run whose event times fall in the window, and size that re-sweep from the source's observed arrival-lag tail rather than a guess. Before that, check coverage - which agents were even reporting during the window, and which were dark - because absence in the SIEM and absence on the host are different claims. Then state the conclusion accordingly: no evidence as of records received up to a stated moment, not clean. On this occasion the flush is where the hypothesis pays off, because an offline laptop is exactly the host an intruder's activity tends to sit on unobserved.

go deeper

for a junior

Be ready to say that a search only sees records that have already arrived, and that an endpoint agent which was offline ships its backlog later, so a negative result can change on its own.

for a middle

Explain the mechanics of late arrival - local buffering, polled APIs, store-and-forward replay - and why re-running the identical event-time query gives you no way to tell what is newly present.

for a senior

Show the working method: an arrival watermark, a re-sweep sized from the source's measured lag tail, a coverage check of which agents were reporting, and a conclusion phrased with those limits rather than the word clean.

for a principal

Own the standard for what a hunt is allowed to assert across the organisation: coverage and arrival caveats stated on every negative finding, per-source lag measured, and dark endpoints treated as an owned gap rather than a footnote.

## The negative result has an expiry date Hunting is the business of testing a hypothesis with no alert behind it, and the answer is very often *no*. That answer carries an implicit clause most analysts never say out loud: **no matching evidence had arrived, from sources that were reporting, at the moment the query ran.** A laptop that spent the afternoon on a train, buffering endpoint telemetry locally and flushing six hours of it when it reconnected to the VPN, breaks every part of that clause at once. The records exist now. Their event times sit inside the window you already swept. Your query saw none of them, and it reported nothing found. This is different from a clock problem. The timestamps here are fine. What moved is *when the data became searchable*, and that is a property of arrival, not of the event. ## Where late arrival comes from - **Offline endpoint agents.** Laptops sleep, travel, sit on hotel networks that block the agent's egress. The agent keeps a local queue and drains it on reconnect, sometimes days later. - **Polled SaaS and cloud audit trails.** Records are fetched from an API that publishes them some minutes after the action, so even a healthy poller is structurally behind. - **Store-and-forward relays.** A site collector that lost its uplink replays the backlog when it returns. - **Backpressure.** An indexer under load, a queue draining slowly, a retry loop. None of these is a fault. They are the normal behaviour of a distributed collection estate, and they mean that *the set of records describing a given hour keeps growing for some time after that hour ends*. (How long a detection rule should wait for stragglers before it fires - the length of a correlation window and its grace period - is a detection-engineering decision made elsewhere. It is not the hunter's lever, and it does not rescue a sweep that has already run.) ## Hunting so that late data cannot fool you Four habits do most of the work. **1. Sweep by arrival on the re-run.** The first pass is naturally bounded by event time - that is the hypothesis. The second pass should ask a different question: *what has arrived since my last pass, whose event times fall in my window?* Arrival is monotonic in your own pipeline, so it partitions cleanly into already-seen and new. Note the arrival watermark you swept up to, so the next pass starts exactly where this one stopped and nothing is double-read or skipped. **2. Size the re-sweep from measured lag.** Every source has an arrival-lag distribution. If 99% of that agent fleet's records land within eight hours, then re-sweeping a day of arrivals covers the realistic tail, and anything arriving beyond that is itself a coverage failure that should be surfaced on its own rather than absorbed into hunts. A source whose lag profile nobody has measured is a source whose negative results you cannot size. **3. Check who was reporting.** Before concluding anything from an absence, look at the agent inventory for the window: which endpoints checked in, which were dark, which were reporting but had a subset of their event channels failing. A hunt across a fleet where forty laptops were offline all afternoon has not covered forty laptops. That is a coverage statement, and it belongs in the finding. **4. Phrase the conclusion so it survives.** *No evidence of this behaviour in records received up to 09:00 UTC, across the 1,240 endpoints reporting during the window; 41 endpoints were offline and will be re-swept on reconnect.* That sentence is defensible a month later. *Clean* is not, and it is the sentence that gets quoted back at you when the flush turns out to have contained the very thing you were looking for. ## When the flush is the find There is a second, sharper reason to care. An intruder's activity is disproportionately likely to be on the host that was *not* being watched in real time - and a machine off the corporate network for six hours is exactly that host. The flush is not just a nuisance that breaks your sweep; it is a delivery of the most interesting telemetry you will get that day. Hunts that re-run on arrival routinely surface their result from a backlog rather than from the original pass, which is the archetypal outcome for this work: found by a hunt, with no alert behind it. The corollary is unpleasant and worth saying in an interview: if that laptop is rebuilt before it ever reconnects, the buffered queue goes with it. Whatever the agent recorded locally and never shipped is gone, and you are left with whatever central sources happened to observe the host from the outside - identity, proxy, mail, network - none of which shows local process activity. That is a concrete reason to collect from a suspect host before rebuilding it, and a concrete reason to make reconnection, not reimaging, the first thing you ask for. ## The claim you can actually make A hunt does not prove a behaviour is absent from the estate. It proves that a specific query, over the records that had arrived from the sources that were reporting, matched nothing. Every honest hunt output names those three limits. The candidate who says *we ran the hunt and it came back clean* has not yet had a backlog land on them the following morning.

  • How far back should the re-sweep reach when a source can arrive hours late?
    Size it from that source's observed arrival lag rather than a guess. If 99% of the fleet's records land within eight hours, re-sweeping a day of arrivals covers the realistic tail, and anything later is a coverage failure that should raise its own alert instead of quietly widening every hunt. A source with no measured lag profile is one whose negative results you cannot size.
  • The laptop was rebuilt before it ever reconnected. What have you lost?
    Everything the agent buffered locally and never shipped, which for a host that was offline is the only record of what ran on it. You are left with central sources that watched the host from outside - identity, proxy, mail, network flow - and none of them shows local process activity. That is why you collect from a suspect host before rebuilding it.
  • Does a large arrival lag also weaken your detections, not just your hunts?
    Yes. A rule evaluated over arriving data will fire when the backlog lands, hours after the behaviour, and a rule scheduled over an event-time window that has already closed may never see those records at all. Arrival lag is a coverage property of the source, so it belongs in the coverage assessment and not just in a hunter's habits.

saying these in an interview costs you the question

  • Reports a hunt as clean without checking which sources were reporting
  • Assumes absence in the SIEM means absence on the host
  • Treats the first negative result as final
  • Counts offline endpoints as covered by the sweep
  • Cannot state the moment their hunt's evidence set was fixed

context