A branch office's endpoint events arrive hours after their host timestamps — what does that break during an intrusion investigation?
answer
- a record carries two clocks
- which one does your rule window use
- late data lands in a closed window
- measure ingest minus event, watch p95
basics
~20 sLate arrival breaks scheduled detections that scan a rolling window of event time, skews any timeline that mixes fast and slow sources, and makes the live picture of a host stale while you are deciding what to do about it.
solid answer
~50 sEvery record has two clocks: the event time written by the host and the ingest time when the platform received it. A saturated site link or a backlogged forwarder pushes them hours apart. Three things break. A rule that runs every five minutes over the previous five minutes of event time never sees the late records at all — their window was already evaluated and closed, and the rule's silence is indistinguishable from a quiet estate. A timeline that mixes a timely source, such as identity sign-in logs, with a lagging one can imply an ordering that never happened. And during live response the host keeps producing evidence that has not landed yet, so my picture at the moment I decide to isolate is stale. I measure lag as ingest minus event time per source, track the p95, and key scheduled rules to ingest time with a lookback that covers it.
code
text · 8 lineshost branch-ws-041
channel Security
event id 4624 (an account was successfully logged on)
logon type 3 (network)
event timestamp 2026-03-10T02:14:07Z written on the host
ingest timestamp 2026-03-10T06:01:52Z received by the collector
...
lag = 3h 47m 45sgo deeper
Know that a record carries both the time the host wrote it and the time the platform received it, and that the two can differ. Being able to say a search over recent time may not show everything that has happened is enough here.
Explain where lag comes from — link saturation, agent backoff, batching, collector backlog — and what it does to a rolling event-time rule window. Be able to write down the lag calculation and say why the tail matters more than the mean.
Show the investigative consequences: a detection that misses silently, a timeline whose ordering you cannot defend, a containment decision made on a partial picture. Say how you measure lag per source and what you change in rule windows because of it.
Own the trade-off between alert timeliness and completeness across sites, what lag a detection programme is engineered to tolerate, and whether a branch link that cannot ship telemetry in time is a network investment or a local-buffering one.
## Two clocks on every record Every record a collection pipeline stores carries at least two timestamps, and confusing them is one of the most expensive mistakes in a SOC. - **Event time** — written on the host, when the thing happened. - **Ingest time** — stamped by the collector or index, when the record arrived. In a healthy pipeline these differ by seconds. Across a saturated branch link, a backlogged agent doing exponential backoff, a batching schedule, or a collector working through a queue, they can differ by hours, and the difference varies by source and by time of day. ## What breaks, in order of how badly it bites **1. Scheduled detections silently miss.** A rule that runs every five minutes over the previous five minutes *of event time* evaluates a window and closes it. Records that arrive three hours later fall into a window already scanned. They are stored and fully searchable, and the rule will never look at them again. The failure produces no error and no alert — and the absence of an alert is indistinguishable from a genuinely quiet estate, which is precisely why this defect survives for months. Two fixes. Key the rule's window to *ingest* time with a lookback long enough to cover the observed skew, which guarantees every record is evaluated exactly once but means an alert can arrive long after the behaviour. Or keep event-time windows and add a grace period, re-evaluating older windows once late data has settled. Either way, the choice is explicit rather than accidental. **2. Timelines built from mixed sources imply false ordering.** Reconstruction pulls from sources with different lag: identity provider sign-ins land in seconds, a branch endpoint's process telemetry lands hours later. If you build the timeline from what is currently *present* rather than from event time across sources you know to be complete, you can conclude that the cloud sign-in preceded the local process launch when in fact it followed it — and an ordering claim is exactly the kind of thing a responder is asked to defend later. Order events by event time, and record for each source how much of the window has actually arrived. **3. Live response works on a stale picture.** You decide to isolate a host at 10:00 on evidence that runs to 06:00 because that is all that has landed. The forty minutes before your decision are still in flight. They will arrive, and they may change the scope. Note also the direction of the claim: pulling a host stops the damage and simultaneously freezes what you can learn — but with a lagging forwarder, isolation can also strand records that were still queued locally and never shipped. ## Lag is not clock skew They look the same on a timeline and have different fixes. **Lag** means the host wrote a correct timestamp and the record travelled slowly; it is measurable as ingest minus event time and correctable by waiting or re-scanning. **Clock skew** means the host wrote the *wrong* time; no amount of waiting fixes it, and the record will be misplaced in every timeline forever. Distinguish them by comparing a host's event times against the collector's receipt times over a long period: a steady offset with tight spread suggests skew, a variable and load-correlated offset suggests lag. Time synchronisation is the fix for one; capacity or buffering is the fix for the other. ## Measuring it Compute `ingest_time - event_time` per record, then aggregate per source, per site and per host class. Track the median and the p95, not an average — lag distributions are long-tailed, and the tail is the part that breaks rules. Alert when a source's p95 crosses the lookback your detections are configured for, because that is the moment the detections start missing. Watch the forwarder's queue depth alongside it; a rising queue predicts the lag before the lag shows up in the index. ## What you tell the investigation When you report on a window covered by a lagging source, say what fraction of it had arrived at the time you drew your conclusions, and re-run the key searches after the backlog drains. `We saw nothing from that site in the hour` is not a finding until you can say the hour had actually arrived.
- Why not simply key every scheduled rule to ingest time and be done with it?It guarantees each record is evaluated once, which is the main thing, but it decouples the alert from when the behaviour happened: an alert can fire at 09:00 for something that occurred at 05:00, and analysts read the timestamp on the alert as now. It also means a burst of backlog replayed at once can produce a flood of stale alerts. Use it, but label the alert with both times.
- How do you separate a lagging source from a host whose clock is wrong?Compare event times against receipt times for that host over days. A steady offset with tight spread points to clock skew, which no waiting fixes and which misplaces the host in every timeline. A variable offset that tracks link utilisation or queue depth points to lag, which drains. The remedies differ: time synchronisation versus capacity or local buffering.
- During triage, how do you know whether a window has finished arriving?Re-run the same search later and compare counts, and check the source's current lag distribution against the age of the window you care about. If the p95 lag exceeds the window's age, the window is still filling and any conclusion about absence is premature.
saying these in an interview costs you the question
- Treats event time and ingest time as interchangeable
- Says late records are re-evaluated by a scheduled rule
- Reads an empty search result as proof of quiet
- Averages lag instead of watching the tail
- Confuses a slow forwarder with a wrong host clock