What is the difference between a log record's event time and its ingest time in a SIEM?
answer
- one record, two clocks
- whose clock wrote which field
- when it happened vs when we heard
- collector stamps arrival, source stamps event
- filter on arrival to see what is new
basics
~20 sEvent time is when the source says the activity happened; ingest time is when your collector received the record. Event time answers when it happened, ingest time answers the earliest moment the SOC could have known about it.
solid answer
~50 sEvery record carries at least two clocks. The event timestamp comes from the source - a host, an appliance, a SaaS service - and claims when the activity occurred. The ingest or receipt timestamp is written by your own collector when the record arrived, so it marks the earliest moment anyone in the SOC could have acted on it. The two drift apart for ordinary reasons (an agent batches, an offline laptop buffers, a forwarder retries, a queue backs up) and for bad reasons (the source's clock has drifted, or it logs local time with no offset). Which one you search matters: filtering on event time asks `what happened in that hour`, filtering on ingest time asks `what has reached us since I last looked`. On a source whose own clock you cannot trust, the collector-stamped field may be the only consistent time you have.
go deeper
Be ready to name both timestamps, say which one the source wrote and which one your platform wrote, and state which field your search filtered on. That last part is what the question is really testing.
Explain why the two drift apart - batching, offline buffering, polled APIs, retries, backlog - and what your parser does when it cannot read the source's own timestamp out of a record.
Show that you check a source's normal arrival lag before calling a gap a failure, and that you choose event or ingest time deliberately depending on whether you are asking what happened or what has newly arrived.
Own the standard rather than the query: every source logs UTC with an explicit offset, the platform keeps both the raw source time and the collector-stamped time, and no pipeline is allowed to normalise one of them away.
## Two timestamps, two different questions A log record in a SIEM is not one moment, it is at least two. The **event time** is whatever the source wrote into the record: the endpoint agent, the firewall, the identity provider, the application. It is a *claim* about when the activity happened, made by a clock you do not control. The **ingest time** (also called receipt, index or collection time depending on the platform) is written by your own collector when the record arrived at your infrastructure. It is a fact about your pipeline, produced by one clock you do run. They answer different questions: - *When did it happen?* - event time, and only as far as the source's clock deserves trust. - *When could we first have known?* - ingest time. Nothing the SOC does could have started earlier than that, no matter how good the detection is. ## Why the gap opens A gap between the two is normal, not automatically suspicious. Common causes: - **Batching.** Many agents ship on an interval rather than per event; a five-minute batch puts up to five minutes between the two fields by design. - **Offline buffering.** An endpoint agent on a laptop that is off the network keeps writing locally and flushes the backlog when it reconnects. That backlog can be hours or days old. - **Pull-based sources.** SaaS and cloud audit trails are usually polled through an API, and many of them publish records some minutes after the action, so the earliest a poller can see them is already late. - **Transport and backpressure.** Forwarder retries, queue depth, an indexer under load. - **Bad clocks.** If the source's clock is wrong, the gap is partly fiction: the event time is simply not where the event was. Because of that last case, a *negative* gap is a red flag on its own. A record whose event time is later than the moment it arrived means the source's clock is running ahead of yours - the event cannot have been recorded after it reached you. ## What each timestamp is good for Searching by **event time** is what investigation needs. When you ask which processes ran on a host between 02:00 and 03:00, you want the hour the activity claims, drawn from every source that observed it. Searching by **ingest time** is what *operating* needs. Two everyday uses: 1. **Seeing what is new.** Re-running a sweep by ingest time returns everything that arrived since the last run, whatever hour those records claim - including a batch of six-hour-old events that landed after your previous query. Re-running the same event-time query gives you no way to tell what is newly present. 2. **Reasoning about latency.** Everything before ingest is collection and transport; everything after it is detection and response. Splitting the two keeps a rule from being blamed for an agent that was offline all afternoon. ## When the parser has to guess If a record arrives without a timestamp the parser can read - an unusual format, a truncated header, a field the sourcetype does not know about - most platforms fall back to the arrival time and stamp the record with it. The event time then silently *becomes* the ingest time. For a source that streams in real time that is harmless. For a source that ships hourly it is badly wrong: every event in the batch collapses onto one instant, and the internal ordering inside that hour is destroyed. Before trusting an ordering, it is worth checking whether the event time you are reading was parsed from the record or invented by your own pipeline. ## Reading a lag honestly A six-hour lag is only meaningful against that source's own profile. Some useful habits: - Know each source's normal arrival lag (a median and a tail). A source that usually lands in ten seconds and is suddenly six hours behind is a collection failure worth paging on. A source that has always been a daily file pull is not. - Alert on the *absence* of arrivals per source, not just on the lag. A silent source and a quiet estate look identical from the SIEM if nobody is watching arrival counts. - Keep both fields. Normalising one away - storing only the source's claim, or only your own arrival stamp - throws away exactly the field you need when the other one turns out to be wrong. ## The claim each field supports Be precise about what you can say. The event time supports *the source recorded this activity at that moment*, which is weaker than *this activity happened at that moment*, because the clock is the source's. The ingest time supports *we held this record from this moment*, which is strong because it is your own clock, but it says nothing at all about when the activity occurred. Interviewers push on exactly this distinction: a candidate who treats the source's timestamp as ground truth, or who cannot say which of the two their query filtered on, has not yet worked a case where the two disagreed.
- What does a SIEM usually do when it cannot parse a timestamp out of an arriving record?It falls back to the arrival time and stamps the record with that, so the event time silently becomes the ingest time. Harmless for a source that streams in real time; badly wrong for one that batches hourly, because every event in the batch collapses onto a single instant and the ordering inside it is gone. Check what the parser did before trusting an order.
- Is a large gap between event time and ingest time always a problem?No. Batching agents, laptops flushing after being offline, and polled SaaS audit APIs all produce hours of lag by design. What matters is whether the lag matches that source's normal profile: a source that usually arrives in ten seconds and is suddenly six hours behind is a collection failure, while a source that has always been a daily pull is just a daily source.
- What does it mean if a record's event time is later than its ingest time?That the source's clock is running ahead of yours, because a record cannot describe an event that happens after it already arrived. It is one of the cheapest skew detectors you have: chart per-source event-time minus ingest-time, and any source sitting persistently on the negative side has a clock problem you should measure before using its ordering.
A postmark and the date written inside the letter. The writer's date can be wrong, missing or in another timezone; the postmark is your own record of when it reached you.
saying these in an interview costs you the question
- Treats the source's own timestamp as ground truth
- Assumes a record exists in the SIEM the moment it happens
- Cannot say which timestamp their query actually filtered on
- Thinks ingest time is meaningless because it is not when it happened
- Normalises away one of the two timestamps at ingest