skip to content

Investigators blame a 40-minute gap in a host's forwarded events on the intruder, but your deployment job stopped the agent — how do you settle it?

level: seniorimportance: nice to knowfreq 32%

answer

  1. a ticket is a claim, not evidence
  2. shipping stopped, recording did not
  3. look for the same gap across the ring
  4. service state changes live in another channel
  5. explained is not the same as clean

basics

~20 s

A change record is a claim, not evidence. Corroborate it with the agent service's own stop and start records, the same gap on other hosts in that deployment ring, and the host's local channel, which kept recording while shipping stopped.

solid answer

~50 s

I would not expect anyone to take my ticket as proof, so I offer independent corroboration and let the responders collect it. Three things settle it. The host's System log holds Service Control Manager state-change records showing the collection agent's service stopping and restarting, timestamps bracketing the gap, written to a channel that never stopped. The deployment ring shows the same minutes missing on the other hosts in that batch; a gap on this one host alone would point the other way. And the deployment system's own job and package-install records show the run independently of my ticket. The crucial technical point is that the forwarder stopping is a shipping gap, not a recording gap: the Event Log service kept writing, so those 40 minutes may still exist locally and can be collected — unless the channel has wrapped since.

go deeper

for a junior

Know that a collection agent stopping halts shipping, not recording, so the missing events may still be on the host. Being able to point at a change record and the agent's service stop and start records is the expected contribution.

for a middle

Explain how each corroborating source is checked: which channel holds service state changes, why the deployment ring is the strongest independent signal, and why the local channel may or may not still hold the window depending on rotation.

for a senior

Demonstrate the conclusion discipline — an explained gap is not a clean window — and show that you hand over access rather than exports, and that you can say what evidence would falsify your own explanation.

for a principal

Own the systemic version: change automation that blinds telemetry mid-window is a design defect. Argue for deployment schedules the SOC knows about, agents that spool across their own restart, and a standing rule that unexplained gaps are investigated rather than assumed benign.

## The situation, honestly stated An intrusion is confirmed on a host you manage. The responders' timeline has a 40-minute hole in the forwarded events, inside the window they care about, and the natural reading of a hole in telemetry during an intrusion is that someone made it. You believe your software-deployment job restarted the collection agent inside a Tuesday change window. You have a ticket. A ticket is an assertion by the person being asked to explain themselves, and no competent responder should accept it as evidence. Your job is to hand over things that can be checked without trusting you. ## Recording and shipping are different gaps This is the distinction that decides the case, and most people conflate the two. | what stopped | still recorded locally? | reaches the collector? | |---|---|---| | the forwarder or collection agent | yes — the Event Log service is a separate service | no, until it restarts and drains | | the Event Log service, or auditing itself | no | nothing to send | A stopped agent produces a **shipping** gap. The host went on writing to its channels the whole time. So the first and best move is not an argument at all: collect those 40 minutes from the host's local channel, where they very likely still are. They are only gone if the channel has since wrapped, which depends on the host's event rate and channel size. If they are there, the gap closes and the timeline is whole, and nobody has to believe anyone. ## Independent corroboration for the change claim If the local copy is gone, the claim still has to stand on evidence other than the ticket: - **The agent service's own state changes.** The Service Control Manager writes service stop and start records into the System log (Event ID 7036 reports a service entering the running or stopped state). Those records are in a different channel from the one that went quiet, and they bracket the gap with timestamps. A stop at the start of the gap and a start at its end is strong corroboration. - **The deployment ring.** If the same package went to two hundred hosts in a batch, the same gap should appear on many of them at overlapping minutes. Fleet-wide correlation is very hard to fake and easy for a responder to check independently. Conversely, a gap on this one host alone, at those exact minutes, is a finding against the benign explanation. - **The deployment system's own records.** Job history and package-install logs on the deployment server, pulled by the responder rather than exported by you, are independent of the ticket. - **The agent's own install or upgrade log** on the host, which typically records the service restart it performed. Offer all of it, and offer access rather than exports. The person under suspicion should not be the sole custodian of their own exculpatory evidence — not because anyone assumes bad faith, but because evidence someone collected about themselves is worth less later, when the claim has to be defended to somebody who was not in the room. ## The conclusion discipline: two claims, not one Even when the change explanation is fully corroborated, keep two statements separate: 1. **The gap has a benign cause.** Corroborated. 2. **Nothing hostile happened during those minutes.** Not established, and not establishable from a change record. An explained gap is still an unobserved gap. Deployment windows are published, predictable and often visible from inside the estate, so an intruder acting during one is not exotic; and an intruder with local rights can stop a service themselves, which would look much the same on this one host but would not reproduce across the ring. Conflating the two claims is how a real intrusion window gets written off as a change artefact. ## How this lands in the report If the local channel yields the missing records, say so and close the gap. If it does not, the report names those 40 minutes as time for which no copy exists anywhere, states the corroborated benign cause of the gap, and explicitly does not claim the window was clean. Scoping treats the interval as unknown, and the watch for re-entry after eradication is written knowing there is a slice of the intrusion nobody ever observed. That is a stronger, more defensible report than one that quietly implies full coverage — and it is the version that survives being read back to you months later. ## The relationship, not just the evidence One last practical note from this side of the table: cooperating fast is worth more than being right. Volunteering the ring correlation and the deployment logs before being asked converts you from a suspect into a source, and it is usually the fastest route to a timeline everyone believes.

  • What would make you doubt the deployment explanation?
    A gap on this host alone while the rest of the ring shipped normally through the same minutes; service stop and start records that do not bracket the gap, or are missing entirely; a stop time that precedes the deployment job's own start; or a gap that is longer than the package's restart takes anywhere else. Any of those turns the benign reading into a finding.
  • If the local channel still holds the 40 minutes, is the gap resolved?
    The evidence gap is, which is the important part — the records exist and can be collected and compared against what the collector holds for adjacent minutes. The reason for the interruption should still be confirmed, because a forwarder stopping for reasons nobody can explain is a durable weakness whether or not it mattered this time.
  • Should you collect the corroborating evidence yourself and send it over?
    Prefer giving the responders access and letting them pull it. Evidence gathered by the person whose actions are in question carries less weight later, and if the claim ever has to be defended outside the incident, provenance matters as much as content. Volunteering where to look is the useful contribution.

saying these in an interview costs you the question

  • Offers the change ticket as proof and stops there
  • Assumes a stopped forwarder means the host stopped recording
  • Treats an explained gap as a clean window
  • Ignores whether the same gap appears across the deployment ring
  • Collects and curates their own exculpatory evidence unsupervised

context