A service account's baseline shows Sunday-night logons only, and last night it authenticated to eighty hosts at 03:00 — how do you reach a verdict?
answer
- four things deviated, not one
- an accepted credential, not a person
- go looking for the disconfirming record
- type 10 needs a named human
- benign true positive, then re-cut
basics
~20 sDecompose the deviation — hour, host count, logon type, source address — then hunt the disconfirming record: change tickets, the patch job's own logs, the source hosts. The owner's explanation is a hypothesis to test, not a close.
solid answer
~50 sStart by decomposing the deviation instead of reacting to it: the hour moved, the host count roughly doubled, the source addresses are the two management servers, and a handful of hosts show logon type 10 (RemoteInteractive) where the baseline had only type 5 and type 3. Each part gets its own explanation. Then go looking for evidence that would close it: a change record moving the patch window, the deployment tool's own job log covering the same hosts and minutes, and the identity of whoever opened those interactive sessions. Ask the platform engineer — but treat what they say as a hypothesis you verify against telemetry, because an owner's confidence is not evidence. If it all lines up, this is a benign true positive: real activity, correctly surfaced, legitimately authorised. The hunt's actual output is then the re-cut baseline, with the new window, the new host count, and a written note of what would invalidate it next.
code
text · 10 linesbaseline for svc-patchmgmt (learned over 12 weeks, cut 2 weeks ago)
4624 LogonType 5 (Service) Computer in {MGMT01, MGMT02}, daily
4624 LogonType 3 (Network) 38-42 distinct Computer, Sundays 22:00-23:00
4624 LogonType 10 (RemoteInteractive) never observed
last night (Wednesday)
03:10-03:48 4624 LogonType 3 TargetUserName=svc-patchmgmt 80 distinct Computer IpAddress=10.10.1.11, 10.10.1.12
03:12 4624 LogonType 10 TargetUserName=svc-patchmgmt Computer=WKS-3311 IpAddress=10.44.9.51
03:19 4624 LogonType 10 TargetUserName=svc-patchmgmt Computer=WKS-3312 IpAddress=10.44.9.51
...go deeper
Be ready to say what a successful logon record does and does not prove, and that the logon type distinguishes a machine-driven network logon from an actual remote-desktop session by a person.
Explain how you decompose a deviation into its axes — hour, population, logon type, source address — and which corroborating source explains each: change records for the schedule, the deployment tool's job log for the hosts, source addresses for the origin.
Demonstrate the discipline of seeking disconfirming evidence and refusing to let one plausible story cover every event. Show that you can name the outcome correctly as a benign true positive and that the deliverable is a re-cut, annotated baseline.
Own the standing arrangement this incident exposes: which teams must notify security before a maintenance window or scope changes, what the SOC commits to in return, and how much analyst time you are prepared to spend re-deriving change you were never told about.
## Decompose before you react The worst move is to treat "the account broke its baseline" as one fact. It is at least four, and each has a different explanation and a different failure mode: 1. **The hour moved** — the baseline said Sunday 22:00–23:00; this was Wednesday 03:10–03:48. 2. **The population roughly doubled** — from around forty hosts to eighty. 3. **The logon type set changed** — the baseline held type 5 (Service) on two management servers and type 3 (Network) on the fleet; last night added type 10 (RemoteInteractive) on three hosts. 4. **The source addresses** — most events sourced from the two management servers; the interactive ones from a different address entirely. A single explanation that covers points 1 and 2 does not automatically cover points 3 and 4, and the commonest analytical error here is letting a plausible story close the whole set. ## What each record actually proves A Windows Security **4624** is a *successful logon*: it proves a credential for that account was accepted by that host. It does not prove a person was present, and it does not prove a process ran. The logon type is where the meaning lives — type **3** is a network logon, which is what remote execution over SMB or WinRM produces and what a fan-out patch run looks like; type **5** is a service logon; type **2** is an interactive console logon and type **10** is RemoteInteractive, the signature of an actual remote-desktop session. Eighty type-3 logons are consistent with a machine doing a job. Three type-10 sessions mean somebody had a desktop, and that somebody needs a name. ## Hunt for the disconfirming evidence A verdict is built by trying to break the benign explanation, not by collecting support for it: - **Change records.** Was there an approved change moving the patch window? Does its date, time and scope match what you observed, or is it merely nearby? - **The tool's own logs.** The deployment or patch platform records which hosts it targeted, at what times, with what result. If its job covers the same eighty hosts within the same thirty-eight minutes, the type-3 logons are explained by a system, not by an assertion. - **The source.** Are all eighty sourced from the two management servers you already baselined? One event sourced from a workstation, a jump host, or an address outside the management range is worth more than the other seventy-nine combined. - **The edges of the window.** Does the activity stop when the job log stops? Activity continuing after the job ends is the classic tell that something rode in alongside it. - **The scope.** Are all eighty hosts in the patch tool's scope, or did the account touch anything it has no business touching — a domain controller, a finance server, a jump host? - **The account itself.** Was its password or key changed recently, is it a member of anything new, and did anything else authenticate as it from elsewhere in the same period? ## Talking to the owner The platform engineer will very likely explain it in one sentence: the maintenance window moved to Wednesday, the fleet grew when a business unit was onboarded, and they logged into three machines by remote desktop to watch the first run of the new schedule. That explanation is exactly what you want — and it is a hypothesis, not a verdict. It is checkable in minutes: the change record, the job log, and their own account's authentication trail during the same minutes. Ask for it in a way that gets the change history rather than a defence, because the person who can explain what your baseline missed is the same person you need to hear from *before* the next window moves. And be honest about what an explanation from an owner is worth on its own. An adversary with the account's credential also produces confident owners, because the owner is describing what they *intended*, not what the estate recorded. ## The verdict and its name If the change record, the job log, the source addresses and the interactive sessions all reconcile, the correct outcome is a **benign true positive**: the activity was real, it genuinely deviated, and it was authorised. That is not a failed hunt and not a false positive — nothing was wrong with the observation, only with the baseline's assumptions. ## The output is a re-cut baseline A hunt lead that closes benign still has to produce something, or the same eighty hosts cost you the same afternoon next month. Re-cut the baseline with the new window and new host count, and record what you learned about its assumptions: the patch schedule is now Wednesday 03:00, scope is eighty hosts and growing, and interactive logons by this account are still *not* baselined — they remain a finding, and a named engineer is the only thing that explains one. Write down what invalidates the new baseline: the window moving again, another onboarding wave, or the tool changing its execution method. Then close the loop with the platform team so the next change arrives as a notification rather than as a 03:00 anomaly. ## What would have flipped this Name these explicitly, because they are what separates a verdict from a rationalisation: a type-3 logon sourced from outside the management pair; a type-10 session no engineer will claim; activity continuing after the job log ends; hosts outside the patch scope; or the same account authenticating from a second source in parallel. Any one of those, and the benign explanation covers most of the events while leaving the interesting ones uncovered — which is exactly how a real intrusion hides inside a legitimate change.
- The platform engineer says the patch window moved and the fleet grew. What do you check before accepting that?The change record's date, time and scope; the patch tool's own job log covering the same eighty hosts and the same thirty-eight minutes; that every network logon sourced from the two management servers; and that activity stopped when the job did. An explanation that matches the timing but not the source addresses is not an explanation.
- Why do the three RemoteInteractive sessions matter more than the eighty network logons?A type-3 network logon is what a fan-out job produces and is fully explained by a job log. Type 10 is a remote desktop session — a human had a desktop as that account. The patch schedule explains machine activity; it does not explain a person, so that part needs a named engineer and their own authentication trail.
- You close it as a benign true positive. What have you actually shipped?A re-cut baseline with the new window, the new host count and its assumptions written down; a statement of what would invalidate it next; the note that interactive logons by this account remain unbaselined and still a finding; and an agreement with the platform team to be told when the window changes again.
- How would an intrusion look different in exactly this data?The legitimate change would explain most of the events and leave a residue: a logon sourced from a workstation rather than the management pair, hosts outside the patch scope, activity continuing after the job log ends, or a parallel authentication from a second address. Riding inside an authorised change is a deliberate technique, so the residue is what you hunt.
saying these in an interview costs you the question
- Accepts the owner's explanation without checking any telemetry
- Lets one explanation close all eighty events at once
- Reads a successful logon as proof a person was present
- Calls it a false positive when the activity genuinely happened
- Closes the lead without re-cutting or annotating the baseline
- Isolates the management servers before establishing what the job was