Why can a UEBA risk score stay flat while a stolen credential is actively being used?
answer
- it scores difference, not intent
- in-envelope use produces no deviation
- activity during learning becomes normal
- silence has many indistinguishable causes
- queue order is not coverage
basics
~20 sBecause a UEBA score measures deviation from a learned normal, not maliciousness. A credential used from the usual country, at usual hours, against the systems that account always touches produces no deviation, and a flat score is evidence of nothing.
solid answer
~50 sThe analytics are built to notice difference, so an adversary who stays inside the account's envelope is invisible by construction: same geography, same working hours, same file shares, volumes within the user's own range. Worse, if the intrusion overlaps the window in which the baseline is learned or refreshed, the adversary's own activity is absorbed into normal and raises the bar for anyone later. The trap is inferring anything from the flat score. Absence of output from a behavioural analytic is indistinguishable from a quiet estate, a broken log feed, or an account excluded from scoring, so I cannot report a false-negative rate from the score distribution. Operationally I treat the ranked queue as a work-prioritisation aid, not coverage: high-value behaviours get their own targeted detections, and I hunt the underlying records against an explicit hypothesis rather than waiting for a number to rise.
go deeper
Remember that these analytics fire on difference from a learned normal, so a stolen credential used the way its owner uses it produces no score, and a low score never means an account was checked.
Explain why in-envelope activity generates no contributing reason, and what baseline poisoning is - activity present during the learning window becoming part of normal for the account.
Show how you operate around the blind spot: monitor telemetry arrival and exclusion lists, write targeted detections for behaviours you fear, and hunt the underlying records against an explicit hypothesis instead of waiting on the queue.
Be able to argue why a false-negative rate cannot be reported from behavioural scoring, what you would substitute as evidence that the capability works, and where analyst capacity truncating a ranked queue is an accepted risk rather than coverage.
## The design assumption UEBA scores deviation. Every contributing reason is a comparison - against the account's own history or against a peer cohort - and points are added when the comparison finds something different. The implicit assumption is that intrusion looks *unlike* the user. That assumption holds for a noisy adversary and fails completely for a careful one. An adversary in possession of valid credentials can choose to operate inside the envelope. Connect from an address in the country the user works in. Work during the user's working hours. Touch only the file shares that account touches every week. Pull data in volumes inside the account's own normal range, spread over days. None of that produces a deviation, so no reason fires, so the score never rises, so the account never enters a queue that is worked top-down. The intrusion is not missed because the tool broke; it is missed because there was nothing for it to be different from. ## Baseline poisoning There is a sharper version. Baselines are learned over a window, and refreshed. If the adversary is present during that window, their behaviour becomes part of what "normal" means for the account. Two consequences follow: the activity that established the foothold never scores, and the widened envelope keeps scoring quiet afterwards - including for the legitimate user, whose own unusual behaviour is now less exceptional. The same effect appears for a newly created account or a newly created cohort. A learned baseline needs history, and during the period before it has any, the comparison is either unavailable or weak - which is exactly the period after an onboarding or a reorganisation when accounts change hands. ## The direction of the claim This is where interviewers separate candidates. Getting these right matters: - A **high score** proves that analytics fired over records. It does not prove malicious activity. - A **flat score** proves that no analytic found a deviation. It does not prove the account is clean. - A **successful authentication record** proves that a credential was accepted, not that its owner was present. And there is a structural problem behind the flat score: a behavioural analytic's non-output has no distinguishable causes. Nothing bad happening, an ingestion feed silently stopping, the account being on an exclusion list, and an adversary operating inside the envelope all render identically as "no reasons fired". That is why you cannot compute a false-negative rate from the score distribution the way you can compute a false-positive rate from cases you worked and dismissed. There is nothing to count. ## What you do about it instead **Verify the inputs, not the silence.** Score health depends on telemetry arriving. Monitor the feeds themselves - VPN session records, logon history, file-access audit - for volume and freshness per source, so that a silent collector failure presents as a data alarm rather than as a calm queue. Also audit the exclusion list: accounts exempted years ago to stop noise are accounts that can no longer score at all. **Do not treat the ranked queue as coverage.** A queue ordered by score is a way of spending finite analyst hours well. It says nothing about which adversary behaviours the estate can see. Behaviours you specifically fear deserve their own targeted detections written against the records that betray them, and their presence is proven by executing the behaviour deliberately - a purple-team exercise - not by observing that no score rose. **Hunt with a hypothesis.** The whole point of hunting is to look for something no alert will raise. "If an adversary is using this credential in-envelope, what would still be different?" is a productive question even when the score is flat: two concurrent sessions for one account from different networks; access to a share the user *has* rights to but has never in fact opened; the user's own workstation idle while their account is authenticating elsewhere; a credential used against a system class the human never uses interactively. **Accept the capacity truncation.** If the queue is worked top-down and staffing reaches only down to some score, then everything below that line is unexamined by policy. That is a defensible resourcing decision, but say it out loud in those terms - "we examine the top N per shift" - rather than letting a low score imply that something was checked. ## The one-line answer UEBA finds users behaving unlike themselves. An adversary who behaves exactly like the user is not an anomaly, and the silence that results is the least informative signal in the SOC.
- The intruder was active during the two weeks the baseline was being learned. What does that do?Their behaviour is absorbed into normal. The foothold activity itself never scores, and the widened envelope keeps the account quiet afterwards, because later access falls inside a baseline the adversary helped define. It is the reason a baseline is not a trust anchor and why re-baselining after a confirmed intrusion is part of eradication.
- Your manager asks for the UEBA false-negative rate this quarter. What do you say?That it is not computable from the tool's output. A non-firing analytic produces nothing to count, and a quiet estate, a dead log feed, an excluded account and an in-envelope adversary all look identical. What can be measured is whether specific behaviours are detected when deliberately executed, and whether every expected telemetry source is still arriving.
- What would you hunt for when the score is flat but you suspect credential misuse?Things that are anomalous in structure rather than in volume: concurrent sessions for one account from two networks, access to shares the account has rights to but has never opened, authentication activity while the user's own device is idle, or interactive use of a system class the human never touches. All are queries over the same records the scoring runs on.
saying these in an interview costs you the question
- Reads a flat risk score as evidence the account is clean
- Treats the ranked queue as a measure of detection coverage
- Forgets that activity during the learning window becomes the baseline
- Quotes a false-negative rate derived from score distribution
- Ignores that a stopped log feed also produces a calm queue