A UEBA console shows a user risk score of 92 - what does that number actually represent?
answer
- an aggregate, not a verdict
- points added per fired reason
- reasons are often correlated
- baseline is learned, not truth
- ordinal, and it decays
basics
~20 sA UEBA risk score is the sum of weighted reasons that fired on one account in a time window, measured against a learned baseline. It orders the queue; it is not a probability that the user is malicious.
solid answer
~50 sTreat it as an aggregate, not a verdict. Products like Exabeam and Securonix stitch an account's logon, VPN and file-access records into a per-user timeline, compare each event against that user's own history and against a peer group, and add points for every contributing reason that finds a deviation. A 92 tells me many reasons fired, or a few heavy ones. It does not tell me which, how independent they were, or that anything malicious happened. So the first thing I open is the contribution breakdown, because several of the reasons are typically the same VPN session counted from different angles, and correlated contributors inflate an additive score. Scores also decay, so 92 describes a window, not a standing property of the person. The underlying records are the evidence; the number is only the queue order.
code
text · 12 linesUser: j.okonkwo Risk score (24h): 92 Queue rank: 1 of 3,412
Contributing reasons (additive):
+15 First VPN session from this ASN for this user (90d)
+15 VPN session outside this user's normal hours (02:10 local)
+12 Logon to file server FS-07 - first for this user (90d)
+10 File-share read volume 8.4x this user's 30d median
+10 File-share read volume above peer-group 95th percentile
... 6 further reasons, +4 to +8 each
Peer group: "Finance - Reporting" (members: 14)
Decays to 0 over 7 days if no further reasons fire.go deeper
Be ready to say plainly that the score is a sum of weighted analytic reasons over a window, that it ranks your queue, and that your first action is opening the breakdown and pulling the underlying records.
Explain the mechanics: which baselines the reasons compare against, why additive scoring double-counts correlated reasons drawn from one event, and why scores decay rather than being cleared by a case closure.
Show triage judgment - collapsing a breakdown to distinct observations, sourcing authorisation from outside the tool, and recording an authorised deviation as a benign true positive rather than a false positive.
Own the question of what the number is allowed to be used for. An uncalibrated ordinal total can rank work; committing to it as a measure of risk, a metric, or grounds for action about a person is a decision with consequences.
## What the number is User and entity behaviour analytics (UEBA) products - Exabeam and Securonix are the two most commonly named in interviews - do not work like a detection rule that emits one alert per match. They ingest an account's records (VPN session starts and stops, Windows logon history, file-share access volumes, application and SaaS audit trails), stitch them into a per-user timeline, and run a library of analytics over that timeline. Each analytic that finds a deviation contributes a **reason** (products variously call them rules, policies or violations) carrying a **weight in points**. The user's risk score is the aggregate of those points over a rolling window. So a 92 is arithmetic, not a judgement. It answers "how many points did this account accumulate", and its only reliable operational meaning is **ordering**: in a 24x7 SOC where the UEBA queue is worked top-down, a 92 gets looked at before a 46. ## What it aggregates Typical contributing reasons on a user timeline compare an event to one of two baselines: - **The user's own history**: first VPN connection from this network in 90 days, first logon to this file server, read volume 8x this account's own median. - **A peer group**: read volume above the 95th percentile for people with the same department, title or manager. That mix matters, because the two baselines fail differently. A self-baseline flags anything new, including a legitimate change of duties. A peer baseline flags anything unusual for the cohort, and is only as good as the cohort's membership. ## What it hides Three things, and interviewers probe all of them. **Which reasons moved it.** An additive total is lossy. A 92 built from eleven small reasons is a different case from a 92 built from two heavy ones, and only the breakdown distinguishes them. A candidate who escalates on the number without opening the breakdown has skipped the actual work. **Correlation between reasons.** The reasons are not independent observations. One 02:10 VPN session from an unfamiliar network can fire "first VPN from this ASN", "session outside the user's normal hours" and "logon from a new geography" - three reasons, one observation. Because the model is additive, correlated contributors compound, and a single unusual event can look like a pattern. Collapsing the breakdown back down to the distinct observations underneath it is the analyst's first move. **Direction of the claim.** A high score proves that analytics fired over records. It does not prove the user did anything wrong, and it does not prove a human was present: a successful authentication record proves only that a credential was accepted. The score is a pointer to evidence, never the evidence. ## Scores are ordinal and they decay Two properties trip candidates up. First, the scale is ordinal within one scoring model: 92 versus 46 tells you which to look at first, but 92 is not "twice as bad" and the numbers are not comparable across tenants, across product versions, or after someone re-weights a reason. Second, scores **decay** - if no further reasons fire, the total falls back toward zero over days. That means a closed case usually leaves the score standing to expire on its own rather than being retracted, and it means a score you look at on Monday is not the score that ranked the queue on Friday night. ## What good triage looks like 1. Open the contribution breakdown before forming any hypothesis. 2. Collapse correlated reasons down to the distinct underlying events. 3. Pull the raw records for those events - the VPN session, the logon, the file-share reads - and read them directly. 4. Establish authorisation from outside the tool: a change record, a project, a ticket, the user's actual role. 5. Record the outcome honestly. If the behaviour genuinely happened, genuinely deviated, and was authorised, that is a **benign true positive**, not a false positive; calling it a false positive misrepresents what the analytics saw and misdirects whoever tunes them later. ## The sentence to have ready "The score ranks my queue; the records decide the case." Everything else in this topic - wrong peer groups, credentials that never move the score, explaining a number to the person it is about - follows from taking that distinction seriously.
- Two of the reasons are 'first VPN from this ASN' and 'VPN outside normal hours' for the same session. Does that matter?Yes. They are one observation scored twice, and an additive model treats them as two pieces of evidence. Before escalating I collapse the breakdown to distinct underlying events - here, one 02:10 VPN session - and judge that event. Correlated contributors are the main reason a single unusual login can present as a top-of-queue case.
- Is a user at 92 twice as suspicious as a user at 46?No. The scale is ordinal within one scoring model and window: it says look at 92 first. It is not calibrated, so it is not a probability, not comparable across products or tenants, and it changes meaning the moment someone re-weights a reason or adds a new one.
- You close the case as authorised activity. What happens to the score?Usually nothing immediate - it decays on its own over the window. Closing a case does not retract points. If the same benign behaviour will recur, the fix is upstream: change the reason's weight, exclude the activity, or repair the peer group. Otherwise the same user tops the queue again next month.
It is a leaderboard position, not a diagnosis. Knowing a runner is first tells you who to watch; it tells you nothing about how they ran.
saying these in an interview costs you the question
- Reads the score as a probability that the account is compromised
- Escalates on the number without opening the contribution breakdown
- Assumes eleven contributing reasons mean eleven independent observations
- Treats the score itself as evidence rather than a pointer to records
- Compares scores across different products or tenants as if the scale were shared