skip to content

A support account viewed 11,000 customer records in one afternoon — which innocent explanations must you eliminate?

level: middleimportance: should knowfreq 57%

answer

  1. hypotheses before verdicts
  2. each one names its own killer observation
  3. cadence and client identity first
  4. look outside the account's own rows
  5. volume prompts, it does not conclude

basics

~10 s

List the benign hypotheses first — a shared credential, an automation using the account's token, a mis-scoped report, a UI that prefetches every result — then find the observation that rules each one out.

solid answer

~40 s

Enumerate before you rank. The realistic alternatives are: an automation or reconciliation job authenticating as that account; a credential shared or stored in a team runbook; a support tool or list view that emits a view row per prefetched result; a report or search whose scope was widened by a recent role change; and authorised bulk work — a data-quality sweep, a migration, a red-team exercise — that nobody told the SOC about. Then pair each with a discriminating observation: request cadence and client identity (token versus browser session), the number of distinct sessions and source addresses, whether the views map onto open tickets, the change record for any role grant, and whether other engineers produced the same pattern historically. Report what remains after that, not what you suspected at the start.

go deeper

for a junior

Be ready to list two or three ordinary reasons an account might touch thousands of records, and to say that a large number by itself is a reason to look, not a conclusion.

for a middle

Explain the method: enumerate hypotheses first, attach to each the observation that would eliminate it, and prefer artefacts from outside the account — change records, tickets, other engineers' history.

for a senior

Demonstrate that you drive to the discriminating test rather than accumulating agreeable evidence, and that you can close the case honestly as a benign true positive when that is what the evidence says.

for a principal

Own the standing arrangement that makes these cases cheap: which teams tell the SOC before bulk work, and what the product must log for authorised and unauthorised bulk access to look different at all.

## Why you enumerate first The expensive mistake in a case like this is not reaching the wrong verdict; it is reaching a verdict first and then collecting evidence that agrees with it. Once "insider taking customer data" is written in the channel, every later artefact gets read in that light and the benign explanations never get tested — they simply go unmentioned. Enumerating the alternatives *before* you rank them forces each one to declare what would kill it, which is the only structure that produces a report someone can attack without it falling apart. ## The hypotheses that are actually alive here All of these produce the same shape in the application's audit log — a large number of `record.view` rows on one account in a short window. | Hypothesis | Observation that would rule it out | |---|---| | An automation or reconciliation job using the account's API token | Requests carry a browser session cookie and a browser user agent, with irregular human-scale gaps | | The credential is shared, or sits in a team runbook | Only one session id and one source address, and it matches the engineer's usual client and network | | A list or search screen that prefetches the full record for every result | Views are not clustered behind a small number of search requests, and each has a distinct navigation path | | A report or export whose scope widened after a recent role grant | No entitlement change in the identity provider's change record before the window | | Authorised bulk work — data-quality sweep, migration, an announced exercise | No ticket, change record, or project asked for it, and no other engineer ran the same pattern | | Ticket-driven support work at unusual volume | The viewed customer ids do not correspond to any open case | Notice that most of the discriminating observations come from *outside* the account's own activity: a change record, a ticket queue, another engineer's history. Evidence drawn only from the same account tends to be circular — it describes the pattern again rather than explaining it. ## Cadence and client identity do the most work If you can only ask one thing, ask how the requests were made. Eleven thousand records in an afternoon is roughly one every second and a half sustained — possible for a person clicking through a list screen that prefetches, and implausible as deliberate reading. Fixed inter-request intervals, a non-browser user agent, an API token, or a source address belonging to infrastructure rather than an office or VPN pool all point at a program. That does not close the case; it *moves* it. The question becomes who owns that program and whether the access was authorised — a different investigation, sometimes a benign one, occasionally worse, because an automation quietly using a human's credentials is its own finding. ## Benign true positive is a real verdict A benign true positive is activity that genuinely happened, genuinely matches the rule, and is genuinely authorised. It is not a false positive: the detection was right about the behaviour and wrong about nothing. Closing this case as "the engineer ran an approved data-quality sweep under a ticket" is a correct and complete outcome, and it should be recorded as such, because the next analyst who sees the same pattern needs to know the difference between "we saw this and it was fine" and "this rule is noisy". ## Volume proves volume The single most common weak answer treats the number as the finding. Eleven thousand is a *prompt*, not a conclusion. It establishes that the access was far outside the account's normal range — which is what makes it worth an hour of someone's time — and nothing about intent, retention, or whether the data left. Intent in particular is almost never visible in telemetry; you can sometimes observe things *consistent* with it (access outside working patterns, records with no ticket, activity that stops the moment a query is asked), but the artefact that shows intent directly is usually a message, a document, or an interview, not a log. ## What you write down End with three separated lists: what was observed, what those observations are consistent with (plural — keep the surviving alternatives visible), and the specific test that would discriminate between the survivors, with how long it takes. That is a report a reviewer can act on, and it is also the version that does not embarrass you when the answer turns out to be a forgotten migration script.

  • If you could gather only one more observation here, which would you choose?
    Whether the requests carried a browser session with human-scale gaps or a token at machine cadence. It splits the case in two: a person paging through a screen, or a program authenticating as the account. Almost every other question you would ask has a different answer depending on which of those is true, so it is the cheapest observation with the largest effect on the conclusion.
  • The engineer says a runbook told them to run a bulk lookup. How do you test that claim?
    Look for artefacts that exist independently of the engineer: the runbook and its edit history, the ticket or request that triggered the work, whether other engineers produced the same pattern before, and whether the set of records touched matches the stated scope. A claim corroborated by contemporaneous records made by other people is worth something; one supported only by the same account's activity is not.
  • How is closing this as a benign true positive different from closing it as a false positive?
    A false positive means the rule was wrong about the behaviour — the thing it claimed to see did not happen. A benign true positive means it happened exactly as described and was authorised. The distinction matters downstream: false positives justify changing the rule, benign true positives usually justify enriching it with context such as ticket linkage, not deleting it.

saying these in an interview costs you the question

  • Starts from insider theft and gathers only confirming evidence
  • Treats record volume alone as proof of intent
  • Forgets automation can authenticate as a human account
  • Never checks whether the access matched open tickets
  • Calls authorised bulk work a false positive

context