First-seen SaaS rules fired 900 alerts overnight after IT migrated every user to a new tenant. What now?
answer
- characterise the flood before working it
- one novel value across everyone means an environment change
- benign true positive, not false positive
- suppress by value with an expiry, not by source
- the durable fix is a learning period
basics
~20 sConfirm the flood is the migration by checking that the novel values are all the same expected new tenant on the change window, then suppress narrowly on that value with an expiry, keep every other detection live, and fix the cause with a cold-start learning period before any first-seen rule may fire for an entity.
solid answer
~50 sFirst establish the cause rather than working the queue: sample the alerts, confirm the novel value is the one new tenant, and match the timing to the change record. If they all share it, this is a benign true positive at scale, not nine hundred investigations. Then suppress precisely — that value, that rule family, a stated expiry — and leave volume, impossible-travel and everything else running, because a blanket mute of the SaaS source hands an adversary a scheduled blind window during exactly the period when everyone's behaviour is unrecognisable. Record it as a time-boxed exception with an owner. The durable fix is a cold-start policy: no first-seen rule fires for an entity or a value until the baseline holds a stated minimum of history, and planned changes seed the baseline in advance. Nine hundred unworked alerts is worse than none, because a real one dies in the pile.
go deeper
Be ready to say that nine hundred alerts sharing one new value is a sign the environment changed, and that the first move is to summarise and group the alerts rather than open them one by one.
Explain the mechanics of the flood — an empty or reset baseline makes every observation novel — and why the verdict is a benign true positive rather than a false positive.
Show the operating judgment: narrow suppression by value with an expiry, other detections left running, an explicit statement that the suppression is an accepted blind window, and a learning period as the durable fix.
Own the coverage argument — that an unworked queue is an undeclared outage of the detection, and that planned change has to reach the SOC as an input rather than as a 03:00 surprise.
## Triage the flood, not the alerts When a detection's output jumps two orders of magnitude overnight, the first hypothesis is that the *rule* changed meaning, not that the estate was overrun. So the first action is a summarising query, not an investigation: group the alerts by the novel value and by the entity, and look at the shape. Nine hundred alerts across nine hundred distinct users all carrying the same previously-unseen destination tenant, all starting within the same hour, is a signature of one environmental change. One user with nine hundred alerts would be a completely different story. Then confirm it externally. The change record for the tenant migration, the identity provider's own configuration history, or a five-minute conversation with the platform team turns a plausible hypothesis into a verified cause. Do not skip that step because the pattern looks obvious — an intruder who has stood up infrastructure that everyone is now being redirected to would produce a similar shape, and the difference is whether the new destination is the one the business asked for. ## Name the verdict correctly These are **benign true positives**. The rule detected precisely what it was written to detect — a value with no precedent — and the cause is legitimate. That matters because the wrong label leads to the wrong fix. Calling them false positives invites someone to weaken the logic, when the logic is fine and what failed was the baseline's assumption that the estate is stable. ## Suppress narrowly, and know what you are giving up The temptation is to disable the rule family or mute the source. Both are over-broad, and in a security context the cost of over-breadth is different from the reliability case: a suppression is a window in which you have chosen not to see, and a migration is exactly when an adversary's activity is most easily explained away as part of the change. Everyone is authenticating from new places, to a new tenant, with re-consented clients — the cover story is free. So scope it: - **By value, not by rule.** Suppress novelty for the one expected destination tenant. Any *other* previously-unseen tenant still alerts, and that is the alert that matters most this month. - **By time.** An explicit expiry that matches the change window, not an open-ended exception. Exceptions without expiries become permanent blind spots that nobody remembers creating. - **By rule family.** Only first-seen logic is broken by the migration. Volume baselines, authentication failure rates, session anomalies and administrative-action rules are unaffected and stay live. - **With an owner and a written record.** Who suppressed it, why, until when, and what compensates for it while it is off. Where the migration runs in cohorts over weeks, re-enable per cohort as each cohort's baseline accumulates, rather than holding one long blanket suppression across the whole programme. ## Fix the cause: a cold-start policy The underlying defect is that a first-seen rule with an empty or reset baseline has no ability to be right. The standing fix is a **learning period**, expressed in the rule rather than in an analyst's memory: - A rule does not fire for an entity until that entity has at least N days, or N observations, in the baseline. - A newly observed value that appears across more than some large share of the population in a short period is treated as an environmental change and raises a single meta-alert rather than one alert per entity — because a value that is new to *everyone at once* is a platform event, and a value that is new to *one account* is the interesting case. - Planned changes are inputs: the change calendar seeds the baseline or triggers the exception before the cutover, so the flood is anticipated rather than discovered at 03:00. That last point is where the SOC's relationship with the platform team does the work. A tenant migration is not a surprise to the organisation, only to the detection. ## Why nine hundred is worse than zero A queue at that size is not worked. Analysts stop reading, the backlog is bulk-closed, and any genuine alert inside it is closed with the rest — the detection is functionally off, but nobody has recorded that it is off. That is the real risk of the flood, and it is why the response is to restore a workable queue quickly and deliberately rather than to grind through the alerts to prove diligence. ## The interview answer Walk it in order: characterise the flood, verify the cause against the change record, label it a benign true positive, suppress by value with an expiry rather than muting the source, keep the other detections live, and then close the loop with a cold-start learning period and a change-calendar input. Say out loud that the suppression is an accepted blind spot, because that sentence is what separates someone silencing a pager from someone managing detection coverage.
- The migration runs in cohorts over six weeks. Do you keep the first-seen rules suppressed for six weeks?No — that is a six-week blind spot on exactly the logic that would catch an unexpected destination. Suppress only the expected tenant value, keep novelty live for every other value, and re-enable per cohort as each cohort's baseline reaches the learning-period minimum. Register the residual gap with an owner and an expiry so it is a decision, not a drift.
- How could you have known this was coming before the alerts arrived?By treating the change calendar as a detection input and by replaying the rule against the planned change beforehand. A cheap safety net helps too: when a rule's firing rate jumps by an order of magnitude, the platform raises one meta-alert saying the rule has stopped discriminating, instead of emitting nine hundred individual claims of maliciousness it cannot support.
- What would make you treat the same alert shape as an intrusion instead of a migration?If the new destination is not the one the change record names, if the timing does not match the approved window, if only a subset of accounts is affected rather than the population, or if the platform team does not recognise the value. The distinguishing evidence is external to the alerts — the alerts themselves look identical either way.
saying these in an interview costs you the question
- Disables the whole SaaS detection family and never re-enables it
- Starts working the 900 alerts one at a time to show diligence
- Calls a legitimate migration a false positive and weakens the rule
- Creates a suppression with no expiry, owner or record
- Assumes the obvious cause without checking the change record