skip to content

300 cloud detections arrive daily and you can work 40. How do you choose which 40?

level: middleimportance: must knowfreq 72%

answer

  1. ordering changes which, not how many
  2. arrival rate against service rate
  3. some evidence dies on a schedule
  4. same call, different account, different meaning
  5. newest-first plus a bounded oldest sweep

basics

~20 s

Work order decides which 260 you never examine, not how many you finish. Order by what a late look costs: evidence about to expire, high blast-radius accounts and identities, and stages of an intrusion you cannot undo.

solid answer

~50 s

Start by saying the arithmetic out loud: at 300 arriving and 40 worked, no ordering policy shrinks the backlog — it only chooses which 260 go unexamined each day. Then order by the **cost of being late**, not by the label the rule stamped on itself. Three levers do most of the work. **Perishability** — items whose supporting telemetry expires soonest, or whose host or container is ephemeral, because a delayed look becomes an impossible one. **Blast radius** — the same API call from an organisation-management or production identity outranks it in a sandbox account. **Stage** — detections that sit late in an intrusion, where data leaves or persistence is created, outrank early reconnaissance-shaped noise, because the damage is already irreversible. In practice that is mostly newest-first, with a bounded oldest-first sweep so the tail cannot silently expire, and the unworked set counted openly rather than allowed to hide.

go deeper

for a junior

Be ready to say that ordering decides which items go unexamined rather than how many get done, and to name two ordering inputs beyond the severity label — for example the account or identity involved and how soon the supporting evidence expires.

for a middle

Expect to explain the mechanics: arrival rate against service rate, newest-first versus oldest-first and each one's failure mode, and why the same API call means different things in a sandbox account and a production one.

for a senior

Demonstrate an ordering policy you have actually run, including the bounded sweep that stops the tail expiring silently, the per-source evidence horizons that drive earliest-expiry-first, and how you keep the unworked residue counted.

for a principal

Own the framing upward: present the daily residue as a structural output of capacity against arrivals, so the choice of what goes unexamined is made deliberately by someone who can fund or change it rather than settled by whoever picks up the queue.

## First, name the arithmetic With 300 arrivals a day and capacity for 40, the queue grows by 260 items every day regardless of what you do. This is worth saying in the first sentence of an interview answer, because it reframes the question honestly: **work order is not a throughput lever, it is a choice about which items are never examined.** A candidate who answers only with a sorting rule has implicitly claimed the backlog is a prioritisation problem. It is a capacity problem with a prioritisation decision inside it, and the structural remedies — automating a class of verdicts, tuning a rule family, deliberately turning something off — are a separate conversation with a separate owner. Your job at the console is to make the daily 40 the *right* 40 and to keep the other 260 countable. ## Order by what a late look costs ### Perishability first Every detection has a date after which it cannot be answered, set by the shortest-retention source you would need. In a multi-account, multi-provider estate those horizons differ per source and per provider — a control-plane trail, a flow log, an endpoint agent's local buffer and a SaaS audit trail rarely expire together. An item whose evidence dies on Friday and an identical item whose evidence dies in eleven weeks are not the same item. Ephemeral compute sharpens this further: if the container or the auto-scaled instance is gone in an hour, the on-host evidence goes with it and only what was already shipped survives. Earliest-expiry-first is the one ordering rule with a hard, external deadline behind it, and it is the one candidates most often omit. ### Blast radius second The same call means different things depending on who made it and where. A permission change made by an organisation-management identity, a role assumable from many accounts, or a principal in the account holding production data outranks the identical event in a sandbox. This is asset and identity context, not rule severity — the rule stamped a label at authoring time without knowing which of your dozens of accounts it would fire in. ### Stage third Behaviour late in an intrusion — data being copied out, a snapshot or key shared outward, a new credential or trust relationship created — deserves the analyst's hour more than behaviour that is early, cheap for an adversary to repeat, and reversible if you are slow. The reason is not that the late behaviour is *scarier*; it is that being late costs more. Something that has already left cannot be un-left, and a persistence mechanism you do not look at today is still there tomorrow. ## Newest-first, oldest-first, or both **Newest-first (LIFO)** maximises the chance of catching an intrusion while it is still running and while the evidence and the live state still exist. Its failure mode is brutal: under permanent overload the oldest items are never reached and quietly expire. **Oldest-first (FIFO)** feels fair and bounds the age of the oldest unopened item, but it spends your scarce hours on the items least likely to still be actionable, and it lets a live intrusion wait behind three-week-old noise. Most working SOCs run mostly newest-first with a **bounded sweep** of the oldest band — a fixed slice of each day or shift spent on items approaching expiry — which is really earliest-expiry-first wearing a simpler uniform. ## The lever candidates over-trust A rule's historical true-positive rate is genuinely useful for ordering, with two caveats an interviewer will probe. It is measured only over items that were worked, so it says nothing about the band nobody opens; and a rule that has never had an alert opened has no measured rate at all, which is not the same as a rate of zero. Sorting the queue purely by rule severity descending is the weaker version of the same mistake: severity was assigned by the rule author, before the rule knew which account, which identity or which asset it would fire on. ## Make the residue visible Whatever you choose, the 260 must not vanish. Count them, know the age of the oldest unopened item, know how many are within a week of their evidence expiring, and be able to say which classes of detection systematically never get opened. A work order you can describe, plus a residue you can measure, is the answer; a sorted queue with an invisible tail is not.

  • If you work newest-first, don't the oldest items simply never get worked?
    Yes, and that is the known failure mode. The fix is a bounded sweep: a fixed share of each shift spent on the items closest to their evidence-expiry date, so the tail ages out on purpose rather than by accident. Without it, newest-first quietly converts the oldest band into permanently unanswerable items while the daily numbers look healthy.
  • Should you order the queue by each rule's historical true-positive rate?
    It is a legitimate input with two caveats. The rate is measured only over alerts somebody worked, so it does not describe the band that is never opened; and a rule whose alerts have never been opened has no measured rate, which is different from a measured zero. Use it as one signal alongside asset context and evidence expiry, never as the sole sort key.
  • How do you defend spending capacity on medium-severity items while criticals wait?
    By naming the cost of lateness rather than the label. A critical whose telemetry retains for a year and whose host is still under your control can wait an hour; a medium on an ephemeral workload, or one whose control-plane evidence expires this week, cannot. Severity is the rule author's prior; expiry and blast radius are facts about your estate.

saying these in an interview costs you the question

  • Claiming a better sort order will clear the backlog
  • Sorting purely by rule severity descending
  • Strict oldest-first because it feels fair
  • Ignoring that evidence for some sources expires sooner
  • Leaving the unworked remainder uncounted and undescribed

context