skip to content

Analysts closed the same nightly cryptomining alert by hand for fourteen months — why did nobody report it?

level: seniorimportance: should knowfreq 44%

answer

  1. the cheap close beats the report
  2. measured on time-to-close, not rules fixed
  3. no route, and no receipt
  4. habituation, not the wasted minutes
  5. make the disposition itself the report

basics

~20 s

Because reporting cost more than closing and produced no visible result. Ninety seconds a night is cheaper than finding an owner two org units away, and analysts are measured on time-to-close, not on rules improved. The real damage is the trained reflex.

solid answer

~50 s

The silence is rational, which is why exhortation will not fix it. Each close costs ninety seconds; reporting means identifying an owner in another org unit, writing the case up, and then hearing nothing, so the expected return on complaining is zero. Analysts are measured on queue depth and time-to-close, both of which the fast close improves. Meanwhile the cost is not the fifteen hours of aggregate handling — it is that the queue has learned this rule means nothing, so a genuine resource hijack on a render node gets the same ninety-second close as the six hundred before it. The fix is structural: make the disposition itself the report by capturing the deciding field, roll those up to a named rule owner automatically, put that owner's name in the alert, and report back to the analysts what changed. A loop that never returns a receipt stops receiving input within weeks.

go deeper

for a junior

Be ready to say why closing a known-bad alert night after night is a problem rather than efficiency, and to name the reflex it builds. If you are the one closing it, saying so out loud is part of the job.

for a middle

Explain the incentives concretely: the close is cheap, the report is expensive, and neither queue depth nor time-to-close rewards the report. Be able to state the habituation cost separately from the wasted-time cost.

for a senior

Show that you fix the route rather than the attitude — the disposition becomes the report, a rollup does the noticing, the alert names an owner, and the loop returns a receipt. Interviewers listen for whether you know where your authority over the rule stops.

for a principal

Own the argument to leadership that a technique can be covered on paper and dead in practice, and the cross-team contract that makes queue evidence land with a detection team you do not manage.

## The economics of the silent close Start by taking the analyst's side seriously. The nightly firing on the render farm takes about ninety seconds to close: recognise the rule, recognise the asset, stamp it, move on. Reporting it properly means working out who owns the rule — a detection team two org units away — writing up a case that explains the pattern, and then waiting. There is no acknowledgement, no ticket status they can see, and no expectation that anything will change. Ninety seconds beats that every night, and it beats it *every* night, which is how you get to fourteen months. Add the measurement layer. Tier-1 analysts are typically measured on queue depth, time-to-triage and time-to-close. Fast-closing a known misfire improves all three. Writing a detection-quality report improves none of them and consumes the time that would have. Nothing in the incentive structure asks for the behaviour that the detection programme depends on. Then add the organisational distance. The queue and the rule sit under different managers with different objectives, and the detection team is defending a false-positive budget of its own. A complaint arriving from outside that structure has no natural landing place. ## What it actually costs The cheap answer is the arithmetic: roughly 612 closures at ninety seconds is about fifteen hours. That number is real, and it is not the reason this matters. The reason it matters is **habituation**. Every close teaches the queue that this rule's output is meaningless. The rule exists to catch resource hijacking (`T1496`) — an intruder monetising your compute, which on a GPU estate is a natural target. When that finally happens on a render node, the alert arrives into a trained reflex: it looks like the six hundred before it, and the analyst closes it in ninety seconds. The rule fired. The telemetry worked. The detection still failed, and the failure is not in the rule at all — it is in the loop around it. A second, quieter cost: the coverage map still says this technique is covered. On paper, a working detection exists. In practice, its output has been neutralised by the people receiving it, and nobody has written that down anywhere. ## Fixing the route rather than the attitude **Make the disposition the report.** If closing already requires recording the deciding field and the asset group, the analyst has filed the report by closing the alert. No separate act of complaining exists to skip. This is the highest-leverage change and it is why the closure schema matters. **Aggregate automatically and route it.** A periodic rollup of closures grouped by rule, verdict and asset group would have surfaced this in the first month: one rule, one asset group, hundreds of identical verdicts. Nobody has to notice; the rollup notices. **Put an owner in the alert.** If the alert itself names who owns the rule, the route is one click rather than an archaeology exercise across org charts. **Return a receipt.** Tell the analysts what happened to the evidence they produced — including "the owner looked at it and the rule stays, here is why". A loop that swallows input without ever reporting back stops receiving input within weeks, and then you are back to silence with better forms. **Measure the contribution.** If detection-quality feedback is genuinely wanted from tier 1, it has to appear somewhere in how tier 1 is judged. Otherwise you are asking people to spend the time they are measured on doing something they are not measured on. ## Where your authority ends What happens to the rule after the evidence lands — narrowing the logic, scoping an exclusion, or dropping it entirely — is the rule owner's decision, made against obligations across the estate that the queue cannot see. The SOC's job is to make sure the decision is made explicitly, by a named person, on real evidence, rather than absorbed silently at ninety seconds a night for fourteen months.

  • The rule's owner acts on the evidence and changes something. What do you owe the analysts?
    A receipt: what was sent, what the owner decided, and what they should expect to see in the queue tomorrow. This is not courtesy, it is what keeps the route alive — a feedback loop whose contributors never learn whether anything happened stops being used within a few weeks, and the next broken rule is absorbed silently again.
  • How would you have found this without an analyst raising it?
    A standing rollup of your own dispositions — closures grouped by rule, verdict and asset group over a rolling window. One rule accounting for hundreds of identical verdicts from a single asset group is visible in the first month and needs nobody to notice or complain. The rollup does the noticing that the incentive structure will not.
  • Is the fifteen hours of analyst time the argument you take to leadership?
    Lead with the coverage argument instead. Fifteen hours over fourteen months is easy to dismiss; a technique the coverage map claims is covered, whose alerts the queue has been trained to dismiss on sight, is not. The time saved is a supporting figure, not the case.

A smoke alarm the kitchen fans a towel at every evening. The waving is not laziness; it is the cheapest available response. It is also exactly what will happen on the night there is a fire.

saying these in an interview costs you the question

  • Blames analysts for not caring enough
  • Says the cost is the wasted ninety seconds
  • Believes a reminder in a team meeting fixes it
  • Adds a reporting form with no route back
  • Assumes the coverage map reflects working detection

context