skip to content

Leadership wants to retire the phish-report button because 97% of reports are benign - what do you argue?

level: principalimportance: nice to knowfreq 30%

answer

  1. a low hit rate is by design
  2. count campaigns caught first by users
  3. delivery-to-report time is the lever
  4. cut the cost, not the sensor
  5. the reports you never get are invisible

basics

~20 s

Hit rate is the wrong measure. The button's value is the campaigns where a human report was the first signal because the gateway had already delivered the mail, and how fast it arrived. Attack the queue's cost, not the sensor.

solid answer

~50 s

Hit rate measures how calibrated thousands of untrained reporters are, not what the sensor is worth, and it is supposed to be low - we asked everyone to report anything odd, so most of it is newsletters, vendors and our own campaigns. The numbers I would bring instead are how many distinct malicious campaigns last year had a user report as their **first** signal, meaning the gateway had already delivered them, and the median time from delivery to report, because that decides how many other recipients still had the message in front of them. If the real complaint is cost, attack cost: dedup by campaign, auto-close known internal and vendor sends with an immediate reply, pre-extract the indicators. Removing the button removes the only detection surface where the target is also the sensor - so if it still goes, it goes into the risk register with the miss data and a named owner.

go deeper

for a junior

Know that most reports are benign by design, that this is not a mark against the people reporting, and that every report still gets an answer.

for a middle

Be able to name the measures that describe the programme - first-signal campaigns, delivery-to-report time, reporter reach, analyst-minutes per report - and why precision is not one of them.

for a senior

Show how you would compress the queue's cost without blinding the sensor, and how you would instrument the programme so the argument is settled with data rather than opinion.

for a principal

Own the framing: fund it, shrink it, or retire it as an accepted risk with a named owner. Be explicit that reporting culture is slow to build, fast to lose, and impossible to measure once it is gone.

## The metric being argued from is the wrong one A reported-message queue in which 97% of reports are benign is not a broken programme; it is the expected output of asking a whole workforce to forward anything that feels wrong. The denominator is every odd-looking mail in the company - marketing campaigns, vendor invoices, recruiters, the finance team's own reminder mails - and the numerator is real attacks, which are rare. A low hit rate is what a **sensitive** sensor with an untrained operator population looks like. Judging it by precision is the same error as judging a smoke detector by how often it goes off because of toast. ## The numbers that would actually settle it Bring four, and be honest about each: 1. **Campaigns where a user report was the first signal.** This is the whole case. It counts only mail the gateway had already delivered - a residual miss the button caught. If it is a handful a year, that is a handful of campaigns that would otherwise have sat in mailboxes until something worse happened. 2. **Median time from delivery to first report.** This decides how much of the campaign is still recoverable. Minutes means most recipients had not read it yet; a day means you are scoping something already acted on. 3. **Reach.** How many distinct people have reported at all, and how concentrated the reporting is in a few enthusiasts. A programme carried by nine people is fragile in a way the raw count hides. 4. **Cost, measured rather than asserted.** Analyst-minutes per report, and the share of reports that are duplicates of a campaign already triaged - usually most of them. If nobody has measured this, the cost complaint and the defence are both anecdote. There is also a value the volume metrics cannot see: a targeted approach to one executive produces exactly one report and no pattern for any volume-based control to notice. The button is often the only way that case ever surfaces. ## Reframe the 97% Those reports are **benign true positives**, not analyst failures and not false positives. The reporters correctly identified mail with phishing-shaped properties; the intent turned out to be legitimate. Recording them as false positives - which many SOCs do - manufactures the exact statistic being used against the programme, and it is worth checking whether the 97% figure in the room was built that way. ## If the cost is real, attack the cost A reported-message queue is unusually compressible, because reports arrive in campaign-shaped clusters: - Dedup at campaign level so thirty-one reports of one newsletter are one unit of work. - Auto-close plus an immediate templated reply for known internal senders and recognised vendor campaigns, with a small sampled review so the auto-close does not become a blind spot. - Pre-extract sender, links and infrastructure so an analyst opens a summary rather than a raw message. - Reserve real analysis for first-seen campaigns. What you must not cut is the reply to the reporter. That is the maintenance the sensor needs, and it is cheap when the case is deduped. ## Why this is not a reversible experiment "Turn it off for a quarter and see" sounds measured and is not. Reporting is a behaviour, and it decays quietly: people who hear nothing back, or who find the button gone, stop looking, and they do not resume on the day you re-enable it. Worse, the thing you would need to measure during the trial - the campaigns nobody reported - is invisible by construction. You cannot count the reports you did not receive. The genuinely reversible experiment is to change the **cost model** for a quarter, keep the sensor, and measure analyst-minutes and reporter satisfaction. ## Owning the decision, not winning the argument The SOC's job here is to make the option set and the evidence visible, not to defend a tool it likes. The honest framing has three options: fund the queue as it stands, shrink its cost by the measures above, or retire it and accept a named risk. If leadership picks the third after seeing the first-signal data, that is their call to make - and it should be recorded as an accepted risk with an owner, the miss data attached, and a review date, so that the first campaign that lands after the button goes is a decision anyone can trace rather than a surprise. And be prepared to be wrong. If a year of data shows no campaign whose first signal was a report, and the gateway detects and pulls everything before users see it, the button really has become redundant as a detection surface. Check whether that is true - or whether nobody was measuring it.

  • What evidence would convince you the button should actually go?
    A year in which no malicious campaign had a user report as its first signal, every reported-malicious message had already been detected and removed by the gateway before the report arrived, and no targeted approach surfaced this way - plus confidence the measurement existed to see those things. Absent that data, the argument is being made from cost alone.
  • How do you halve the queue's cost without discouraging reporters?
    Dedup at campaign level, auto-close with an immediate templated reply for known internal and vendor sends, pre-extract sender and link details so triage starts from a summary, and reserve deep analysis for first-seen campaigns. Keep a sampled review of auto-closed reports. The one thing you never cut is the answer back to the reporter.
  • Leadership goes ahead and retires it anyway - what do you do?
    Record it as an accepted risk with a named owner, the first-signal data attached and a review date, and say plainly what detection is lost - delivered mail nobody else will look at, and targeted approaches with no volume pattern. Then instrument what remains, so the next campaign that lands is traceable to a decision rather than a surprise.

Nobody scraps a smoke alarm because it mostly detects toast. You ask instead how many fires it caught first, and how early.

saying these in an interview costs you the question

  • Defends the programme on hit rate alone
  • Records benign reports as false positives and then argues from that number
  • Treats switching the button off as a reversible experiment
  • Accepts the cost complaint without ever measuring analyst-minutes
  • Argues the case with no miss data at all
  • Refuses the decision rather than framing options and a risk owner

context