How do you set a log-retention horizon from a right-skewed distribution of observed intrusion dwell times?
answer
- the tail is where the cases live
- half of cases exceed the median
- add discovery lag and case duration
- internally found cases are short by construction
- your data is truncated at current retention
basics
~10 sTake a high percentile of dwell, never the median, then add the lag from discovery to first search plus the weeks the case runs. Treat your own dwell data as truncated at current retention.
solid answer
~50 sThe median is the wrong statistic because half of intrusions exceed it by definition, and the mass that matters is the tail. I take a high percentile — 90th or 95th of confirmed cases — and treat it as a floor. Then I add what raw dwell omits: the lag between discovery and an analyst actually searching, and the duration of the case, because the archive keeps ageing while you work. I segment by discovery method, since internally detected cases are short by construction and externally notified ones are long, and argue the horizon from the external branch where deep history is needed. Finally I state the bias: my measured distribution is censored at roughly the current retention, because an intrusion older than the window is hard to confirm at all, so the derived number is an underestimate.
go deeper
Know what dwell time means — from earliest adversary activity to detection — and that a retention window has to reach back past it to be useful at all.
Be ready to explain why a right-skewed distribution makes the median useless here, and to name the operational slack that gets added on top of a percentile.
Demonstrate the censoring argument: your own dwell data is truncated at your current horizon, so every derived number understates the requirement, and say so before someone else does.
Be prepared to defend a horizon you know is an estimate with wide error bars, and to explain what evidence would move it in either direction.
## The quantity you are trying to cover Dwell time is the elapsed time between the earliest adversary activity in the estate and detection of the intrusion. A retention horizon is useful only if it reaches back past that point for the cases you will actually have to investigate. So the derivation is: look at the distribution of dwell across confirmed intrusions, choose a coverage target, and add the operational slack that the raw dwell number leaves out. ## Why the median is the wrong number Dwell distributions are strongly right-skewed: many short cases, a thin tail running to many months. Three consequences: - The median is, by definition, exceeded by half of cases. A horizon set at the median fails half the time. - The arithmetic mean is dragged by the tail but still sits far below it, so it fails a large share of cases while sounding more sophisticated. - The tail is where the expensive cases live. Long-dwell intrusions are the ones with the most systems touched and the most to reconstruct — precisely the cases you cannot afford to be blind on. So the input is a high percentile — commonly the 90th or 95th of confirmed cases — and it is a floor rather than the final number. ## A worked shape Suppose your last twelve confirmed intrusions gave these dwell figures in days: | statistic | value | | --- | --- | | median | 9 | | 75th percentile | 41 | | 90th percentile | 190 | | longest two | 250 and 310 | A horizon set at the median of 9 days is absurd, and one at 41 still misses the two cases that mattered most. The 90th percentile at 190 days is the honest starting point — and note how violently the number moves with the tail. Removing one long case can halve it. That sensitivity is itself the finding: a horizon derived from a dozen cases is an estimate with wide error bars, which is an argument for rounding generously rather than for precision. ## What the dwell number leaves out Retention has to cover more than dwell: - **Discovery-to-search lag.** You learn of the intrusion, then triage, then open a case. Days pass, and the far edge of your window slides forward the whole time. - **Case duration.** A large investigation runs for weeks. Records that were inside the window on day one of the case can expire during it — a genuinely nasty failure mode, and one reason to snapshot or hold the relevant data at case opening rather than relying on the live index. - **Re-opening.** Cases get reopened when new intelligence arrives or a partner extends their timeline. Adding a margin for these on top of the percentile is not padding; it is the difference between a horizon that covers the intrusion and one that covers the intrusion only if you work fast. ## The bias you must declare Your own dwell distribution is censored by the very number you are trying to set. An intrusion that started before your retention window is hard to confirm and hard to date, so it either never becomes a confirmed case or is recorded with an understated dwell. The measured distribution is therefore truncated at roughly the current horizon, and every statistic derived from it is an underestimate. Published industry dwell figures help as a prior, but they carry the same censoring plus a population that does not match your estate. A second selection effect points the same way: internally detected cases are short because your detections fire on early behaviour, while externally notified cases are long because an outsider only sees a downstream consequence. If you pool the two branches, the internal cases pull the percentile down and hide the cases the horizon exists for. Segment by discovery method and argue the horizon from the external branch. ## What the answer looks like A defensible derivation states four things: the percentile chosen and why, the operational margin added, the segmentation by discovery method, and the direction of the bias. It ends with a number you can defend as a floor, plus an explicit statement that the number is an underestimate — which is the honest position and, incidentally, the strongest argument available when the horizon goes to a budget conversation.
- You only have a dozen confirmed intrusions to work from. Is that enough to derive a percentile?Not with any precision — a 90th percentile over twelve cases is essentially the second-longest case, and removing it moves the answer by months. Treat it as an order-of-magnitude estimate: it tells you whether the honest horizon is weeks, months or a year, not whether it is 180 or 200 days. Round generously, blend in a published prior, and say out loud that the estimate is unstable rather than presenting a false decimal.
- Why does a case that runs for six weeks create a retention problem of its own?Because the window slides forward while you work. Records that were inside the horizon when the case opened can expire mid-investigation, and you lose the ability to re-derive an earlier conclusion or answer a new question about the same period. The fix is to preserve the relevant slice at case opening — export or hold it out of the rotation — rather than treating the live index as your evidence store.
- Should you use a published industry dwell median instead of your own numbers?Use it as a prior, not as the answer. It is drawn from a population of organisations that mostly do not look like yours, it carries the same censoring your own data does, and a median is the wrong statistic either way. Where it helps is sanity-checking direction: if your measured dwell is far below the published figure, the likelier explanation is that you are not finding long-dwell intrusions, not that you have none.
saying these in an interview costs you the question
- Sets the horizon at the median or the mean dwell
- Presents a percentile from a dozen cases as a precise number
- Ignores that the window keeps sliding during a long case
- Treats measured dwell as unbiased despite the current retention limit
- Pools internally detected and externally notified cases into one statistic