skip to content

How do you measure a paid intel feed's unique contribution against the free sources you already run?

level: middleimportance: should knowfreq 47%

answer

  1. the vendor's count is a volume claim
  2. normalise before you compare anything
  3. delivered, unique, matched, acted on
  4. lead measured against your exposure time
  5. sample indicators later for staleness

basics

~20 s

Normalise every source into one store, then measure a funnel: how many indicators only this feed carried, how many of those matched your own telemetry, and how many changed a verdict. Also measure how early it carried them.

solid answer

~50 s

I build an overlap harness. Every source lands in one store with the artefact normalised - defanged forms unwrapped, case folded, URLs separated from domains separated from addresses - and stamped with the source and the time we first received it. Then I compute a funnel over ninety days: delivered, unique to this feed, unique and matched against our own telemetry, and matched and changed a verdict. Volume dies at the second step and most feeds die at the third. Alongside that I measure first-seen lead time: for artefacts several sources carried, how far ahead was this one, and did that lead land before or after the thing touched us. Finally I sample the unique indicators weeks later to see how fast they go stale, because a feed whose value is a forty-minute lead is worthless consumed as a nightly file.

go deeper

for a junior

Be able to say why a feed's headline indicator count tells you nothing, and name the one comparison that matters: how much of it no other source you already have was carrying.

for a middle

Expect to describe the mechanics - one normalised store, receipt timestamps per source, and a funnel from delivered through unique to matched - and to explain why normalisation changes the answer dramatically.

for a senior

Demonstrate the judgment that turns numbers into a recommendation: lead time measured against your own exposure, staleness sampling driving consumption cadence, and honesty about small-sample estates.

for a principal

Be ready to defend the measurement itself to a vendor who disputes it, and to decide how much analyst time an evidence-gathering exercise like this deserves against everything else the team is not doing.

## Why the vendor's number is not the measurement An account manager quoting ten million indicators under management is making a volume claim. Volume is the one property of a feed that is free to inflate: reprint public blocklists, expand every URL into its domain and its address, keep everything forever. The question that decides a renewal is not how much a feed carries but **how much of it you would not otherwise have had, and whether any of that changed what you did**. ## Build one store, then normalise Overlap analysis is only meaningful if the comparison is fair, so everything - paid feeds, free lists, sharing-community output, your own internal observations - lands in a single store with three stamps: the artefact, the source, and the time *you received it*. Then normalise, because most apparent uniqueness is formatting: - unwrap defanged forms (`hxxp`, `[.]`), fold case, strip trailing dots on domains; - keep artefact **types** apart - a URL, its domain and its resolved address are three different observations and collapsing them manufactures both false overlap and false uniqueness; - deduplicate by artefact plus type, keeping the earliest receipt per source. Skip this and you will "discover" that a vendor contributes eighty percent unique indicators, when in truth it reprints the same lists with a different escaping convention. ## The funnel that actually decides Measure four numbers over a fixed window, ninety days is a common choice: | Stage | What it counts | | --- | --- | | Delivered | Everything the feed sent | | Unique | Artefacts no other source you run carried in the window | | Matched | Unique artefacts that appear anywhere in your telemetry | | Acted on | Matches that produced or changed a verdict | The drop from delivered to unique kills the reprint feeds. The drop from unique to matched is where most of the rest die: an indicator about infrastructure that never touched your estate is not intelligence about you. The final column is the only one that survives contact with a budget conversation, and it is usually a very small number - sometimes one. ## First-seen lead time A feed that carries an artefact three days after a free source is not a duplicate in any operationally useful sense; it is late. For every artefact carried by more than one source, compute the receipt-time delta between this feed and the earliest other source. Then ask the question that matters: **did the lead land before the thing touched us?** Being forty minutes ahead of everyone else is decisive if the artefact arrived before the first phishing message was delivered, and worthless if it arrived a day after the session had already been stolen. Lead time is only value when it is measured against your own exposure time, not against other vendors. ## The age curve Sample the feed's unique indicators and re-check them at intervals - seven, thirty, ninety days - for whether they are still true: does the domain still resolve to the same infrastructure, is the address still serving the same thing, has the name been re-registered by an ordinary business. Two things fall out. First, a shape: some feeds are almost entirely fast-moving phishing and command-and-control infrastructure with a half-life measured in days, others are slow, long-lived artefacts. Second, a consumption requirement: a feed whose value is concentrated in the first hours must be consumed by streaming API or frequent poll and delivered into a control in near real time. Buying that feed and ingesting it nightly discards the property you paid for. ## Traps that ruin the analysis - **Counting indicators instead of matches.** The headline number is the vendor's marketing metric; borrow theirs and you have run their analysis, not yours. - **Vendors reselling each other.** Two subscriptions can share an upstream. Uniqueness is measured against *your actual source set*, whoever owns it. - **Small-estate sample size.** A small identity-and-SaaS estate may legitimately see zero matches in ninety days from a genuinely good feed. That is evidence about your exposure as much as about the feed - so pair the funnel with the qualitative question of whether the feed covers the threats that plausibly reach you. - **Retrospective bias.** Only counting matches on alerts that already fired hides the feed's contribution to sweeps of historic telemetry, so run the match stage against stored telemetry rather than against the alert queue alone. - **Confusing this with detection quality.** This measures a *source*, not a rule. A feed can be excellent and the detection built on it noisy, or the reverse. ## What good looks like A one-page result per source: delivered, unique, matched, acted on; median lead time against your other sources and against your own exposure; a staleness curve; and, in a sentence, the verdicts that came out differently. That page is the artefact you take into a renewal, and it is also the honest answer when the number in the last column is zero.

  • Your ninety-day funnel shows the feed contributed hundreds of unique indicators and zero matches. Is it worthless?
    It is unproven rather than proven worthless, and I would say so. Zero matches is evidence about my exposure as well as the feed - a small estate can genuinely go a quarter without touching any of it. I would extend the window, run the unique set against stored telemetry rather than only against fired alerts, and ask whether the feed's collection covers threats that plausibly reach us. If it still shows nothing, the honest recommendation is to stop paying for it.
  • Two vendors both claim unique coverage. How do you tell whether they share an upstream?
    Look at the receipt-time deltas on the artefacts they share: an upstream relationship usually shows up as a tight, repeatable lag rather than a random spread, and often as identical field values or identical typos in descriptions. I also ask each vendor directly which of their collections are original and which are aggregated, and I measure uniqueness against my actual source set rather than against a market.

saying these in an interview costs you the question

  • Accepts indicator count as a measure of feed value
  • Compares sources without normalising artefact types or formats
  • Calls a late-arriving duplicate the same as a duplicate
  • Stops at uniqueness and never joins against own telemetry
  • Treats zero matches as proof the feed is bad
  • Measures lead time only against other vendors, never against exposure

context