Why can't a 1.8-million-indicator threat feed be pushed straight into a proxy block list?
answer
- two costs, not one
- the list runs in the traffic path
- object groups have entry ceilings
- malicious addresses are rarely dedicated
- block narrowly, match widely
basics
~20 sEvery entry costs something to evaluate and something to be wrong about. Enforcement lists have size and push limits, and adversary infrastructure usually sits on shared or CDN-fronted addresses that also carry your own business traffic.
solid answer
~40 sThere are two costs, and volume ignores both. The first is match cost: an enforcement point evaluates its list on traffic in the path, and proxy destination lists and firewall object groups have vendor ceilings on entries, plus real compile-and-push time across every region. The second, and the one that hurts, is collateral: adversary infrastructure is mostly rented and mostly shared, so a large share of malicious addresses are shared hosting, CDN edges or big SaaS platforms carrying your payroll or CRM traffic too. So I select. I enforce specific, high-confidence indicators — a URL or a hostname, rarely a bare address, never a prefix — and I send the rest to matching in the SIEM, where being wrong produces an alert an analyst closes rather than an outage nobody can attribute to me.
go deeper
Be ready to say that a block list costs both match capacity and collateral damage, and that a feed's size is not evidence of protection. Name the safest indicator types to enforce: URLs and hostnames.
Explain where the limits actually bite — entries per object group, compile and distribution time across regions — and how indicator type and confidence drive which entries go inline versus which go to matching.
Show a written selection policy: what you enforce, what you refuse to enforce, what stays on a documented allow-list, and how you keep visibility on everything you chose not to block.
Own the trade the business is actually making: enforcement buys prevention and buys outage risk in the same purchase, and your team should be able to state in advance which classes of indicator it will ever push into a traffic path.
## Two places an indicator can land An indicator is an artefact somebody observed — an address, a domain, a URL, a file hash. A feed delivers them in bulk, and the important decision is not whether to take them but *where to put them*. A **detection** point matches indicators against records that have already been written: proxy access logs, DNS query logs, flow records, endpoint telemetry. A hit produces an alert, and a human decides what it means. An **enforcement** point sits in the traffic path: a forward proxy's destination or category list, an egress firewall object group, a DNS sinkhole. A hit stops a connection, immediately, with no human in the loop. Bulk-loading a feed into the second kind is the mistake. The question is which small subset of indicators earns a place there. ## Cost one: match cost and list mechanics An inline control evaluates its lists against traffic, so list size is not free. Concretely: - Firewall object groups and access lists have vendor-specific ceilings on entries, and large ones consume constrained hardware matching resources; you can hit the ceiling and have the push simply fail. - Proxy destination lists must be compiled and distributed to every node in every region; a very large list turns a routine push into a slow, heavyweight operation. - A list push is a **production change**, and in most estates there is no staged copy of the proxy fleet to try it on first. The first place the change runs is the place users are. None of this makes a few thousand entries a problem. It makes *the feed's whole volume* a problem, and it makes the push itself something you plan rather than something you fire. ## Cost two: collateral, which is the real one Adversary infrastructure is rented, borrowed or stolen. Command-and-control lives on a VPS, on a compromised site on shared hosting, on a bucket or app inside a large SaaS platform, behind a CDN edge that fronts thousands of tenants. The address you were handed is very often carrying other people's traffic — including yours. Granularity is therefore the whole game, and it runs from safest to most dangerous: | Indicator | What a block hits | |---|---| | Full URL | one path on one host | | Hostname | one service | | Single address | every tenant on that address | | Prefix (a /24, a /16) | every tenant in the range, including yours | A feed that hands you netblocks is handing you outages, not prevention. ## The errors are not symmetrical Be explicit about the direction of each mistake. A wrong **match** costs an analyst a few minutes closing an alert. A wrong **block** costs users a failure that carries no message naming its cause: the help desk hears about it before the SOC does, the symptom looks like a vendor outage, and the loss is business revenue rather than security. That asymmetry, not the feed's confidence score, is what should drive how much of a feed you are willing to enforce. ## A selection policy you can defend 1. Enforce URLs and hostnames from high-confidence, recently observed sources — and prefer indicators your own incidents produced. 2. Enforce a bare address only when you have reason to believe the host is dedicated; refuse to enforce prefixes at all. 3. Keep an explicit, owned allow-list of CDN, shared-hosting and SaaS ranges you will never block at the address layer, with a reason recorded against each entry so nobody later tidies it away. 4. Everything you decline to enforce still goes somewhere: match it against your logs so a hit raises an alert. Declining to block is not declining to look. 5. Run any new list in alert-before-block first, and turn enforcement on one region at a time. ## What the volume claim actually says 'We block 1.8 million indicators' describes a subscription, not a defence. The defensible statements are narrower and much harder to produce: which indicators are enforced, what they hit last month, what business traffic they nearly hit, and which of them stopped something the rest of the stack would have missed. Volume is what a feed sells; specificity is what an enforcement path can afford.
- Which is the cheaper thing to be wrong about — an indicator you blocked or one you only matched?The one you matched. A false match costs an analyst a few minutes and leaves the user unaffected. A false block produces a failure with no error message that names the SOC, so it is usually diagnosed by somebody else, late, as a vendor problem. That asymmetry is why most of a feed belongs in detection rather than enforcement.
- Where do the indicators you decline to enforce actually go?Into matching against the records you already collect — proxy, DNS, flow and endpoint logs — so a hit raises an alert with context instead of dropping traffic. You keep full visibility of the feed and give up only the automatic block, which is the part that can take a business system down.
- Does blocking a domain cost the same as blocking an address?The evaluation cost is broadly similar; the collateral cost is not. One hostname is one service, so a wrong hostname block breaks one thing. One address can be thousands of tenants behind a CDN or shared host, so a wrong address block breaks whatever else lives there — and you will not know what that is until somebody complains.
A block list is a bouncer with a photo album, not a filing cabinet. Every extra photo slows the door, and one blurry photo turns away a paying customer.
saying these in an interview costs you the question
- Treats the number of blocked indicators as a security metric
- Assumes every address on a feed is dedicated to the adversary
- Calls a block list free because lookups are fast
- Forgets that pushing a list is a production change
- Blocks a prefix to catch one host on it