skip to content

How do you prove what share of your endpoints still resolve through the sanctioned filtering resolver?

level: seniorimportance: nice to knowfreq 28%

answer

  1. you cannot count what you never received
  2. join the askers to the inventory
  3. a second source splits silence apart
  4. make the endpoint prove the path

basics

~20 s

Not from resolver logs alone — they cannot show a host that asked somebody else. Join the querying sources to the asset inventory, add egress records for other resolver destinations, and make the endpoint prove the path.

solid answer

~50 s

Three sources, none sufficient alone. The resolver's own logs give the set of source addresses that did ask you; intersect that with the asset inventory and the gap is your unknown population, not your compliant one. Egress records from the branch — sessions to port 53 or 853 destinations that are not you, and 443 to addresses on your list of public providers — recover part of that gap. Then an active check: publish a name only your filter can answer and have the managed endpoint resolve it on a schedule, which proves the path for that process at that moment. The caveats are the answer's real content: NAT at the branch collapses sources so per-host attribution needs a DHCP or inventory join, one process's resolver does not certify the host, and an implant carrying its own resolver never appears in any of the three. What you can publish is a floor.

go deeper

for a junior

Understand the basic asymmetry: a resolver can log every query it receives, and has no record whatsoever of a host that sent its queries somewhere else.

for a middle

Explain the join — querying sources against the asset inventory — and what the three resulting buckets mean, including sources that appear but are not in the inventory at all.

for a senior

Demonstrate the second and third sources: egress records showing resolution to other destinations, and an active check from the endpoint that survives address translation, plus the limits of each.

for a principal

Be prepared to state a coverage floor with its residual to an auditor, and to defend why a true percentage is unobtainable rather than promising a number the estate cannot support.

## Why the obvious answer is wrong The instinct is to count distinct source addresses in the filtering resolver's query logs and call that coverage. It cannot be: the control only records the clients that chose to ask it. A host resolving through the ISP, through a public encrypted endpoint, or through an implant's own resolver produces exactly one artefact in your logs — nothing. Counting the hosts that asked you tells you the size of the population you can see, and says nothing about the size of the population you cannot. This is the vantage problem the whole leaf turns on, and it is the reason the question is asked of the person who owns endpoint resolver policy: they will be asked, by an auditor or by their own management, what share of the estate resolves through the sanctioned path, and "all of it, we deployed the filter" is not an answer anyone can defend. ## Source one: who did ask Start with the querying sources over a window — a week is usually enough to catch machines that are only used occasionally. Join them to the asset inventory. Three buckets fall out: assets that asked (seen), assets that did not (unknown), and querying sources that are not in the inventory at all (which is a separate and often more interesting finding at a branch with unmanaged devices). The join is where the branch estate bites. If the site NATs to a single public address, every query arrives from one source and per-host attribution from the resolver alone is impossible. You then need either records that preserve the internal address, or a DHCP-lease and inventory join by time, which carries real error, or you push the identity down to the endpoint — see source three. ## Source two: who asked somebody else Egress records at the branch router recover part of the unknown bucket: - sessions to port 53 at any destination that is not your resolver — clear-text resolution elsewhere; - sessions to port 853 — DNS-over-TLS to some provider; - sessions to 443 at addresses on your list of known public DoH providers — suggestive, not conclusive. Note precisely what these evidence. A flow record carries the five-tuple, byte and packet counts and timestamps, and no payload at all, so it shows that a host talked to a resolver-shaped destination and never which names it asked for. And an endpoint the adversary hosts on 443 at an unlisted name produces a record that looks like any other web session, so this source has a hard ceiling. ## Source three: ask the endpoint to prove it The strongest per-device evidence is an active check. Publish a name that only your filtering resolver can answer — an internal-only zone, or a name your filter synthesises — and have a managed endpoint resolve it on a schedule and report the result with its own identity attached. A successful resolution proves that this process, on this host, at this moment, reached the sanctioned resolver. It survives NAT, because identity comes from the endpoint rather than the source address. Its limit is the same precision that makes it useful: it proves the path for the process that ran the check. On a host where the browser holds its own DoH configuration, the operating-system stub and the browser can disagree completely, and a check running as a system service will report a healthy path while the browser resolves elsewhere. If the browser population is what you care about, the check has to run where the browser resolves. ## What you can honestly claim Put the three together and you can state: this many inventory assets were observed resolving through the filter in the window; this many were observed resolving elsewhere, by destination; this many were not observed at all, and here is why that number is not zero. That is a floor on coverage with a named residual, and it is defensible. What you cannot claim is a percentage of resolution, because software that brings its own resolver and its own destination is invisible to all three sources by construction — which is the same reason the filter did not stop it. ## The failure to avoid The common mistake is reporting the absence of an event as compliance: "no blocked queries from that segment" or "no host in that VLAN appears with a bad name". Absence from a control that only sees what it is asked is consistent with a clean host, a switched-off host, a host resolving elsewhere, and a broken log pipeline. Splitting those apart is the entire job here, and it takes a second source every time.

  • The branch NATs behind one address, so every query arrives from the same source. How do you attribute?
    Not from the resolver alone. Either capture records that preserve the internal address before translation, or join by time to DHCP leases and inventory and accept the error that introduces, or move the evidence to the endpoint: have a managed agent resolve a name only your filter answers and report the result with the device identity attached. The last one is the only option that survives translation cleanly.
  • A host has never appeared in the filtering resolver's logs. What can you conclude?
    Nothing, until you add a second source. It is equally consistent with a machine that is switched off, one resolving through another resolver, one whose queries are all answered from a local cache, and a logging path that quietly stopped delivering. Each has a different remedy, so the useful move is to check the egress records for that address and, if it is a managed device, run the active check against it.
  • Your active check reports a healthy path on a laptop whose browser is resolving through a public endpoint. Is the check wrong?
    No, it is answering a narrower question than the one being asked of it. It proves that the process running the check reached the sanctioned resolver, and applications resolve independently, so a system-level check says nothing about the browser. If browser resolution is what you are reporting on, the check must exercise the same resolver the browser uses, or be paired with a policy read of the browser's own configuration.

saying these in an interview costs you the question

  • Reports coverage from resolver query logs alone
  • Treats absence of queries as proof of compliance
  • Assumes one process's resolver represents the whole host
  • Counts a NATed branch as a single endpoint
  • Claims a percentage of resolution rather than a floor

context