skip to content

Two dependency scanners report 61 and 9 findings for one JVM service - how do you adjudicate?

level: middleimportance: must knowfreq 68%

answer

  1. same inputs before comparing outputs
  2. diff, do not argue
  3. key on the normalised identifier
  4. feed, granularity, duplicates, freshness
  5. more findings is not more coverage

basics

~20 s

Confirm both read the same resolved dependency set, then diff the lists by alias-normalised identifier. Every difference resolves to a feed gap, a coarser range, an un-collapsed duplicate, or stale data. The bigger number is not the safer one.

solid answer

~50 s

First equalise the inputs: both tools must have seen the same resolved graph and the same artifact, otherwise you are comparing two different questions. Then export both lists, key them on the alias-normalised flaw identifier plus package and version, and diff. Almost every delta lands in one of four buckets. One feed carries an advisory the other does not, usually because one reads an ecosystem-native database and the other an enriched central one. The same flaw is filed against a whole module in one feed and the single vulnerable artifact in the other, so a coarse range multiplies one issue into dozens of rows. Aliases are not collapsed, so one flaw is counted once per database. Or one tool's advisory cache is stale. Decide which source is authoritative per ecosystem, record it, and present a normalised union with dispositions rather than either raw count.

go deeper

for a junior

Know that a finding count depends on which advisory feed the tool reads, and that the same flaw can appear under more than one identifier. Do not treat a bigger number as a better scan.

for a middle

Be ready to walk the diff method: equalise inputs, key rows on the normalised identifier, then classify each difference as feed coverage, advisory granularity, a duplicate, a range edge case, or stale data.

for a senior

Demonstrate that you settle this by choosing an authoritative source per ecosystem and recording the decision, and that you hand a release meeting a normalised list with dispositions instead of a vendor headline number.

for a principal

Own the estate-level position: which feeds are canonical, who pays for the normalisation layer, and how you answer an auditor or customer who ran a different tool and got a different number.

A finding count is not a measurement of a service. It is the output of three things: the dependency set that was scanned, the advisory data the tool was reading, and the matching rules it applied. Two numbers as far apart as 61 and 9 almost always come from the second and third, and the way to settle it is a diff, not an argument about which tool is better. ## Step one: equalise the inputs Before comparing outputs, confirm both tools answered the same question. Did both resolve the same dependency set, from the same lockfile state, on the same artifact? A tool reading a build manifest and a tool reading a built artifact will legitimately see different component lists. If the inputs differ, stop - the finding gap is explained by the input gap and the rest of the analysis is noise. ## Step two: normalise and diff Export both result sets. For every row, produce a key: the alias-normalised flaw identifier, plus the package coordinate and the resolved version. Alias normalisation matters because one flaw commonly appears under a central identifier and an ecosystem one, and a tool that reports both without collapsing them doubles its own count. Now diff the two keyed sets. What you want is not a number but a classified list of differences. ## Step three: classify every delta **Feed coverage.** One tool is reading an ecosystem-native database and the other an enriched central one. The ecosystem database often has the package-level record days before central enrichment produces structured applicability data, and it may carry advisories that never got a central identifier at all. Conversely a central record may exist for a component the ecosystem feed does not cover. This bucket is usually the largest one. **Advisory granularity.** The same flaw can be filed at very different resolutions. One database records that a specific module of a large project is affected in a narrow range; another files the same flaw against the whole project, so every artifact published under it matches. If the coarse record touches fifty artifacts in your graph, one flaw becomes fifty rows. Neither database is lying - one is more precise. Note the granularity difference explicitly, because it will recur every release. **Duplicate identifiers.** Beyond aliases, databases sometimes carry two records for one flaw that were never linked, or a record that has since been retracted as a duplicate. These inflate one side only. **Range interpretation.** Ranges are evaluated against a version scheme, and the edges are where tools disagree: pre-release and build-metadata suffixes, an unbounded introduced event with no fixed version yet, or a version string that does not sort the way the range author assumed. Two tools evaluating the same range can land on different sides of the boundary for the same installed version. **Data freshness.** A local advisory database that has not been refreshed produces both false negatives, for advisories published since, and stale positives, for records corrected or retracted since. Always check when each side last updated, before you conclude anything about matching logic. ## Step four: adjudicate and record Once every delta is classified, the decision is straightforward and is about sources, not tools: which database is authoritative for each ecosystem you build in, and which is a secondary safety net. Write it down with the reasoning, because this argument recurs every time a tool is added or a feed changes. Then present the release meeting a normalised union - the deduplicated flaw list with a disposition per item - rather than either vendor's headline number. A count without a source and a normalisation rule attached is not evidence of anything. ## The trap The instinctive read is that 61 is the thorough tool and 9 is the one missing things, so pick 61 and feel safe. Frequently the reverse is closer to the truth: the larger number is one flaw multiplied by a coarse range, plus alias duplicates, plus rows for components that are not in the shipped artifact at all. Meanwhile the smaller list may contain a real finding the larger one missed entirely, because its feed does not cover that ecosystem. Volume is not coverage. The only defensible position is the classified diff, and the only durable fix is to fix the source configuration so the gap does not reopen next quarter.

  • Which number do you take into the release meeting?
    Neither raw count. I take the normalised union - one row per distinct flaw, with the source that reported it and a disposition - plus a short note on why the vendor numbers differ. A raw count with no source or de-duplication rule attached cannot support a release decision.
  • One coarse advisory adds fifty rows for a single flaw. How do you handle that?
    Treat it as one flaw with fifty affected coordinates, not fifty findings. Record that the finer-grained feed is authoritative for that ecosystem, keep the coarse feed as a secondary signal, and make sure the collapsing rule lives in the pipeline rather than in someone's head.
  • How do you stop the same disagreement recurring next quarter?
    Fix the configuration, not the report: pin which advisory source is authoritative per ecosystem, make cache refresh a scheduled step with a visible last-updated timestamp, and keep the cross-tool diff running so a new gap surfaces as a change rather than as a surprise in a release meeting.

saying these in an interview costs you the question

  • Assumes the tool with more findings is more thorough
  • Counts one flaw twice because two databases named it
  • Blames the tool without diffing the two lists
  • Never checks when each advisory cache was refreshed
  • Compares outputs without confirming the same dependency set was scanned

context