What does an SCA scanner actually match on when it flags a vulnerable dependency?
answer
- it never reads your source
- package identity plus a version range
- resolved version, not declared range
- presence is not exposure
basics
~20 sIt matches each resolved package version in your dependency tree against an advisory's affected version range for that same package identity. That is version arithmetic. It never inspects whether your code calls the vulnerable function.
solid answer
~40 sSoftware composition analysis resolves your dependency tree to concrete components — ecosystem, name and version, often expressed as a package URL such as `pkg:npm/[email protected]` — and asks one question per component: does this version fall inside a range that some advisory marks as affected for this package? If yes, it emits a finding. Everything else is inference the scanner did not make: it does not know whether the vulnerable function is ever called, whether any attacker-controlled data reaches it, or whether your platform already blocks the attack. So a finding means "a known-vulnerable component is present", not "this application is vulnerable". That gap is exactly why a triage step exists, and why a queue of three hundred criticals is mostly a statement about presence rather than exposure.
code
json · 12 lines{
"advisory": {
"package": "pkg:npm/example-parser",
"affected": [{ "introduced": "0.0.0", "fixed": "4.2.1" }]
},
"resolved": {
"name": "example-parser",
"version": "4.1.9",
"pulledInBy": "some-framework"
}
...
}go deeper
Be ready to say in one sentence that the scanner compares a resolved package version against an advisory's affected range, and that this is not the same as your application being exploitable.
Explain the mechanics: how the graph is resolved, purl versus CPE identity, and the concrete false-positive and false-negative sources such as distribution backports and vendored copies.
Show you design around the limits — scan the resolved artifact, keep the matched identity in the record, and treat the queue as an inventory statement that still needs a triage pass before anyone is paged.
Own the framing with leadership: finding counts track dependency growth, not risk, so pick metrics that survive that fact and set expectations before a board slide turns a presence count into an incident count.
## The mechanic, stated plainly A software composition analysis (SCA) scanner does two things. First it builds an **inventory**: it resolves your project's dependency graph — direct and transitive — down to concrete components, each identified by ecosystem, name and version. That identity is commonly written as a package URL (purl), e.g. `pkg:maven/org.example/[email protected]`. Second it consults an **advisory database** (public advisory feeds carrying CVE, GHSA or OSV entries, sometimes plus a vendor's own data). Each advisory says, for one package identity, which version ranges are affected and which version fixed the flaw. The match is then trivial: for every resolved component, does its version fall inside an affected range for that package? If yes, emit a finding. That is the whole of it. ## The two identifier schemes, and why one is noisier Ecosystem packages (npm, PyPI, Maven Central, crates.io, NuGet) are matched on their native coordinates, which are unambiguous — a name in a registry plus a version. Operating-system and vendor components are frequently matched on CPE, a vendor/product/version string. CPE matching is fuzzier: two products can share a name, a vendor string can vary between advisories, and a distribution routinely **backports** a security patch into an older version number. The upstream version string still looks affected even though the shipped binary is fixed, so the scanner reports a match that is not a real exposure. This is one of the largest sources of false positives on container image scans. ## What the match cannot tell you Four separate things, and conflating them is the classic wrong answer in this domain: 1. **Whether the vulnerable code executes.** The flaw may live in a module your application never loads. Version matching has no view of call paths. 2. **Whether an attacker can drive it.** Even code that does execute may only ever process data your own operators control. 3. **Whether the component ships.** A component pulled in for a local toolchain is not in the artifact you deploy, and the scanner's report will happily list it beside the ones that are. 4. **What the deployment looks like.** Network position, authentication in front of the endpoint, a runtime guard — none of that is in the version comparison. So the direction of the claim matters: **presence of a known-vulnerable component is evidence, not a verdict.** ## Where the match is also wrong in the other direction False negatives are less discussed and worth naming in an interview: - **Components installed outside the package manager.** A binary copied into an image, a JAR downloaded by a build step, a statically linked library — no manifest entry, no match. - **Vendored or shaded copies.** Source pasted into your repository or a dependency repackaged inside another artifact keeps the flaw and loses the version string that would have matched it. - **Advisory lag.** A flaw that has no advisory yet, or an advisory whose affected range was published incompletely, produces no finding at all. - **Resolution drift.** If the scanner reads the declared range rather than what the build actually resolved, it can analyse a different version from the one you shipped. Scanning a lockfile or the built artifact, not just the manifest, is what closes this. ## Why this is the first question of triage Once you accept that the scanner performed a version comparison, the rest of the triage discipline follows naturally. The scanner has told you *what is inside*. The next questions — is the vulnerable code ever executed, and can anyone hostile drive it — are questions about *your application*, and no advisory database contains their answers. A candidate who can state that boundary crisply is already ahead of one who reads the queue as a list of live exploits. ## What good practice looks like Scan the resolved graph (a lockfile or the built artifact), not the declared ranges. Record the component identity you matched on, because a finding without a resolved version and a source is unreviewable later. Expect and account for backport false positives on OS packages. And treat the count of findings as a measure of inventory size, not of risk — the number goes up when you add dependencies and goes down when you delete them, which is not the same as being safer.
- Why does scanning the lockfile or built artifact beat scanning the manifest?A manifest declares ranges; the build resolves them to exact versions, and transitive dependencies rarely appear in the manifest at all. Scanning declared ranges can analyse a version you never shipped, in either direction. The lockfile or the built artifact tells you what actually resolved, which is the only thing an advisory range can be honestly compared against.
- Your image scan reports a critical in an OS package the distribution says it patched. What happened?The distribution backported the fix into its own build while keeping the upstream version number, so a version-range match still fires. The shipped binary is fixed; the identifier is not. Confirm against the distribution's own advisory data for that package build, record the evidence, and prefer scanners that consume distro advisory feeds rather than upstream ranges alone.
- Does a clean scan mean the artifact has no known vulnerabilities?No. It means nothing in the resolved inventory matched a published advisory range. Components installed outside the package manager, vendored or shaded copies, and flaws with no advisory yet all produce silence rather than a finding. A clean report is bounded by the inventory's completeness and the advisory feed's freshness.
It is a librarian checking whether any book on your shelf appears on a recall list. It never opens the book to see whether you ever read the recalled chapter.
saying these in an interview costs you the question
- Says a flagged dependency means the application is exploitable
- Thinks the scanner analyses source code call paths
- Treats the finding count as a measure of risk
- Assumes a clean scan proves the artifact has no known flaws
- Believes a distribution backport removes the version match