Why does transitive depth dominate a dependency scanner's finding count?
answer
- count the closure, not the manifest
- two direct imports, hundreds of modules
- deduplicate rows into distinct components
- depth predicts volume, not severity
- the fix path runs through the parent
basics
~10 sFindings are counted against the resolved graph, not the manifest. Each direct dependency drags in its own closure, so a two-import project can resolve to hundreds of modules. Depth predicts volume, not severity.
solid answer
~50 sTake a Go log-shipper CLI that customers install on their own servers: the manifest has two direct imports, the resolved module graph has around four hundred modules, and most of the scanner queue sits three or four levels down. That is the normal shape. Scanners match advisories against every resolved component, and each direct dependency contributes its whole closure, so the count tracks the graph rather than the manifest. Two consequences follow. First, depth is a volume signal, not a risk signal - a deep package is not safer, it is just more numerous. Second, depth changes who can act: you usually cannot bump a package you never declared, so the fix path runs through the intermediate maintainer's release, which makes time-to-fix a distribution rather than a number. Report distinct affected components, not raw rows.
go deeper
Know that the packages you declared are a small fraction of what gets installed, and that a scanner reports against everything that resolved, not against the file you edited.
Be able to explain why the closure compounds with depth, why rows and components are different numbers, and why a fix for a deep package usually means bumping the parent that pulled it in.
Show how you turn a large queue into work: deduplicate to components, split deployed from build-only, batch by parent bump, and set metrics on decisions made rather than on findings closed, since upstream controls part of the clock.
Own the reporting contract. Decide which number the organisation commits to externally, given that a chunk of remediation time belongs to maintainers you do not employ, and defend it when a customer asks for a fixed remediation window.
### The counting unit is the resolved graph A scanner does not read your manifest and stop. It matches advisories against the **resolved** set of components - everything the resolution actually produced. Your manifest is the seed; the closure is the inventory. This is why the two numbers people quote about a project are almost never the same number: 'we have twelve dependencies' describes a file, and 'we have four hundred components' describes what was installed. The compounding is mechanical. Each direct dependency brings its own direct dependencies, which bring theirs. Even modest fan-out compounds quickly over a few levels, and the deeper levels contain far more distinct packages than the top one. So the *population* of the queue is dominated by depth, purely by arithmetic. A useful worked shape: a Go module log-shipper distributed as a CLI that customers install on their own servers, with two direct imports and roughly four hundred modules in the resolved graph. Nearly the whole scanner queue for that binary is packages the authors never typed. ### Rows are not components Before drawing any conclusion from a queue size, deduplicate. The same vulnerable package reached by six different paths is one component to upgrade, not six findings, and one package with four advisories is one upgrade too. Teams that report raw rows get numbers that swing wildly for reasons that have nothing to do with risk - a new service onboarded, a scanner changing how it enumerates paths - and lose the ability to say whether things are getting better. Report: - distinct affected components, - distinct advisories, - how many are in the deployed artifact versus build or test only, - age of the oldest unresolved one. ### Depth does not mean safe, and it does not mean dangerous The common wrong instinct is to sort by depth and treat level four as background noise. Depth is a statement about how the package entered your project, not about what it does. A deep package can be the exact library parsing untrusted input on the request path, because a thin wrapper you depend on may be delegating all of the real work downward. Equally, depth does not make something scarier. Severity comes from the flaw, presence comes from scope, and reachability comes from what executes; depth contributes to none of the three. ### The resolved graph can also over-state what ships The graph is a superset of what runs, and by different amounts in different ecosystems. In Go, the module graph includes modules needed to build the tests of packages you import, while the linker includes only packages actually imported transitively by the binary you built. A module can appear in the requirements and contribute nothing to the shipped binary. Scanning at module granularity therefore over-reports for a compiled binary, and analysis of the built artifact itself gives a tighter answer. In interpreted ecosystems, the opposite exposure applies: what is installed on disk is present and loadable whether you import it or not. ### Depth changes who has to act This is the practically important consequence and the one interviewers push on. You do not usually own a package four levels down. The clean fix is that the intermediate maintainer releases a version depending on the fixed one and you take the parent bump. That is out of your control, so: - Time-to-fix becomes a distribution. Some findings close in an hour with a parent bump; some sit until an upstream release exists; a few require your ecosystem's override mechanism, which is a resolver decision with its own trade-offs. - The right metric is not 'time from advisory to zero findings' but something you actually control, such as time from advisory to a decision recorded on each affected service. - Batching matters. Bumping one parent frequently closes several deep findings at once, so scheduling by parent rather than by finding is what shrinks the queue. ### What to say in an interview Get the sequence right: the count is a property of the resolved graph; deduplicate to components; separate what is in the deployed artifact from what is build- or test-only; then rank what remains by severity, exposure and what the affected code actually does. Depth is a fact about the shape of the graph, and it belongs in your capacity planning, not in your risk ranking.
- Your queue shows 900 rows but only 60 distinct packages. Which number do you report, and why?Report the 60, alongside the count of distinct advisories and how many sit in the deployed artifact. Rows count paths and advisory duplicates, so they move for reasons unrelated to risk and are not actionable; components map one-to-one onto upgrades, which is the unit of work and the unit of progress.
- A module appears in a Go project's requirements but its packages are never imported. Is it in the shipped binary?No. The module graph includes requirements needed to build tests of imported packages, while the linker only includes packages that are transitively imported by what you built. Module-level scanning therefore over-reports for a compiled binary, and analysing the built artifact narrows the queue with evidence rather than assertion.
- Why can 'upgrade the direct parent' be unavailable for a deep finding?Because there may be no release yet that depends on the fixed version, the parent may be unmaintained, or the fix may only exist behind a major version the parent cannot adopt without breaking its own consumers. That is when teams reach for an ecosystem override, accepting a resolution the maintainers never tested.
saying these in an interview costs you the question
- Treats the manifest as the dependency inventory
- Sorts by depth and calls deep findings background noise
- Reports raw finding rows instead of distinct affected components
- Assumes every resolved module ends up in the shipped artifact
- Promises a fix SLA that depends entirely on upstream releasing