A security team wants to know, for every service in a company's portfolio, whether any dependency anywhere in the resolved graph carries a copyleft license (like GPL) or a known CVE. Why is this materially harder than just checking each service's manifest file, and what has to be built to answer it reliably?
answer
- manifest = curated subset, not what ships
- lockfile has full resolved graph
- need per-node license+CVE metadata
- SBOM aggregates across services
- reachability != presence
basics
~20 sThe manifest only lists the top-level libraries a team chose - it says nothing about the hundreds of other libraries those libraries secretly depend on. To really check licenses or vulnerabilities you have to look at the entire expanded graph, not just the short list a human wrote.
solid answer
~40 sManifests only capture direct dependencies with version ranges, not the fully resolved set of concrete packages actually shipped. A CVE or a restrictive license can live several hops deep in a transitive dependency that no one on the team chose or even knows exists, so manifest-only scanning produces false negatives. Reliable auditing requires generating the fully resolved graph (from the lockfile, which pins exact versions) for every service, then scanning every node - not just declared ones - against vulnerability databases and license metadata, aggregated into something like a software bill of materials (SBOM) so the whole portfolio can be queried centrally instead of re-derived per service on demand.
go deeper
Should recognize that a project uses more code than what's listed in its manifest, so problems can hide in packages nobody named.
Should know that the lockfile, not the manifest, is the source of truth for what's actually installed, and that tools exist to scan it.
Should describe the full pipeline - lockfile to per-node metadata to SBOM to aggregated scanning - and the operational cost of keeping it current.
Should reason about portfolio-scale tooling investment, reachability-based triage prioritization, and the organizational trade-off between audit thoroughness and the resulting finding-volume/triage burden.
## Why a manifest audit is a sample, not a census Checking a manifest file for licensing or vulnerability problems only tells you about the packages a team consciously chose - the direct dependencies with their name and a version range. That is a small, human-curated list, typically single digits to low dozens of entries. But what actually gets compiled into a binary, bundled into a JS build, or loaded at runtime is the fully resolved transitive graph: every direct dependency's own dependencies, and theirs, recursively, often numbering in the hundreds. A manifest-only audit is therefore auditing a small, cherry-picked sample of what's actually running, and anything problematic that lives purely in the transitive portion of the graph - which, by volume, is most of it - is invisible to that approach. This is the core reason manifest-level scanning produces **false negatives**: it can report "clean" for a service that is, in reality, shipping a GPL-licensed utility three hops down, or a package with a disclosed critical CVE that nobody on the team has ever directly interacted with. ## What reliable auditing needs instead To audit reliably you need three things the manifest alone doesn't give you. 1. **First, the fully resolved graph, not the declared one** - this comes from the lockfile (`package-lock.json`, Gradle's resolved configuration, `Cargo.lock`, etc.), which records every concrete package-and-version actually selected during resolution, including every transitive node. 2. **Second, metadata for every node in that graph, not just the top-level ones**: license identifiers, publish provenance, and known-vulnerability records, which typically means cross-referencing each resolved package/version pair against a vulnerability database and a license-detection tool that can read the license declared in each package's own metadata (which is itself sometimes wrong, ambiguous, or missing, compounding the difficulty). 3. **Third**, because this has to happen per-service and the same transitive package can appear differently pinned across dozens of services in a portfolio, you need **aggregation** - a way to ask "across our entire portfolio, does anything anywhere depend on vulnerable-package X at version Y" without re-resolving every service's graph by hand each time. ## Where the SBOM fits This is exactly the problem a **software bill of materials (SBOM)** is built to solve: it's a structured, machine-readable manifest of the actual fully resolved graph - typically in a standard format like CycloneDX or SPDX - generated once per build from the lockfile, that downstream tooling can scan and diff without re-running dependency resolution. Generating one SBOM per service and feeding all of them into a central scanner turns "walk 40 different services' transitive graphs by hand" into "query one aggregated index," which is the only way this scales past a handful of services. ## What each choice costs The trade-offs here run in both directions. - **Doing full transitive-graph auditing costs real infrastructure**: SBOM generation has to run in CI on every build (because the resolved graph can change between builds even without a manifest change, if version ranges are loose or a re-resolution picks up a newer transitive version), vulnerability and license databases have to be kept current, and someone has to own triage for the inevitable flood of findings in deeply transitive packages that may not even be reachable at runtime (a CVE in a code path that's never actually invoked is a real finding but a much lower-priority one, and determining reachability is its own hard, separate problem). - **Skipping it** and relying on manifest-only or even top-level-lockfile-only spot checks is cheaper but systematically blind to the majority of the graph, which is precisely where risk accumulates because those are the packages nobody chose and nobody is watching. ## How this fails in production The failure mode this produces in production is well documented: license and security incidents routinely originate several hops into the transitive graph, not from a team's direct choices. - A widely cited example is the 2018 `event-stream` npm incident, where a popular package's maintainer handed off control to an unknown party who added a small, targeted malicious dependency; because `event-stream` was itself often a transitive dependency of other tools rather than something projects declared directly, most affected teams had no direct-dependency signal at all pointing to the problem - only full-graph scanning could surface it. - Similarly, license-compliance incidents (a company accidentally shipping GPL code inside a proprietary product) very often trace back to a small utility package buried three or four levels deep that nobody on the engineering team ever reviewed, because reviewing a direct dependency's license is standard practice but reviewing every transitive one by hand simply doesn't scale without tooling that does the graph traversal automatically.
- Why can two builds of the same service, with an unchanged manifest, produce different SBOMs?If the manifest uses loose version ranges and there's no lockfile (or the lockfile gets regenerated), a newly published transitive package version can be picked up between builds, changing the resolved graph and therefore the SBOM even though nothing in the manifest itself changed. This is why CI should generate the SBOM from the actual lockfile used in that build, not re-derive it from the manifest.
- What does 'reachability' mean in this context, and why does it matter for prioritizing a vulnerability finding?Reachability is whether the vulnerable code path in a transitive package is actually invoked by the application at runtime, as opposed to merely being present in the resolved graph. A CVE in a function your build never calls is a lower operational risk than one in code that executes on every request, so mature vulnerability-management programs try to distinguish 'present' from 'reachable' to prioritize triage, though determining reachability accurately is itself a hard static/dynamic-analysis problem.
- Why isn't scanning just the top-level lockfile entries (ignoring deeper transitive ones) good enough?The lockfile's top-level entries are still only the direct dependencies restated with pinned versions; the vast majority of resolved packages are nested transitive entries deeper in the file, so skipping those misses most of the actual attack surface - the whole point of using the lockfile is to scan every resolved node, not just the ones that happen to appear at the top.
Like inspecting a restaurant by only checking the ingredients on the menu description, while ignoring everything the chef's pre-made sauces and stocks were made from - the food safety risk lives in the whole supply chain, not just the dishes you can read about.
saying these in an interview costs you the question
- Believes scanning the manifest file is sufficient for a real audit
- Doesn't know what an SBOM is or why one would be generated per build
- Assumes the resolved dependency set never changes if the manifest hasn't changed
- Conflates 'a vulnerability is present in the graph' with 'the vulnerable code path is actually executed'
- Has no answer for how findings would be aggregated across many services/repos