Why does an SBOM built from a project's lockfile differ from one built by scanning the shipped image?
answer
- declared intent versus shipped bytes
- pre-build, in-build, post-build, runtime
- a lockfile knows one ecosystem only
- image scans also see OS packages
- multi-stage builds drop the builder stage
basics
~20 sEach vantage observes something different. A lockfile lists what the package manager resolved for one ecosystem before the build; scanning the shipped image lists what actually landed on disk, including operating-system packages and files no manifest ever mentioned.
solid answer
~50 sAn SBOM is only ever a report of what one observation point could see. Parsing a lockfile is a **pre-build** vantage: it reads declared intent for a single ecosystem, so you get exact pinned versions and hashes for application dependencies — and nothing about the base image's OS packages, nothing about native libraries, and it still includes dev and test dependencies that never ship. Scanning the built image is a **post-build** vantage: it inspects real files and package databases in the layers, so it sees the OS half and whatever actually shipped, but it identifies components by evidence rather than by declaration, so anything without a package record or manifest is invisible to it. The two documents describe the same service and legitimately disagree. The useful question in an interview is never "which one is correct" but "what is each one structurally blind to".
go deeper
Be ready to name where the list came from — a lockfile, the build, the built artifact, or a running process — and say one thing each one cannot see. Naming the base image's OS packages as the lockfile's gap is enough at this level.
Explain the mechanics: a lockfile records declarations for one ecosystem, an artifact scan infers components from package databases and embedded metadata. Be able to walk through a multi-stage build and predict which components each vantage reports.
Show that you design for the gap rather than arguing about which document is right. Expect to describe merging a build-time graph with artifact-level analysis, and to explain to a responder how much to trust a given document based on how it was produced.
Own the tradeoff between reach and fidelity. Cheap manifest parsing covers the whole estate immediately and lies quietly; build-time generation is accurate and requires touching every pipeline. Be ready to argue which you buy first and what you tell auditors in the meantime.
### What a vantage point is An SBOM (Software Bill of Materials) is a list of the components inside an artifact. It is not discovered by magic: some process stands at a particular point in the software's life, looks at whatever evidence is available there, and writes down what it can see. That standing point is the *vantage*, and it is the single biggest determinant of what ends up in the document — bigger than the format you chose, bigger than which generator you ran. There are four vantages worth naming. **1. Manifest and lockfile parsing (pre-build).** The generator reads declaration files: the manifest that says what you asked for, and the lockfile that says what the resolver settled on. This is cheap, needs no build, and yields exact versions plus integrity hashes because a lockfile exists precisely to pin them. Its shape is also its limit: it covers exactly one ecosystem's declarations. Operating-system packages in the base image are not declared there. Native libraries linked by a compiler are not declared there. Development and test dependencies *are* declared there and will appear even though they never reach production. And anything copied into the repository by hand is, by definition, not a declaration. **2. Build-time generation.** The generator runs inside the build, with access to the fully resolved dependency graph as the build system actually used it, before that graph is flattened into an opaque artifact. This is the highest-fidelity vantage for the application half: it knows the graph, it knows direct versus transitive, and it can be tied to the exact build that produced the exact artifact. Its cost is reach — you must touch every build, and for anything you do not build yourself you cannot stand here at all. **3. Filesystem and binary analysis of the built artifact (post-build).** The generator opens the image layers, the archive, or the binary and infers components from evidence: OS package databases, embedded manifests inside packaged archives, file names, and version strings. This is the vantage that sees what *actually ships*, including the base image, and it works on artifacts you did not build. But identification here is inference, not declaration. Where evidence is thin, the component silently does not appear. **4. Runtime observation.** Watch a running process and record what it actually loads and opens. This tells you what is live rather than what is present, which is genuinely useful for prioritisation, but it structurally undercounts: a code path never exercised during the observation window looks like it is not there, and by definition you only get this after deployment, which is too late to be the document you ship with the release. ### Why the two documents in the question disagree Put a concrete case beside it. A Python model-training service is built in a multi-stage container build: the builder stage installs a compiler toolchain and downloads wheels; the final stage copies in only the installed package tree and the application code. - The lockfile SBOM lists every Python package the resolver chose, including the test-only ones, and nothing else. - The image SBOM lists the base image's OS packages, plus the Python packages it can still find records for in the final layers, and *not* the compilers — because those stayed in the builder stage and were never copied forward. - A build-time SBOM taken in the builder stage would list the compilers, which is accurate for that stage and misleading as a description of what you deployed. None of these is a bug. Each is a faithful report from where it stood. ### What to do with that First, know which vantage produced any document you are handed, and read its blind spots off that fact rather than trusting the component count. Both major SBOM formats carry creation metadata naming what produced the document, and a mature process records the vantage explicitly. Second, treat vantages as complementary rather than competing. The common production pairing is build-time generation for the application dependency graph plus artifact analysis for the OS and packaged layers, merged and attached to the release. Runtime observation is a prioritisation input on top, not a replacement. Third, expect the counts to differ and be able to explain the difference. "The lockfile document has 340 entries and the image document has 190" is not evidence of a broken generator — it is dev dependencies plus a multi-stage build behaving exactly as designed.
- Your lockfile-based document lists more components than your image scan. Is that a defect?Usually not. A lockfile includes development and test dependencies that are excluded from the runtime image, and a multi-stage build discards builder-stage tooling. Both are expected. It becomes a defect only when a component you know ships is missing from the artifact-level document — that points at a real evidence gap, such as a bundled or statically linked component the scan could not identify.
- Where does runtime observation fit alongside these two?It answers a different question: not what is present, but what is actually loaded. That is valuable for prioritising which findings to chase, because a component never loaded is a weaker candidate for urgent work. It is a poor SBOM of record though — it only exists after deployment, and any code path not exercised during the observation window is silently absent.
- If you can only run one, which vantage do you attach to the released artifact?Build-time generation, because it can be tied to the exact build that produced the exact artifact and it sees the resolved graph before it is flattened. In practice you supplement it with artifact analysis so the base-image and OS layer is covered too, since the build's own view of application dependencies says nothing about them.
A packing list written before you pack, an inventory taken of the sealed box, and a note of which items you actually used on the trip all describe the same luggage and none of them match.
saying these in an interview costs you the question
- Assumes any generator over any input yields the same list
- Thinks a lockfile document covers the base image's OS packages
- Expects build-only compilers to appear in the shipped image
- Calls a component-count difference an automatic generator bug
- Treats an SBOM as a statement of exploitability rather than contents