skip to content

What does Syft do when you point it at a container image, and what does it not report?

level: juniorimportance: must knowfreq 68%

answer

  1. inventory, not risk
  2. layers unpacked, then parsed
  3. one parser per ecosystem
  4. no metadata, no component
  5. SPDX or CycloneDX on request

basics

~10 s

Syft unpacks the image's layers, runs per-ecosystem catalogers over the resulting filesystem to list installed packages, then serialises that list as SPDX or CycloneDX. It reports what is present, not whether anything is vulnerable.

solid answer

~40 s

Syft resolves the image reference, pulls it, and unpacks the layers into a filesystem view. Over that view it runs catalogers, one per ecosystem: the `dpkg`/`apk`/`rpm` databases for OS packages, `node_modules` and lockfiles for npm, `dist-info`/`egg-info` for Python, `.gemspec` files for RubyGems, jar manifests and `pom.properties` for Java, embedded build info for Go binaries. Each hit becomes a component with a name, version, a package URL and the file locations that evidenced it, and the whole set is written out as a table by default or as `spdx-json`/`cyclonedx-json` on request. What it does not do: match those components against advisory data, say how the image was built, say who vouches for it, or see anything that left no package metadata behind, such as a binary downloaded into `/usr/local/bin` during the build.

go deeper

for a junior

Be ready to say plainly what comes out of the run: a list of components with versions and identifiers, in SPDX or CycloneDX, and that vulnerabilities are a separate step.

for a middle

Explain the mechanics: layers unpacked into a filesystem view, per-ecosystem catalogers matched against evidence files, purl assigned per component. Name two kinds of evidence.

for a senior

Demonstrate that you reason about the blind spots before you trust the document, and that you know which artifact you scanned - image, source tree or build stage - changes the answer.

for a principal

Own the framing that an inventory is only as good as the evidence it was derived from, and that publishing one creates an obligation to keep it attached to a specific artifact rather than a project.

## The job in one line An SBOM generator answers one question about one artifact: **what is inside it**. Syft is the common example for container images and filesystem trees. It does not answer *how it came to be* (that is provenance) or *who vouches for it* (that is a signature). Mixing those three up is the canonical wrong answer in this domain, and an interviewer will listen for it. ## What the run actually does **1. Resolve and fetch the input.** You give Syft an image reference, a local OCI layout, or a directory. For a registry reference it pulls the manifest and the layer blobs; nothing is executed, and the image is never started. This matters: the inventory is a *static* read of files, not an observation of a running process. **2. Build a filesystem view.** Layers are applied in order to produce the filesystem the container would see. Syft's `--scope` controls this: the default squashed view is the final filesystem, while an all-layers view also includes files that a later layer deleted. A package removed in a cleanup layer is absent from one view and present in the other — which is the first reason two people's SBOMs of "the same image" can disagree. **3. Run catalogers.** A cataloger is a small parser that knows one ecosystem's evidence: | Evidence found on disk | Component it yields | | --- | --- | | `/var/lib/dpkg/status`, `apk` db, rpm db | OS packages | | `node_modules/*/package.json`, lockfiles | npm packages | | `*.dist-info/METADATA`, `egg-info` | Python packages | | `specifications/*.gemspec` | Ruby gems | | jar/war `MANIFEST.MF`, `pom.properties` | Java libraries | | Go binaries' embedded build info | Go modules | Which catalogers fire depends entirely on what evidence exists. Nothing declares "this is a Ruby image"; the Ruby cataloger simply finds gemspecs or it does not. **4. Emit identifiers.** Each component carries a package URL (`purl`) such as `pkg:gem/[email protected]` or `pkg:deb/debian/[email protected]~deb12u2?arch=amd64`, plus evidence of where it was found. A CPE may also be attached, sometimes synthesised by pattern rather than looked up — useful for matching against older advisory data, but not authoritative. **5. Serialise.** Default output is a human table; `-o spdx-json` and `-o cyclonedx-json` produce the machine formats, and `-o syft-json` produces the tool's own richer model. ## What it does not report - **Vulnerabilities.** Inventory and vulnerability matching are separate steps: something else takes the component list and matches it against advisory ranges. An SBOM that is a year old lists the same components while the advisories about them change daily. - **Exploitability or reachability.** Even after matching, the document says nothing about whether the vulnerable code is called or whether an attacker can reach it. - **Provenance.** Which builder produced the image, from which source revision, under which parameters — none of that is in an SBOM. - **Anything without package metadata.** A statically linked C library, source vendored into the tree, a tarball extracted into `/opt`, a binary fetched with `curl` in a build step: no dpkg entry, no gemspec, no component. The document will not flag the gap; the component simply is not there. ## Why juniors get asked this Because the failure that follows is expensive: teams publish an SBOM, tick the box, and believe they have coverage of things the generator never saw. Being able to say *what evidence the tool relies on* is what separates "I ran the command" from "I know what the output means".

  • You run Syft against an image and again against that app's source checkout and get very different component counts. Why?
    Different catalogers fire because the evidence differs. The source checkout has manifests and lockfiles, so you get declared dependencies including dev-only ones, and no OS layer at all. The image has an OS package database and only the runtime dependencies that survived installation, plus anything the base image contributed. Neither is wrong; they describe different artifacts, and only the image inventory describes what you actually ship.
  • Does an SBOM from Syft tell you whether the image is exploitable?
    No. It is an inventory. Turning it into risk needs a second step that matches components against advisory data, and even a confirmed match says only that a vulnerable version is present. Whether the vulnerable code path is reachable, and whether an attacker can trigger it in your deployment, are further questions the document does not address.
  • Why does the output list a package URL and sometimes a CPE for the same component?
    A package URL identifies a component precisely within its ecosystem, so `pkg:npm/[email protected]` is unambiguous. CPE is the older identifier style used by much advisory data, and generators often synthesise candidate CPEs by pattern rather than resolving them. That is why matching against CPE-keyed data produces both misses and false hits, while purl-keyed matching is far tighter.

It is a stocktake of a warehouse done by reading the labels on the shelves. Anything someone carried in without a label is simply not on the list.

saying these in an interview costs you the question

  • Says the generator scans for CVEs itself
  • Assumes every file on disk becomes a component
  • Calls the SBOM proof the image was not tampered with
  • Confuses the inventory with a provenance statement
  • Thinks an image SBOM includes the repo's dev dependencies

context