skip to content

What is an SBOM (Software Bill of Materials), and how does SBOM-based scanning differ from pointing a scanner directly at a running or stored image?

level: middleimportance: should knowfreq 38%

answer

  1. SBOM = machine-readable ingredient list (SPDX / CycloneDX)
  2. generators: Syft, Trivy, BuildKit --sbom
  3. split inventory (build-time, once) from matching (anytime)
  4. re-scan stored SBOM against fresh feeds, no image needed
  5. sign/attest for provenance (SLSA); regenerate per build

basics

~20 s

An SBOM is a machine-readable inventory of every component and version in an image (formats: SPDX, CycloneDX). You generate it once at build time; then you can re-scan that list against fresh CVE feeds anytime without the image, and detect newly-disclosed CVEs in components you already shipped. Direct scanning re-inventories the image each run.

solid answer

~60 s

An **SBOM** is a structured, machine-readable list of everything an image contains: each component, its version, and often its license and origin. The two dominant formats are **SPDX** and **CycloneDX**, and tools like Syft, Trivy, and Docker Scout can produce one. Direct scanning does inventory-plus-match in a single pass over the image. **SBOM-based scanning splits those steps.** You generate the SBOM **once**, at build time when you have the full context (build tools, lockfiles, exact layers), store it as an artifact alongside the image, and then run CVE matching against the **SBOM** whenever you like. The payoff: - **Re-scan without the image.** New CVEs are disclosed daily. You can re-match yesterday's SBOM against today's feed and learn that an image you shipped last month now has a known-critical flaw, without pulling or rebuilding it. - **Consistency and provenance.** The inventory is captured at build time, so it reflects exactly what shipped, and can be signed/attested for supply-chain integrity. - **Speed and reuse.** Inventory is the expensive part; do it once, match many times.

code

bash · 8 lines
bash
# Generate an SBOM at build time
syft myapp:1.4.0 -o cyclonedx-json > sbom.json
# or with Trivy
trivy image --format cyclonedx --output sbom.json myapp:1.4.0

# Later, re-match the stored SBOM against today's CVE feed (no image needed)
grype sbom:sbom.json
trivy sbom sbom.json

go deeper

for a junior

Know an SBOM is a machine-readable list of what is in the image, in SPDX or CycloneDX format.

for a middle

Explain the inventory-vs-match split and why re-scanning stored SBOMs catches newly-disclosed CVEs without the image.

for a senior

Cover build-time accuracy, attestation/provenance, staleness risks, and running both SBOM and live scans.

for a principal

Position SBOMs within a supply-chain program (SLSA/provenance, fleet-wide continuous re-matching, compliance), not just a scan input.

## What an SBOM is A **Software Bill of Materials** is to software what an ingredients list is to food: a complete, machine-readable manifest of every component that went into an artifact. For a container image that means each OS package and application dependency, with **name, version, and often license, supplier, and a package URL (purl)**. It is emitted in a standard format so any tool can consume it. The two you must know: - **SPDX** (Linux Foundation / ISO standard), broad, license-focused heritage. - **CycloneDX** (OWASP), security-focused, compact, popular for vuln workflows. Generators include **Syft** (from Anchore), **Trivy** (`trivy image --format cyclonedx`), and **Docker Scout** / BuildKit, which can attach an SBOM as an attestation during `docker buildx build --sbom=true`. ## The key distinction: inventory vs matching Every scan is really two phases: 1. **Inventory** the artifact (walk layers, read package DBs and lockfiles). This is the slow, context-heavy part and is most accurate at **build time**. 2. **Match** the inventory against a vulnerability feed. This is fast and only as fresh as the feed. **Direct image scanning** fuses both every run: it re-inventories the image and matches. **SBOM-based scanning** decouples them: inventory once into an SBOM, then match that SBOM against feeds repeatedly. ## Why the decoupling matters - **Continuous re-scanning of shipped images.** CVEs are disclosed constantly. A component that was 'clean' when you built is not clean forever. With stored SBOMs you can re-run matching every night across your whole fleet and flag images that have *newly* become vulnerable, no rebuild, no registry pull, even for images no longer easy to reproduce. - **Build-time accuracy.** At build you have the richest context: exact lockfiles, transitive resolution, layers before they are squashed or stripped. An SBOM captured then can be more complete than re-deriving inventory from a slim published image. - **Provenance and attestation.** An SBOM can be **signed** and attached as an attestation (e.g. via Sigstore/cosign, or BuildKit SBOM attestations) so downstream consumers get a verifiable statement of what an image contains, a core supply-chain (SLSA) practice and increasingly a compliance/regulatory expectation. - **Interoperability.** One SBOM feeds many tools: vuln scanners, license-compliance checks, policy engines. You are not re-parsing the image for each. ## Tradeoffs and caveats - **Staleness of the SBOM itself.** An SBOM describes the image at build time. If someone mutates a running container or rebuilds the base without regenerating the SBOM, it drifts. Regenerate on every build and tie it to the image digest. - **Only as good as the inventory.** If the SBOM missed a hand-installed binary, every SBOM-based scan inherits that blind spot. Direct scanning of the actual image can occasionally catch things a stale SBOM misses. - **Not a replacement, a complement.** Mature pipelines do both: generate+scan at build, attach the SBOM, and re-scan stored SBOMs continuously, while still periodically scanning live images/registries. ## Practical flow 1. In CI, build the image and generate an SBOM (`syft`, or `trivy --format cyclonedx`, or BuildKit `--sbom`). 2. Scan it immediately and gate on severity (next question). 3. Store/attest the SBOM keyed to the image **digest**. 4. Run a scheduled job that re-matches all stored SBOMs against the latest feed and alerts on new criticals, driving the rebuild cadence (the patch-and-rebuild question). The interview line: an SBOM turns 'what is in this image' into a durable, portable artifact, so vulnerability *matching* can happen anytime, anywhere, against ever-fresh data, independent of the image itself.

  • Why is generating the SBOM at build time better than deriving it from the final published image?
    At build time you have the full context: complete lockfiles, transitive dependency resolution, and layers before they are stripped or squashed. A slim final image may hide or omit metadata, so a build-time SBOM is typically more complete and accurately reflects exactly what shipped, keyed to the image digest.
  • How does an SBOM enable finding vulnerabilities in images you already deployed?
    Because matching is decoupled from the image, a scheduled job re-runs CVE matching over your stored SBOMs against the latest feed every day. When a new CVE is published against a component you shipped weeks ago, it surfaces immediately without pulling or rebuilding the image, driving your rebuild priorities.
  • What are the two common SBOM formats and who backs them?
    SPDX, a Linux Foundation and ISO standard with strong license-compliance heritage, and CycloneDX, an OWASP project focused on security use cases. Most tooling (Syft, Trivy, Scout) can emit either; CycloneDX is common in vulnerability workflows, SPDX in compliance.

saying these in an interview costs you the question

  • Thinking an SBOM itself finds vulnerabilities (it is the inventory; a scanner matches it)
  • Believing one SBOM stays accurate after a rebuild or container mutation
  • Not knowing SPDX and CycloneDX as the standard formats
  • Assuming SBOM-based scanning replaces rather than complements image scanning
  • Forgetting to key/attest the SBOM to the image digest

context